Data storage method and related equipment

By modifying the storage nodes associated with the data before data storage, the problem of high data overhead during capacity balancing is solved, thereby improving the performance of the storage system.

CN121635781APending Publication Date: 2026-03-10CHENGDU HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

In storage systems, the data overhead generated during capacity balancing is significant, impacting system performance.

Method used

By modifying the storage nodes associated with the data before data storage, the first data is moved from a node with less remaining storage capacity to a node with more remaining storage capacity, thus avoiding data migration and achieving capacity balancing among storage nodes.

Benefits of technology

It reduces data overhead during capacity balancing and improves the performance of the storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121635781A_ABST
    Figure CN121635781A_ABST
Patent Text Reader

Abstract

The invention provides a data storage method and related equipment, which are applied to the field of communication and can reduce data overhead in a capacity balancing process so as to improve the performance of a storage system. In the method, a first device obtains to-be-stored first data; the first device determines whether a storage node associated with the first data needs to be modified into a second node; when it is determined that modification is required, the first device stores the first data to the second node.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of communication, and in particular, to a data storage method and related equipment. BACKGROUND

[0002] With the rapid development of Internet technology, the size of data to be stored also increases explosively. In order to store a large amount of data, a capacity balancing technology can be used. The capacity balancing technology can, on the one hand, reasonably allocate the size of data stored by different storage nodes. On the other hand, as the overall system design of the storage system is becoming increasingly complex, the capacity balancing technology also plays a crucial role in ensuring the performance and available capacity of the storage system.

[0003] In a storage system, the capacity balancing technology can be used to balance the storage capacity of different storage nodes. For example, in the process of writing stored data, the size of data received by each disk can not be completely consistent, resulting in an unbalanced size of data stored between disks, and ultimately leading to the fact that some disks in the storage system still have a large amount of available space, while some disks are full. At this time, the capacity balancing technology can be used to balance the storage capacity. For example, data A is stored in disk A, and the storage capacity of disk A is full, while the storage capacity of disk B is larger. At this time, the disk storing data A can be adjusted to disk B, thereby improving the first parameter of the storage capacity of disk A and disk B. However, in the process of capacity balancing, a large amount of data overhead is generated, which affects the performance of the storage system.

[0004] Therefore, how to reduce the data overhead in the process of capacity balancing is a technical problem to be solved. SUMMARY

[0005] Embodiments of the present application provide a data storage method and related equipment, which can reduce the data overhead in the process of capacity balancing, thereby improving the performance of the storage system.

[0006] The first aspect of the present application provides a data storage method, which is executed by a first device, or executed by part of components (such as processors, chips or chip systems, etc.) in the first device, or the method can also be implemented by a logic module or software that can realize all or part of the functions of the first device. In the first aspect and its possible implementation manners, the data storage method is taken as an example to be executed by the first device, the first device obtains first data to be stored; the first device determines whether a storage node associated with the first data needs to be modified to a second node; when it is determined that the modification is needed, the first device stores the first data to the second node.

[0007] The first aspect provides a capacity balancing method, which involves changing the storage node associated with the first data from the first node to the second node, where the remaining storage capacity of the second node is greater than that of the first node. After the modification, the first data will be stored on the second node. The principle behind this application's capacity balancing is that when the storage node associated with the first data is changed to the second node, the first data will be stored on the second node, thus reducing the difference in remaining storage capacity between the first and second nodes. It is evident that this method of modifying the associated node can promote capacity balancing among storage nodes. Furthermore, in this application, the modification of the associated node occurs before the data is stored; therefore, this application does not require data migration, such as migrating the first data from the first node to the second node. This reduces the overhead of the capacity balancing process, thereby improving the performance of the storage system.

[0008] Optionally, "the storage node associated with the first data is the first node" can mean that the first data belongs to the first node. In this case, the first node serves as a preparatory node for storing the first data, that is, the first data is prepared to be stored in the first node, but has not yet been stored in the first node.

[0009] Optionally, this application may perform two data ownership determinations, i.e., determine the node associated with the first data twice. The first data ownership determination refers to the software pre-determining the storage node of the data after the first device receives the first data. For example, the software can determine the storage node associated with the first data based on the traffic volume. This first determination is a simple one. At this point, some problems may arise. For example, the first determination may determine that the storage node associated with the first data is the first node, but the storage space in the first node is full. In this case, the first data is not suitable to be stored in the first node. Therefore, a second determination is required. The second determination occurs after the first determination, i.e., "determining whether it is necessary to change the storage node associated with the first data to the second node" in the first aspect mentioned above.

[0010] In one optional implementation of the first aspect, the first device acquires a first dataset, which includes the first data; when a first condition is met, the first device modifies the storage node associated with the first data to a second node, wherein the first condition includes: the number of data in the first dataset with a similarity higher than a first threshold to the first data is less than a second threshold; when the first condition is not met, the first device does not modify the storage node associated with the first data to a second node.

[0011] In the implementation manner, the number of data in the first data set that is similar to the first data and is higher than the first threshold is counted, and the deduplication count of the first data is obtained. The deduplication count is the number of data in the first data set that is highly similar or identical to the first data, so as to determine the number of data in the first data set and the data duplicated with the first data that needs to be deleted. For example, when there are 6 data in the first data set that is similar to the first data and is higher than the first threshold, the deduplication count of the first data is 6. When the deduplication count is less than the second threshold, that is, when the deduplication count is small enough, the modification of the ownership of the first data has little effect on the deduplication of the data. It can be seen that the implementation manner has little effect on the deduplication of the data, and the deduplication characteristics of the data are retained.

[0012] In an optional implementation manner of the first aspect, the first condition further includes that the first parameter of the first node is higher than a third threshold, and the first parameter is related to the size of the data stored by the first node.

[0013] Based on the implementation manner, when the first parameter of the first node is higher than the third threshold, the first device is likely to modify the storage node associated with the first data to the second node, and the first parameter is related to the size of the data stored by the first node, that is, the first device determines the ownership of the first node, that is, the storage node associated with the first node, in combination with the capacity dimension of the node, which is beneficial to improve the utilization rate of each storage node.

[0014] In an optional implementation manner of the first aspect, the first parameter is a ratio of the size of the data stored by the first node to a first average value, and the first average value is an average value of the size of the data stored by all nodes in a first storage system, and the first storage system includes the first node and the second node.

[0015] Based on the implementation manner, the first parameter can determine the balance degree of the first node relative to the first storage system. For example, when the first parameter is too high, for example, 200%, it indicates that the size of the data stored by the first node is much larger than the average size of the data stored by the nodes in the first storage system, and the first node is unbalanced relative to other nodes in the first storage system. By taking the first parameter as an index of the first condition, that is, taking the balance degree as an index of whether to modify the ownership of the first data associated with the first node, that is, as an index of modifying the ownership of part of the data in the first node, the size of the data stored in the first node can be adjusted according to the balance degree of the first node, so as to balance the first node, which is beneficial to balance the first storage system and improve the utilization rate of the storage space of the first storage system.

[0016] In an optional implementation of the first aspect, the first parameter is a ratio of a size of data stored by the first node to a first average value, the first average value being an average value of sizes of data stored by all nodes in a first storage system, the first storage system including the first node and the second node.

[0017] Based on the above implementation, the first parameter can determine the balance degree of the first node relative to the first storage system. For example, when the first parameter is too high, such as 200%, it indicates that the size of data stored by the first node is much larger than the average size of data stored by the nodes in the first storage system, and the first node is unbalanced relative to other nodes in the first storage system. By taking the first parameter as an index of the first condition, i.e., taking the balance degree as an index of whether to modify the ownership of the first data associated with the first node, i.e., as an index of modifying the ownership of part of the data in the first node, the size of the data stored in the first node can be regulated according to the balance degree of the first node, so as to realize the balance of the first node, which is beneficial to the balance of the first storage system and improves the utilization rate of the storage space of the first storage system.

[0018] Optionally, the first parameter of the second node being lower than the third threshold value and the first parameter of the first node being higher than the third threshold value, the first parameter being related to the size of data stored by the first node, can be collectively used as the first condition. The first device can set a weight for the difference between the first parameter of the first node and the third threshold value and a weight for the difference between the deduplication count of the first node and the second threshold value, respectively. If the sum of the two after weighting exceeds a certain threshold value, it is considered that the storage node associated with the first data needs to be modified. That is, the capacity balance is controlled by means of weighted summation.

[0019] In an optional implementation of the first aspect, the first device modifies the storage node associated with the first data to the second node by modifying the first parameter in the metadata of the first data to a second parameter, where the first parameter is used to indicate that the storage location of the first data is the first node, and the second parameter is used to indicate that the storage location of the first data is the second node.

[0020] Based on the above implementation, the first device can modify the information in the metadata for identifying the storage location of the first data by modifying the first parameter in the metadata of the first data to the second parameter, so as to modify the storage location of the first data and modify the ownership of the first data without data migration, so that the first data can be stored in the second node.

[0021] In an optional implementation of the first aspect, when the second condition is met, the first device determines whether the storage node associated with the first data needs to be modified to the second node, the second condition comprising: the first control interface being in an open state, wherein the first control interface is in the open state when an average read-write rate of the first storage system is greater than a fourth threshold value and / or a data amount of the first data set is less than a fifth threshold value, and the first storage system comprises the first node and the second node. When the first condition is not met, the first data is stored to the first node.

[0022] Based on the above implementation, a precondition-second condition is further set before determining whether to modify the node associated with the first data. The second condition can be determined from the read-write rate of the first storage system and the data amount of the obtained data set to determine whether the node associated with the first data needs to be modified, which can reduce the impact on the performance of the first storage system and is also conducive to the capacity balancing of the first storage system.

[0023] In an optional implementation of the first aspect, the second condition further comprises: a size of data stored by the first storage system is greater than a sixth threshold value.

[0024] Based on the above implementation, the second condition is a precondition for determining whether to modify the node associated with the first data. The second condition can be determined from the capacity of the storage system to determine whether to perform the judgment process of modifying the node associated with the first data. Therefore, when the size of the data stored by the first storage system exceeds the sixth threshold value, the capacity balancing of the first storage system can be achieved by adjusting part of the data, such as the node associated with the first data.

[0025] The second aspect of the present application provides a communication device, which comprises a transceiver unit and a processing unit, and is used to execute all or part of the operations of the first aspect. The communication device can be a computing node, a management node device, or a part of a component used to perform related operations, such as a line card, an interface board, etc. It can also be a chip system used to perform related operations, which can include one or more chips. When the communication device is a chip system, the receiving module and the sending module can be, for example, the interface circuit of the chip, and the processing unit can be, for example, the processing circuit of the chip.

[0026] For example, when the method of the first aspect is executed, the transceiver unit is used to obtain first data to be stored, and the storage node associated with the first data is a first node. The processing unit is used to determine whether the storage node associated with the first data needs to be modified to a second node, wherein the remaining storage capacity of the second node is greater than the remaining storage capacity of the first node. When it is determined that the modification is needed, the first data is stored to the second node.

[0027] In an optional implementation of the second aspect, the transceiving unit is configured to: obtain a first data set, the first data set comprising the first data; and the processing unit is configured to: modify the storage node associated with the first data to the second node when a first condition is met, wherein the first condition comprises: a number of data in the first data set that is similar to the first data is less than a second threshold; and not modify the storage node associated with the first data to the second node when the first condition is not met.

[0028] In an optional implementation of the second aspect, the first condition further comprises: a first parameter of the first node is higher than a third threshold, the first parameter being related to a size of data stored in the first node.

[0029] In an optional implementation of the second aspect, the first parameter is a ratio of the size of data stored in the first node to a first average value, the first average value being an average of sizes of data stored in all nodes in a first storage system, the first storage system comprising the first node and the second node.

[0030] In an optional implementation of the second aspect, the first parameter of the second node is lower than the third threshold.

[0031] In an optional implementation of the second aspect, the processing unit is configured to: modify a second parameter in metadata of the first data to a third parameter, wherein the second parameter is used to indicate that the first data is stored in the first node, and the third parameter is used to indicate that the first data is stored in the second node.

[0032] In an optional implementation of the second aspect, the determination of whether the storage node associated with the first data needs to be modified to the second node is performed when a second condition is met, the second condition comprising: the first control interface is in an open state, wherein the first control interface is in the open state when an average read-write speed of a first storage system is greater than a fourth threshold and / or an amount of data in the first data set is less than a fifth threshold, wherein the first storage system comprises the first node and the second node; and the first data is stored in the first node when the second condition is not met.

[0033] In an optional implementation of the second aspect, the second condition further comprises: a size of data stored in the first storage system is greater than a sixth threshold.

[0034] The third aspect of the present application provides a communication system, comprising a first node, a second node, and a first device configured to perform the method in the first aspect or any possible implementation of the first aspect.

[0035] The fourth aspect of the present application provides a communication device, comprising a processor and a communication interface. The processor and the communication interface are configured to perform the method of the first aspect and any possible implementation manner thereof.

[0036] Optionally, the processor is coupled with a memory, for example, the memory is configured to store programs or instructions. The at least one processor is configured to execute the programs or instructions, so that the device implements all or part of the operations in the first aspect and any possible implementation manner thereof.

[0037] The fifth aspect of the present application provides a computer readable storage medium, which stores programs or instructions. When the programs or instructions are run on a processor, the method of the first aspect and any possible implementation manner thereof is executed.

[0038] The sixth aspect of the present application provides a computer program product, comprising programs or instructions. When the programs or instructions are run on a processor, all or part of the operations in the first aspect and any possible implementation manner thereof are implemented.

[0039] In a specific design, the computer program product can be the computer readable storage medium mentioned in the eighth aspect.

[0040] The seventh aspect of the present application provides a chip system, comprising at least one processor, which is configured to support all or part of the functions of the method of the first aspect and any possible implementation manner thereof.

[0041] The technical effects brought by any one of the second aspect to the seventh aspect can be referred to the technical effects brought by the first aspect and any possible implementation manner thereof, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1A An architecture schematic diagram of a distributed system provided by an embodiment of the present application;

[0043] Figure 1B Another architecture schematic diagram of a distributed system provided by an embodiment of the present application;

[0044] Figure 2 An architecture schematic diagram of a centralized system provided by an embodiment of the present application;

[0045] Figure 3 A flow schematic diagram of a data storage method provided by an embodiment of the present application;

[0046] Figure 4 A schematic diagram of the first condition provided by an embodiment of the present application;

[0047] Figure 5 A diagram of the sixth threshold provided for the embodiments of the present application;

[0048] Figure 6 A flow diagram of an application example of the data storage method provided in the embodiments of the present application;

[0049] Figure 7 A structural diagram of a communication device provided for the embodiments of the present application;

[0050] Figure 8 Another structural diagram of a communication device provided for the embodiments of the present application. DETAILED DESCRIPTION

[0051] To make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0052] Some terms related to the embodiments of the present application are explained below.

[0053] 1. De-duplication

[0054] De-duplication, the full name of which is data deduplication, is a data reduction technology. In the storage process, there may be a large amount of duplicate data, which not only occupies valuable transmission bandwidth, but also consumes a large amount of storage space. The emergence of de-duplication technology is to solve this problem.

[0055] The principle of de-duplication is: when the system finds that multiple data blocks are the same, it will only keep one instance of the data, and create a unique identifier for it. For other duplicate data blocks, the system only stores a reference pointing to this unique instance. In this way, even if there is a large amount of duplicate content in the data, the actual storage space occupied will be greatly reduced.

[0056] Block-level de-duplication refers to slicing (also called chunking) data streams or files according to a certain way, performing hash calculation on data blocks, and finding the same data blocks for deletion in this way. Block-level de-duplication is divided into fixed-length de-duplication and variable-length de-duplication.

[0057] Deduplication techniques can be divided into two types: fixed-length deduplication and offline deduplication. Fixed-length deduplication involves slicing the data to a fixed length, performing a hash calculation on each slice, and then writing the data. Non-duplicate data is written separately, while duplicate data is simply written as a reference. Variable-length deduplication works by using a content-based variable-length data segmentation algorithm. This algorithm intelligently identifies modified and unmodified data, avoiding the problem of unmodified data being split into new data blocks due to data displacement caused by modifications. This technique maximizes deduplication performance and deduplication rate, aiming to avoid backing up any redundant data.

[0058] 2. In the embodiments of this application, information, data and data stream can be interchanged with each other. Information, data and data stream are exemplary names and can be replaced with any possible names, such as message, signaling, data packet, data block or information stream, etc.

[0059] 3. In this application's embodiments, "including" means "including but not limited to." When A includes multiple elements or situations, A can be one or more of those elements or situations. For example, if A includes B or C, then A can be B, A can be C, and A can also be B and C. "At least one" means one or more, and "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, "at least one of A, B, and C" includes A, B, C, AB, AC, BC, or ABC. Furthermore, unless otherwise specified, the ordinal numbers such as "first" and "second" mentioned in the embodiments of this application are used to distinguish multiple objects and are not used to limit the order, sequence, priority or importance of multiple objects.

[0060] The system architecture on which the embodiments of this application are based is illustrated below.

[0061] Optionally, this application can be applied to distributed storage systems or centralized storage systems.

[0062] To facilitate understanding of the embodiments of this application, Figure 1A and Figure 1B A schematic diagram of a possible, non-limiting distributed storage system is shown.

[0063] See Figure 1AThe diagram shows the system architecture of a distributed storage system. This system employs a compute-storage separation architecture, comprising a compute node cluster and a storage node cluster. The compute node cluster includes one or more compute nodes 110 ( Figure 1A The diagram shows two compute nodes 110 (but is not limited to two compute nodes 110), and these compute nodes 110 can communicate with each other. A compute node 110 is a computing device, such as a server, desktop computer, or the controller of a storage array. In terms of hardware, such as... Figure 1A As shown, the compute node 110 includes at least a central processing unit (CPU) 112, memory 113, and a network interface card (NIC) 114. The CPU 112 is an example of a processor, but it can be replaced with other processors. The processor handles data access requests from outside the compute node 110 or requests generated internally within the compute node 110. For example, when the processor 112 receives write data requests from users, it temporarily stores the data in these write data requests in memory 113. When the total amount of data in memory 113 reaches a certain threshold, the processor 112 sends the data stored in memory 113 to the storage node 100 for persistent storage. In addition, the processor 112 is also used for data computation or processing, such as metadata management, deduplication, data compression, virtualization of storage space, and address translation. Figure 1A Only one CPU 112 is shown in the illustration. In practical applications, there may be one or more CPUs 112, and each CPU 112 may have one or more CPU cores. This application does not limit the number of CPUs or the number of CPU cores.

[0064] The memory 113 refers to an internal memory that exchanges data directly with the processor, which can read and write data and serve as temporary data storage for the operating system or other programs running. The memory includes at least two types of memory, for example, the memory can be a random access memory or a read-only memory (Read Only Memory, ROM). For example, the random access memory is a dynamic random access memory (Dynamic Random Access Memory, DRAM) or a storage class memory (Storage Class Memory, SCM). The DRAM is a semiconductor memory, like most random access memories (Random Access Memory, RAM), which is a kind of volatile memory device. The SCM is a composite storage technology that combines the characteristics of traditional storage devices and memories. The storage class memory can provide faster read and write speeds than hard disks, but the access speed is slower than DRAM, and the cost is also cheaper than DRAM. However, the DRAM and SCM are only exemplary in the embodiments of the present application, and the memory can also include other random access memories, such as static random access memory (Static Random Access Memory, SRAM), etc. For read-only memory, for example, it can be a programmable read-only memory (Programmable Read Only Memory, PROM), an erasable programmable read-only memory (Erasable Programmable Read Only Memory, EPROM), etc. In addition, the memory 113 can also be a dual in-line memory module or a dual-line memory module (Dual In-line Memory Module, DIMM), i.e. a module composed of dynamic random access memory (DRAM), and can also be a solid state disk (Solid State Disk, SSD). In practical applications, multiple memories 113 can be configured in the computing node 110, and different types of memories 113 can be configured. The number and type of the memory 113 are not limited in the embodiments of the present application. Optionally, the memory 113 can be configured to have a power retention function. The power retention function refers to that when the system is powered off and then powered on again, the data stored in the memory 113 will not be lost. The memory with the power retention function is called a non-volatile memory.

[0065] The network card 114 is used for communication with the storage node 100. For example, when the total amount of data in the memory 113 reaches a certain threshold, the computing node 110 can send a request to the storage node 100 through the network card 114 to persistently store the data. Optionally, the computing node 110 can also include a bus for communication between the components inside the computing node 110. In terms of function, since Figure 1A The main function of the computing node 110 in the system is to perform computing services, and when storing data, the computing node 110 can use remote memory to achieve persistent storage, so it has less local memory than a conventional server, thereby achieving cost and space savings. However, this does not mean that the computing node 110 cannot have local memory. In actual implementation, the computing node 110 can also be internally provided with a small amount of hard disk or externally connected with a small amount of hard disk.

[0066] Any computing node 110 can access any storage node 100 in the storage node cluster through the network. The storage node cluster includes a plurality of storage nodes 100 Figure 1A Three storage nodes 100 are shown in the system, but the number of storage nodes 100 is not limited to three. One storage node 100 includes one or more control units 101, a network card 104, and a plurality of hard disks 105. The network card 104 is used for communication with the computing node 110. The hard disk 105 is used for storing data and can be a magnetic disk or other types of storage medium, such as a solid state disk or a shingled magnetic recording hard disk. The control unit 101 is used to write data into the hard disk 105 or read data from the hard disk 105 according to the read / write data request sent by the computing node 110. During the process of reading and writing data, the control unit 101 needs to convert the address carried in the read / write data request into an address that can be recognized by the hard disk. That is, the control unit 101 also has a certain computing function.

[0067] Figure 1B Another system architecture diagram of a distributed storage system applied by the embodiment of the present application is shown. The system is a storage-computing integrated architecture, and the system includes a storage cluster. The storage cluster includes one or more servers 110 Figure 1B Three servers 110 are shown in the system, but the number of servers 110 is not limited to three. The servers 110 can communicate with each other. The server 110 is a device that has both computing and storage capabilities, such as a server, a desktop computer, etc. Exemplarily, the server 110 includes an ARM server, a Linux server, or an X86 server.

[0068] In hardware, for example, Figure 1BAs shown, the server 110 at least includes a processor 112, a memory 113, a network card 114 and a hard disk 105. The processor 112, the memory 113, the network card 114 and the hard disk 105 are connected through a bus. Among them, the processor 112 and the memory 113 are used to provide computing resources. Specifically, the processor 112 is a central processing unit CPU, which is used to process data access requests from outside the server 110 (application server or other servers 110), and is also used to process requests generated inside the server 110. The memory 113 refers to the internal memory that exchanges data directly with the processor, which can read and write data as temporary data storage for the operating system or other programs running. In practical applications, multiple memories 113 can be configured in the server 110, and different types of memories 113 can be configured. The number and type of the memory 113 are not limited in the embodiments of the present application. Optionally, the memory 113 can be configured to have a power retention function. The hard disk 105 is used to provide storage resources, such as storing data. It can be a magnetic disk or other types of storage media, such as a solid state disk or a stacked magnetic recording hard disk, etc. The network card 114 is used for communication with other servers 110.

[0069] It should be noted that the above Figure 1A , Figure 1B is only a schematic architecture of the distributed storage system, and in other possible implementation manners of the present application, the distributed storage system can also use other architectures, for example, the distributed storage system can also adopt a full fusion architecture, etc.

[0070] Further, the above distributed storage system can be used to provide storage services, for example, the application server can be provided with storage services, and the application server can directly access the distributed storage system or access the distributed storage system through a switch to store user data through the distributed storage system.

[0071] Figure 2 A possible, non-limiting centralized storage system schematic diagram is shown, Figure 2 The centralized storage system is characterized by a unified portal through which all data from external devices such as application servers pass. For example Figure 2 As shown, the portal of the centralized storage system can be an engine 221 of the centralized storage system. Among them, the engine 221 can include one or more controllers, Figure 2 Take a controller 222 as an example for illustration.

[0072] Optionally, the engine 221 can further include a front-end interface 225 and a back-end interface 226, wherein the front-end interface 225 is configured to communicate with the application server 200, thereby providing storage services for the application server 200. The back-end interface 226 is configured to communicate with the hard disk 234, thereby expanding the capacity of the storage system. Through the back-end interface 226, the engine 221 can connect more hard disks 234, thereby forming a storage resource pool.

[0073] Optionally, in the controller 222, a central processing unit (CPU) 223 and a memory 224 can be included. The CPU 223 is configured to process data access requests from outside the storage system (such as an application server or other storage systems), and is also configured to process requests occurring inside the storage system. It should be noted that the CPU 223 is only an example of a processor in the controller 222, and the CPU 223 can be replaced by other processors. When the CPU 223 receives a write data request sent by the application server 200 through the front-end interface 225, the CPU 223 temporarily saves user data in the write data request in the memory 224. When the total amount of user data in the memory 224 reaches a certain threshold, the CPU 223 sends the user data stored in the memory 224 to the hard disk 234 through the back-end interface for persistent storage.

[0074] It should be noted that, Figure 2 Only one engine 221 is shown in the figure, but in actual applications, two or more engines 221 can be included in the storage system, and the multiple engines 221 can be redundant or load balanced. In addition, in an implementation, the engine 221 can further include a hard disk slot. In this case, the hard disk 234 can be directly deployed in the engine 221, and the back-end interface 226 is an optional configuration. When the storage space of the system is insufficient, more hard disks or hard disk frames can be connected through the back-end interface 226.

[0075] In addition, it should be noted that, Figure 2 Only a structural schematic diagram of a centralized storage system is provided by way of example. In other application scenarios, the storage system 220 can be composed of multiple independent storage servers, and the storage servers can communicate with each other. Each storage server can include a processor, a memory, a network card, and a hard disk, etc.

[0076] With the rapid development of Internet technology, the size of data to be stored also increases explosively. In order to store a large amount of data, a capacity balancing technology can be used. On the one hand, the capacity balancing technology can reasonably allocate the size of data stored by different storage nodes. On the other hand, as the overall system design of the storage system becomes more and more complex, the capacity balancing technology also plays a crucial role in ensuring the performance and available capacity of the storage system.

[0077] In a storage system, for example Figure 1A The capacity balancing technology can be used to balance the storage capacity of different storage nodes, for example, in the process of writing storage data, due to the size of data received by each disk may not be completely consistent, the size of the stored data between disks is not balanced, eventually leading to some disks in the storage system still have a large amount of available space, and some disks are full. At this time, with the help of capacity balancing technology, the balance of storage capacity can be achieved. For example, data A is stored in disk A, and the storage capacity of disk A is full, while the storage capacity of disk B is more. At this time, the disk storing data A can be adjusted to disk B, so as to improve the first parameter of the storage capacity of disk A and disk B. However, during the capacity balancing process, a large amount of data overhead is generated, which affects the performance of the storage system.

[0078] Therefore, how to reduce the data overhead in the capacity balancing process is a technical problem to be solved.

[0079] The data storage method provided by the embodiments of the present application will be described in detail below. Optionally, Figure 3 The method is illustrated by taking the first device as the execution subject of the interaction in Figure 3 The present application does not limit the execution subject of the interaction. For example, when Figure 3 The method is applied to Figure 1A The distributed storage system in Figure 1A The first device can be the computing node 110, and when Figure 3 The method is applied to Figure 1B The first device can be the server 110, and when Figure 3 The method is applied to Figure 2 The storage system 220 in Figure 2 The first device can be the controller 222. The execution subject of S301-S303 and related implementation manners can be the first device, or a chip, chip system, or processor supporting the first device to implement the method, and can also be a logic module or software capable of implementing all or part of the function of the first device. The first device in S301-S303 and related implementation manners can also be replaced by a chip, chip system, or processor supporting the first device to implement the method, and can also be replaced by a logic module or software capable of implementing all or part of the function of the controller.

[0080] As shown in Figure 3 The data storage method provided by the embodiments of the present application comprises the following steps:

[0081] S301, the first device acquires first data to be stored;

[0082] The first data is associated with a first node.

[0083] It should be noted that the data in the present application is an example of naming, for example, it can also be replaced by Data Block, information or data packet. Wherein, the data block is the basic unit of stored data, and is also the smallest physical storage unit. Alternatively, in the present application, the first data can be a data block, or a whole composed of multiple data blocks.

[0084] Optionally, the first device acquires a first data set, and the first data set includes the first data. For example, the first data can include n data blocks, n≥2, n is a positive integer, wherein the first data is one or more data blocks in the n data blocks. Alternatively, each data in the first data set, each data, or each data block can be fixed length or variable length.

[0085] Optionally, the storage node can be a server, and the server has at least one hard disk for storing data, for example, the hard disk can be a magnetic disk or other types of storage medium, such as a solid state disk or a shingled magnetic recording hard disk. When Figure 3 The method shown in the method is applied to Figure 1A The system shown, the storage node can be a storage node 110; when Figure 3 The method shown in the method is applied to Figure 1B The system shown, the storage node can be a module or device containing a hard disk 105, when Figure 3 The method shown in the method is applied to Figure 2 The system shown, the storage node can be a module or device containing a hard disk 234.

[0086] Optionally, the first device can be located in the same storage system as the first node, at this time the first device is a control module, controller, computing node or management node in the storage system. Or the first device constitutes a whole storage system, and the first node is a storage node in the storage system.

[0087] Optionally, the client can send the first data to the first device through a network, for example, the client can be a smart phone, a tablet PC, a mobile phone, a video phone, an electronic book reader, a desktop PC, a laptop PC, a notebook computer, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical instrument, a camera, and a wearable device (for example, a head-mounted device (HMD) such as electronic glasses, electronic clothes, an electronic bracelet, an electronic necklace, an electronic accessory (that is, an application accessory), an electronic tattoo, or a smart watch), and the like. Alternatively, the server can send the first data to the first device, for example, the server can be an application server, or the server can send the first data to the first device through a switch.

[0088] Optionally, the first data has not been stored to any storage node, or written to any storage node, or in other words, the data has not been written to a hard disk or stored in a hard disk in the storage system before S301 and the following S301.

[0089] Optionally, the "storage node associated with the first data is the first node" can mean that the first data belongs to the first node, and at this time, the first node serves as a preliminary node for storing the first data, that is, the first data is preliminarily stored in the first node, but has not been stored in the first node.

[0090] Optionally, two data attribution determinations can be made in the present application, that is, two nodes associated with the first data are determined. The first data attribution determination means that after the first device receives the first data, the software will preliminarily determine the storage node of the data, for example, the software can determine the storage node associated with the first data according to the size of the traffic, and the first determination is a simple determination. At this time, some problems can occur, for example, the first determination determines that the storage node associated with the first data is the first node, but the storage space in the first node is full, and at this time, the first data is not suitable for being stored in the first node, so the second determination is needed, which occurs after the first determination, that is, S302 described below, which will be described in detail below:

[0091] S302, the first device determines whether the storage node associated with the first data needs to be modified to a second node.

[0092] Among them, the second node has a greater remaining storage capacity than the first node or a greater first parameter than the first node.

[0093] Optionally, the first device, the first node, and the second node can be located in the same storage system, and the first node and the second node are both storage nodes in the storage system.

[0094] Optionally, the second node remaining storage capacity is greater than the first node remaining storage capacity can be replaced by the second node currently stored data size is less than the size of the first node stored data, for example, the second node has stored 100M of data, and the first node stores 500M of data, it can be considered that the size of the data stored by the first node is less than the first node, or the second node expects to store data less than the size of the data stored by the first node, for example, the first node has currently stored 500M of data, but there are 5M of data before the first data to be stored in the first node, the first node expects to store 505M of data, and the second node has currently stored 500M of data, but there are 1M of data before the first data to be stored in the second node, the second node expects to store 501M of data, and it is considered that the second node expects to store data less than the size of the data stored by the first node.

[0095] As described above, the present application can make two judgments of data attribution, and S302 can be the second judgment. Optionally, S302 can be a judgment made by the hardware layer or in the background, that is, a background judgment. At this time, since S302 occurs in the background, the user does not perceive, and the impact on the storage data delay and the like is small. In addition, S302 can determine the storage node associated with the first data according to more conditions, more comprehensively. Specifically, in S302, the storage node associated with the first data can be determined again according to the following conditions:

[0096] In an optional implementation, the first device obtains a first data set, the first data set including the first data; when a first condition is met, the first device modifies the storage node associated with the first data to a second node, wherein the first condition includes: the number of data in the first data set that is similar to the first data above a first threshold is less than a second threshold; when the first condition is not met, the first device does not modify the storage node associated with the first data to the second node.

[0097] In the above implementation, the number of data in the first data set that is similar to the first data above a first threshold is counted, the first data duplication count is obtained, the duplication count is counted, and the first data in the first data set and the data that is highly similar or identical to the first data is counted, so as to determine the number of data that needs to be deleted in the first data and the data that is duplicated with the first data in the first data set. For example, when there are 6 data in the first data set that is similar to the first data above a first threshold, the first data duplication count is 6. When the duplication count is less than the second threshold, that is, when the duplication count is small enough, modifying the attribution of the first data has less impact on data duplication. It can be seen that the above implementation has less impact on data duplication and preserves the duplication characteristics of the data.

[0098] Optionally, the data with the similarity higher than the first threshold value to the first data can be considered as the data that is duplicated or nearly duplicated with the first data. According to the principle of deduplication, part of the data that is duplicated or nearly duplicated can be deleted, and only one instance of the data is kept, and a unique identifier is created for the instance. For other data, the system only stores a reference pointing to the unique instance.

[0099] Optionally, the first threshold value can be preset by the system or input by the user. For example, the first threshold value can be 99%, that is, when there are 3 data in the first data set with the similarity higher than 99% to the first data, the deduplication count of the first data is 3.

[0100] Optionally, the second threshold value should be small enough to reduce the impact on data deduplication. For example, the second threshold value can be 2, when there are 6 data in the first data set with the similarity higher than the first threshold value to the first data, the storage node associated with the first data is not modified to the second node. When there is 1 data in the first data set with the similarity higher than the first threshold value to the first data, the storage node associated with the first data is modified to the second node.

[0101] Optionally, the second threshold value can be set by the system or according to the deduplication level in the deduplication application or software used by the user, wherein the deduplication application or software is used for deduplication of user data, and the user data includes the first data. The deduplication level is related to the size of the data that is deduplicated. For example, there is 1G of storage space in the storage system, and the user needs to store 2G of data, so the deduplication level is set to be relatively high. For example, there is 1G of storage space in the storage system, and the user needs to store 1.2G of data, so the deduplication level is set to be relatively low.

[0102] In an optional implementation, the first condition further includes that the first parameter of the first node is higher than a third threshold value, and the first parameter is related to the size of the data stored by the first node.

[0103] Based on the above implementation, when the first parameter of the first node is higher than the third threshold value, the first device can modify the storage node associated with the first data to the second node, and the first parameter is related to the size of the data stored by the first node, that is, the first device determines the ownership of the first node and the storage node associated with the first node in combination with the capacity dimension of the node, which is beneficial to improve the utilization rate of each storage node.

[0104] For example, the first parameter can be a size of the stored data of the first node, and the first threshold can be a storage capacity threshold of the first node, i.e., a maximum storage capacity, such as 1 G; for example, the first parameter can be a ratio of the size of the stored data of the first node to the storage capacity threshold of the first node, and the first threshold can be 100% or 99%, etc.; for example, the first parameter is a ratio of the size of the stored data of the first node to a first average value, and the first average value is an average value of the sizes of the stored data of all nodes in the first storage system, and the first threshold can be 120% or 110%, etc.; for example, the first parameter can be a size of a remaining available storage space of the first node, and the first threshold can be 20 M or 10 M, etc.

[0105] In an optional implementation, the first parameter is a ratio of the size of the stored data of the first node to a first average value, and the first average value is an average value of the sizes of the stored data of all nodes in the first storage system, and the first storage system includes the first node and the second node.

[0106] Based on the above implementation, the first parameter can determine the balance degree of the first node relative to the first storage system. For example, when the first parameter is too high, such as 200%, it indicates that the size of the stored data of the first node is much larger than the average size of the stored data of the nodes in the first storage system, and the first node is unbalanced relative to other nodes in the first storage system. By taking the first parameter as an index of the first condition, i.e., taking the balance degree as an index of whether to modify the ownership of the first data associated with the first node, i.e., an index of modifying the ownership of part of the data in the first node, the size of the stored data in the first node can be adjusted according to the balance degree of the first node, so as to realize the balance of the first node, which is beneficial to the balance of the first storage system and improves the utilization rate of the storage space of the first storage system.

[0107] In an optional implementation, the first parameter of the second node is lower than the third threshold.

[0108] Based on the above implementation, the first parameter is a ratio of the size of the stored data of the first node to a first average value, and when the first parameter of the second node is lower than the third threshold and the first parameter of the first node is higher than the threshold, it indicates that the size of the stored data of the second node is smaller than the size of the stored data of the first node, and the remaining storage space of the second node is larger. At this time, the first data associated with the first node is modified to be associated with the second node, and is subsequently stored to the second node, so as to improve the utilization rate of the storage space of the second node and improve the balance degree of the first storage system.

[0109] Optionally, as Figure 4The first parameter of the second node being lower than the third threshold value and the first parameter of the first node being higher than the third threshold value, the first parameter being related to the size of the data stored by the first node, can be collectively taken as the first condition. The first device can set a weight for the difference between the first parameter of the first node and the third threshold value and a weight for the difference between the duplicate count of the first node and the second threshold value, and if the sum of the two after weighting exceeds a certain threshold value, it is considered that the storage node associated with the first data needs to be modified. That is, the capacity balance is controlled by means of weighted summation. For example, the difference between the first parameter of the first node and the first average value can be set to 40% of the weight, and the difference between the duplicate count of the first node and the first threshold value can be set to 60% of the weight, and if the sum of the two after weighting exceeds 3, it is considered that the storage node associated with the first data needs to be modified. For example, the ratio of the size of the data stored by the first node to the first average value is 200%, the third threshold value is 120%, the difference between the first parameter of the first node and the third threshold value is 0.8, and for example, the second threshold value can be 1, when there are 5 data in the first data set whose similarity to the first data is higher than the first threshold value, that is, the duplicate count of the first node is 5, the difference between the duplicate count of the first node and the second threshold value is 4, and at this time, the weighted sum of the two is: 0.8 x 40% + 4 x 60% = 2.72 < 3, so the storage node associated with the first node is not modified.

[0110] In an optional implementation, the first device modifies the storage node associated with the first data to the second node by modifying the first parameter in the metadata of the first data to a second parameter, wherein the first parameter is used to indicate that the storage location of the first data is the first node, and the second parameter is used to indicate that the storage location of the first data is the second node.

[0111] Based on the above implementation, the first device can modify the information in the metadata for identifying the storage location of the first data by modifying the first parameter in the metadata of the first data to a second parameter, thereby modifying the storage location of the first data, modifying the ownership of the first data without data migration, and enabling the first data to be stored in the second node.

[0112] Optionally, the above metadata can be located in the metadata file of the first device, and there is a field in the metadata for identifying the storage location of the data, the field of the first data originally being the first parameter and being modified to the second parameter, thereby modifying the storage location of the first data to the second node.

[0113] In an optional implementation, when the second condition is met, the first device determines whether the storage node associated with the first data needs to be modified to the second node, the second condition comprising: the first control interface being in an open state, wherein the first control interface is in the open state when an average read-write rate of the first storage system is greater than a fourth threshold value and / or a data amount of the first data set is less than a fifth threshold value, and the first storage system comprises the first node and the second node. When the first condition is not met, the first data is stored to the first node.

[0114] Based on the above implementation, before determining whether to modify the node associated with the first data, a precondition-second condition is further set, the second condition can determine whether the node associated with the first data needs to be modified from two dimensions of the read-write rate of the first storage system and the data amount of the obtained data set, which can reduce the influence on the performance of the first storage system and is also beneficial to the capacity balance of the first storage system.

[0115] In an optional implementation, the second condition further comprises: a size of data stored by the first storage system is greater than a sixth threshold value.

[0116] Based on the above implementation, the second condition is a precondition for determining whether to modify the node associated with the first data, and the second condition can determine whether to perform the judgment process of modifying the node associated with the first data from the capacity of the storage system. Therefore, when the size of data stored by the first storage system exceeds the sixth threshold value, the capacity balance of the first storage system can be realized by adjusting part of the data, for example, the node associated with the first data.

[0117] Optionally, the sixth threshold value can be a capacity threshold value of the first storage system, or a maximum storage data amount of the first storage system, or a maximum storage capacity of the first storage system. For example, the sixth threshold value can be 1 TB. As shown in the figure, the sixth threshold value can be dynamically configured according to the total capacity of the first storage system. Figure 5

[0118] S303, when it is determined that the modification is needed, the first device stores the first data to the second node.

[0119] Optionally, before S303, the first data has not been stored to the first storage system, so that when the first device stores the first data to the second node, data migration is not needed, that is, the first data does not need to be migrated from any node in the first storage system to the second node.

[0120] ​Optionally, S303 can be automatically completed, that is, once the storage node associated with the first node is modified from the first node to the second node, the first data is automatically stored to the second node. As shown above, the storage location of the first data can be modified to the second node by modifying the parameter in the metadata for identifying the storage location of the first data, and thus the first data is automatically stored to the second node.

[0121] Optionally, S303 can also be replaced by that the first device writes the first data to the first node or the first device stores the first data to the hard disk of the first node.

[0122] For the convenience of understanding, the above data storage method will be introduced below in combination with specific examples.

[0123] The embodiment of the present application is an example of the data storage method. Please refer to Figure 6 , Figure 6 A flowchart of an embodiment provided by the embodiment of the present application is shown in the following. Figure 6 The specific flow in the above embodiment includes:

[0124] S601: The first device receives a first data set;

[0125] For example, the first data set can be a plurality of data blocks sent by a server to the first device, and the first device needs to store the plurality of data blocks to the storage node of the first storage system.

[0126] S602: The first device determines the storage node associated with the first data;

[0127] It can be understood that the above S602 is the first time to determine the storage node associated with the first data. For example, the software configured in the first device can determine the storage node associated with the first data according to the traffic size. The following describes the first node as the storage node associated with the first data determined in S602.

[0128] S603: The first device determines whether the first control interface is opened;

[0129] It should be noted that S603 is an optional step. When S603 is not executed, if the first control interface is opened, S604 is directly executed, and if the first control interface is closed, S609 is directly executed.

[0130] For example, the opening and closing of the first control interface can be determined by the first device according to the read-write performance of the first storage system, such as read-write speed and the size of the first data set. When the read-write performance of the first storage system is poor and / or the size of the first data set is large, the read-write performance of the first storage system cannot quickly write the first data, and the first control interface is closed. When the read-write performance of the first storage system is good and / or the size of the first data set is small, the read-write performance of the first storage system supports quickly writing the first data, and the first control interface is opened.

[0131] S604: The first device determines whether a capacity threshold of the first storage system is reached.

[0132] For example, the first device can count the size of the data already stored in the first storage system, and determine whether the size reaches the capacity threshold of the first storage system. If yes, S605 is executed, and if no, S609 is executed.

[0133] It can be understood that S605 to S608 below provide a capacity balancing method. By modifying the nodes associated with the data, that is, modifying the ownership of the data, the storage data of the allocation of the storage nodes in the first storage system is adjusted, so that the capacity of the first storage system is balanced.

[0134] S605: The first device traverses each node in the first storage system.

[0135] For example, the first device can traverse each node in the first storage system to determine the size of the first parameter of each node in the first storage system. On the one hand, the size of the first parameter of the first node can be determined, and the node in the first storage system whose first parameter is greater than the third threshold can also be found.

[0136] S606: The first device determines whether the number of data in the first data set that is similar to the first data and is higher than the first threshold is less than the second threshold.

[0137] When it is determined in S606 whether the number of data in the first data set that is similar to the first data and is higher than the first threshold is less than the second threshold, S607 is executed, otherwise S609 is executed.

[0138] S607: The first device determines whether the first parameter of the first node is greater than the third threshold.

[0139] The first parameter is the ratio of the size of the data stored in the first node to the first average value, and the first average value is the average of the size of the data stored in all nodes in the first storage system.

[0140] It can be understood that when the first parameter is too large, for example, greater than the third threshold, the size of the stored data in the first node is large, at this time, the node associated with the first data can be modified so that the first data is stored to the second node instead of the first node, thereby avoiding further increasing the size of the stored data in the first device, which is beneficial to the balance of the first storage system.

[0141] When the first parameter is greater than the third threshold in S607, S608 is performed, and when the first parameter is greater than the third threshold, S609 is performed.

[0142] S608: The first device modifies the storage node associated with the first data.

[0143] Optionally, when the first device S605 traverses each storage node in the first storage system, if a node with a first parameter less than the third threshold or a stored data size less than the size of the stored data in the first node is found, assuming that the node is a second node, at this time, the first device can modify the storage node associated with the first data from the first node to the second node.

[0144] S609: The first device stores the first data to the first node.

[0145] Optionally, S609 can also be replaced by the first device writing the first data to the first node or the first device storing the first data to the hard disk of the first node.

[0146] From the above introduction, it can be seen that the present application has the following technical effects:

[0147] Firstly, the present application provides a capacity balancing method, that is, modifying the storage node associated with the first data from the first node to the second node, wherein the remaining storage capacity of the second node is greater than that of the first node, and after the modification, the first data will be stored to the second node. The principle of the present application for realizing capacity balancing is that when the storage node associated with the first data is modified to the second node, the first data will be stored to the second node, at this time, the difference between the remaining storage capacities of the first node and the second node will decrease. When not modified, the first data will be stored to the first node with smaller remaining storage capacity, at this time, the difference between the remaining storage capacities of the first node and the second node will increase. It can be seen that by modifying the node associated with the data in this way, the capacity balance between the storage nodes can be promoted. Moreover, the modification of the node associated with the data in the present application occurs before storing the data, so the present application does not need to perform data migration, for example, does not need to migrate the first data from the first node to the second node, so the overhead of the capacity balancing process can be reduced.

[0148] Furthermore, this application counts the number of data in the first dataset whose similarity to the first data exceeds a first threshold, thus obtaining the deduplication count of the first data. The deduplication count is calculated by counting data in the first dataset that are highly similar to or identical to the first data, thereby determining the number of data to be deleted from the first data and its duplicates in the first dataset. This application also limits the modification of the storage node associated with the first data (i.e., changing the ownership of the first data) only when the deduplication count is less than a second threshold (i.e., when the deduplication count is sufficiently small). When the deduplication count is sufficiently small, changing the ownership of the data has a minimal impact on data deduplication. Therefore, the capacity balancing method provided in this application has a minimal impact on data deduplication and can preserve the data's deduplication characteristics. Moreover, this application modifies the storage node associated with the data by modifying corresponding parameters in the metadata. This allows for the early determination of data ownership by balancing the data's deduplication characteristics and the system's capacity characteristics without requiring the creation of new data structures to store capacity distribution information. This achieves capacity balancing of the storage system while minimizing the impact on data deduplication.

[0149] Furthermore, this application can determine the balance of the first node relative to the first storage system through the first parameter. By using the first parameter as an indicator of the first condition—that is, using the balance as an indicator of whether to modify the ownership of the first data associated with the first node, i.e., modifying the ownership of some data in the first node—it is possible to adjust the size of the data stored in the first node based on the balance of the first node, thereby achieving the balance of the first node. This is beneficial to the balance of the first storage system and improves the utilization rate of the storage space of the first storage system. Moreover, by using the opening of the first control interface and the ratio of the size of the data stored in the first node to the average size of the first storage system exceeding a threshold as preconditions for modifying the storage node associated with the first data, this application can jointly determine whether to modify the storage node associated with the first data from two dimensions: the capacity of the storage system and the balance of the storage system. This makes the method provided by this application have a smaller impact on the performance of the storage system and is beneficial to the capacity balance of the storage system.

[0150] The embodiments of this application have been described above from the perspective of methodology. The communication devices in the embodiments of this application will be introduced below from the perspective of specific device implementation.

[0151] Please see Figure 7 This application provides a schematic diagram of a communication device 700, for example, in Figure 1A In the system shown, the distributed storage system as a whole can function as a communication device 700, or the communication device 700 can also function as a computing node 110. Figure 1B In the system shown, the communication device 700 can be the server 110. Figure 2In the illustrated system, the storage system 220 as a whole can be the communication device 700, or the communication device 700 can also be the controller 222. The communication device 700 at least includes a processing unit 701 and a transceiver unit 702.

[0152] As an example, the communication device 700 can implement the above-mentioned Figure 3 The functions of the first device in the illustrated method can also be implemented by the above-mentioned Figure 3 The beneficial effects of the above-mentioned

[0153] The transceiver unit 702 is configured to obtain first data to be stored, and the storage node associated with the first data is a first node.

[0154] The processing unit 701 is configured to determine whether the storage node associated with the first data needs to be modified to a second node, wherein the remaining storage capacity of the second node is greater than that of the first node; and when it is determined that the modification is needed, store the first data to the second node.

[0155] In an optional implementation, the transceiver unit 702 is specifically configured to obtain a first data set, and the first data set includes the first data; and the processing unit 701 is specifically configured to modify the storage node associated with the first data to the second node when a first condition is met, wherein the first condition includes that the number of data in the first data set that has a similarity to the first data higher than a first threshold is less than a second threshold; and not modify the storage node associated with the first data to the second node when the first condition is not met.

[0156] In an optional implementation, the first condition further includes that a first parameter of the first node is higher than a third threshold, and the first parameter is related to the size of the data stored in the first node.

[0157] In an optional implementation, the first parameter is a ratio of the size of the data stored in the first node to a first average value, and the first average value is an average value of the size of the data stored in all nodes in a first storage system, and the first storage system includes the first node and the second node.

[0158] In an optional implementation, the first parameter of the second node is lower than the third threshold.

[0159] In an optional implementation, the processing unit 701 is specifically configured to modify a second parameter in the metadata of the first data to a third parameter, wherein the second parameter is used to indicate that the storage location of the first data is the first node, and the third parameter is used to indicate that the storage location of the first data is the second node.

[0160] In an optional implementation, when the second condition is met, it is determined whether the storage node of the first data association needs to be modified to the second node, the second condition comprising: the first control interface is in an open state, wherein the first control interface is in the open state when an average read-write rate of the first storage system is greater than a fourth threshold value and / or a data amount of the first data set is less than a fifth threshold value, wherein the first storage system comprises the first node and the second node; and the first data is stored to the first node when the second condition is not met.

[0161] In an optional implementation, the second condition further comprises: a size of data stored by the first storage system is greater than a sixth threshold value.

[0162] It should be noted that the information execution process and the like of the units of the communication apparatus 700 described above can be specifically refer to the descriptions in the method embodiments described above, and will not be described here.

[0163] Please refer to Figure 8 , Figure 8 The structure of the communication apparatus 800 involved in the above embodiments provided by the embodiments of the present application, for example, in the system shown in Figure 1A , the distributed storage system as a whole can be the communication apparatus 800, or the communication apparatus 800 can be the computing node 110, in the system shown in Figure 1B , the communication apparatus 800 can be the server 110, in the system shown in Figure 2 , the storage system 220 as a whole can be the communication apparatus 800, or the communication apparatus 800 can be the controller 222. The structure of the communication apparatus 800 can refer to the structure shown in Figure 8 .

[0164] The communication apparatus 800 comprises at least one processor 801, at least one communication port 802, at least one memory 803, and one or more antennas 804. The processor 801, the memory 803 and the communication port 802 are connected, for example, through a bus, in the embodiments of the present application, the connection can comprise various interfaces, transmission lines or buses, etc., and the present embodiment does not limit this. The antenna 804 is connected with the communication port 802.

[0165] As an implementation example, Figure 8 the communication apparatus is the foregoing Figure 3In the first device in the related embodiments, the communication port 802 is configured to obtain first data to be stored, and a storage node associated with the first data is a first node; the processor 801 is configured to determine whether the storage node associated with the first data needs to be modified to a second node, where a remaining storage capacity of the second node is greater than a remaining storage capacity of the first node; and when it is determined that the modification is needed, the first data is stored in the second node.

[0166] It should be noted that the above Figure 8 The execution process of each device in the communication device shown in the above method embodiments of the present application is described in detail, and the description is not repeated here.

[0167] The processor 801 is mainly used for processing communication protocols and communication data, and controlling the entire communication device, executing software programs, and processing data of the software programs, for example, for supporting the communication device to perform the actions described in the embodiments. The communication device can include a baseband processor and a central processor. The baseband processor is mainly used for processing communication protocols and communication data, and the central processor is mainly used for controlling the entire first device, executing software programs, and processing data of the software programs. Figure 8 The processor 801 in the above embodiments can integrate the functions of the baseband processor and the central processor. Those skilled in the art can understand that the baseband processor and the central processor can also be independent processors interconnected by bus technology. Those skilled in the art can understand that the first device can include multiple baseband processors to adapt to different network modes, and the first device can include multiple central processors to enhance its processing capability. Various components of the first device can be connected by various buses. The baseband processor can also be referred to as a baseband processing circuit or a baseband processing chip. The central processor can also be referred to as a central processing circuit or a central processing chip. The function of processing communication protocols and communication data can be built into the processor, or stored in the memory in the form of software programs, and the processor executes the software programs to realize the baseband processing function.

[0168] The memory 803 is mainly used for storing software programs and data. The memory 803 can exist independently and be connected to the processor 801. Alternatively, the memory 803 can be integrated with the processor 801, for example, integrated in a chip. The memory 803 can store program codes for executing the technical solutions of the embodiments of the present application, and the processor 801 controls the execution. Various computer programs executed can also be regarded as a driver of the processor 801.

[0169] Figure 8Only one memory and one processor are shown. In an actual first device, there can be multiple processors and multiple memories. The memory can also be referred to as a storage medium or a storage device, etc. The memory can be a storage element on the same chip as the processor, i.e., an on-chip storage element, or an independent storage element, and the embodiments of the present application do not limit this.

[0170] The communication port 802 can be configured to support the receiving or transmitting of radio frequency signals between the communication device and a terminal. The communication port 802 can be connected to the antenna 804. The communication port 802 includes a transmitter Tx and a receiver Rx. Specifically, one or more antennas 804 can receive radio frequency signals, the receiver Rx of the communication port 802 is configured to receive the radio frequency signals from the antenna and convert the radio frequency signals into digital baseband signals or digital intermediate frequency signals, and provide the digital baseband signals or digital intermediate frequency signals to the processor 801, so that the processor 801 further processes the digital baseband signals or digital intermediate frequency signals, such as demodulation processing and decoding processing. In addition, the transmitter Tx in the communication port 802 is also configured to receive modulated digital baseband signals or digital intermediate frequency signals from the processor 801, and convert the modulated digital baseband signals or digital intermediate frequency signals into radio frequency signals, and transmit the radio frequency signals through one or more antennas 804. Specifically, the receiver Rx can selectively perform one or more levels of down-mixing processing and analog-to-digital conversion processing on the radio frequency signals to obtain digital baseband signals or digital intermediate frequency signals, and the order of the down-mixing processing and the analog-to-digital conversion processing can be adjusted. The transmitter Tx can selectively perform one or more levels of up-mixing processing and digital-to-analog conversion processing on the modulated digital baseband signals or digital intermediate frequency signals to obtain radio frequency signals, and the order of the up-mixing processing and the digital-to-analog conversion processing can be adjusted. The digital baseband signals and the digital intermediate frequency signals can be collectively referred to as digital signals.

[0171] The transceiver can also be referred to as a transceiving unit, a transceiver, a transceiving device, etc. Optionally, the devices in the transceiving unit for implementing the receiving function can be regarded as a receiving unit, and the devices in the transceiving unit for implementing the transmitting function can be regarded as a transmitting unit, i.e., the transceiving unit includes the receiving unit and the transmitting unit, the receiving unit can also be referred to as a receiver, an input port, a receiving circuit, etc., and the transmitting unit can be referred to as a transmitter, a transmitter, or a transmitting circuit, etc.

[0172] It should be noted that, Figure 8 The communication device shown can be specifically configured to implement the steps implemented by the first device in any of the preceding method embodiments, and achieve the corresponding technical effects of the first device, Figure 8 The specific implementation of the communication device shown can be referred to the description in any of the preceding method embodiments, which will not be repeated here.

[0173] The embodiment of the present application further provides a computer readable storage medium storing one or more computer-executable instructions that, when executed by a processor, cause the processor to perform the method of any possible implementation of the communication apparatus as described in the foregoing embodiments, wherein the communication apparatus can be specifically the first device in the foregoing embodiments.

[0174] The embodiment of the present application further provides a computer program product (or computer program) storing one or more computers that, when executed by the processor, cause the processor to perform the method of any possible implementation of the communication apparatus as described above, wherein the communication apparatus can be specifically the first device in the foregoing embodiments.

[0175] The embodiment of the present application further provides a chip system including a processor for supporting the communication apparatus to implement the functions involved in the possible implementation of the communication apparatus described above. In a possible design, the chip system can further include a memory for storing the necessary program instructions and data of the communication apparatus. The chip system can be composed of a chip, or can include a chip and other discrete devices, wherein the communication apparatus can be specifically the first device in the foregoing embodiments.

[0176] The embodiment of the present application provides a communication system including a first node, a second node and a processor for executing the method of any possible implementation of the first node, the second node and the processor. Figure 3 The first device for executing the method of any possible implementation of the first node, the second node and the processor.

[0177] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented by other means. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0178] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0179] In addition, each of the functional units in the various embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0180] When the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application essentially make contributions to the part or the whole or part of the technical solutions, which can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a first device, etc.) to execute all or part of the steps of the method according to the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0181] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A data storage method, characterized by, The method comprises: acquiring first data to be stored, the storage node associated with the first data being a first node; determining whether the storage node associated with the first data needs to be modified to a second node, wherein the second node has a larger remaining storage capacity than the first node; when it is determined that the modification is needed, storing the first data to the second node.

2. The data storage method of claim 1, wherein, The acquiring of the first data to be stored comprises: acquiring a first data set, the first data set comprising the first data; The determining of whether the storage node associated with the first data needs to be modified to the second node comprises: when a first condition is met, modifying the storage node associated with the first data to the second node, wherein the first condition comprises that the number of data in the first data set that has a high similarity to the first data is less than a second threshold value; when the first condition is not met, not modifying the storage node associated with the first data to the second node.

3. The data storage method of claim 2, wherein, The first condition further comprises that a first parameter of the first node is higher than a third threshold value, the first parameter being related to the size of the data stored in the first node.

4. The data storage method of claim 3, wherein, The first parameter is a ratio of the size of the data stored in the first node to a first average value, the first average value being an average value of the size of the data stored in all nodes in a first storage system, the first storage system comprising the first node and the second node.

5. The data storage method according to claim 3 or 4, characterized by, The first parameter of the second node is lower than the third threshold value.

6. The data storage method according to any one of claims 1 to 5, characterized in that, The storage node associated with the first data is modified to the second node in the following manner: modifying a second parameter in metadata of the first data to a third parameter, wherein the second parameter is used to indicate that the storage location of the first data is the first node, and the third parameter is used to indicate that the storage location of the first data is the second node.

7. The data storage method according to any one of claims 1 to 6, characterized by, After the acquiring of the first data set, the method further comprises: when a second condition is met, determining whether the storage node associated with the first data needs to be modified to the second node, the second condition comprising that a first control interface is in an open state, wherein the first control interface is in the open state when an average read-write speed of a first storage system is greater than a fourth threshold value and / or the data amount of the first data set is less than a fifth threshold value, wherein the first storage system comprises the first node and the second node; when the second condition is not met, storing the first data to the first node.

8. The data storage method of claim 7, wherein, The second condition further comprises that the size of the data stored in the first storage system is greater than a sixth threshold value.

9. A data storage device, characterized by The method comprises: a transceiving unit, configured to acquire first data to be stored, the storage node associated with the first data being a first node; a processing unit, configured to determine whether the storage node associated with the first data needs to be modified to a second node, wherein the second node has a larger remaining storage capacity than the first node; and when it is determined that the modification is needed, store the first data to the second node.

10. The data storage device of claim 9, wherein, The transceiving unit is specifically configured to: acquire a first data set, the first data set comprising the first data; The processing unit is specifically configured to: modifying a storage node associated with the first data to a second node when a first condition is satisfied, wherein the first condition comprises: a number of data in the first data set that is similar to the first data with a similarity higher than a first threshold is less than a second threshold; not modifying the storage node associated with the first data to the second node when the first condition is not satisfied.

11. The data storage device of claim 10, wherein, The first condition further comprises: a first parameter of the first node is higher than a third threshold, the first parameter being related to a size of data stored in the first node.

12. The data storage device of claim 11, wherein, The first parameter is a ratio of the size of data stored in the first node to a first average value, the first average value being an average of sizes of data stored in all nodes in a first storage system, the first storage system comprising the first node and the second node.

13. The data storage device of claim 11 or 12, wherein, The first parameter of the second node is lower than the third threshold.

14. The data storage device of any of claims 9 to 13, wherein, The processing unit is specifically configured to: modify a second parameter in metadata of the first data to a third parameter, wherein the second parameter is used to indicate that a storage location of the first data is the first node, and the third parameter is used to indicate that the storage location of the first data is the second node.

15. The data storage device of any of claims 9 to 14, wherein, The processing unit is further configured to: determine whether the storage node associated with the first data needs to be modified to the second node when a second condition is satisfied, the second condition comprising: a first control interface being in an open state, wherein the first control interface is in the open state when an average read-write speed of a first storage system is greater than a fourth threshold and / or an amount of data in the first data set is less than a fifth threshold, wherein the first storage system comprises the first node and the second node; store the first data to the first node when the second condition is not satisfied.

16. The data storage device of claim 15, wherein, The second condition further comprises: a size of data stored in the first storage system is greater than a sixth threshold.

17. A communications device, characterized by comprising: a communication interface and a processor; The communication interface and the processor perform the method according to any one of claims 1 to 8.

18. A communication system, characterized by comprising: a first node, a second node, and a first device configured to perform the method according to any one of claims 1 to 8.

19. A computer-readable storage medium, characterized in that, The medium stores instructions, when the instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

20. A computer program product, characterised in that, comprising instructions, when the instructions are executed on a processor, the method according to any one of claims 1 to 8 is performed.

21. A chip, characterized by comprising at least one processing unit and an interface circuit, the interface circuit being configured to provide program instructions or data for the at least one processing unit, the at least one processing unit being configured to execute the program instructions to implement the method according to any one of claims 1 to 8.