Multi-replica data processing method and apparatus, and electronic device

By calculating the request processing capability corresponding to the unit data amount of data nodes in the distributed storage system, determining the request probability and selecting the target replica, the problem of mismatch in the processing requests of data nodes is solved, and the overall service capability of the system is improved.

WO2025149835A1PCT designated stage expired Publication Date: 2025-07-17CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/063271
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-12-30
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

In a distributed storage system, the processing requests of the data nodes do not match their processing capabilities, resulting in excessive requests received on data nodes with large disk capacity, exceeding the processing capabilities and affecting the overall service capabilities of the system.

Method used

By determining the request processing capability corresponding to the unit data amount of each data node, the request probability of each replica is calculated, and the target replica is selected for processing requests, ensuring that the request matches the node's capabilities.

Benefits of technology

It improves the overall service capabilities of the distributed system, avoids the problem that nodes with large data volumes cannot handle due to excessive requests, and improves system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024063271_17072025_PF_FP_ABST
    Figure IB2024063271_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a multi-replica data processing method and apparatus, and an electronic device. The method comprises: in response to a processing instruction for multi-replica data, determining data nodes where a plurality of replicas of the multi-replica data are located, wherein each of the plurality of replicas is located at a different data node; on the basis of a request processing capability corresponding to a unit data volume of the data node where each replica is located, determining a request probability for each replica; and on the basis of the request probability for each replica, selecting a target replica from the plurality of replicas, and sending a processing request for the target replica to the data node where the target replica is located. The invention ensures that a processing request received by a data node matches the processing capability provided by said data node, thereby improving the overall service capability of distributed systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field of Processing Method, Apparatus and Electronic Device for Multi-copy Data

[0001] This application relates to the field of data storage, and more particularly, to a processing method, apparatus and electronic device for multi-copy data. Background Art

[0002] A distributed storage system usually consists of a metadata node and multiple data nodes. The metadata node mainly records the metadata information in the system, while user data is stored in the data nodes in the form of files or data blocks.

[0003] The multi-copy mechanism is the simplest way to maintain data redundancy in a distributed system. When a user writes data, the system will replicate the data block to generate multiple copies and store these copies on multiple data nodes in the distributed system according to a preset placement strategy. In this way, when a copy on a certain data node is abnormal, the data can still be obtained from other data nodes, thus ensuring the high availability of the data. When the states of all data nodes are normal, the client may send a processing request to a single data node or send the processing request to each data node in turn. From the system level, the processing requests received by each data node are proportional to the amount of data stored on it, which easily leads to a large number of requests being received on the data node with a large disk capacity, exceeding the processing capacity of the data node. Summary of the Invention

[0004] This application provides a processing method, apparatus and electronic device for multi-copy data to ensure that the processing requests received by each data node match the processing capabilities provided by the data node, thereby improving the overall service capacity of the distributed system.

[0005] In a first aspect, this application provides a processing method for multi-copy data, including: in response to a processing instruction for multi-copy data, determining the data nodes where multiple copies of the multi-copy data are located, where each of the multiple copies is located on a different data node; determining the request probability for each of the copies based on the request processing capabilities corresponding to the unit data amount of the data nodes where the copies are located; based on the request probabilities for each of the copies, selecting a target copy from the multiple copies, and sending a processing request for the target copy to the data node where the target copy is located.

[0006] In one implementation, the method further includes: periodically obtaining the request processing capabilities and the amount of stored data of each data node from the metadata node; determining the request processing capabilities corresponding to the unit data amount of each data node according to the request processing capabilities and the amount of stored data of each data node.

[0007] In one implementation, determining the request processing capability corresponding to the unit data volume of each data node according to the request processing capability and the stored data volume of each data node includes: dividing the request processing capability of each data node by the stored data volume to obtain the request processing capability corresponding to the unit data volume of each data node.

[0008] In one implementation, determining the request probability for each replica based on the request processing capability corresponding to the unit data volume of the data node where each replica is located includes: performing a normalization process on the request processing capability corresponding to the unit data volume of the data node where each replica is located, and determining the request processing capability corresponding to the unit data volume of the data node where each replica is located after normalization as the request probability for each replica.

[0009] In one implementation, the request processing capability is the query rate per second.

[0010] In one implementation, selecting a target replica from the multiple replicas based on the request probability for each replica includes: determining the random number range corresponding to each replica within a preset random number range according to the request probability for each replica, where the ratio of the random number ranges corresponding to each replica is the same as the ratio of the request probabilities; generating a random number within the preset random number range, and determining the replica corresponding to the random number range where the random number is located as the target replica.

[0011] In one implementation, responding to a processing instruction for multi-replica data and determining the data nodes where the multiple replicas of the multi-replica data are located includes: responding to a processing instruction triggered by an open operation on the multi-replica data, and obtaining the data nodes where the multiple replicas of the multi-replica data are located from the metadata node.

[0012] In a second aspect, the present application provides a multi-replica data processing apparatus, including: a node determination module, configured to respond to a processing instruction for multi-replica data and determine the data nodes where the multiple replicas of the multi-replica data are located, where each of the multiple replicas is located in a different data node; a probability calculation module, configured to determine the request probability for each replica based on the request processing capability corresponding to the unit data volume of the data node where each replica is located; and a request module, configured to select a target replica from the multiple replicas based on the request probability for each replica, and send a processing request for the target replica to the data node where the target replica is located.

[0013] In one implementation, the device further includes: an acquisition module, configured to periodically obtain the request processing capabilities and the stored data volumes of the respective data nodes from the metadata node; a determination module, configured to determine the request processing capabilities corresponding to the unit data volumes of the respective data nodes according to the request processing capabilities and the stored data volumes of the respective data nodes.

[0014] In one implementation, the determination module is configured to: divide the request processing capabilities of the respective data nodes by the stored data volumes to obtain the request processing capabilities corresponding to the unit data volumes of the respective data nodes.

[0015] In one implementation, the probability calculation module is configured to: perform a normalization process on the request processing capabilities corresponding to the unit data volumes of the data nodes where the respective replicas are located, and determine the request probabilities for the respective replicas by using the request processing capabilities corresponding to the unit data volumes of the data nodes where the respective replicas are located after the normalization. In one implementation, the request module is configured to: determine, according to the request probabilities for the respective replicas, the random number ranges corresponding to the respective replicas from within a preset random number range, where the ratio of the random number ranges corresponding to the respective replicas is the same as the ratio of the request probabilities; generate a random number within the preset random number range, and determine the replica corresponding to the random number range where the random number is located as the target replica.

[0016] In one implementation, the request processing capability is the query rate per second.

[0017] In one implementation, the request module is configured to: determine, according to the request probabilities for the respective replicas, the random number ranges corresponding to the respective replicas from within a preset random number range, where the ratio of the random number ranges corresponding to the respective replicas is the same as the ratio of the request probabilities; generate a random number within the preset random number range, and determine the replica corresponding to the random number range where the random number is located as the target replica.

[0018] In one implementation, the node determination module is configured to: in response to a processing instruction triggered by an open operation on the multi-copy data, obtain the data nodes where the multiple replicas of the multi-copy data are located from the metadata node.

[0019] In a third aspect, the present application provides an electronic device, including: a memory and a processor; the memory is configured to store a computer program; the processor is configured to execute the computer program stored in the memory, and when the computer program runs, the processor executes the method described in the first aspect.

[0020] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the method described in the first aspect.

[0021] Fifth aspect, the present application provides a computer program product, including a computer program which, when executed by a processor, implements the method described in the first aspect.

[0022] In the method, apparatus and electronic device for processing multi-copy data provided by the present application, the request probability for each copy is determined according to the request processing capacity corresponding to the unit data volume of the data node where the copy is located, so as to ensure that the processing requests received by the data node match the request processing capacity corresponding to the unit data volume of the data node, and avoid an excessive number of processing requests being received on the data node with a large data volume, resulting in its inability to process, thereby improving the overall service capacity of the distributed system. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0024] FIG. 1 is a schematic diagram of a distributed system provided by an embodiment of the present application;

[0025] FIG. 2 is a schematic flowchart of a method for processing multi-copy data provided by an embodiment of the present application;

[0026] FIG. 3 is a schematic structural diagram of an apparatus for processing multi-copy data provided by an embodiment of the present application;

[0027] FIG. 4 is a schematic block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0029] Introduce the professional terms related to the embodiments of the present application.

[0030] Metadata node: The node in a distributed system where the metadata of a file is stored, usually used to store the status information of the file and the location information of data blocks, etc.

[0031] Data node: A node in a distributed system where data blocks are stored. It is usually responsible for operations such as writing, storing, reading, and deleting data block copies.

[0032] Machine model: The model of the machine in a distributed system. Machines of the same model have the same hardware configuration, such as disk size, CPU model, memory size, etc.

[0033] In many cases, data nodes in a distributed system cannot be guaranteed to use hardware that provides consistent performance. For example, the disk size and CPU performance of different data nodes may be different. As shown in FIG1 , the distributed system includes a metadata node 101 and multiple data nodes 102, wherein the data nodes 102 include two types of machines, A and B. The number of disks of the two types of machines is the same, and the size of the HDD disk on the A type machine is larger than that on the B type machine. Taking the case where the client 103 reads data from the distributed system as an example, since the read performance of the HDD disk is independent of the disk size, the read performance of the machines of model A and model B is consistent, that is, the processing capabilities of the machines of model A and model B for the read requests received from the client 103 are the same. However, in the case where the disk on the machine of model A is larger, if more user data is stored on it, the read requests received by the machine of model A will be higher than the read requests received by the machine of model B. In the case where the read performance of the two is consistent, the machine of model A receives more read requests, and these read requests may exceed the processing capability of the machine of model A, thereby affecting the overall service capability of the distributed system.

[0034] For example, the read performance of machine type A and machine type B can both process 100 read requests per second. The disk size of the A-type machine is 8TB, and the disk size of the B-type machine is 4TB. Assuming that the amount of stored data of the A-type machine and the B-type machine is 4TB, that is, the amount of user data stored in both is the same, then the number of read requests received by both is roughly the same. However, if the amount of stored data of the A-type machine continues to increase, for example, reaching 8TB, then the number of read requests received on the A-type machine may be much higher than the number of read requests received on the B-type machine. Since the read performance of the two is the same, the number of read requests received on the B-type machine may be lower than the read performance, while the number of read requests received on the A-type machine may far exceed the read performance, which may cause serious processing timeouts and other situations, affecting the overall service capabilities of the distributed system.

[0035] In view of this, in the embodiments of the present application, a method for processing multi-copy data is proposed. When selecting replica nodes of multi-copy data for data processing, the request processing capacity corresponding to the unit data volume of the data node where the replica is located is considered, and the processing requests are sent to the data nodes where different replicas are located with a certain access probability, so as to ensure that the processing requests received by the data nodes match the capabilities provided by the data nodes, thereby improving the overall service capacity of the system.

[0036] FIG. 2 is a schematic flowchart of a method for processing multi-copy data provided by an embodiment of the present application. The execution subject of the embodiment of the present application is a client. As shown in FIG. 2, the method includes steps S201 to S203.

[0037] S20K In response to a processing instruction for multi-copy data, determine the data nodes where multiple replicas of the multi-copy data are located, where each replica among the multiple replicas is located in a different data node.

[0038] The triggering manner of the processing instruction for multi-copy data can be triggered by a user operation or a command. Optionally, the processing instruction can be an instruction that triggers the client to read data from a distributed system. Taking the processing instruction as a read instruction as an example, for example, the multi-copy data is a multi-copy file, and the user triggers the read instruction for the multi-copy data by performing an open operation on the file through the client. Another example is that the read instruction for the multi-copy data is triggered in the form of a program command during the execution of the program on the client. The triggering manner of the processing instruction in the embodiments of the present application is not limited.

[0039] In response to a processing instruction for multi-copy data, the client first determines the data nodes where each replica of the multi-copy data is located. By way of example, in response to a processing instruction triggered by an open operation on multi-copy data, the client obtains the data nodes where multiple replicas of the multi-copy data are located from the metadata node. Assume that there are four replicas of the multi-copy data, which are located in four different data nodes respectively. Then the client determines the four different data nodes where the four replicas are located from the metadata node. For example, the client obtains the node identifiers of the four data nodes where the four replicas of the multi-data copy are located from the metadata node.

[0040] S202. Based on the request processing capacity corresponding to the unit data volume of the data node where each replica is located, determine the request probability for each replica.

[0041] For a specific data node, its processing capacity for the received processing requests, that is, this data node If the request processing capacity of a data node is fixed, then the more data stored on the data node, the lower the request processing capacity per unit data volume, and the less data stored on the data node, the higher the request processing capacity per unit data volume. The request processing capacity per unit data volume reflects the actual ability of the data node to process requests. Therefore, based on the request processing capacity per unit data volume of the data nodes where each replica is located, the request probability for each replica is determined. The higher the request processing capacity per unit data volume of a data node, the higher the request probability for the replica on that data node. Conversely, the lower the request processing capacity per unit data volume of a data node, the lower the request probability for the replica on that data node, so that the number of processing requests received by each data node matches its actual processing ability.

[0042] S203. Based on the request probability for each replica, select a target replica from multiple replicas and send a processing request for the target replica to the data node where the target replica is located.

[0043] Based on the request probability for each replica, select a target replica from multiple replicas. The higher the request probability, the higher the probability that the corresponding replica is selected as the target replica. Conversely, the lower the request probability, the lower the probability that the corresponding replica is selected as the target replica. That is, the higher the request processing capacity per unit data volume of a data node, the higher the probability that the replica on that data node is selected as the target replica, and the lower the request processing capacity per unit data volume of a data node, the lower the probability that the replica on that data node is selected as the target replica. The client sends a processing request for the target replica to the data node where the target replica is located. For example, if the processing instruction is a read instruction, the client sends a read request for the target replica to the data node where the target replica is located to obtain the corresponding data from the data node where the target replica is located and complete the reading of the multi-replica data.

[0044] The method for processing multi-replica data provided by the embodiments of the present application, when processing multi-replica data, determines the request probability for each replica according to the request processing capacity per unit data volume of the data node where the replica is located, that is, the probability of sending a processing request to each data node, so as to ensure that the processing requests received by the data node match the request processing capacity per unit data volume of the data node, avoid receiving too many processing requests on the data node with a large amount of data and causing it to be unable to process, thereby improving the overall service ability of the distributed system.

[0045] Based on the above embodiments, an explanation is given on how to determine the request processing capacity per unit data volume of each data node.

[0046] The metadata node of the distributed system stores the relevant information of each data node, including the request processing capacity and the stored data volume of each data node. The client periodically obtains the request processing capacity and the stored data volume of each data node from the metadata node; according to the request processing capacity and the stored data volume of each data node, determine the request processing capacity corresponding to the unit data volume of each data node.

[0047] Among them, the request processing capacity of each data node can be fixed or pre-configured, which is used to characterize the processing capacity of each data node for the received requests. The request processing capacity of each data node can be reported by each data node to the metadata node actively, or the metadata node can obtain it from each data node actively. Similarly, the stored data volume of each data node can be reported by each data node to the metadata node actively, or the metadata node can obtain it from each data node actively.

[0048] Based on the request processing capacity and the stored data volume of each data node obtained from the metadata node, the client determines the request processing capacity corresponding to the unit data volume of each data node. Optionally, divide the request processing capacity of each data node by the stored data volume to obtain the request processing capacity corresponding to the unit data volume of each data node.

[0049] Optionally, the request processing capacity is the Queries Per Second (QPS), which characterizes the number of requests that the data node can respond to per second.

[0050] Exemplarily, the QPS of data node a and data node B is both 100. The disk size of data node a is 4TB, and the disk size of data node b is 8TB. Assume that the stored data volume of data node a is 4TB, and the stored data volume of data node b is also 4TB. Then, in the case of taking 1TB as the unit data volume, the QPS corresponding to the unit data volume of data node a is 100 / 4, and the QPS corresponding to the unit data volume of data node b is also 100 / 4. Assume that the user data stored in the disk of data node b continues to increase later and reaches a stored data volume of 8TB. Then, the QPS corresponding to the unit data volume of data node b is 100 / 8. It can be seen that as the stored data volume increases, the QPS corresponding to the unit data volume of the data node decreases.

[0051] After the client determines the data nodes where multiple copies of multi-copy data are located, it determines the request probability for each copy based on the request processing capabilities corresponding to the unit data volume of the data nodes where each copy is located, including: performing normalization processing on the request processing capabilities corresponding to the unit data volume of the data nodes where each copy is located, and determining the request probability for each copy as the request processing capabilities corresponding to the unit data volume of the data nodes where each copy is located after normalization.

[0052] Referring to the above example, when the stored data volume of data node a is 4TB and the stored data volume of data node b is also 4TB, the QPS corresponding to the unit data volume of data node a and data node b is both 100 / 4. After normalizing the QPS corresponding to the unit data volume of data node a and data node b, the results are both 0.5, that is, the request probabilities for the copies on data node a and data node b are both 0.5.

[0053] In the case where the stored data volume of data node a is 4TB and the stored data volume of data node b is 8TB, the QPS corresponding to the unit data volume of data node a is 100 / 4, and the QPS corresponding to the unit data volume of data node b is 100 / 8. After normalizing the QPS corresponding to the unit data volume of data node a and data node b, the results are 2 / 3 and 1 / 3 respectively, that is, the request probabilities for the copies on data node a and data node b are 2 / 3 and 1 / 3.

[0054] After determining the request probability for each copy, according to the request probability for each copy, determine the random number range corresponding to each copy from the preset random number range, where the ratio of the random number ranges corresponding to each copy is the same as the ratio of the request probabilities; generate a random number within the preset random number range, and determine the copy corresponding to the random number range where the random number is located as the target copy.

[0055] Taking the request probabilities of the replicas on data node a and data node b as 2 / 3 and 1 / 3 respectively as an example, since the request probability of the replica on data node a is twice that of the replica on data node b, the range of random numbers corresponding to the replica on data node a should also be twice that of the range of random numbers corresponding to the replica on data node b. Assuming the preset random number range is 1 - 300, the range of random numbers corresponding to the replica on data node a is 1 - 200, and the range of random numbers corresponding to the replica on data node b is 201 - 300. Or, the range of random numbers corresponding to the replica on data node a is 101 - 300, and the range of random numbers corresponding to the replica on data node b is 1 - 100. As long as the ratio of the random number ranges corresponding to each replica meets the requirements, the specific partitioning method is not limited in the embodiments of this application.

[0056] Taking the range of random numbers corresponding to the replica on data node a as 1 - 200 and the range of random numbers corresponding to the replica on data node b as 201 - 300 as an example, when the client selects a replica, it generates a random number between 1 and 300. Assuming the generated random number is 165, and this random number falls within the range of random numbers corresponding to the replica on data node a, then the replica on data node a is the target replica, and the client sends a read request for the target replica to data node a, thereby obtaining the corresponding data from data node a. Assuming the generated random number is 265, and this random number falls within the range of random numbers corresponding to the replica on data node b, then the replica on data node b is the target replica, and the client sends a read request for the target replica to data node b, thereby obtaining the corresponding data from data node b.

[0057] Since the range of random numbers corresponding to the replica on data node a is twice that of the range of random numbers corresponding to the replica on data node b, when generating random numbers, the probability that the random number falls within the range of random numbers corresponding to the replica on data node a is approximately twice that of data node b, so that the number of read requests received on data node a is approximately twice that of data node b. In this way, it matches the request processing capabilities corresponding to the unit data volume of data nodes a and b, enabling the read requests to be reasonably allocated according to the processing capabilities of the data nodes, rather than the data node with a larger data volume receiving more requests, thus ensuring that the performance of each data node in the distributed system is reasonably utilized and improving the overall service ability of the distributed system.

[0058] 3 is a schematic diagram of the structure of a multi-copy data processing device provided in an embodiment of the present application. As shown in FIG3 , the multi-copy data processing device 300 includes a node determination module 301 , a probability calculation module 302 and a request module 303 .

[0059] The node determination module 301 is used to determine multiple nodes of the multiple copies of data in response to a processing instruction for the multiple copies of data. The data nodes where the replicas are located, where each of the multiple replicas is located on a different data node.

[0060] The probability calculation module 302 is used to determine the request probability for each replica based on the request processing capacity corresponding to the unit data volume of the data node where each replica is located.

[0061] The request module 303 is used to select a target replica from multiple replicas based on the request probability of each replica, and send a processing request for the target replica to the data node where the target replica is located.

[0062] In one implementation, the multi-copy data processing device 300 also includes: an acquisition module, which is used to periodically obtain the request processing capability and the amount of stored data of each data node from the metadata node; and a determination module, which is used to determine the request processing capability corresponding to the unit data amount of each data node based on the request processing capability and the amount of stored data of each data node.

[0063] In one implementation, the determination module is used to: divide the request processing capacity of each data node by the amount of stored data to obtain the request processing capacity corresponding to the unit data amount of each data node.

[0064] In one implementation, the probability calculation module 302 is used to: normalize the request processing capacity corresponding to the unit data volume of the data node where each replica is located, and determine the normalized request processing capacity corresponding to the unit data volume of the data node where each replica is located as the request probability for each replica.

[0065] In one implementation, the request processing capacity is the query rate per second.

[0066] In one implementation, the request module 303 is used to: determine the random number range corresponding to each replica from a preset random number range according to the request probability of each replica, wherein the ratio of the random number range corresponding to each replica is the same as the ratio of the request probability; generate a random number within the preset random number range, and determine the replica corresponding to the random number range where the random number is located as the target replica.

[0067] In one implementation, a node determination module 301 is configured to: in response to a processing instruction triggered by an open operation on multi-copy data, obtain data nodes where multiple copies of the multi-copy data are located from a metadata node.

[0068] The device according to the embodiment of the present application can be used to execute the method for processing multi-copy data in the foregoing embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0069] FIG. 4 is a schematic block diagram of an electronic device provided by an embodiment of the present application. As shown in FIG. 4, the electronic device 400 may include at least one processor 401, which is configured to implement the method for processing multi-copy data provided by the embodiment of the present application.

[0070] Optionally, the electronic device 400 further includes at least one memory 402, which is configured to store program instructions and / or data. The memory 402 is coupled to the processor 401. The coupling in the embodiment of the present application is an indirect coupling or communication connection between devices, units or modules, and may be electrical, mechanical or other forms, and is used for information interaction between devices, units or modules. The processor 401 may cooperate with the memory 402. The processor 401 may execute program instructions stored in the memory 402. At least one of the at least one memories may be included in the processor.

[0071] Optionally, the electronic device 400 further includes a communication interface 403, which is configured to communicate with other devices through a transmission medium, so that the electronic device 400 can communicate with other devices. The communication interface 403 may be, for example, a transceiver, an interface, a bus, a circuit or a device capable of implementing a transceiver function. The processor 401 may use the communication interface 403 to send and receive data and / or information, and is configured to implement the method provided by the embodiment of the present application. For specific details, refer to the detailed description in the foregoing embodiment, which will not be elaborated here.

[0072] In the embodiment of the present application, the specific connection medium between the foregoing processor 401, memory 402 and communication interface 403 is not limited. In the embodiment of the present application, in FIG. 4, the processor 401, memory 402 and communication interface 403 are connected through a bus 404. The bus 404 is represented by a thick line in FIG. 4, and the connection manners between other components are only for illustrative purposes and are not limited thereto. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only one thick line is used in FIG. 4, but it does not mean that there is only one bus or one type of bus.

[0073] It should be understood that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments may be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by the hardware decoding processor, or executed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0074] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static Static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM). It should be noted that the memories of the systems and methods described herein are intended to include, but are not limited to, these and any other suitable types of memories.

[0075] This application also provides a computer-readable storage medium that stores a computer program (which can also be referred to as code or instructions). When the computer program is run, it causes the computer to execute the method in any of the foregoing embodiments.

[0076] This application also provides a computer program product that includes a computer program, and when the computer program is executed by a processor, it implements the method in any of the foregoing embodiments.

[0077] The terms "unit", "module", etc. used in this specification can be used to represent computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution.

[0078] Those of ordinary skill in the art can realize that the various illustrative logical blocks and steps described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application. In several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0079] The unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0080] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0081] In the above embodiments, the functions of each functional unit can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0082] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs, etc., which can store program codes.

[0083] The user information involved in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0084] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art in the technical field disclosed by this application can easily think of changes or substitutions within the technical scope disclosed by this application, and all should be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

Claims 1. A method for processing multi-copy data, comprising: In response to a processing instruction for multi-copy data, determine the data nodes where multiple copies of the multi-copy data are located, where each of the multiple copies is located in a different data node; based on the request processing capabilities corresponding to the unit data volume of the data nodes where each of the copies is located, determine the request probabilities for each of the copies; based on the request probabilities for each of the copies, select a target copy from the multiple copies, and send a processing request for the target copy to the data node where the target copy is located.

2. The method according to claim 1 further comprises: Periodically obtain the request processing capabilities and the stored data volumes of each data node from the metadata node; based on the request processing capabilities and the stored data volumes of each data node, determine the request processing capabilities corresponding to the unit data volume of each data node.

3. The method according to claim 2, wherein, The determining, based on the request processing capabilities and the stored data volumes of each data node, the request processing capabilities corresponding to the unit data volume of each data node includes: dividing the request processing capabilities of each data node by the stored data volume to obtain the request processing capabilities corresponding to the unit data volume of each data node.

4. The method according to claim 1, wherein, The determining, based on the request processing capabilities corresponding to the unit data volume of the data nodes where each of the copies is located, the request probabilities for each of the copies includes: performing normalization processing on the request processing capabilities corresponding to the unit data volume of the data nodes where each of the copies is located, and determining the request processing capabilities corresponding to the unit data volume of the data nodes where each of the copies is located after normalization as the request probabilities for each of the copies.

5. The method according to any one of claims 1-4, wherein, The request processing capability is the queries per second rate.

6. The method according to any one of claims 1-4, wherein The selecting, based on the request probabilities for each of the copies, a target copy from the multiple copies includes: determining, according to the request probabilities for each of the copies, the random number ranges corresponding to each of the copies within a preset random number range, where the ratio of the random number ranges corresponding to each of the copies is the same as the ratio of the request probabilities; generating a random number within the preset random number range, and determining the copy corresponding to the random number range where the random number is located as the target copy.

7. The method according to any one of claims 1 - 4, wherein, The responding to a processing instruction for multi-copy data and determining the data nodes where multiple copies of the multi-copy data are located includes: responding to a processing instruction triggered by an open operation on the multi-copy data, and obtaining from the metadata node the data nodes where multiple copies of the multi-copy data are located.

8. A processing device for multi-copy data, comprising: A node determination module, configured to respond to a processing instruction for multi-copy data and determine the data nodes where multiple copies of the multi-copy data are located, where each of the multiple copies is located in a different data node; a probability calculation module, configured to determine the request probabilities for each of the copies based on the request processing capabilities corresponding to the unit data volume of the data nodes where each of the copies is located; a request module, configured to select a target copy from the multiple copies based on the request probabilities for each of the copies, and send a processing request for the target copy to the data node where the target copy is located.

9. An electronic device, comprising: A memory and a processor; The memory is used to store a computer program; The processor is used to execute the computer program stored in the memory, and when the computer program runs, it causes the processor to execute the method described in any one of claims 1-7.

10. A computer-readable storage medium, wherein, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, it causes the processor to execute the method described in any one of claims 1-7.

11. A computer program product, comprising a computer program, and when the computer program is executed by a processor, it implements the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Load balancing method based on object storage device

    CN101013387A

  • Hotspot detection method and device, monitoring server and storage medium

    CN113489776A

  • Probability-based load balancing method and device, electronic equipment and storage medium

    CN114079656A

  • Distributed database system load balancing method and device

    CN116991580A