Multi-copy data processing method and device and electronic equipment
By calculating the request probability of multiple replica data nodes in a distributed storage system, the problem of mismatch between data node processing requests and capabilities is solved, and the overall service capability of the system is improved.
Patent Information
- Application Number
- CN202410042029.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-11
AI Technical Summary
In a distributed storage system, the processing requests of the data nodes do not match their processing capabilities, resulting in data nodes with large disk capacity receiving too many requests, exceeding their processing capabilities, affecting the overall service capabilities of the system.
By determining the request processing capability corresponding to the unit data amount of the data node where the multiple replica data is located, the request probability is calculated, and the target replica is selected for processing requests based on the probability, ensuring that the request matches the processing capability of the data node.
The overall service capability of the distributed system is improved, and the problem that data nodes with large amounts of data cannot be processed due to excessive requests is improved, which improves system performance.
Smart Images

Figure CN120295746A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data storage, and more specifically, to a method, apparatus, and electronic device for processing multi-copy data. Background Art
[0002] A distributed storage system usually consists of a metadata node and multiple data nodes. The metadata node mainly records the metadata information in the system, while user data is stored in the data nodes in the form of files or data blocks.
[0003] The multi-copy mechanism is the simplest way to maintain data redundancy in a distributed system. When a user writes data, the system will generate multiple copies of the data block and store these copies on multiple data nodes in the distributed system according to a preset placement policy. In this way, when a copy on a certain data node is abnormal, the data can still be obtained from other data nodes, thus ensuring the high availability of the data. When the states of all data nodes are normal, the client may send a processing request to a single data node or send the processing request to each data node in sequence. From a system level perspective, the processing requests received by each data node are proportional to the amount of data stored on it, which easily leads to a large number of requests being received on a data node with a large disk capacity, exceeding the processing capacity of the data node. Summary of the Invention
[0004] This application provides a method, apparatus, and electronic device for processing multi-copy data to ensure that the processing requests received by each data node match the processing capabilities provided by the data node, thereby improving the overall service capacity of the distributed system.
[0005] In a first aspect, this application provides a method for processing multi-copy data, including:
[0006] In response to a processing instruction for multi-copy data, determining the data nodes where multiple copies of the multi-copy data are located, where each copy of the multiple copies is located on a different data node;
[0007] Based on the request processing capabilities corresponding to the unit data amounts of the data nodes where each copy is located, determining the request probabilities for each copy;
[0008] Based on the request probabilities for each copy, selecting a target copy from the multiple copies and sending a processing request for the target copy to the data node where the target copy is located.
[0009] In one implementation, the method further includes:
[0010] Periodically obtaining the request processing capabilities and the amounts of stored data of each data node from the metadata node;
[0011] Determine the request processing capability corresponding to the unit data volume of each data node according to the request processing capability and the stored data volume of each data node.
[0012] In one implementation, the determining the request processing capability corresponding to the unit data volume of each data node according to the request processing capability and the stored data volume of each data node includes:
[0013] Divide the request processing capability of each data node by the stored data volume to obtain the request processing capability corresponding to the unit data volume of each data node.
[0014] In one implementation, the determining the request probability for each replica based on the request processing capability corresponding to the unit data volume of the data node where each replica is located includes:
[0015] Perform normalization processing on the request processing capability corresponding to the unit data volume of the data node where each replica is located, and determine the request probability for each replica as the request processing capability corresponding to the unit data volume of the data node where each replica is located after normalization.
[0016] In one implementation, the request processing capability is the query rate per second.
[0017] In one implementation, the selecting a target replica from the multiple replicas based on the request probability for each replica includes:
[0018] Determine the random number range corresponding to each replica within a preset random number range according to the request probability for each replica, where the ratio of the random number ranges corresponding to each replica is the same as the ratio of the request probabilities;
[0019] Generate a random number within the preset random number range, and determine the replica corresponding to the random number range where the random number is located as the target replica.
[0020] In one implementation, the determining the data nodes where the multiple replicas of the multi-replica data are located in response to a processing instruction for the multi-replica data includes:
[0021] In response to a processing instruction triggered by an open operation on the multi-replica data, obtain the data nodes where the multiple replicas of the multi-replica data are located from the metadata node.
[0022] In a second aspect, the present application provides a processing device for multi-replica data, including:
[0023] A node determination module, configured to determine data nodes where multiple copies of the multi-copy data are located in response to a processing instruction for the multi-copy data, wherein each of the multiple copies is located in a different data node;
[0024] A probability calculation module, configured to determine a request probability for each of the copies based on a request processing capability corresponding to a unit data volume of the data node where each copy is located;
[0025] A request module, configured to select a target copy from the multiple copies based on the request probability for each of the copies, and send a processing request for the target copy to the data node where the target copy is located.
[0026] In one implementation, the apparatus further includes:
[0027] An acquisition module, configured to periodically obtain a request processing capability and a stored data volume of each data node from a metadata node;
[0028] A determination module, configured to determine a request processing capability corresponding to a unit data volume of each data node according to the request processing capability and the stored data volume of each data node.
[0029] In one implementation, the determination module is configured to:
[0030] Divide the request processing capability of each data node by the stored data volume to obtain a request processing capability corresponding to a unit data volume of each data node.
[0031] In one implementation, the probability calculation module is configured to:
[0032] Perform normalization processing on the request processing capabilities corresponding to the unit data volumes of the data nodes where each copy is located, and determine the request processing capabilities corresponding to the unit data volumes of the data nodes where each copy is located after normalization as the request probabilities for each of the copies.
[0033] In one implementation, the request processing capability is the queries per second rate.
[0034] In one implementation, the request module is configured to:
[0035] Determine a random number range corresponding to each of the copies from a preset random number range according to the request probability for each of the copies, wherein the ratio of the random number ranges corresponding to each of the copies is the same as the ratio of the request probabilities;
[0036] Generate a random number within the preset random number range, and determine the copy corresponding to the random number range where the random number is located as the target copy.
[0037] In one implementation, the node determination module is configured to:
[0038] In response to a processing instruction triggered by an open operation on the multi-copy data, obtain data nodes where multiple copies of the multi-copy data are located from the metadata node.
[0039] In a third aspect, the present application provides an electronic device, including: a memory and a processor;
[0040] The memory is used to store a computer program;
[0041] The processor is configured to execute the computer program stored in the memory, and when the computer program runs, the processor executes the method described in the first aspect.
[0042] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the method described in the first aspect.
[0043] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0044] In the method, apparatus, and electronic device for processing multi-copy data provided by the present application, the request probability for each copy is determined according to the request processing capacity corresponding to the unit data volume of the data node where the copy is located, so as to ensure that the processing requests received by the data node match the request processing capacity corresponding to the unit data volume of the data node, and avoid receiving too many processing requests on the data node with a large data volume, resulting in its inability to process, thereby improving the overall service capacity of the distributed system. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0046] Figure 1 is a schematic diagram of a distributed system provided by an embodiment of the present application;
[0047] Figure 2 is a flowchart of a method for processing multi-copy data provided by an embodiment of the present application;
[0048] Figure 3It is a schematic structural diagram of a multi-copy data processing device provided by an embodiment of the present application;
[0049] Figure 4 It is a schematic block diagram of an electronic device provided by an embodiment of the present application. Specific embodiments
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0051] Introduce the professional terms related to the embodiments of the present application.
[0052] Metadata node: The node where the metadata of files in a distributed system is stored, usually used to store the status information of files and the location information of data blocks, etc.
[0053] Data node: The node where data blocks are stored in a distributed system, usually responsible for operations such as writing, storing, reading, and deleting data block copies.
[0054] Machine type: The model of machines in a distributed system. Machines of the same model have the same hardware configuration, such as disk size, CPU model, memory size, etc.
[0055] In many cases, data nodes in a distributed system cannot guarantee the use of hardware with consistent performance. For example, there may be certain differences in the disk size, CPU performance, etc. of different data nodes. As Figure 1 shown, the distributed system includes a metadata node 101 and multiple data nodes 102. Among them, there are machines of two types A and B in the data node 102. The number of disks of machines of types A and B is the same, and the size of the HDD disk on the machines of type A is larger than that of type B. Taking the example of the client 103 reading data from the distributed system, since the reading performance of the HDD disk is independent of the disk size, the reading performance of the machines of types A and B remains the same, that is, the processing capabilities of the machines of types A and B for the reading requests received by the client 103 are the same. However, in the case where the disk on the machine of type A is larger, if more user data is stored on it, this will cause the reading requests received on the machine of type A to be higher than those received on the machine of type B. With the same reading performance of the two, more reading requests are received on the machine of type A, and these reading requests may exceed the processing capabilities of the machine of type A, thus affecting the overall service capabilities of the distributed system.
[0056] Exemplarily, the read performance of both Model A machines and Model B machines is 100 read requests that can be processed per second. The disk size of Model A machines is 8TB, and the disk size of Model B machines is 4TB. Assuming that the stored data volume of both Model A machines and Model B machines is 4TB, that is, the user data volumes stored by both are the same, then the number of read requests received by both is also roughly the same. However, if the stored data volume of Model A machines continues to increase, for example, reaches 8TB, then the number of read requests received by Model A machines may be much higher than the number of read requests received by Model B machines. Since the read performance of both is the same, the number of read requests received by Model B machines may be lower than the read performance, while the number of read requests received by Model A machines may far exceed the read performance, which may lead to serious processing timeouts and other situations, affecting the overall service ability of the distributed system.
[0057] In view of this, in the embodiments of the present application, a method for processing multi-copy data is proposed. When selecting replica nodes of multi-copy data for data processing, the request processing ability corresponding to the unit data volume of the data node where the replica is located is considered, and the processing requests are sent to the data nodes where different replicas are located with a certain access probability, so as to ensure that the processing requests received by the data nodes match the capabilities provided by the data nodes, thereby improving the overall service ability of the system.
[0058] Figure 2 It is a flowchart of a method for processing multi-copy data provided by an embodiment of the present application. The execution subject of the embodiment of the present application is a client. As Figure 2 shown, the method includes:
[0059] S201. In response to a processing instruction for multi-copy data, determine the data nodes where multiple replicas of the multi-copy data are located, where each of the multiple replicas is located in a different data node.
[0060] The triggering method of the processing instruction for multi-copy data can be triggered by user operations or commands. Optionally, the processing instruction can be an instruction that triggers the client to read data from the distributed system. Taking the processing instruction as a read instruction as an example, for example, the multi-copy data is a multi-copy file, and the user triggers the read instruction for the multi-copy data by performing an open operation on the file through the client. Another example is that the read instruction for the multi-copy data is triggered in the form of a program command during the execution of the program on the client. The triggering method of the processing instruction in the embodiments of the present application is not limited.
[0061] In response to a processing instruction for multi-copy data, the client first determines the data nodes where each copy of the multi-copy data is located. By way of example, in response to a processing instruction triggered by an open operation on multi-copy data, the client obtains from the metadata node the data nodes where multiple copies of the multi-copy data are located. Assuming that the multi-copy data has a total of four copies, each located in a different data node, the client determines from the metadata node the four different data nodes where the four copies are located. For example, the client obtains from the metadata node the node identifiers of the four data nodes where the four copies of the multi-data copy are located respectively.
[0062] S202. Determine the request probability for each copy based on the request processing capacity corresponding to the unit data volume of the data node where each copy is located.
[0063] For a specific data node, its processing capacity for the received processing requests, that is, the request processing capacity of the data node is fixed. Then, the more data stored on the data node, the lower the request processing capacity corresponding to the unit data volume, and the less data stored on the data node, the higher the request processing capacity corresponding to the unit data volume. The request processing capacity corresponding to the unit data volume reflects the actual capacity of the data node to process requests. Therefore, by determining the request probability for each copy based on the request processing capacity corresponding to the unit data volume of the data node where each copy is located, the higher the request processing capacity corresponding to the unit data volume of the data node, the higher the request probability for the copy on that data node. On the contrary, the lower the request processing capacity corresponding to the unit data volume of the data node, the lower the request probability for the copy on that data node, so that the number of processing requests received by each data node matches its actual processing capacity.
[0064] S203. Select a target copy from multiple copies based on the request probability for each copy, and send a processing request for the target copy to the data node where the target copy is located.
[0065] Based on the request probabilities of each replica, a target replica is selected from multiple replicas. The higher the request probability, the higher the probability that the corresponding replica is selected as the target replica. On the contrary, the lower the request probability, the lower the probability that the corresponding replica is selected as the target replica. That is, the higher the request processing capacity corresponding to the unit data volume of a data node, the higher the probability that the replica on that data node is selected as the target replica; the lower the request processing capacity corresponding to the unit data volume of a data node, the lower the probability that the replica on that data node is selected as the target replica. The client sends a processing request for the target replica to the data node where the target replica is located. For example, if the processing instruction is a read instruction, the client sends a read request for the target replica to the data node where the target replica is located to obtain the corresponding data from the data node where the target replica is located, and complete the reading of the multi-replica data.
[0066] The method for processing multi-replica data provided by the embodiments of the present application determines the request probabilities for each replica according to the request processing capacity corresponding to the unit data volume of the data node where the replica is located when processing multi-replica data, that is, the probability of sending a processing request to each data node, so as to ensure that the processing requests received by the data node match the request processing capacity corresponding to the unit data volume of the data node, and avoid receiving too many processing requests on the data node with a large amount of data and causing it to be unable to process, thereby improving the overall service capacity of the distributed system.
[0067] Based on the above embodiments, an explanation is given on how to determine the request processing capacity corresponding to the unit data volume of each data node.
[0068] The metadata node of the distributed system stores the relevant information of each data node, including the request processing capacity and the stored data volume of each data node. The client periodically obtains the request processing capacity and the stored data volume of each data node from the metadata node; according to the request processing capacity and the stored data volume of each data node, the request processing capacity corresponding to the unit data volume of each data node is determined.
[0069] Among them, the request processing capacity of each data node can be fixed or pre-configured, and is used to characterize the processing capacity of each data node for received requests. The request processing capacity of each data node can be reported by each data node to the metadata node actively, or the metadata node can obtain it from each data node actively. Similarly, the stored data volume of each data node can be reported by each data node to the metadata node actively, or the metadata node can obtain it from each data node actively.
[0070] The client determines the request processing capacity corresponding to the unit data volume of each data node based on the request processing capacity and the stored data volume of each data node obtained from the metadata node. Optionally, divide the request processing capacity of each data node by the stored data volume to obtain the request processing capacity corresponding to the unit data volume of each data node.
[0071] Optionally, the request processing capacity is the Queries Per Second (QPS), which represents the number of requests that a data node can respond to per second.
[0072] For example, the QPS of data node a and data node B is both 100. The disk size of data node a is 4TB, and the disk size of data node b is 8TB. Assume that the stored data volume of data node a is 4TB, and the stored data volume of data node b is also 4TB. Then, in the case of taking 1TB as the unit data volume, the QPS corresponding to the unit data volume of data node a is 100 / 4, and the QPS corresponding to the unit data volume of data node b is also 100 / 4. Assume that the user data stored in the disk of data node b continues to increase later and reaches a stored data volume of 8TB. Then, the QPS corresponding to the unit data volume of data node b is 100 / 8. It can be seen that as the stored data volume increases, the QPS corresponding to the unit data volume of the data node decreases.
[0073] After the client determines the data nodes where multiple replicas of the multi-replica data are located, based on the request processing capacity corresponding to the unit data volume of the data nodes where each replica is located, it determines the request probability for each replica, including: performing a normalization process on the request processing capacity corresponding to the unit data volume of the data nodes where each replica is located, and determining the request probability for each replica as the request processing capacity corresponding to the unit data volume of the data nodes where each replica is located after normalization.
[0074] Referring to the above example, when the stored data volume of data node a is 4TB and the stored data volume of data node b is also 4TB, the QPS corresponding to the unit data volume of data node a and data node b is both 100 / 4. After normalizing the QPS corresponding to the unit data volume of data node a and data node b, the result is 0.5 for both, that is, the request probability for the replicas on data node a and data node b is 0.5.
[0075] When the stored data volume of data node a is 4TB and the stored data volume of data node b is 8TB, the QPS corresponding to the unit data volume of data node a is 100 / 4, and the QPS corresponding to the unit data volume of data node b is 100 / 8. After normalizing the QPS corresponding to the unit data volume of data node a and data node b, the obtained results are 2 / 3 and 1 / 3 respectively, that is, the request probabilities for the replicas on data node a and data node b are 2 / 3 and 1 / 3 respectively.
[0076] After determining the request probabilities for each replica, according to the request probabilities for each replica, determine the random number range corresponding to each replica from the preset random number range, where the ratio of the random number ranges corresponding to each replica is the same as the ratio of the request probabilities; generate a random number within the preset random number range, and determine the replica corresponding to the random number range where the random number is located as the target replica.
[0077] Taking the request probabilities for the replicas on data node a and data node b as 2 / 3 and 1 / 3 respectively as an example, since the request probability for the replica on data node a is twice that of the replica on data node b, the random number range corresponding to the replica on data node a should also be twice that of the random number range corresponding to the replica on data node b. Assume that the preset random number range is 1 - 300, the random number range corresponding to the replica on data node a is 1 - 200, and the random number range corresponding to the replica on data node b is 201 - 300. Or, the random number range corresponding to the replica on data node a is 101 - 300, and the random number range corresponding to the replica on data node b is 1 - 100. As long as the ratio of the random number ranges corresponding to each replica meets the requirements, the specific division method is not limited in the embodiments of the present application.
[0078] Taking the random number range corresponding to the replica on data node a as 1 - 200 and the random number range corresponding to the replica on data node b as 201 - 300 as an example, when the client selects a replica, a random number is generated between 1 - 300. Assume that the generated random number is 165, and this random number falls within the random number range corresponding to the replica on data node a, then the replica on data node a is the target replica, and the client sends a read request for the target replica to data node a, thereby obtaining the corresponding data from data node a. Assume that the generated random number is 265, and this random number falls within the random number range corresponding to the replica on data node b, then the replica on data node b is the target replica, and the client sends a read request for the target replica to data node b, thereby obtaining the corresponding data from data node b.
[0079] Since the range of random numbers corresponding to the replicas on data node a is twice that of the replicas on data node b, when generating random numbers, the probability that the random numbers fall within the range of random numbers corresponding to the replicas on data node a is approximately twice that of data node b. As a result, the number of read requests received on data node a is approximately twice that of data node b. In this way, it matches the request processing capabilities corresponding to the unit data volume of data nodes a and b, enabling read requests to be reasonably allocated according to the processing capabilities of the data nodes, rather than having more requests for data nodes with larger data volumes. This ensures that the performance of each data node in the distributed system is reasonably utilized, improving the overall service capacity of the distributed system.
[0080] Figure 3 It is a schematic structural diagram of a multi-copy data processing device provided by an embodiment of the present application. As Figure 3 shown, the multi-copy data processing device 300 includes:
[0081] A node determination module 301, configured to determine data nodes where multiple replicas of multi-copy data are located in response to a processing instruction for the multi-copy data, where each of the multiple replicas is located on a different data node;
[0082] A probability calculation module 302, configured to determine the request probability for each replica based on the request processing capabilities corresponding to the unit data volume of the data nodes where each replica is located;
[0083] A request module 303, configured to select a target replica from the multiple replicas based on the request probability for each replica, and send a processing request for the target replica to the data node where the target replica is located.
[0084] In one implementation, the multi-copy data processing device 300 further includes:
[0085] An acquisition module, configured to periodically obtain the request processing capabilities and the stored data volumes of each data node from the metadata node;
[0086] A determination module, configured to determine the request processing capabilities corresponding to the unit data volume of each data node according to the request processing capabilities and the stored data volumes of each data node.
[0087] In one implementation, the determination module is configured to:
[0088] Divide the request processing capabilities of each data node by the stored data volume to obtain the request processing capabilities corresponding to the unit data volume of each data node.
[0089] In one implementation, the probability calculation module 302 is configured to:
[0090] Normalize the request processing capabilities corresponding to the unit data volume of the data nodes where each copy is located, and determine the request processing capabilities corresponding to the unit data volume of the data nodes where each copy is located after normalization as the request probabilities for each copy.
[0091] In one implementation, the request processing capability is the queries per second rate.
[0092] In one implementation, the request module 303 is used for:
[0093] Determine the random number ranges corresponding to each copy from a preset random number range according to the request probabilities for each copy, where the ratio of the random number ranges corresponding to each copy is the same as the ratio of the request probabilities;
[0094] Generate a random number within the preset random number range, and determine the copy corresponding to the random number range where the random number is located as the target copy.
[0095] In one implementation, the node determination module 301 is used for:
[0096] In response to a processing instruction triggered by an open operation on multi-copy data, obtain the data nodes where multiple copies of the multi-copy data are located from the metadata node.
[0097] The device according to the embodiments of the present application can be used to execute the multi-copy data processing method in the foregoing embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0098] Figure 4 It is a schematic block diagram of an electronic device provided by the embodiments of the present application. As Figure 4 shown, the electronic device 400 may include at least one processor 401, which is used to implement the multi-copy data processing method provided by the embodiments of the present application.
[0099] Optionally, the electronic device 400 further includes at least one memory 402, which is used to store program instructions and / or data. The memory 402 and the processor 401 are coupled. The coupling in the embodiments of the present application is an indirect coupling or communication connection between devices, units or modules, which can be electrical, mechanical or other forms, and is used for information interaction between devices, units or modules. The processor 401 may cooperate with the memory 402. The processor 401 may execute the program instructions stored in the memory 402. At least one of the at least one memory may be included in the processor.
[0100] Optionally, the electronic device 400 further includes a communication interface 403, configured to communicate with other devices via a transmission medium, so that the electronic device 400 can communicate with other devices. The communication interface 403 may be, for example, a transceiver, an interface, a bus, a circuit, or a device capable of implementing a transceiver function. The processor 401 can use the communication interface 403 to send and receive data and / or information, and is configured to implement the method provided in the embodiments of the present application. For specific details, refer to the detailed description in the foregoing embodiments, and details are not described herein again.
[0101] In the embodiments of the present application, the specific connection medium between the foregoing processor 401, memory 402, and communication interface 403 is not limited. In the embodiments of the present application Figure 4 it is assumed that the processor 401, memory 402, and communication interface 403 are connected via a bus 404. The bus 404 is Figure 4 represented by a thick line in [the figure]. The connection manners between other components are only schematically illustrated and are not to be taken as limiting. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 4 only one thick line is used to represent it in [the figure], but it does not mean that there is only one bus or one type of bus.
[0102] It should be understood that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the foregoing method embodiments may be completed by the integrated logic circuit in the hardware of the processor or by instructions in the form of software. The foregoing processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware decoding processor, or may be executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the foregoing method.
[0103] It should also be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.
[0104] The present application also provides a computer-readable storage medium, which stores a computer program (which can also be referred to as code, or instructions). When the computer program is run, it causes the computer to execute the method in any of the foregoing embodiments.
[0105] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method in any of the foregoing embodiments.
[0106] The terms "unit", "module", etc. used in this specification can be used to represent computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution.
[0107] Those of ordinary skill in the art will realize that the various illustrative logical blocks and steps described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application. In several embodiments provided in this application, it should be understood that the disclosed apparatus, device, and method can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.
[0108] The unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0109] In addition, each functional unit in the various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0110] In the above embodiments, the functions of each functional unit can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid-state disk (SSD)), etc.
[0111] If this function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs, etc., which can store program codes.
[0112] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0113] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.
Claims
1. A method for processing multi-copy data, characterized in that, including: responding to a processing instruction for multi-copy data, determining data nodes where multiple copies of the multi-copy data are located, wherein each of the multiple copies is located in a different data node; determining a request probability for each of the copies based on a request processing capacity corresponding to a unit data amount of the data node where each copy is located; selecting a target copy from the multiple copies based on the request probability for each of the copies, and sending a processing request for the target copy to the data node where the target copy is located.
2. The method according to claim 1, characterized in that, further including: periodically obtaining the request processing capacity and the stored data amount of each data node from a metadata node; determining a request processing capacity corresponding to a unit data amount of each data node according to the request processing capacity and the stored data amount of each data node.
3. The method according to claim 2, wherein The determining a request processing capacity corresponding to a unit data amount of each data node according to the request processing capacity and the stored data amount of each data node includes: dividing the request processing capacity of each data node by the stored data amount to obtain a request processing capacity corresponding to a unit data amount of each data node.
4. The method according to claim 1, characterized in that The determining a request probability for each of the copies based on a request processing capacity corresponding to a unit data amount of the data node where each copy is located includes: performing normalization processing on the request processing capacity corresponding to a unit data amount of the data node where each copy is located, and determining the request processing capacity corresponding to a unit data amount of the data node where each copy is located after normalization as the request probability for each of the copies.
5. The method according to any one of claims 1-4, characterized in that, The request processing capacity is the queries per second rate.
6. The method according to any one of claims 1-4, characterized in that, The selecting a target copy from the multiple copies based on the request probability for each of the copies includes: determining a random number range corresponding to each of the copies within a preset random number range according to the request probability for each of the copies, wherein the ratio of the random number ranges corresponding to each of the copies is the same as the ratio of the request probabilities; generating a random number within the preset random number range, and determining the copy corresponding to the random number range where the random number is located as the target copy.
7. The method according to any one of claims 1 to 4, characterized in that, The responding to a processing instruction for multi-copy data, determining data nodes where multiple copies of the multi-copy data are located includes: responding to a processing instruction triggered by an open operation on the multi-copy data, and obtaining data nodes where multiple copies of the multi-copy data are located from a metadata node.
8. A processing device for multi-copy data, characterized in that, including: a node determination module, configured to respond to a processing instruction for multi-copy data, and determine data nodes where multiple copies of the multi-copy data are located, wherein each of the multiple copies is located in a different data node; a probability calculation module, configured to determine a request probability for each of the copies based on a request processing capacity corresponding to a unit data amount of the data node where each copy is located; a request module, configured to select a target copy from the multiple copies based on the request probability for each of the copies, and send a processing request for the target copy to the data node where the target copy is located.
9. An electronic device, characterized in that, including: a memory and a processor; the memory is used for storing a computer program; The processor is configured to execute the computer program stored in the memory, and when the computer program runs, it causes the processor to execute the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, it causes the processor to execute the method according to any one of claims 1-7.
11. A computer program product, characterized in that, It includes a computer program which, when executed by the processor, implements the method according to any one of claims 1-7.