A business data storage method and apparatus
By determining and balancing the storage efficiency across nodes in a distributed storage cluster, the method addresses the issue of low data reliability and effectiveness in existing storage methods, enhancing user experience through efficient data replication.
Patent Information
- Application Number
- CN202210393314.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-04-15
AI Technical Summary
In the existing service data storage method, the master node does not consider the performance and load capacity of different nodes when sending service data to the slave node, resulting in low data reliability, poor storage effect and poor user experience.
By determining the storage efficiency of multiple slave nodes in a distributed storage cluster and storing service data in the master node, and sending a copy of service data to multiple slave nodes, the number and amount of data sent are controlled to ensure that the difference between the slave node with the highest storage efficiency and the lowest slave node is within the threshold range, data consistency and load balancing are achieved.
It improves data reliability and storage effect, improves user experience, ensures that multiple slave nodes can evenly complete data storage within similar time periods, and avoids node overload.
Smart Images

Figure CN114756172B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and apparatus for storing service data. Background Art
[0002] Distributed storage clusters are usually applied in the field of cloud storage technology. To prevent service data loss, in addition to storing service data in the master node of the distributed storage cluster, service data replicas are also stored in each slave node.
[0003] Taking a Raft (a distributed consensus protocol) cluster as an example, if the Raft cluster includes one master node (leader) and two slave nodes (FollowerA, FollowerB), the service data written by the user will be first sent to the leader of the distributed storage cluster for storage, and then the leader will distribute the service data to FollowerA and FollowerB to store service data replicas in the slave nodes to ensure distributed consistency.
[0004] There are at least the following problems in the prior art:
[0005] In the existing service data storage method, when the master node sends service data to the slave nodes, the performance and load capacity of different nodes are not considered, resulting in low data reliability, poor storage effect, and poor user experience. Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a method and apparatus for storing service data, which can store service data replicas in multiple slave nodes by combining the storage efficiency of different slave nodes, ensuring data reliability, improving the storage effect, and enhancing the user experience.
[0007] To achieve the above object, according to the first aspect of the embodiments of the present invention, a method for storing service data is provided, including:
[0008] Receiving service data to be stored, and storing the service data in the master node of the distributed storage cluster;
[0009] Determining the storage efficiency of multiple slave nodes in the distributed storage cluster;
[0010] Using the master node to send service data to multiple slave nodes to store service data replicas in multiple slave nodes respectively; wherein, the quantity of service data to be stored sent to any slave node does not exceed a first quantity threshold, and the difference between the quantity of service data to be stored in the slave node with the highest storage efficiency and the quantity of service data to be stored in the slave node with the lowest storage efficiency does not exceed a second quantity threshold.
[0011] Further, the step of using the master node to send service data to multiple slave nodes for separately storing the service data in the multiple slave nodes includes:
[0012] Number the service data stored in the master node;
[0013] Obtain the slave node storage schedule, and determine candidate slave nodes according to the slave node storage schedule and the first quantity threshold; wherein, the slave node storage schedule indicates the current quantity of service data to be stored in multiple slave nodes;
[0014] Calculate the quantity difference between the current quantity of service data to be stored in the candidate slave nodes and the current quantity of service data to be stored in the slave node with the lowest storage efficiency, and determine whether the quantity difference does not exceed the second quantity threshold. If so, determine the candidate slave nodes as the target slave nodes, and use the master node to send the target service data to the target slave nodes; wherein, the sum of the quantity of the target service data and the current quantity of service data to be stored in the target slave nodes does not exceed the first quantity threshold.
[0015] Further, the step of using the master node to send service data to multiple slave nodes for separately storing the service data in the multiple slave nodes further includes:
[0016] Determine the data volume corresponding to the numbered service data;
[0017] Obtain the slave node storage schedule, and determine candidate slave nodes according to the slave node storage schedule and the first data volume threshold; wherein, the slave node storage schedule indicates the data volume corresponding to the current service data to be stored in multiple slave nodes;
[0018] Calculate the data volume difference between the data volume of the current service data to be stored in the candidate slave nodes and the data volume of the current service data to be stored in the slave node with the lowest storage efficiency, and determine whether the data volume difference does not exceed the second data volume threshold. If so, determine the candidate slave nodes as the target slave nodes, and use the master node to send the target service data to the target slave nodes, wherein, the sum of the data volume of the target service data and the data volume of the current service data to be stored in the target slave nodes does not exceed the first data volume threshold.
[0019] Further, before the step of obtaining the slave node storage schedule, it further includes:
[0020] Obtain in real time the storage status of the service data to be stored in multiple slave nodes;
[0021] Construct a slave node storage schedule according to the storage status.
[0022] Further, it further includes:
[0023] Set a storage efficiency threshold, and determine the duration when the storage efficiency of the slave node is less than the storage efficiency threshold;
[0024] If the duration exceeds the duration threshold, adjust the slave nodes of the distributed storage cluster.
[0025] Further, the steps of determining the storage efficiency of multiple slave nodes in the distributed storage cluster include:
[0026] Determine the storage efficiency of multiple slave nodes according to the heartbeats between the master node and multiple slave nodes respectively.
[0027] Further, it also includes:
[0028] Update the storage efficiency according to the storage progress schedule of the slave nodes.
[0029] According to the second aspect of the embodiments of the present invention, a business data storage device is provided, including:
[0030] A receiving module, configured to: receive the business data to be stored and store the business data in the master node of the distributed storage cluster;
[0031] A storage efficiency determination module, configured to: determine the storage efficiency of multiple slave nodes in the distributed storage cluster;
[0032] A business data storage module, configured to: use the master node to send the business data to multiple slave nodes to store business data copies in multiple slave nodes respectively; wherein, the quantity of the business data to be stored sent to any slave node does not exceed the first quantity threshold, and the difference between the quantity of the business data to be stored in the slave node with the highest storage efficiency and the quantity of the business data to be stored in the slave node with the lowest storage efficiency does not exceed the second quantity threshold.
[0033] According to the third aspect of the embodiments of the present invention, an electronic device for business data storage is provided, including:
[0034] One or more processors;
[0035] A storage device, configured to store one or more programs,
[0036] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the above-mentioned business data storage methods.
[0037] According to the fourth aspect of the embodiments of the present invention, a computer-readable medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements any of the above-mentioned business data storage methods.
[0038] One embodiment of the above invention has the following advantages or beneficial effects: By receiving service data to be stored and storing the service data in the master node of the distributed storage cluster; determining the storage efficiency of multiple slave nodes in the distributed storage cluster; using the master node to send the service data to multiple slave nodes to store service data copies in the multiple slave nodes respectively; wherein, the quantity of service data to be stored sent to any slave node does not exceed the first quantity threshold, and the difference between the quantity of service data to be stored in the slave node with the highest storage efficiency and the quantity of service data to be stored in the slave node with the lowest storage efficiency does not exceed the second quantity threshold. Therefore, it overcomes the technical problems in the existing service data storage methods, such as low data reliability, poor storage effect, and poor user experience due to the failure to consider the performance and load capacity of different nodes when the master node sends service data to the slave nodes. Furthermore, it achieves the technical effect of combining the storage efficiency of different slave nodes, storing service data copies in multiple slave nodes, ensuring data reliability, improving the storage effect, and enhancing the user experience.
[0039] The further effects of the above non-conventional optional methods will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0041] Figure 1 is a schematic diagram of the main process of a service data storage method provided according to an embodiment of the present invention;
[0042] Figure 2a is a schematic diagram of the main process of a service data storage method provided according to another embodiment of the present invention;
[0043] Figure 2b is a schematic diagram of the main process of a service data storage method provided according to still another embodiment of the present invention;
[0044] Figure 3 is a schematic diagram of the main modules of a service data storage device provided according to another embodiment of the present invention;
[0045] Figure 4 is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;
[0046] Figure 5 is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0048] Figure 1 is a schematic diagram of the main process of a service data storage method provided according to an embodiment of the present invention; as Figure 1 shown, the service data storage method provided by the embodiment of the present invention is mainly applied to a distributed storage cluster including multiple nodes, and mainly includes:
[0049] Step S101, receive the service data to be stored and store the service data in the master node of the distributed storage cluster.
[0050] Specifically, according to the embodiment of the present invention, the above service data may be service data to be stored directly written by a user, or service data to be stored written via a service system. The master node of the distributed storage cluster receives the above service data to be stored, and then the master node distributes the service data to the slave nodes of the distributed storage cluster to store service data copies in the slave nodes, so as to store a copy of the service data in each of the full nodes of the distributed storage cluster, avoiding data loss.
[0051] Step S102, determine the storage efficiency of multiple slave nodes in the distributed storage cluster.
[0052] Specifically, according to the embodiment of the present invention, the step of determining the storage efficiency of multiple slave nodes in the distributed storage cluster includes:
[0053] Determine the storage efficiency of multiple slave nodes according to the heartbeats between the master node and the multiple slave nodes.
[0054] In specific implementation, if there is no service data copy stored in the slave node and the storage efficiency cannot be determined according to the actual storage situation, it can be determined through the heartbeat between the master node and the slave node.
[0055] Preferably, according to the embodiment of the present invention, the above method further includes:
[0056] Update the storage efficiency according to the storage progress schedule of the slave nodes.
[0057] After the standby slave nodes store all or part of the business data copies, the storage efficiency of the slave nodes can be determined according to the storage status recorded in the slave node storage schedule (recording the number of stored data copies within a certain time interval). According to the magnitude of the storage efficiency of the slave nodes, controlling the quantity or data volume of the business data to be stored sent to different slave nodes can ensure that the efficiency of storing business data copies in multiple slave nodes is as consistent as possible, thereby achieving data consistency between different nodes and improving the storage efficiency.
[0058] Step S103: Use the master node to send business data to multiple slave nodes to store business data copies in the multiple slave nodes respectively; wherein, the quantity of the business data to be stored sent to any slave node does not exceed the first quantity threshold, and the difference between the quantity of the business data to be stored in the slave node with the highest storage efficiency and the quantity of the business data to be stored in the slave node with the lowest storage efficiency does not exceed the second quantity threshold. Among them, the first quantity threshold can be set to 1000 as the limiting condition for restricting the maximum storage quantity per time; the second quantity threshold can be set to 256 as the limiting condition for judging the maximum storage progress difference between nodes. It should be noted that the above values are only examples, and the specific values of the first quantity threshold and the second quantity threshold can be adjusted according to the storage performance of the actual nodes and the requirements for data consistency.
[0059] Specifically, according to the embodiments of the present invention, the step of using the master node to send business data to multiple slave nodes to store business data in the multiple slave nodes respectively includes:
[0060] Number the business data stored in the master node;
[0061] Obtain the slave node storage schedule, and determine candidate slave nodes according to the slave node storage schedule and the first quantity threshold; wherein, the slave node storage schedule indicates the quantity of the business data to be stored currently in multiple slave nodes;
[0062] Calculate the quantity difference between the quantity of the business data to be stored currently in the candidate slave nodes and the quantity of the business data to be stored currently in the slave node with the lowest storage efficiency, and judge whether the quantity difference does not exceed the second quantity threshold. If so, determine the candidate slave node as the target slave node, and use the master node to send the target business data to the target slave node; wherein, the sum of the quantity of the target business data and the quantity of the business data to be stored currently in the target slave node does not exceed the first quantity threshold.
[0063] As described above, the storage efficiency of a slave node can be determined based on the storage status recorded in the slave node storage schedule. At the same time, since the slave node storage schedule also indicates the current number of service data to be stored by multiple slave nodes, a slave node with the highest storage efficiency can be selected as a candidate slave node from the slave nodes whose current number of service data to be stored does not exceed the first quantity threshold. Through the above settings, the service data stored in the master node is numbered, so as to control, from the perspective of the quantity of service data, the quantity difference between the number of service data to be stored in any slave node and the number of service data to be stored in the slave node with the lowest storage efficiency to be less than the second data threshold, ensuring that all service data is stored by multiple slave nodes within a similar time period to ensure data consistency. At the same time, it is controlled that the sum of the quantity of target service data sent from the master node to the slave node and the current number of service data to be stored does not exceed the first quantity threshold, avoiding overloading of the slave node.
[0064] Optionally, according to an embodiment of the present invention, the step of using the master node to send service data to multiple slave nodes to store service data in multiple slave nodes respectively further includes:
[0065] Determine the data volume corresponding to the numbered service data;
[0066] Obtain the slave node storage schedule, and determine candidate slave nodes according to the slave node storage schedule and the first data volume threshold; wherein, the slave node storage schedule indicates the data volumes corresponding to the current service data to be stored by multiple slave nodes;
[0067] Calculate the data volume difference between the data volume of the current service data to be stored in the candidate slave node and the data volume of the current service data to be stored in the slave node with the lowest storage efficiency, and determine whether the data volume difference does not exceed the second data volume threshold. If so, determine the candidate slave node as the target slave node, and use the master node to send the target service data to the target slave node, wherein the sum of the data volume of the target service data and the data volume of the current service data to be stored in the target slave node does not exceed the first data volume threshold.
[0068] Through the above settings, the data volume corresponding to the numbered service data is determined, so as to control, from the perspective of the data volume of service data, the data volume difference between the data volume of the service data to be stored in any slave node and the data volume of the service data to be stored in the slave node with the lowest storage efficiency to be less than the second data volume threshold, ensuring that all service data is stored by multiple slave nodes within a similar time period to ensure data consistency. At the same time, it is controlled that the sum of the data volume of the target service data sent from the master node to the slave node and the data volume of the current service data to be stored does not exceed the first data volume threshold, avoiding overloading of the slave node.
[0069] Further, according to the embodiments of the present invention, before the step of obtaining the storage schedule of the slave nodes, the above method further includes:
[0070] Obtain the storage status of the service data to be stored in multiple slave nodes in real time;
[0071] Construct a storage schedule for the slave nodes according to the storage status.
[0072] Through the above settings, by constructing a storage schedule for the slave nodes, the current number or data volume of the service data to be stored in the slave nodes can be learned in real time, which helps the master node to adjust the number or data volume of the service data sent to the slave nodes subsequently to ensure data consistency.
[0073] Exemplarily, according to the embodiments of the present invention, the above method further includes:
[0074] Set a storage efficiency threshold and determine the duration during which the storage efficiency of the slave node is less than the storage efficiency threshold;
[0075] If the duration exceeds the duration threshold, perform adjustment processing on the slave nodes of the distributed storage cluster.
[0076] Through the above settings, in order to further ensure the storage effect of the distributed storage cluster, if the storage efficiency of one or several slave nodes is lower than the storage efficiency threshold within a certain duration (the above duration threshold), the slave node can be maintained or new slave nodes can be added. That is, a storage efficiency threshold is set as a trigger condition for replacing or maintaining the slave nodes. The numerical value of the storage efficiency threshold can be set according to the actual situation. If the requirement for the storage efficiency of the slave nodes is not high, a lower storage efficiency threshold can be set; if the requirement for the storage efficiency of the slave nodes is high, a higher storage efficiency threshold needs to be set.
[0077] According to the technical solution of the embodiments of the present invention, since the method of receiving the service data to be stored and storing the service data in the master node of the distributed storage cluster; determining the storage efficiency of multiple slave nodes in the distributed storage cluster; using the master node to send the service data to multiple slave nodes to store service data copies in multiple slave nodes respectively; wherein, the number of service data to be stored sent to any slave node does not exceed the first quantity threshold, and the difference between the number of service data to be stored in the slave node with the highest storage efficiency and the number of service data to be stored in the slave node with the lowest storage efficiency does not exceed the second quantity threshold, the technical problems in the existing service data storage methods, such as low data reliability, poor storage effect, and poor user experience caused by the master node not considering the performance and load capacity of different nodes when sending service data to the slave nodes, are overcome. Furthermore, the technical effect of storing service data copies in multiple slave nodes by combining the storage efficiency of different slave nodes, ensuring data reliability, improving the storage effect, and enhancing the user experience is achieved.
[0078] Figure 2a is a schematic diagram of the main process of a service data storage method provided according to another embodiment of the present invention; as Figure 2a shown, the service data storage method provided by the embodiment of the present invention is applied to a distributed storage cluster, and mainly includes:
[0079] Step S201: Receive the service data to be stored, and store the service data in the master node of the distributed storage cluster.
[0080] Specifically, according to the embodiment of the present invention, the above service data may be service data to be stored directly written by a user, or service data to be stored written via a service system. The master node of the distributed storage cluster receives the above service data to be stored, and then the master node distributes the service data to the slave nodes of the distributed storage cluster to store service data copies in the slave nodes, so as to store a copy of service data in each of the full nodes of the distributed storage cluster, avoiding data loss.
[0081] Step S202: Determine the storage efficiency of multiple slave nodes according to the heartbeats between the master node and the multiple slave nodes respectively.
[0082] In specific implementation, if there is no service data copy stored in the slave node and the storage efficiency cannot be determined according to the actual storage situation, it can be determined through the heartbeat between the master node and the slave node.
[0083] Step S203: Number the service data stored in the master node; obtain the slave node storage progress table, and determine candidate slave nodes according to the slave node storage progress table and the first quantity threshold; wherein, the slave node storage progress table indicates the current quantity of service data to be stored of multiple slave nodes.
[0084] Step S204: Calculate the quantity difference between the current quantity of service data to be stored in the candidate slave nodes and the current quantity of service data to be stored in the slave node with the lowest storage efficiency.
[0085] Step S205: Determine whether the quantity difference does not exceed the second quantity threshold; if so, that is, the quantity difference does not exceed the second quantity threshold, then execute Step S206; if not, that is, the quantity difference exceeds the second quantity threshold, then go to Step S203.
[0086] Step S206: Determine the candidate slave node as the target slave node, and use the master node to send the target service data to the target slave node; wherein, the quantity of the target service data and the current quantity of service data to be stored in the target slave node do not exceed the first quantity threshold.
[0087] Through the above settings, the business data stored in the master node is numbered, so as to control, from the perspective of the quantity of business data in the subsequent process, the quantity difference between the quantity of business data to be stored in any slave node and the quantity of business data to be stored in the slave node with the lowest storage efficiency to be less than the second data threshold, ensuring that multiple slave nodes all complete the storage of all business data within a similar time period to ensure data consistency. At the same time, it is controlled that the sum of the quantity of target business data sent from the master node to the slave node and the quantity of current business data to be stored does not exceed the first quantity threshold, avoiding overloading of the slave nodes.
[0088] Further, according to the embodiment of the present invention, before the step of obtaining the storage progress schedule of the slave nodes, the above method further includes:
[0089] Obtaining in real time the storage status of the business data to be stored in multiple slave nodes;
[0090] Constructing a storage progress schedule of the slave nodes according to the storage status.
[0091] Through the above settings, by constructing a storage progress schedule of the slave nodes, the current quantity of business data to be stored or the data volume of the current business data to be stored in the slave nodes can be learned in real time, which helps the master node to adjust the quantity or data volume of business data sent to the slave nodes subsequently to ensure data consistency.
[0092] Exemplarily, according to the embodiment of the present invention, the above method further includes:
[0093] Setting a storage efficiency threshold and determining the duration during which the storage efficiency of the slave node is less than the storage efficiency threshold;
[0094] If the duration exceeds the duration threshold, performing an adjustment process on the slave nodes of the distributed storage cluster.
[0095] Through the above settings, in order to further ensure the storage effect of the distributed storage cluster, if the storage efficiency of one or several slave nodes is lower than the storage efficiency threshold within a certain duration (the above duration threshold), the slave nodes can be maintained or new slave nodes can be added.
[0096] See Figure 2b , the embodiment of the present invention provides another business data storage method, which mainly includes:
[0097] Step S211: Receiving the business data to be stored and storing the business data in the master node of the distributed storage cluster.
[0098] Step S212: Determining the storage efficiency of multiple slave nodes according to the heartbeats between the master node and multiple slave nodes respectively and the storage progress schedule of the slave nodes.
[0099] Step S213, number the service data stored in the master node; obtain the slave node storage schedule, and determine candidate slave nodes according to the slave node storage schedule and the first quantity threshold; wherein, the slave node storage schedule indicates the current quantity of service data to be stored by multiple slave nodes.
[0100] Step S214, calculate the quantity difference between the current quantity of service data to be stored in the candidate slave nodes and the current quantity of service data to be stored in the slave node with the lowest storage efficiency.
[0101] Step S215, determine whether the quantity difference does not exceed the second quantity threshold; if so, that is, the quantity difference does not exceed the second quantity threshold, then execute Step S216; if not, that is, the quantity difference exceeds the second quantity threshold, then go to Step S213.
[0102] Step S216, determine the data volume corresponding to the numbered service data; obtain the slave node storage schedule, and determine candidate slave nodes according to the slave node storage schedule and the first data volume threshold; wherein, the slave node storage schedule indicates the data volumes of the current service data to be stored by multiple slave nodes.
[0103] Step S217, calculate the data volume difference between the data volumes of the current service data to be stored in the candidate slave nodes and the data volumes of the current service data to be stored in the slave node with the lowest storage efficiency.
[0104] Step S218, determine whether the data volume difference does not exceed the second data volume threshold, if so, execute Step S219; if not, go to Step S216;
[0105] Step S219, determine the candidate slave nodes as the target slave nodes, and use the master node to send the target service data to the target slave nodes, wherein the sum of the data volume of the target service data and the data volume of the current service data to be stored in the target slave node does not exceed the first data volume threshold; the sum of the quantity of the target service data and the quantity of the current service data to be stored in the target slave node does not exceed the first quantity threshold.
[0106] Through the above settings, number the service data stored in the master node, and determine the data volume corresponding to the numbered service data, so as to control, from the two perspectives of the quantity and data volume of the service data, that is, to control the quantity difference between the quantity of service data to be stored in any slave node and the quantity of service data to be stored in the slave node with the lowest storage efficiency to be less than the second data threshold, and also to control the data volume difference between the data volume of the service data to be stored in any slave node and the data volume of the service data to be stored in the slave node with the lowest storage efficiency to be less than the second data volume threshold, ensuring that all service data is stored in multiple slave nodes within a similar time period to ensure data consistency.
[0107] Meanwhile, control the sum of the quantity of target service data sent from the master node to the slave node and the quantity of service data to be stored currently so as not to exceed the first quantity threshold, and control the sum of the data volume of the target service data sent from the master node to the slave node and the data volume of the service data to be stored currently so as not to exceed the first data volume threshold, significantly avoiding overloading of the slave node.
[0108] According to the embodiments of the present invention, it can be understood that, from the perspective of only the data volume of the slave service data, it is also possible to ensure that all service data is stored in multiple slave nodes within a similar time period to achieve data consistency.
[0109] According to the technical solution of the embodiments of the present invention, by adopting the technical means of receiving service data to be stored, storing the service data in the master node of the distributed storage cluster; determining the storage efficiency of multiple slave nodes in the distributed storage cluster; using the master node to send service data to multiple slave nodes to store service data copies in the multiple slave nodes respectively; wherein, the quantity of service data to be stored sent to any slave node does not exceed the first quantity threshold, and the difference between the quantity of service data to be stored in the slave node with the highest storage efficiency and the quantity of service data to be stored in the slave node with the lowest storage efficiency does not exceed the second quantity threshold, the technical problems in the existing service data storage method are overcome. When the master node sends service data to the slave node, the performance and load capacity of different nodes are not considered, resulting in low data reliability, poor storage effect, and poor user experience. Furthermore, the technical effect of combining the storage efficiency of different slave nodes, storing service data copies in multiple slave nodes, ensuring data reliability, improving the storage effect, and enhancing the user experience is achieved.
[0110] Figure 3 It is a schematic diagram of the main modules of a service data storage device provided by another embodiment of the present invention; as Figure 3 shown, the service data storage device 300 provided by the embodiments of the present invention is mainly arranged in a distributed storage cluster including multiple nodes, and mainly includes:
[0111] A receiving module 301, configured to: receive service data to be stored and store the service data in the master node of the distributed storage cluster.
[0112] Specifically, according to the embodiments of the present invention, the above service data may be service data to be stored directly written by a user, or service data to be stored written via a service system. The master node of the distributed storage cluster receives the above service data to be stored, and then the master node distributes the service data to the slave nodes of the distributed storage cluster to store service data copies in the slave nodes, realizing that a copy of service data is stored in each node of the distributed storage cluster to avoid data loss.
[0113] A storage efficiency determination module 302, configured to: determine the storage efficiencies of multiple slave nodes in a distributed storage cluster.
[0114] Specifically, according to an embodiment of the present invention, the above-mentioned storage efficiency determination module 302 is further configured to:
[0115] Determine the storage efficiencies of multiple slave nodes according to the heartbeats between the master node and the multiple slave nodes respectively.
[0116] In specific implementation, if the slave nodes have not stored business data replicas and the storage efficiency cannot be determined according to the actual storage situation, it can be determined through the heartbeats between the master node and the slave nodes.
[0117] Preferably, according to an embodiment of the present invention, the above-mentioned storage efficiency determination module 302 is further configured to:
[0118] Update the storage efficiency according to the storage progress schedule of the slave nodes.
[0119] After all or part of the business data replicas have been stored in the slave nodes, the storage efficiency of the slave nodes can be determined according to the storage situation recorded in the storage progress schedule of the slave nodes (recording the number of stored data replicas within a certain time interval). According to the magnitudes of the storage efficiencies of the slave nodes, controlling the quantity or data volume of the business data to be stored sent to different slave nodes can ensure that the efficiency of storing business data replicas in multiple slave nodes is kept as consistent as possible, thereby achieving data consistency between different nodes and improving storage efficiency.
[0120] A business data storage module 303, configured to: use the master node to send business data to multiple slave nodes to store business data replicas in the multiple slave nodes respectively; wherein, the quantity of the business data to be stored sent to any slave node does not exceed a first quantity threshold, and the difference between the quantity of the business data to be stored in the slave node with the highest storage efficiency and the quantity of the business data to be stored in the slave node with the lowest storage efficiency does not exceed a second quantity threshold. Among them, the first quantity threshold can be set to 1000 as a limiting condition for restricting the maximum storage quantity per time; the second quantity threshold can be set to 256 as a limiting condition for judging the maximum storage progress difference between nodes. It should be noted that the above values are only examples, and the specific values of the first quantity threshold and the second quantity threshold can be adjusted according to the storage performance of the actual nodes and the requirements for data consistency.
[0121] Specifically, according to an embodiment of the present invention, the above-mentioned business data storage module 303 is further configured to:
[0122] Number the business data stored in the master node;
[0123] Obtain the storage schedule of the slave nodes, and determine candidate slave nodes according to the storage schedule of the slave nodes and the first quantity threshold; wherein, the storage schedule of the slave nodes indicates the current quantity of service data to be stored by multiple slave nodes;
[0124] Calculate the quantity difference between the current quantity of service data to be stored in the candidate slave nodes and the current quantity of service data to be stored in the slave node with the lowest storage efficiency, and determine whether the quantity difference does not exceed the second quantity threshold. If so, determine the candidate slave node as the target slave node, and use the master node to send the target service data to the target slave node; wherein, the sum of the quantity of the target service data and the current quantity of service data to be stored in the target slave node does not exceed the first quantity threshold.
[0125] As mentioned in the previous description, the storage efficiency of the slave nodes can be determined according to the storage status recorded in the storage schedule of the slave nodes. At the same time, since the storage schedule of the slave nodes also indicates the current quantity of service data to be stored by multiple slave nodes, therefore, the slave node with the highest storage efficiency can be selected as the candidate slave node from the slave nodes whose current quantity of service data to be stored does not exceed the first quantity threshold. Through the above settings, the service data stored in the master node is numbered, so as to control, from the perspective of the quantity of service data in the subsequent process, the quantity difference between the quantity of service data to be stored in any slave node and the quantity of service data to be stored in the slave node with the lowest storage efficiency to be less than the second data threshold, ensuring that multiple slave nodes complete the storage of all service data within a similar time period to ensure data consistency. At the same time, control the sum of the quantity of the target service data sent from the master node to the slave node and the current quantity of service data to be stored not to exceed the first quantity threshold, avoiding overloading of the slave nodes.
[0126] Optionally, according to the embodiment of the present invention, the above service data storage module 303 is further configured to:
[0127] Determine the data volume corresponding to the numbered service data;
[0128] Obtain the storage schedule of the slave nodes, and determine candidate slave nodes according to the storage schedule of the slave nodes and the first data volume threshold; wherein, the storage schedule of the slave nodes indicates the data volume corresponding to the current service data to be stored by multiple slave nodes;
[0129] Calculate the data volume difference between the data volume of the current service data to be stored in the candidate slave nodes and the data volume of the current service data to be stored in the slave node with the lowest storage efficiency, and determine whether the data volume difference does not exceed the second data volume threshold. If so, determine the candidate slave node as the target slave node, and use the master node to send the target service data to the target slave node, wherein, the sum of the data volume of the target service data and the data volume of the current service data to be stored in the target slave node does not exceed the first data volume threshold.
[0130] Through the above settings, the data volume corresponding to the numbered service data is determined, so as to control, from the perspective of the data volume of the service data, the data volume difference between the data volume of the service data to be stored in any slave node and the data volume of the service data to be stored in the slave node with the lowest storage efficiency to be less than the second data volume threshold, ensuring that all service data is stored in multiple slave nodes within a similar time period to ensure data consistency. At the same time, it is controlled that the sum of the data volume of the target service data sent from the master node to the slave node and the data volume of the currently to-be-stored service data does not exceed the first data volume threshold, avoiding overloading of the slave nodes.
[0131] Further, according to an embodiment of the present invention, the above service data storage device 300 further includes a slave node storage schedule construction module for:
[0132] Before the step of obtaining the slave node storage schedule, the above method further includes:
[0133] Obtain the storage status of the service data to be stored in multiple slave nodes in real time;
[0134] Construct a slave node storage schedule according to the storage status.
[0135] Through the above settings, by constructing a slave node storage schedule, the current number or data volume of the service data to be stored in the slave node can be learned in real time, which helps the master node to adjust the number or data volume of the service data sent to different slave nodes subsequently to ensure data consistency.
[0136] Exemplarily, according to an embodiment of the present invention, the above service data storage device 300 further includes a slave node adjustment module for:
[0137] Set a storage efficiency threshold and determine the duration during which the storage efficiency of the slave node is less than the storage efficiency threshold;
[0138] If the duration exceeds the duration threshold, perform adjustment processing on the slave nodes of the distributed storage cluster.
[0139] Through the above settings, in order to further ensure the storage effect of the distributed storage cluster, if the storage efficiency of one or several slave nodes is lower than the storage efficiency threshold within a certain duration (the above duration threshold), the slave node can be maintained or new slave nodes can be added. That is, a storage efficiency threshold is set as a trigger condition for replacing or maintaining the slave node, and the numerical value of the storage efficiency threshold can be set according to the actual situation. If the requirement for the storage efficiency of the slave node is not high, a lower storage efficiency threshold can be set; if the requirement for the storage efficiency of the slave node is high, a higher storage efficiency threshold needs to be set.
[0140] According to the technical solution of the embodiment of the present invention, by receiving service data to be stored and storing the service data in the master node of the distributed storage cluster; determining the storage efficiency of multiple slave nodes in the distributed storage cluster; and using the master node to send the service data to the multiple slave nodes to store service data copies in the multiple slave nodes respectively, wherein the quantity of service data to be stored sent to any slave node does not exceed a first quantity threshold, and the difference between the quantity of service data to be stored in the slave node with the highest storage efficiency and the quantity of service data to be stored in the slave node with the lowest storage efficiency does not exceed a second quantity threshold, the technical problems in the existing service data storage method, such as low data reliability, poor storage effect, and poor user experience caused by not considering the performance and load capacity of different nodes when the master node sends service data to the slave nodes, are overcome. Furthermore, the technical effect of combining the storage efficiency of different slave nodes to store service data copies in multiple slave nodes, ensuring data reliability, improving the storage effect, and enhancing the user experience is achieved.
[0141] Figure 4 An exemplary system architecture 400 to which the service data storage method or service data storage device according to the embodiment of the present invention can be applied is shown.
[0142] As Figure 4 shown, the system architecture 400 may include terminal devices 401, 402, 403, a network 404, and a server 405 (this architecture is only an example, and the components included in the specific architecture can be adjusted according to the specific situation of the application). The network 404 is used to provide a medium for communication links between the terminal devices 401, 402, 403 and the server 405. The network 404 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0143] Users can use the terminal devices 401, 402, 403 to interact with the server 405 through the network 404 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 401, 402, 403, such as service data storage applications, web browser applications, search applications, instant messaging tools, database clusters, social platform software, etc. (only examples).
[0144] The terminal devices 401, 402, 403 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0145] The server 405 may be a server that provides various services, such as a server (only for example) that stores business data / processes data for a user using the terminal devices 401, 402, and 403. The server may analyze and process data such as the business data to be stored received, and feedback the processing results (such as storage efficiency, business data - only for example) to the terminal devices.
[0146] It should be noted that the business data storage method provided by the embodiments of the present invention is generally executed by the server 405. Correspondingly, the business data storage device is generally arranged in the server 405.
[0147] It should be understood that Figure 4 the numbers of the terminal devices, the network, and the server in
[0148] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 5 Figure 5 The terminal device or server shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0149] As Figure 5 shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the system 500 are also stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0150] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as required. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as required so that the computer program read from it can be installed into the storage section 508 as required.
[0151] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment disclosed in the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above functions defined in the system of the present invention are executed.
[0152] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0153] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0154] The modules described in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a receiving module, a storage efficiency determination module, and a service data storage module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases. For example, the receiving module can also be described as "a module for receiving service data to be stored and storing the service data in the master node of a distributed storage cluster".
[0155] On the other hand, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or can exist separately and not be assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: receiving service data to be stored and storing the service data in the master node of a distributed storage cluster; determining the storage efficiency of multiple slave nodes in the distributed storage cluster; using the master node to send the service data to multiple slave nodes to store service data copies in the multiple slave nodes respectively; where the amount of service data to be stored sent to any slave node does not exceed a first quantity threshold, and the difference between the amount of service data to be stored in the slave node with the highest storage efficiency and the amount of service data to be stored in the slave node with the lowest storage efficiency does not exceed a second quantity threshold.
[0156] According to the technical solution of the embodiment of the present invention, the service data to be stored is received and stored in the master node of the distributed storage cluster; the storage efficiency of multiple slave nodes in the distributed storage cluster is determined; the master node is used to send the service data to the multiple slave nodes to store service data copies in the multiple slave nodes respectively; wherein, the quantity of the service data to be stored sent to any slave node does not exceed the first quantity threshold, and the difference between the quantity of the service data to be stored in the slave node with the highest storage efficiency and the quantity of the service data to be stored in the slave node with the lowest storage efficiency does not exceed the second quantity threshold. Therefore, the technical problem in the existing service data storage method that due to the fact that when the master node sends service data to the slave nodes, the performance and load capacity of different nodes are not considered, resulting in low data reliability, poor storage effect and poor user experience is overcome. Furthermore, the technical effect of combining the storage efficiency of different slave nodes, storing service data copies in multiple slave nodes, ensuring data reliability, improving the storage effect and enhancing the user experience is achieved.
[0157] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for storing service data, characterized in that, Including: Receiving service data to be stored and storing the service data in the master node of the distributed storage cluster; Determining the storage efficiency of multiple slave nodes in the distributed storage cluster; Using the master node to send service data to the multiple slave nodes to store service data copies in the multiple slave nodes respectively; wherein, the quantity of service data to be stored sent to any slave node does not exceed a first quantity threshold, and the difference between the quantity of service data to be stored in the slave node with the highest storage efficiency and the quantity of service data to be stored in the slave node with the lowest storage efficiency does not exceed a second quantity threshold.
2. The service data storage method according to claim 1, wherein The step of using the master node to send service data to the multiple slave nodes to store service data copies in the multiple slave nodes respectively includes: Numbering the service data stored in the master node; Obtaining the slave node storage schedule, and determining candidate slave nodes according to the slave node storage schedule and the first quantity threshold; wherein, the slave node storage schedule indicates the current quantity of service data to be stored of multiple slave nodes; Calculating the quantity difference between the current quantity of service data to be stored in the candidate slave nodes and the current quantity of service data to be stored in the slave node with the lowest storage efficiency, and determining whether the quantity difference does not exceed the second quantity threshold. If so, determining the candidate slave nodes as target slave nodes, and using the master node to send target service data to the target slave nodes; wherein, the quantity of the target service data and the current quantity of service data to be stored in the target slave nodes do not exceed the first quantity threshold.
3. The service data storage method according to claim 2, wherein The step of using the master node to send service data to the multiple slave nodes to store service data copies in the multiple slave nodes respectively further includes: Determining the data volume corresponding to the numbered service data; Obtaining the slave node storage schedule, and determining candidate slave nodes according to the slave node storage schedule and the first data volume threshold; wherein, the slave node storage schedule indicates the data volume corresponding to the current service data to be stored of multiple slave nodes; Calculating the data volume difference between the data volume of the current service data to be stored in the candidate slave nodes and the data volume of the current service data to be stored in the slave node with the lowest storage efficiency, and determining whether the data volume difference does not exceed the second data volume threshold. If so, determining the candidate slave nodes as target slave nodes, and using the master node to send target service data to the target slave nodes, wherein, the data volume of the target service data and the data volume of the current service data to be stored in the target slave nodes do not exceed the first data volume threshold.
4. The service data storage method according to claim 2, wherein Before the step of obtaining the slave node storage schedule, it further includes: Obtaining in real time the storage status of the service data to be stored in the multiple slave nodes; Constructing the slave node storage schedule according to the storage status.
5. The service data storage method according to claim 2, wherein It further includes: Setting a storage efficiency threshold, and determining the duration when the storage efficiency of the slave node is less than the storage efficiency threshold; If the duration exceeds the duration threshold, performing adjustment processing on the slave nodes of the distributed storage cluster.
6. The service data storage method according to claim 1, characterized in that, The step of determining the storage efficiency of multiple slave nodes in the distributed storage cluster includes: Determine the storage efficiency of the multiple slave nodes according to the heartbeats between the master node and the multiple slave nodes respectively.
7. The service data storage method according to claim 6, wherein Further comprising: Update the storage efficiency according to the storage progress schedule of the slave nodes.
8. A business data storage device, characterized in that, Comprising: A receiving module, configured to: receive service data to be stored, and store the service data in the master node of the distributed storage cluster; A storage efficiency determination module, configured to: determine the storage efficiency of multiple slave nodes in the distributed storage cluster; A service data storage module, configured to: use the master node to send service data to the multiple slave nodes to store service data copies in the multiple slave nodes respectively; wherein, the quantity of service data to be stored sent to any slave node does not exceed a first quantity threshold, and the difference between the quantity of service data to be stored in the slave node with the highest storage efficiency and the quantity of service data to be stored in the slave node with the lowest storage efficiency does not exceed a second quantity threshold.
9. An electronic device for storing service data, characterized in that, Comprising: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Data storage method and device, main node and distributed system
CN111552441A
Apparatus and method for facilitating a transfer of container between slave nodes
KR101626067B1