Method and device for optimizing performance of cache server cluster
By dynamically adjusting the processing weight and data replica distribution of the cache server, the performance bottleneck of the distributed cache server cluster when processing massive data requests is solved, and more efficient load balancing and data access efficiency are achieved.
Patent Information
- Application Number
- CN202411979147.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional distributed cache server clusters will face performance bottlenecks when processing massive data requests. How to improve their performance to serve more users is an urgent problem.
By obtaining the service data of each cache server, its processing weight at the next moment is calculated, and its requested number is adjusted according to that weight. In addition, the number and location of its data replicas are adjusted based on the number of requests of the cache server and the preset data replica strategy.
Dynamic load balancing is realized, the concurrent processing capability and data access efficiency of the cluster are improved, network delay is reduced, and data availability and system stability are ensured through adaptive data replica strategies.
Smart Images

Figure CN119996517A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of distributed cache technology, and in particular to a performance optimization method and device for a cache server cluster. Background Art
[0002] As the amount of data increases, the system needs to be able to process and access this data quickly and efficiently. The cache server acts as an intermediary server in the network. It stores frequently accessed data in the local network. In this way, when users access data, they do not need to go through the Internet, which greatly improves the access speed of data.
[0003] Distributed cache server clusters are widely used in the business of large Internet companies due to their high availability, high performance, and good scalability. They provide data services to users through the collaboration of multiple cache server nodes. However, with the rapid growth of business volume and the sharp expansion of data scale, traditional distributed cache server clusters will face performance bottlenecks when processing massive data requests. How to improve the performance of distributed cache server clusters and serve more users is an urgent problem to be solved. Summary of the invention
[0004] The present disclosure provides a method and device for optimizing the performance of a cache server cluster, so as to at least solve the above technical problems existing in the prior art.
[0005] According to a first aspect of the present disclosure, a method for optimizing the performance of a cache server cluster is provided, the method comprising:
[0006] Get the service data of each cache server in the cache server cluster;
[0007] Calculating a processing weight of the corresponding cache server at a next moment according to the service data, and adjusting the number of requests of the corresponding cache server based on the processing weight;
[0008] Based on the number of requests to the cache server and a preset data copy strategy, the number and location of the data copies of the corresponding cache server are adjusted.
[0009] In one possible implementation, a monitoring device is provided in the cache server cluster, and the monitoring device is used to obtain service data of each cache server.
[0010] In one possible implementation manner, the calculating the processing weight of the corresponding cache server according to the service data includes:
[0011] Calculating a comprehensive load index of the corresponding cache server according to the server data; the comprehensive load index is used to characterize the load condition of the cache server;
[0012] The processing weight of the cache server at the next moment is calculated based on the comprehensive load index.
[0013] In one possible implementation, the comprehensive load index of the corresponding cache server is calculated in the following manner:
[0014] Load_i=w_cpu*cpu_i+w_mem*mem_i+w_net*net_i
[0015] The processing weight of the cache server at the next moment is calculated in the following manner:
[0016]
[0017] Among them, cpu_i is the controller utilization rate of server i; mem_i is the memory utilization rate; net_i is the network bandwidth utilization rate; Load_i is the comprehensive load index; w_cpu is the weight parameter of the controller; w_mem is the weight parameter of the memory; w_net is the weight parameter of the network bandwidth; K is a constant used to adjust the size of the processing weight; ε is a minimum value; Weight_i is the processing weight of server i at the next moment.
[0018] In one possible implementation, based on the number of requests of the cache server and a preset data copy strategy, adjusting the number and location of the data copies of the corresponding cache server includes:
[0019] Obtaining the access frequency of data and the comprehensive load index of the cache server from the access log of the cache server;
[0020] Calculating the expected number of copies of the data based on the access frequency;
[0021] The number of replicas allocated to each cache server is calculated according to the comprehensive load index and the number of replicas.
[0022] In one possible implementation, the expected number of replicas of data is calculated in the following manner:
[0023] TargetReplica_j=round(M*log(access_j+ε))
[0024] The number of replicas assigned to each cache server is calculated as follows:
[0025]
[0026] Among them, TargetReplica_j is the expected number of replicas of data j; M is a normal number used to adjust the number of replicas; log() ensures that the relationship between the number of replicas and the access frequency is logarithmic; ε is the minimum value; Expect_i is the expected number of replicas of server i; access_j is the number of accesses to data item j within a certain period of time.
[0027] In one possible implementation, the cache server cluster uses a lock-free data structure.
[0028] In one possible implementation, when the request includes large-scale data, before adjusting the number and location of the data copies of the corresponding cache server based on the number of requests of the cache server and the preset data copy strategy, the method further includes:
[0029] Divide large-scale data into small blocks and distribute each small block to each cache server.
[0030] In one possible implementation, dividing the large-scale data into small blocks of data and distributing each small block of data to each cache server includes:
[0031] Dividing the large-scale data into a plurality of small blocks of data;
[0032] Performing hash function calculation on each of the small blocks of data to obtain a hash value for each small block of data;
[0033] Determine the cache server where each small piece of data is stored based on the number of cache servers and the hash value;
[0034] Distribute each small block of data to the corresponding cache server.
[0035] According to a second aspect of the present disclosure, a performance optimization device for a cache server cluster is provided, the device comprising:
[0036] A data acquisition module, used to acquire service data of each cache server in the cache server cluster;
[0037] A first adjustment module, configured to calculate a processing weight of a corresponding cache server at a next moment according to the service data, and adjust the number of requests of the corresponding cache server based on the processing weight;
[0038] The second adjustment module is used to adjust the number and location of the data copies of the corresponding cache server based on the number of requests of the cache server and a preset data copy strategy.
[0039] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0040] at least one processor; and
[0041] a memory communicatively connected to the at least one processor; wherein,
[0042] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present disclosure.
[0043] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described in the present disclosure.
[0044] The performance optimization method and device of the cache server cluster disclosed in the present invention first obtain the service data of each cache server in the server cluster, thereby calculating the processing weight of the cache server at the next moment, and then adjust the request data of the corresponding cache server, and then adjust the number and location of the data copies of the cache server according to the preset data copy strategy. The present invention improves the concurrent processing capability and data access efficiency of the cluster through dynamic load balancing, effectively reduces network latency, and improves the performance of the cluster. By introducing an adaptive data copy strategy, the distribution of copies is dynamically adjusted according to the load of the server and the access frequency of the data. Even if some servers fail, the data availability and stability of the cluster system can still be guaranteed, and the concurrent processing capability and data access efficiency of the cluster can also be improved.
[0045] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which:
[0047] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.
[0048] Figure 1 The schematic diagram shows the implementation process of the performance optimization method of the cache server cluster in the embodiment of the present disclosure. Figure 1 ;
[0049] Figure 2 The schematic diagram shows the implementation process of the performance optimization method of the cache server cluster in the embodiment of the present disclosure. Figure 2 ;
[0050] Figure 3 The schematic diagram shows the implementation process of the performance optimization method of the cache server cluster in the embodiment of the present disclosure. Figure 3 ;
[0051] Figure 4 The schematic diagram shows the implementation process of the performance optimization method of the cache server cluster in the embodiment of the present disclosure. Figure 4 ;
[0052] Figure 5 A schematic diagram showing the structure of a performance optimization device for a cache server cluster according to an embodiment of the present disclosure is shown;
[0053] Figure 6 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0054] In order to make the purpose, features, and advantages of the present disclosure more obvious and easy to understand, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0055] In a distributed cache server cluster, differences in hardware configuration and network environment of each server node will lead to load imbalance between server nodes. Once load imbalance occurs, some servers will be overloaded while other servers will be idle, and the overall processing speed and performance will be affected. Therefore, how to achieve load balancing in a distributed environment is an important issue to improve the performance of a distributed cache server cluster.
[0056] In addition, for distributed situations, data storage and access strategies are also very important. In traditional situations, data is usually replicated and stored on different servers to improve data availability. However, although this method can improve data availability, it will take up additional storage space, and data consistency is also a problem.
[0057] In a distributed cache server cluster, different servers may undertake different tasks, and the corresponding network configuration needs to be performed according to the nature of the tasks. Among them, data storage and reading servers are mainly responsible for storing and reading data, and need to have a large storage space and high-speed reading and writing capabilities. In terms of network configuration, it is necessary to ensure the communication speed with other servers in order to efficiently transmit data. For request processing and distribution servers, they are mainly responsible for processing client requests, including distributing requests to the corresponding data storage servers and aggregating the processing results of each server and returning them to the client. Since this type of server needs to process a large number of network requests, the network configuration must have a high network bandwidth and ensure low latency. For computing and analysis servers, a large number of computing tasks may need to be performed for analyzing and processing data. In terms of network configuration, this type of server may need to configure more advanced network hardware in addition to ensuring the communication speed with other servers to support more complex parallel computing and distributed processing. For backup and recovery servers, they may undertake the task of data backup and recovery. In terms of network configuration, this type of server needs to maintain a stable network connection and a large network bandwidth for large amounts of data transmission.
[0058] A performance optimization method and device for a cache server cluster provided by the present disclosure are described below in conjunction with the accompanying drawings.
[0059] like Figure 1 As shown, the present disclosure provides a performance optimization method for a cache server cluster, the method comprising:
[0060] S101, obtaining service data of each cache server in the cache server cluster;
[0061] It is understandable that the present disclosure first builds a cluster. Specifically, the present disclosure builds and starts multiple cache server nodes according to application and business requirements and hardware resources to form a cache server cluster. Configure the network according to different performance and configuration to undertake different tasks: For the cache server cluster, network configuration is a very important part. It is necessary to ensure that the network communication between servers is unimpeded, and to ensure low latency and high bandwidth to achieve the goal of fast read and write cache.
[0062] In the present disclosure, a monitoring device is installed in the constructed cache server cluster to collect server data of each cache server to analyze the load and performance data of each cache node. Specifically, the real-time monitoring device regularly collects the performance status data of each cache server, such as the controller CPU usage, memory usage, network bandwidth usage, etc. Then the collected monitoring data is summarized and analyzed to obtain the load status of each server node.
[0063] S102, calculating a processing weight of the corresponding cache server at the next moment according to the service data, and adjusting the number of requests of the corresponding cache server based on the processing weight;
[0064] It should be noted that if the load of a server is too high, its weight will be reduced to allow other servers to take over more requests, so as to achieve load balancing for the entire system. In the present disclosure, the processing weight of each cache server at the next moment can be calculated through service data, in order to adjust its processing weight for requests according to the load conditions of the server. The load index value of a server with a higher load is larger, and the calculated processing weight will be smaller, and fewer requests will be allocated to servers with higher loads as needed; and vice versa. This enables the present disclosure to dynamically adjust the processing weight of the server according to the real-time load conditions of the server, so that requests can be evenly distributed on each server, improving the stability and scalability of the system.
[0065] S103: Based on the number of requests to the cache server and a preset data copy strategy, adjust the number and location of the data copies of the corresponding cache server.
[0066] In the present disclosure, for data with high access frequency, the number of its copies can be increased, and the copies can be distributed as much as possible on servers with lower loads to improve the reading speed and reliability of these data. The present disclosure can obtain the access frequency of data from the access log of the cache server. Then the target number of copies is calculated, so as to calculate the expected number of copies of a single cache server, and the data copies are created and deleted according to the expected number of copies. In the actual operation of the server machine damage, the target number of copies TargetReplica_j can be regularly compared with the actual number of copies of each server. If the actual number of copies is less than the target number of copies, a new copy is created on the server with low load; if the actual number of copies is more than the target number of copies, the redundant copies are deleted.
[0067] The performance optimization method of the cache server cluster provided in the present invention realizes dynamic load balancing, optimizes data copy strategy to improve data access efficiency, adopts lock-free data structure to avoid lock contention under concurrent conditions, and utilizes big data sharding technology to improve the system's processing power and cache hit rate, thereby enabling the distributed cache server cluster to process large-scale data more efficiently and meet the needs of the big data era.
[0068] In some embodiments, a monitoring device is provided in the cache server cluster, and the monitoring device is used to obtain service data of each cache server.
[0069] It is understandable that the cache server can be installed and configured with appropriate cache service software for each server, such as Redis, Memcached, etc. At the same time, a real-time monitoring device, such as Zabbix, Prometheus, etc., is deployed in the cluster to collect and analyze the load and performance data of each cache node. The present disclosure monitors the CPU usage, memory usage, network bandwidth usage, etc. through a real-time monitoring device. Then the collected monitoring data is summarized and analyzed to obtain the load status of each server node.
[0070] In some embodiments, such as Figure 2 As shown, the calculating the processing weight of the corresponding cache server according to the service data includes:
[0071] S201, calculating a comprehensive load index of a corresponding cache server according to the server data; the comprehensive load index is used to characterize the load condition of the cache server;
[0072] S202: Calculate the processing weight of the cache server at the next moment based on the comprehensive load index.
[0073] It can be understood that the comprehensive load index is calculated using the obtained CPU usage, memory usage, and network bandwidth usage. The comprehensive load index is used to describe the load of server i. In the present disclosure, the comprehensive load index of the corresponding cache server is calculated in the following manner:
[0074] Load_i=w_cpu*cpu_i+w_mem*mem_i+w_net*net_i
[0075] Among them, cpu_i is the controller usage of server i; mem_i is the memory usage; net_i is the network bandwidth usage; Load_i is the comprehensive load index; w_cpu is the weight parameter of the controller; w_mem is the weight parameter of the memory; w_net is the weight parameter of the network bandwidth; the weight parameter is used to indicate the importance of the controller, memory or network bandwidth in the total load, and the value of the weight parameter can be adjusted according to the specific situation of the application. In general, the higher the load of the server, the higher the Load_i value.
[0076] After obtaining the comprehensive load index, the processing weight of the cache server at the next moment can be calculated in the following manner:
[0077]
[0078] Among them, K is a constant used to adjust the size of the processing weight, and ε is a minimum value set to prevent the denominator from being 0. Weight_i is the processing weight of server i at the next moment
[0079] Specifically, the load index Load_i value of a server with a higher load is larger, and the calculated processing weight Weight_i will be smaller, which means that fewer requests will be assigned to servers with a higher load, and vice versa. The method is to distribute the requests to each server according to the weight. The adjustment method in the present disclosure is based on the load status of the server. The weight of the server with a high load will be reduced, and the weight of the server with a low load will be increased. For example, if the load of a server suddenly increases, its weight will be lowered according to this formula to reduce the number of requests assigned to this server. The performance optimization method of the cache server cluster provided in the present disclosure can dynamically adjust its processing weight according to the real-time load condition of the server, so that the requests can be evenly distributed on each server, thereby improving the stability and scalability of the system.
[0080] In some embodiments, Figure 3 As shown, based on the number of requests of the cache server and the preset data copy strategy, the number and location of the data copies of the corresponding cache server are adjusted, including:
[0081] S301, obtaining the access frequency of data and the comprehensive load index of the cache server from the access log of the cache server;
[0082] S302, calculating the expected number of copies of the data according to the access frequency;
[0083] S303: Calculate the number of replicas allocated to each cache server according to the comprehensive load index and the number of replicas.
[0084] Specifically, a replica is an exact copy of data. Replicas are widely used in many scenarios. For example, in distributed systems, in order to improve the reliability and availability of data, multiple copies of data are usually created and stored on different servers. If a server fails, other copies of the data can still be used, thus avoiding data loss. At the same time, by storing copies of data on multiple servers, the reading speed of the system can be increased because a server with a lower load can be selected from multiple replicas for reading.
[0085] On the basis of load balancing, the present disclosure introduces an adaptive data copy strategy. The data copy strategy dynamically adjusts the number and location of data copies according to the access frequency and distribution status of the data. For example, for data with high access frequency, the number of its copies can be increased, and the copies can be distributed as much as possible on servers with low loads to improve the reading speed and reliability of these data.
[0086] First, the present disclosure needs to obtain the access frequency of data from the access log of the cache server. At the same time, the monitoring system needs to obtain the load value Load_i of each server as in step 1. Assume that the number of accesses of data item j in a certain time period is access_j, and the load value of server i is Load_i. Then the expected number of copies of the data is calculated in the following way:
[0087] TargetReplica_j=round(M*log(access_j+ε))
[0088] Then, according to the server load, an expected number of copies is assigned to each server. If the expected number of copies of server i is Expect_i, the number of copies assigned to each cache server is calculated in the following way:
[0089]
[0090] Among them, TargetReplica_j is the expected number of replicas of data j; M is a normal number used to adjust the number of replicas; log() ensures that the relationship between the number of replicas and the access frequency is logarithmic and the growth is relatively smooth; ε is a minimum value to ensure that the internal value of log is greater than 0; Expect_i is the expected number of replicas of server i; access_j is the number of accesses to data item j within a certain period of time.
[0091] In the present disclosure, the target number of replicas TargetReplica_j can be compared with the actual number of replicas of each server regularly. If the actual number of replicas is less than the target number of replicas, a new replica is created on a server with low load; if the actual number of replicas is more than the target number of replicas, the redundant replicas are deleted.
[0092] In some embodiments, the cache server cluster uses a lock-free data structure.
[0093] It is understandable that cache servers often need to handle a large number of concurrent requests, so special care needs to be taken with atomic operations and concurrency control. Locks are the traditional synchronization mechanism, but they can cause problems such as deadlocks and priority inversions. To avoid these problems and improve performance, lock-free or wait-free data structures can be selected. These data structures take advantage of the hardware's atomic operation instructions, such as "compare-and-swap" or "atomic-add", so that they can work effectively under high concurrency conditions.
[0094] The lock - free data structure is applied in the present disclosure: The selected lock - free data structure is applied to the appropriate position of the cache server. For example, a lock - free queue can be used to implement the request queue, and a lock - free hash table can be used to store cache data. Since the lock - free data structure does not require locking, the performance degradation problem caused by multiple threads competing for the same lock can be avoided.
[0095] The present disclosure can further optimize the lock - free data structure. For example, a lock - free data structure suitable for the cache access pattern can be selected. If the data access has strong locality, then a data structure such as a skip list that can provide fast lookup capabilities can be used.
[0096] In some embodiments, when the request includes a large amount of data, before adjusting the number and location of the data replicas of the corresponding cache server based on the number of requests of the cache server and the preset data replica policy, the method further includes:
[0097] Dividing the large - scale data into small - block data and distributing each small - block data to each cache server respectively.
[0098] In some embodiments, as Figure 4 shown, the dividing the large - scale data into small - block data and distributing each small - block data to each cache server respectively includes:
[0099] S401, dividing the large - scale data into multiple small - block data;
[0100] S402, performing a hash function calculation on each small - block data to obtain the hash value of each small - block data;
[0101] S403, determining the cache server where each small - block data is stored based on the number of cache servers and the hash value;
[0102] S404, distributing each small - block data to the corresponding cache server respectively.
[0103] Specifically, in the present disclosure, the data is fragmented using a defined fragmentation rule. During the specific operation, a hash function is executed on the large data, and the obtained hash function value is used to select the fragment where the data is located. For example, assume there are N pieces of data and M cache servers. For the i - th piece of data, we can calculate its hash value through the hash function, and then take the remainder of the hash value with respect to M (this operation can ensure that all data is evenly distributed on all cache servers). The obtained result j (0 <= j < M) indicates that this piece of data should be stored on the j - th server.
[0104] Specifically, the hash function is
[0105] H_i=hash(data_i)
[0106] j=H_i mod M
[0107] Among them, H_i represents the hash value of the i-th data, hash() is the hash function, data_i is the i-th data, j represents the result of the hash function calculation, indicating which server the data should be stored on, and M is the total number of servers.
[0108] Through the above method, the big data is sharded to different cache servers. Next, through the previous steps, the number and location of data copies are dynamically adjusted according to the data access frequency and server load to continue to optimize the system's read performance.
[0109] The present invention can quickly locate the shard where the data is located through the same hash function and modulo operation each time the data is accessed. At the same time, because the hash function and sharding technology are used, big data can be more evenly distributed on different servers, which will effectively improve the cache hit rate and the processing power of the entire system.
[0110] The present disclosure can evenly distribute requests according to the load status of the server, and can also add copies of data with high access frequency and distribute the copies on servers with lower loads, thereby improving the speed of data reading and the overall performance of the system.
[0111] As a specific implementation, assume that there is a three-node Redis distributed cache server cluster, the node names are Node1, Node2 and Node3, which are deployed on three different servers respectively, and the service data monitored by the real-time monitoring device are as follows:
[0112] Node1: CPU usage 35%, memory usage 50%, network bandwidth usage 40%;
[0113] Node2: CPU usage 70%, memory usage 70%, network bandwidth usage 60%;
[0114] Node3: CPU usage 30%, memory usage 40%, network bandwidth usage 20%;
[0115] First, calculate the load index Load and processing weight Weight of each node based on the load data:
[0116] Assume that the weight parameters w_cpu = 0.4, w_mem = 0.3, w_net = 0.3, the weight adjustment parameter K = 10^3, and the minimum value ε = 10^-3, then calculate,
[0117] Node1: Load_1=35*0.4+50*0.3+40*0.3=22+15+12=49, Weight_1=K / (Load_1+ε)=10^3 / (49+10^-3)
[0118] Node2: Load_2=70*0.4+70*0.3+60*0.3=28+21+18=67, Weight_2=K / (Load_2+ε)=10^3 / (67+10^-3)
[0119] Node3: Load_3=30*0.4+40*0.3+20*0.3=12+12+6=30, Weight_3=K / (Load_3+ε)=10^3 / (30+10^-3)
[0120] The following form will be more intuitive:
[0121] Table 1 Load index and processing weight of each node
[0122] node Load Weight Node1 49 20.40816327 Node2 67 14.92537313 Node3 30 33.33333333
[0123] It can be seen that Node3 has the smallest load, so the allocated request weight is the highest.
[0124] Next, suppose there is a batch of data in the cache, one of which is data D, which has been accessed 500 times in the past period of time, and currently has 1 copy on Node1, and 0 copies on Node2 and Node3. The calculation parameter M = 100, and the target number of copies and the expected number of copies are calculated as:
[0125] TargetReplica_D=round(M*log(500+ε))=round(100*log(500))≈300
[0126] Expect_1=TargetReplica_D / (Load_1+ε)≈6.11
[0127] Expect_2=TargetReplica_D / (Load_2+ε)≈4.47
[0128] Expect_3=TargetReplica_D / (Load_3+ε)≈10
[0129] It can be seen that Node3 expects the largest number of replicas. We can create more replicas on Node3 to optimize the reading speed.
[0130] Finally, for data access and storage, we use lock-free data structures and sharding technology. Taking data D as an example, assuming that the sharding rule is the hash function H(D) mod 3, and H(D) = 7, then the shard where D is located is 7 mod 3 = 1, that is, data D should be stored on Node1. If a new access request requires reading data D, we can directly locate Node1 and read from the copy of Node3, which greatly improves processing efficiency and data reliability.
[0131] The present invention provides a performance optimization method for a cache server cluster, which realizes load balancing and performance optimization by monitoring the load of each cache server in real time and dynamically adjusting the processing weight of the request according to its performance status. An adaptive data copy strategy is introduced to dynamically adjust the number and distribution location of data copies according to the access frequency and distribution of the data, further improving the reading speed of the system and the reliability of the data. At the same time, by utilizing advanced lock-free data structures and optimized big data sharding technology, the cache hit rate and processing capacity are greatly improved, and the network delay is effectively reduced. The method of the present invention provides an efficient, reliable, and easily scalable cache service solution for large-scale, high-concurrency Internet application scenarios.
[0132] like Figure 5 As shown, the present disclosure provides a performance optimization device for a cache server cluster, the device comprising:
[0133] The data acquisition module 501 is used to acquire service data of each cache server in the cache server cluster;
[0134] A first adjustment module 502, configured to calculate a processing weight of a corresponding cache server at a next moment according to the service data, and adjust the number of requests of the corresponding cache server based on the processing weight;
[0135] The second adjustment module 503 is used to adjust the number and location of the data copies of the corresponding cache server based on the number of requests of the cache server and a preset data copy strategy.
[0136] The present disclosure provides a performance optimization device for a cache server cluster, wherein a data acquisition module 501 acquires service data of each cache server in the cache server cluster; a first adjustment module 502 calculates a processing weight of the corresponding cache server at the next moment according to the service data, and adjusts the number of requests of the corresponding cache server based on the processing weight; a second adjustment module 503 adjusts the number and location of data copies of the corresponding cache server based on the number of requests of the cache server and a preset data copy strategy.
[0137] It should be noted that the performance optimization device for the cache server cluster of the embodiment of the present invention solves the problem based on a principle similar to that of the aforementioned performance optimization method for the cache server cluster. Therefore, the implementation process, implementation principle, and beneficial effects of the performance optimization device for the cache server cluster can all be referred to the description of the implementation process, implementation principle, and beneficial effects of the aforementioned method, and the repeated parts will not be repeated.
[0138] An embodiment of the present disclosure provides an electronic device, including:
[0139] at least one processor; and
[0140] a memory communicatively connected to the at least one processor; wherein,
[0141] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any one of the above embodiments.
[0142] An embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method described in any of the above embodiments.
[0143] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0144] Figure 6 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0145] like Figure 6As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0146] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0147] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as a performance optimization method for a cache server cluster. For example, in some embodiments, the performance optimization method for a cache server cluster may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the performance optimization method for the cache server cluster described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the performance optimization method for the cache server cluster in any other appropriate manner (e.g., by means of firmware).
[0148] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0149] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0150] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0151] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0152] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0153] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0154] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0155] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0156] The above is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present disclosure, which should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. A performance optimization method for a cache server cluster, characterized in that: The method comprises: Get the service data of each cache server in the cache server cluster; Calculating a processing weight of the corresponding cache server at a next moment according to the service data, and adjusting the number of requests of the corresponding cache server based on the processing weight; Based on the number of requests to the cache server and a preset data copy strategy, the number and location of the data copies of the corresponding cache server are adjusted.
2. The method according to claim 1, characterized in that A monitoring device is provided in the cache server cluster, and the monitoring device is used to obtain service data of each cache server.
3. The method according to claim 1, characterized in that The calculating the processing weight of the corresponding cache server according to the service data includes: Calculating a comprehensive load index of the corresponding cache server according to the server data; the comprehensive load index is used to characterize the load condition of the cache server; The processing weight of the cache server at the next moment is calculated based on the comprehensive load index.
4. The method according to claim 3, characterized in that The comprehensive load index of the corresponding cache server is calculated in the following way: Load_i=w_cpu*cpu_i+w_mem*mem_i+w_net*net_i The processing weight of the cache server at the next moment is calculated in the following manner: Among them, cpu_i is the controller utilization rate of server i; mem_i is the memory utilization rate; net_i is the network bandwidth utilization rate; Load_i is the comprehensive load index; w_cpu is the weight parameter of the controller; w_mem is the weight parameter of the memory; w_net is the weight parameter of the network bandwidth; K is a constant used to adjust the size of the processing weight; ε is a minimum value; Weight_i is the processing weight of server i at the next moment.
5. The method according to claim 1, characterized in that Based on the number of requests to the cache server and a preset data copy strategy, the adjusting the number and location of the data copies of the corresponding cache server includes: Obtaining the access frequency of data and the comprehensive load index of the cache server from the access log of the cache server; Calculating the expected number of copies of the data based on the access frequency; The number of replicas allocated to each cache server is calculated according to the comprehensive load index and the number of replicas.
6. The method according to claim 5, characterized in that The expected number of copies of the data is calculated as follows: TargetReplica_j=round(M*log(access_j+ε)) The number of replicas assigned to each cache server is calculated as follows: Among them, TargetReplica_j is the expected number of replicas of data j; M is a normal number used to adjust the number of replicas; log() ensures that the relationship between the number of replicas and the access frequency is logarithmic; ε is the minimum value; Expect_i is the expected number of replicas of server i; access_j is the number of accesses to data item j within a certain period of time.
7. The method according to claim 1, characterized in that The cache server cluster adopts a lock-free data structure.
8. The method according to claim 1, characterized in that: When the request includes large-scale data, before adjusting the number and location of the data copies of the corresponding cache server based on the number of requests of the cache server and the preset data copy strategy, the method further includes: Divide large-scale data into small blocks and distribute each small block to each cache server.
9. The method according to claim 8, characterized in that The method of dividing the large-scale data into small blocks of data and distributing each small block of data to each cache server includes: Dividing the large-scale data into a plurality of small blocks of data; Performing hash function calculation on each of the small blocks of data to obtain a hash value for each small block of data; Determine the cache server where each small piece of data is stored based on the number of cache servers and the hash value; Distribute each small block of data to the corresponding cache server.
10. A performance optimization device for a cache server cluster, characterized in that: The device comprises: A data acquisition module, used to acquire service data of each cache server in the cache server cluster; A first adjustment module, configured to calculate a processing weight of a corresponding cache server at a next moment according to the service data, and adjust the number of requests of the corresponding cache server based on the processing weight; The second adjustment module is used to adjust the number and location of the data copies of the corresponding cache server based on the number of requests of the cache server and a preset data copy strategy.