Target address information acquisition method and device, equipment and storage medium
By obtaining the historical load data of slave nodes at the master node and using the load prediction model, dynamically selecting the target slave nodes, solving the problem of lack of forward-looking load algorithms in the existing technology, achieving more accurate load balancing and system stability improvement.
Patent Information
- Application Number
- CN202510702731.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-05
AI Technical Summary
The existing load algorithms lack prospects and cannot accurately predict node load changes, resulting in huge business traffic in a certain node, serious business processing delays, and unstable system.
The master node obtains the historical load evaluation index data of the slave node, uses the load prediction model to predict future load conditions, and determines the single-value load index based on historical and predicted load evaluation indicators, dynamically selects the target slave node and feedbacks the address information.
Improve the accuracy and system stability of load balancing, reduce response time and service interruptions, reduce operation and maintenance costs, and maximize resource utilization.
Smart Images

Figure CN120602490A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network communication technology, and in particular to a method, device, equipment and storage medium for obtaining target address information. Background Art
[0002] Distributed storage systems have become a core component of enterprise IT infrastructure in cloud computing and big data scenarios. With the diversification of business models, the same storage cluster often simultaneously handles multiple protocol services. On the one hand, it is necessary to provide second-level responses for low-latency small file reads and writes, while on the other hand, it must meet offline throughput requirements for terabyte (TB)-level sequential writes. This "hybrid deployment" greatly amplifies the fluctuations in node load over time. To adapt to these changes, operations and maintenance personnel typically use Domain Name System (DNS) load balancing to direct client requests to multiple nodes, thereby achieving "ingress traffic balancing" at the logical level.
[0003] In the open source community, PowerDNS (PDNS) is widely used for its support for piped / remote-backend dynamic responses. Developers simply write an external process that, upon receiving a query, calculates and returns a set of A / AAAA records based on existing load balancing algorithms. PDNS then transparently transmits these records to the client. However, existing load balancing algorithms lack foresight, are inflexible, or fail to identify true bottlenecks. They can only make decisions based on the current load situation, which can easily lead to excessive traffic on a single node and severe processing delays. Summary of the Invention
[0004] The present application provides a method, apparatus, device and storage medium for obtaining target address information, in order to at least solve the problems that the existing load algorithms in the relevant technologies are not forward-looking, are not flexible enough or cannot characterize the real bottleneck, and can only make decisions based on the load conditions at the current moment, etc., which may ultimately lead to huge business traffic at a certain node and serious business processing delays due to certain situations.
[0005] The present application provides a method for obtaining target address information, which is applied to a distributed storage system. The system includes a master node and multiple slave nodes. The method is executed by the master node and includes:
[0006] After obtaining the address query request sent by the client, the multiple historical load evaluation indicator data fed back by each slave node are processed separately to obtain the historical load evaluation tensor corresponding to each slave node;
[0007] Input the historical load evaluation index tensor corresponding to the i-th slave node into the pre-built load prediction model to obtain the predicted load evaluation index tensor corresponding to the i-th slave node in the future preset time period, where i is a positive integer;
[0008] Determine a single-value load index corresponding to the i-th slave node according to the historical load evaluation index tensor and the predicted load evaluation index tensor corresponding to the i-th slave node;
[0009] Selecting a preset number of target slave nodes from the plurality of slave nodes according to the single-value load index corresponding to each slave node, and obtaining target address information corresponding to each target slave node;
[0010] Feedback the target address information to the client.
[0011] The present application also provides a target address information acquisition device, comprising:
[0012] The acquisition module is used to obtain the address query request sent by the client;
[0013] A processing module is used to process the multiple historical load evaluation indicator data fed back by each slave node in advance after determining that the acquisition module obtains the address query request, and obtain the historical load evaluation tensor corresponding to each slave node;
[0014] A prediction module is used to input the historical load evaluation index tensor corresponding to the i-th slave node into a pre-built load prediction model to obtain the predicted load evaluation index tensor corresponding to the i-th slave node in a preset time period in the future, where i is a positive integer;
[0015] A determination module, configured to determine a single-value load index corresponding to the i-th slave node based on a historical load evaluation index tensor and a predicted load evaluation index tensor corresponding to the i-th slave node;
[0016] a selection module, configured to select a preset number of target slave nodes from a plurality of slave nodes according to a single-value load index corresponding to each slave node, and obtain target address information corresponding to each target slave node;
[0017] The sending module is used to feed back the target address information to the client.
[0018] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods for obtaining target address information are implemented.
[0019] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned target address information acquisition methods when executed by a processor.
[0020] Through this application, after obtaining the address query request, the multiple historical load evaluation index data fed back by each slave node are obtained for processing, and the historical load evaluation tensor is obtained, and then input into the load prediction model, so that the predicted load evaluation index tensor in the future preset time period can be obtained, thereby more accurately predicting the future load of the node, and then according to the historical load evaluation index tensor and the predicted load evaluation index tensor corresponding to the i-th slave node, the single-value load index corresponding to the i-th slave node is determined. The single-value load index can simultaneously reflect the "busyness at this moment" and the "trend in the short term". Therefore, according to the single-value load index corresponding to each slave node, a preset number of target slave nodes can be selected from multiple slave nodes, and the target address information corresponding to the target slave nodes can be obtained, and transmitted to the client through the PowerDNS main process in the master node, so that the client can access the target slave node according to the destination address information. In this method, because a variety of historical load evaluation index data are obtained, the health status of the node can be more comprehensively evaluated. By predicting the load conditions during a preset time period in the future, the solution can anticipate changes in node loads in advance, avoiding directing traffic to nodes that are about to face high loads or nodes that are about to be occupied by temporary activities such as backup tasks, thereby reducing queuing peaks and improving overall system performance. By optimizing the load balancing strategy, system instability factors caused by uneven node loads can be reduced, and the overall stability and reliability of the system can be improved. Response time and service interruptions can be reduced, and user access speed and experience can be improved. Moreover, because the system can manage resources more effectively, operation and maintenance personnel can reduce manual intervention, thereby reducing operation and maintenance costs. Furthermore, by dynamically allocating traffic, system resources can be maximized and resource waste can be avoided. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] Figure 1 A flowchart of a method for obtaining target address information provided in an embodiment of the present application;
[0023] Figure 2 A flowchart of another method for obtaining target address information provided in an embodiment of the present application;
[0024] Figure 3 A flowchart of another method for obtaining target address information provided in an embodiment of the present application;
[0025] Figure 4A simplified flowchart of a partial method in another method for obtaining target address information provided in an embodiment of the present application;
[0026] Figure 5 A schematic diagram of the overall process structure of a method for obtaining target address information provided in an embodiment of the present application;
[0027] Figure 6 A schematic diagram of the structure of a target address information acquisition device provided in an embodiment of the present application;
[0028] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0030] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0031] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0032] Distributed storage systems have become a core component of enterprise IT infrastructure in cloud computing and big data scenarios. In recent years, Ceph-based object and file storage, due to its excellent consistency model and horizontal scalability, has been widely deployed in industries such as finance, energy, and the internet. However, with the diversification of business models, the same storage cluster often simultaneously hosts multiple protocols, including Network File System (NFS), Common Internet File System (CIFS), and Hadoop Distributed File System (HDFS). On the one hand, they must provide low-latency, sub-second response times for small file reads and writes, while on the other hand, they must meet offline throughput requirements for terabyte-level sequential writes. This "hybrid deployment" significantly amplifies node load fluctuations over time. To accommodate these changes, operations personnel typically use Domain Name System (DNS) load balancing to direct client requests to multiple nodes, achieving "ingress traffic evenness" at the logical level.
[0033] PowerDNS (PDNS) is widely used in the open source community for its support for piped / remote-backend dynamic responses. Developers simply write an external process that, upon receiving a query, calculates and returns a set of A / AAAA records (resource record sets, or RRsets) based on existing load balancing algorithms. PDNS then transparently transmits these to the client.
[0034] Currently, existing load balancing algorithms can be roughly divided into the following categories:
[0035] 1. Static round-robin or weighted round-robin: The earliest approach was to assign a fixed weight to each node IP address, and then PDNS would return them in turn during each query. The disadvantage is that the weights don't reflect the instantaneous status of CPU, memory, and network congestion in real time, leading to an imbalance during peak business hours, with hot nodes often queuing and cold nodes being idle.
[0036] 2. Single-dimensional health detection + minimum value optimization. Many systems regularly capture the CPU or connection number indicators of nodes and place the lowest value at the top of the RRset. Although this method is more flexible than polling, it still has two drawbacks: First, only looking at a single resource dimension cannot depict the true bottleneck (for example, disk IO and network card bandwidth are the real sources of speed limit); second, the algorithm only reflects the "current" load and has no forward-looking capabilities. If a node is currently idle but will be occupied by backup tasks in 30 seconds, the latest DNS response will still direct traffic to the node, causing the "queue peak" to be magnified several times within a few minutes.
[0037] To solve the above problems, the embodiment of the present application provides a method for obtaining target address information. Figure 1 As shown, the method is applied to a distributed storage system, the system includes a master node and multiple slave nodes, and the method is executed by the master node, including the following method steps:
[0038] Step S101: After obtaining the address query request sent by the client, the various historical load evaluation indicator data fed back by each slave node are processed separately to obtain the historical load evaluation tensor corresponding to each slave node.
[0039] Specifically, historical evaluation indicator data includes, for example, CPU utilization, memory utilization, network card export rate, and the number of active client connections, where active clients are clients that send address requests to the master node within a short period of time. The reason for collecting multiple historical evaluation indicator data here is to avoid considering only single dimensions such as CPU or number of connections. Multi-dimensional indicators such as disk IO and network card bandwidth may also be considered to more comprehensively evaluate the health status of the node. Of course, historical load evaluation indicator data includes but is not limited to the above-mentioned ones. Here, the above-mentioned indicators are used as examples for ease of understanding.
[0040] All slave nodes collect information such as CPU utilization, memory usage, network interface card ingress and egress rates, and the number of active client connections in real time, and report it to the master node using SNMP. This reporting can be done in real time or periodically, for example, every 5 seconds.
[0041] Therefore, after obtaining the address query request sent by the client, the master node can process the various historical load assessment indicator data fed back by each pre-acquired slave node separately. The specific processing process includes, for example, pulling the data of each node at a fixed step, completing missing value interpolation, outlier truncation, zero mean normalization and other operations, and injecting statistical features such as timestamp (time stamp of the time when the data was collected), day-week (collection time) period position code (fixed), moving average (average of the change value between the previous second data and the current second data), moving variance (variance of the change value between the previous second data and the current second data), and finally generating a tensor based on the Long Short-Term Memory (LSTM) algorithm. The specific method of generating tensors using the LSTM algorithm is a relatively mature technology in this field and will not be elaborated here.
[0042] It should be noted that in order to save storage space, the master node writes the message into the ring cache immediately after the listening thread receives it. The cache length can cover the data of the last ten minutes.
[0043] Step S102: input the historical load evaluation index tensor corresponding to the i-th slave node into a pre-built load prediction model to obtain the predicted load evaluation index tensor corresponding to the i-th slave node in a preset time period in the future.
[0044] Specifically, the load prediction model is a load preset model generated by training a tensor generated using sample load evaluation index data, which can predict the load evaluation index tensor in a future preset time period, where i is a positive integer.
[0045] For each slave node's corresponding historical load evaluation index tensor, the load evaluation index tensor corresponding to the node in a future preset time period is predicted. In a specific example, the future preset time period is preferably the next 2 minutes or 5 minutes, which can be set according to actual conditions.
[0046] Step S103 : determining a single-value load index corresponding to the i-th slave node based on the historical load evaluation index tensor and the predicted load evaluation index tensor corresponding to the i-th slave node.
[0047] Specifically, the historical load evaluation index tensor and the predicted load evaluation index tensor corresponding to the i-th slave node both include multiple elements, and the specific number of elements is the same as the type of historical load evaluation index data.
[0048] The same type of elements in the historical load evaluation index tensor and the predicted load evaluation index tensor are fused to obtain a fused tensor, and then a linear projection process is performed on multiple elements in the fused tensor to obtain a single-value load index.
[0049] Among them, the single-value load index reflects both the "current busyness" and the "short-term trend".
[0050] Step S104 : selecting a preset number of target slave nodes from the plurality of slave nodes according to the single-value load index corresponding to each slave node, and obtaining target address information corresponding to each target slave node.
[0051] Specifically, as previously described, the single-value load index corresponding to each slave node is used to reflect the "current busyness" and "short-term trend" of the slave node. Therefore, the single-value load index can be used to evaluate the load pressure of each node at the moment and the load change trend in the short term, that is, within a preset time period in the future, so as to evaluate which slave nodes can serve as the load processing nodes corresponding to the current access request. For example, the single-value load indexes are sorted from small to large, and the slave nodes corresponding to the preset number of single-value load indices at the top of the ranking are selected as target slave nodes, and the target address information corresponding to each target slave node is obtained.
[0052] Step S105: Feedback the target address information to the client.
[0053] Among them, after the master node obtains the destination address information corresponding to multiple target slave nodes, it first sorts the slave nodes according to the current load pressure of each slave node, and then configures the sorting of the corresponding destination address information according to the sorting order, and assembles a string that complies with the PDNS pipe specification and writes it to the standard output. After receiving it, the PDNS master process in the master node immediately transmits it to the client. The specific load pressure can be determined according to the number of connected clients.
[0054] The embodiment of the present application provides a method for obtaining target address information. After obtaining the address query request, the method obtains a plurality of historical load evaluation index data fed back by each slave node for processing, obtains a historical load evaluation tensor, and then inputs it into the load prediction model. The predicted load evaluation index tensor in the future preset time period can be obtained, thereby more accurately predicting the future load of the node. Then, based on the historical load evaluation index tensor and the predicted load evaluation index tensor corresponding to the i-th slave node, the single-value load index corresponding to the i-th slave node is determined. The single-value load index can simultaneously reflect the "busyness at the moment" and the "trend in the short term". Therefore, a preset number of target slave nodes can be selected from multiple slave nodes according to the single-value load index corresponding to each slave node, and the target address information corresponding to the target slave nodes can be obtained. The target address information is transmitted to the client through the PowerDNS main process in the master node, so that the client can access the target slave node according to the destination address information. In this method, because a plurality of historical load evaluation index data are obtained, the health status of the node can be more comprehensively evaluated. By predicting the load conditions during a preset time period in the future, the solution can anticipate changes in node loads in advance, avoiding directing traffic to nodes that are about to face high loads or nodes that are about to be occupied by temporary activities such as backup tasks, thereby reducing queuing peaks and improving overall system performance. By optimizing the load balancing strategy, system instability factors caused by uneven node loads can be reduced, and the overall stability and reliability of the system can be improved. Response time and service interruptions can be reduced, and user access speed and experience can be improved. Moreover, because the system can manage resources more effectively, operation and maintenance personnel can reduce manual intervention, thereby reducing operation and maintenance costs. Furthermore, by dynamically allocating traffic, system resources can be maximized and resource waste can be avoided.
[0055] In an optional example, based on the above embodiment, as described above, the load evaluation index tensor includes multiple elements. According to the historical load evaluation index tensor and the predicted load evaluation index tensor corresponding to the i-th slave node, the single-value load index corresponding to the i-th slave node is determined. The specific implementation process can be referred to the following method steps, specifically see Figure 2 Shown, including:
[0056] Step S201 : performing load fusion processing based on the historical load assessment index tensor and the predicted load assessment index tensor to obtain a comprehensive vector.
[0057] Specifically, the number of elements included in the comprehensive vector is the same as the number of elements included in the load evaluation indicator tensor.
[0058] During specific execution, the historical load evaluation index tensor and the predicted load evaluation index tensor can be load fused based on the pre-configured first weight for the historical load evaluation index tensor and the pre-configured second weight for the predicted load evaluation index tensor to obtain a comprehensive vector, where the sum of the first weight and the second weight is 1.
[0059] The specific expression can be found in the following formula:
[0060] Z=λX+(1-λ)X ′ (Formula 1)
[0061] Among them, X is the historical load evaluation index tensor, X ′ is the predicted load evaluation index tensor, λ is the first weight value, (1-λ) is the second weight value, and Z is the comprehensive vector.
[0062] Step S202 , performing linear projection processing on multiple elements in the comprehensive vector to obtain a single-value load index.
[0063] Specifically, Z is a multidimensional comprehensive vector. For example, in the embodiment of the present application, four load assessment index data are included, so the corresponding Z is also a four-dimensional comprehensive vector. Therefore, when linear projection processing is performed on multiple elements in the comprehensive vector, it can be expressed by an expression similar to the following:
[0064] y=ω1z1+ω2z3+…+ω n z n (Formula 2)
[0065] Among them, ω n is the nth weight vector, which is used to indicate the contribution ratio of each dimension, z n is the n-th dimension vector in the comprehensive vector, y is a single-value load index, n is a positive integer, for example, n is 4 in this embodiment, ω1 to ω n It can be configured according to actual conditions.
[0066] In the above method steps, by fusing the historical load assessment indicator tensor and the predicted load assessment indicator tensor, a comprehensive vector is obtained that comprehensively considers the node's historical load and future predicted load, thereby providing a more comprehensive and dynamic load assessment. By assigning different weights to the historical and predicted load assessment indicators, the degree of reliance on historical data and future predictions can be flexibly adjusted, allowing the system to adjust the relative importance of predictions and history based on actual conditions. Linear projection of the comprehensive vector simplifies the data dimension and fuses multiple load indicators into a single-value load index. This single-value index can be more intuitively used for load balancing decisions. This load fusion and projection method can achieve more effective load balancing, reducing performance bottlenecks and service interruptions caused by uneven load. By fusing historical and predicted data, the accuracy of load predictions can be improved, reducing resource misallocation caused by prediction bias. More accurate load assessment and balancing help improve system stability and reliability, reducing performance fluctuations caused by irrational resource allocation. In the above method, by dynamically adjusting weights and fusing, resource utilization can be improved, resource waste can be avoided, and sudden load increases can be better handled.
[0067] Based on any of the aforementioned embodiments, the master node may further continuously record the mean absolute percentage error (MAPE) between the predicted vector and the actual sampled vector. The master node may adjust the first weight value and the second weight value based on the MAPE obtained in each of the preset rounds, thereby ensuring the self-healing capability of the aforementioned fusion algorithm. This ensures that even if the model experiences a temporary inaccuracy, the parsing result will not significantly deviate from the actual cluster load.
[0068] Therefore, the method may further include the following steps, see Figure 3 Shown, including:
[0069] Step S301, after feeding back the target address information to the client according to the address query request in each round, obtain the actual load assessment index tensor that matches one or more predicted load assessment index tensors.
[0070] Specifically, as described above, the master node stores data that is traced back 10 minutes at each moment, and predicts data for a preset time period in the future, such as the next 2 minutes. Then, after another 2 minutes, the actual load assessment index tensor corresponding to the predicted load assessment index tensor can be obtained. The method for obtaining the actual load assessment index tensor is the same or similar to the method for obtaining the historical load assessment index tensor described above, and will not be repeated here.
[0071] Step S302: determining a mean absolute percentage error based on the actual load assessment index tensor and the predicted load assessment index tensor corresponding to the actual load assessment index tensor.
[0072] Specifically, please refer to the following expression:
[0073]
[0074] Among them, n is the number of actual load evaluation index tensors obtained after each round of execution, y i is the actual load evaluation index tensor of the i-th, is the predicted load evaluation metric tensor corresponding to the actual load evaluation metric tensor, and MAPE is the mean absolute percentage error.
[0075] Step S303 : adjusting the weight value of the first weight and the weight value of the second weight respectively according to the mean absolute percentage errors respectively obtained in preset rounds.
[0076] Specifically, for example, if the MAPE obtained in multiple rounds (exceeding a preset number of rounds) is less than a certain lower limit, it means that the aforementioned load forecast model is relatively reliable, and the system can automatically adjust the first weight value and increase the second weight value. Otherwise, it means that the reliability of the load forecast model is slightly low, and the specific value of the first weight value needs to be increased, and the fusion processing is based on the historical load evaluation index tensor.
[0077] Of course, you can also specify specific weight adjustment rules, for example, see the following:
[0078] When the mean absolute percentage errors obtained in the preset rounds are all less than or equal to the first preset threshold, reducing the weight value of the first weight according to a preset step size;
[0079] Alternatively, when there is an average absolute percentage error greater than a first preset threshold and less than or equal to a second preset threshold among the average absolute percentage errors obtained in the preset rounds, or when the number of average absolute percentage errors greater than the second preset threshold is less than a preset number, the weight value of the first weight is increased according to a preset step size, wherein the second preset threshold is greater than the first preset threshold.
[0080] Specifically, for example, the initial values of the first weight value and the second weight value are both 0.5, and then the first weight value and the second weight value are continuously adjusted according to actual conditions, for example, with a preset step size of 0.01. The first preset threshold value is, for example, 30%, and the second preset threshold value is, for example, 70%. The preset number can be determined based on the ratio of the preset number of rounds. For example, if the preset number of rounds is 10, then the preset number is 6.
[0081] Further optionally, the method may further include:
[0082] When there is an average absolute percentage error greater than or equal to a preset number of average absolute percentage errors in the preset rounds that is greater than a second preset threshold, the asynchronous training process is triggered. The asynchronous training process is used to instruct the load prediction model to be retrained based on the latest obtained multiple historical load evaluation indicator data within a preset time period based on the current moment.
[0083] Specifically, if the MAPE exceeds a higher threshold (70%) for a set number of times, asynchronous retraining (retraining the load forecasting model) will be triggered. The master node can push the data pushed by each slave node in the last 24 hours to the training server, which completes incremental learning and uploads the new model to the model repository. The master node's model hot-swap thread detects the new version and immediately switches to it, without stopping the DNS service.
[0084] For the specific execution process, see Figure 4 As shown, an error calculator is designed in the master node to obtain the generated actual load evaluation index tensor in real time, and calculate the mean absolute percentage error based on the predicted load evaluation index tensor and the actual load evaluation index tensor corresponding to the preset load evaluation index tensor.
[0085] The mean absolute percentage error is used to evaluate whether weights should be adjusted according to any of the above adjustment rules, including the first and second weights. Furthermore, a determination is made as to whether the retraining trigger threshold has been reached; otherwise, the current round ends. If so, the data is sent to the training server, a new load forecast model is generated and archived, and the master node hot-swaps the model.
[0086] The specific implementation process has been introduced in detail in the previous article, so I will not go into details here.
[0087] In an optional example, based on any of the foregoing embodiments, after determining the single-value load index corresponding to the i-th slave node based on the historical load evaluation index tensor and the predicted load evaluation index tensor corresponding to the i-th slave node, the method may further include the following method steps:
[0088] Step a1: Select a target survival time according to the single-value load index corresponding to each slave node.
[0089] Step a2: Feedback the target lifetime to the client.
[0090] Specifically, the master node can calculate the variance based on the single-value load index corresponding to all slave nodes. The load fluctuation of each node in the cluster is determined based on the variance. If the overall fluctuation is small, it means that the cluster is in a stable period. At this time, the DNS TTL can be boldly extended (for example, sixty seconds or even hundreds of seconds) to allow the client's local cache to take full effect to reduce authoritative queries; conversely, if the variance increases significantly, the TTL is shortened to a few seconds so that the new scheduling results can quickly replace the client cache and achieve traffic migration in seconds. Specifically, the TTL will be sent to the client through PDNS. The new scheduling results can quickly replace the client cache and achieve traffic migration in seconds. The decision on the length of TTL is entirely left to the algorithm rather than manual configuration, so that "steady-state consumption saving and peak flexibility" can be met at the same time.
[0091] In this method, the algorithm automatically adjusts the TTL based on node load variance, eliminating the need for manual intervention and improving system automation and responsiveness. When cluster load is stable, extending the DNS TTL reduces the number of client queries to the DNS server, thereby reducing network traffic and server load. Reducing DNS queries saves bandwidth and network resources, a particularly important benefit in high-latency or high-cost network environments. When load fluctuates significantly, shortening the DNS TTL allows for faster updates of client cache records, ensuring that traffic can be quickly migrated to less-loaded nodes, enabling rapid response to peak loads. Adjusting the TTL enables traffic migration within seconds, which is crucial for business scenarios requiring rapid response. Dynamic TTL adjustment helps maintain system stability under varying load conditions and mitigates performance issues caused by uneven load. Rapid traffic migration and stable performance enhance user experience and minimize service interruptions and latency. Extending the TTL saves resources during stable load periods, while shortening it ensures efficient resource utilization during peak load periods. This strategy achieves a balance between resource conservation during stable load periods and flexibility during peak load periods, meeting the needs of the system in diverse states.
[0092] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0093] In an optional embodiment, the method may also include: when a node is removed from the DNS response due to hardware failure or background rebalancing, the master node does not immediately put it back to the optimal position upon recovery, but assigns a recovery coefficient to the node that increases exponentially over time. The coefficient is extremely small at the initial stage of recovery, and the node only carries very little traffic; if no abnormalities occur during the smooth transition, the coefficient will slowly rise to one. This can not only utilize the capabilities of the newly recovered node, but also avoid jitter due to instantaneous "hot and cold switching". A smooth grayscale recovery curve is thus formed. Therefore, when selecting the target node according to the aforementioned method steps, in addition to selecting the target node based on the single-value load index, the recovery coefficient corresponding to each node must also be considered. When selecting the target node, the master node gives priority to the slave node with a recovery coefficient of 1. If the single-value load index of a slave node with a recovery coefficient of 1 is higher than the preset load index threshold, it means that the load on the node is too heavy, and it will no longer be selected as the target slave node. As a fallback, a slave node with a recovery coefficient of at least close to 1 and a single-value load index less than the preset load index threshold will be selected as the target slave node, so as to maintain a certain degree of load balance as much as possible without increasing the load pressure of each node too much.
[0094] Figure 5 The overall flow chart of the method for obtaining target address information is shown in FIG. Figure 5 The diagram shows multiple slave nodes, including nodes 1 to 3, and a master node. Data is transmitted between the slave nodes and the master node in the form of SNMP encoding. The master node executes the corresponding prediction decision method steps and sends them to the PDNS master process in the master node in the form of a pipe. The PDNS master process then sends them to the client. The client's domain name request is the address query request described above. Figure 5 As mentioned above, all monitoring and decision-making links are completed in a closed loop within the master node, and the client only perceives changes in the resolved IP and TTL.
[0095] Furthermore, this method can also include one or more of the following expansion ideas, such as: 1. Introducing reinforcement learning to dynamically adjust scheduling weights based on business and energy prices; 2. Deploying lightweight prediction models to edge gateway nodes to achieve "local self-judgment + central calibration" in cross-regional multi-active scenarios; 3. Using a federated learning framework, data centers only exchange gradients without sharing raw monitoring data, balancing privacy and accuracy; 4. Incorporating protocol identification, traffic classification, and QoS order preservation into a unified weighting system to achieve differentiated services. Through these derivative approaches, this method can evolve from a single storage gateway scheduling solution to a full-stack intelligent control platform covering computing power, networking, and energy.
[0096] The present invention combines short-term time series prediction with real-time monitoring, introducing foresight into the distributed file system DNS load balancing. Theoretically, this method can divert traffic in advance before the business peak arrives, significantly reducing the node queue depth and user-perceived delay; in background rebalancing, Scrub or instantaneous hardware speed reduction scenarios, it can also quickly migrate clients to healthy nodes with a shorter TTL, shortening the fault impact window. In addition, the error adaptive weight ensures that the system can still operate reliably when the model is slightly inaccurate; grayscale recovery makes the node regression process smooth, avoiding jitter and hot and cold spot flips. Since all new logic is encapsulated in the back-end program of the main node, there is zero intrusion into the existing collection and service processes, so the operation and maintenance upgrade cost is extremely low, and it can be quickly switched and rolled back online, which greatly improves the feasibility and maintainability of the solution.
[0097] The embodiment of the present application also provides a device for obtaining target address information, see Figure 6 As shown, the device includes: an acquisition module 601, a processing module 602, a prediction module 603, a determination module 604, a selection module 605, and a sending module 606.
[0098] Acquisition module 601, used to obtain the address query request sent by the client;
[0099] The processing module 602 is configured to process the multiple historical load evaluation indicator data fed back by each slave node after the acquisition module 601 acquires the address query request, and acquire a historical load evaluation tensor corresponding to each slave node;
[0100] The prediction module 603 is configured to input the historical load evaluation index tensor corresponding to the i-th slave node into a pre-built load prediction model to obtain a predicted load evaluation index tensor corresponding to the i-th slave node in a preset future time period, where i is a positive integer;
[0101] A determination module 604 is configured to determine a single-value load index corresponding to the i-th slave node based on the historical load evaluation index tensor and the predicted load evaluation index tensor corresponding to the i-th slave node;
[0102] A selection module 605 is configured to select a preset number of target slave nodes from a plurality of slave nodes according to a single-value load index corresponding to each slave node, and obtain target address information corresponding to each target slave node;
[0103] The sending module 606 is used to feed back the target address information to the client.
[0104] In one optional example, the load assessment indicator tensor includes multiple elements. The determination module 604 is specifically configured to perform load fusion processing based on the historical load assessment indicator tensor and the predicted load assessment indicator tensor to obtain a comprehensive vector, wherein the number of elements included in the comprehensive vector is the same as the number of elements included in the load assessment indicator tensor;
[0105] Perform linear projection on multiple elements in the integrated vector, with a single-value load index.
[0106] In an optional example, the determination module 604 is specifically used to perform load fusion processing on the historical load assessment index tensor and the predicted load assessment index tensor based on pre-configuring a first weight for the historical load assessment index tensor and pre-configuring a second weight for the predicted load assessment index tensor to obtain a comprehensive vector, wherein the sum of the first weight and the second weight is 1.
[0107] In an optional example, the acquisition module 601 is further configured to obtain an actual load assessment indicator tensor that matches one or more predicted load assessment indicator tensors after feeding back the target address information to the client according to the address query request in each round;
[0108] The processing module 602 is further configured to determine a mean absolute percentage error based on the actual load evaluation index tensor and a predicted load evaluation index tensor corresponding to the actual load evaluation index tensor;
[0109] The weight value of the first weight and the weight value of the second weight are adjusted according to the mean absolute percentage errors obtained in the preset rounds.
[0110] In an optional example, the processing module 602 is specifically configured to reduce the weight value of the first weight according to a preset step size when the mean absolute percentage errors respectively obtained in the preset rounds are less than or equal to the first preset threshold;
[0111] Alternatively, when there is an average absolute percentage error greater than a first preset threshold and less than or equal to a second preset threshold among the average absolute percentage errors obtained in the preset rounds, or when the number of average absolute percentage errors greater than the second preset threshold is less than a preset number, the weight value of the first weight is increased according to a preset step size, wherein the second preset threshold is greater than the first preset threshold.
[0112] In an optional example, processing module 602 is specifically used to trigger an asynchronous training process when there is an average absolute percentage error greater than or equal to a preset number of average absolute percentage errors obtained in the preset rounds that is greater than a second preset threshold. The asynchronous training process is used to instruct the load prediction model to be retrained based on the latest obtained multiple historical load evaluation indicator data within a preset time period based on the current moment.
[0113] In an optional example, the processing module 602 is further configured to select a target survival time according to a single-value load index corresponding to each slave node;
[0114] The sending module 606 is further configured to feed back the target lifetime to the client.
[0115] For the description of the features in the embodiment corresponding to the target address information acquisition device provided in the embodiment of the present application, please refer to the relevant description of the embodiment corresponding to the target address information acquisition method, and no further details will be given here.
[0116] The embodiment of the present application provides a target address information acquisition device. After obtaining the address query request, it obtains a variety of historical load evaluation index data fed back by each slave node for processing, obtains the historical load evaluation tensor, and then inputs it into the load prediction model. It can obtain the predicted load evaluation index tensor in the future preset time period, thereby more accurately predicting the future load of the node. Then, based on the historical load evaluation index tensor and the predicted load evaluation index tensor corresponding to the i-th slave node, the single-value load index corresponding to the i-th slave node is determined. The single-value load index can simultaneously reflect the "busyness at the moment" and the "trend in the short term". Therefore, according to the single-value load index corresponding to each slave node, a preset number of target slave nodes can be selected from multiple slave nodes, and the target address information corresponding to the target slave nodes can be obtained. The target address information is transmitted to the client through the PowerDNS main process in the master node, so that the client can access the target slave node according to the destination address information. In this method, because a variety of historical load evaluation index data are obtained, the health status of the node can be more comprehensively evaluated. By predicting the load conditions during a preset time period in the future, the solution can anticipate changes in node loads in advance, avoiding directing traffic to nodes that are about to face high loads or nodes that are about to be occupied by temporary activities such as backup tasks, thereby reducing queuing peaks and improving overall system performance. By optimizing the load balancing strategy, system instability factors caused by uneven node loads can be reduced, and the overall stability and reliability of the system can be improved. Response time and service interruptions can be reduced, and user access speed and experience can be improved. Moreover, because the system can manage resources more effectively, operation and maintenance personnel can reduce manual intervention, thereby reducing operation and maintenance costs. Furthermore, by dynamically allocating traffic, system resources can be maximized and resource waste can be avoided.
[0117] The embodiment of the present application also provides an electronic device, such as Figure 7 As shown, it includes a memory 10 and a processor 20. The memory 10 stores a computer program, and the processor 20 is configured to run the computer program to execute the steps in any of the above-mentioned target address information acquisition method embodiments.
[0118] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned target address information acquisition method embodiments, or execute the steps of any of the above-mentioned data reading method embodiments when running.
[0119] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0120] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned target address information acquisition method embodiments are implemented.
[0121] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned data reading method embodiments are implemented.
[0122] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0123] The above is a detailed introduction to the target address information acquisition method, device, equipment and storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for obtaining target address information, characterized in that: The method is applied to a distributed storage system, the system including a master node and multiple slave nodes. The method is executed by the master node and includes: After obtaining the address query request sent by the client, the multiple historical load evaluation indicator data fed back by each slave node are processed separately to obtain the historical load evaluation tensor corresponding to each slave node; Input the historical load evaluation index tensor corresponding to the i-th slave node into the pre-built load prediction model to obtain the predicted load evaluation index tensor corresponding to the i-th slave node in a preset time period in the future, where i is a positive integer; Determining a single-value load index corresponding to the i-th slave node according to the historical load evaluation index tensor corresponding to the i-th slave node and the predicted load evaluation index tensor; Selecting a preset number of target slave nodes from the plurality of slave nodes according to the single-value load index corresponding to each of the slave nodes, and obtaining target address information corresponding to each of the target slave nodes; Feedback the target address information to the client.
2. The method according to claim 1, characterized in that The load assessment index tensor includes multiple elements, and determining the single-value load index corresponding to the i-th slave node according to the historical load assessment index tensor corresponding to the i-th slave node and the predicted load assessment index tensor specifically includes: Performing load fusion processing according to the historical load assessment index tensor and the predicted load assessment index tensor to obtain a comprehensive vector, wherein the number of elements included in the comprehensive vector is the same as the number of elements included in the load assessment index tensor; A linear projection process is performed on a plurality of elements in the integrated vector, and the single-value load index is obtained.
3. The method according to claim 2, characterized in that The performing load fusion processing according to the historical load evaluation index tensor and the predicted load evaluation index tensor to obtain a comprehensive vector specifically includes: Based on pre-configuring a first weight for the historical load assessment index tensor and pre-configuring a second weight for the predicted load assessment index tensor, load fusion processing is performed on the historical load assessment index tensor and the predicted load assessment index tensor to obtain the comprehensive vector, wherein the sum of the first weight and the second weight is 1.
4. The method according to claim 3, characterized in that The method further comprises: After feeding back the target address information to the client according to the address query request in each round, obtaining an actual load assessment indicator tensor that matches one or more of the predicted load assessment indicator tensors; determining a mean absolute percentage error based on the actual load assessment index tensor and a predicted load assessment index tensor corresponding to the actual load assessment index tensor; The weight value of the first weight and the weight value of the second weight are adjusted respectively according to the mean absolute percentage errors obtained in preset rounds.
5. The method according to claim 4, characterized in that The adjusting the weight value of the first weight and the weight value of the second weight respectively according to the mean absolute percentage errors respectively obtained in the preset rounds specifically includes: When the mean absolute percentage errors obtained in the preset rounds are all less than or equal to a first preset threshold, reducing the weight value of the first weight according to a preset step size; Alternatively, when the average absolute percentage errors obtained in the preset rounds include an average absolute percentage error greater than the first preset threshold and less than or equal to the second preset threshold, or when the number of average absolute percentage errors greater than the second preset threshold is less than a preset number, the weight value of the first weight is increased according to a preset step size, wherein the second preset threshold is greater than the first preset threshold.
6. The method according to claim 5, characterized in that The method further comprises: When the average absolute percentage errors obtained in the preset rounds are greater than or equal to the preset number and are greater than the second preset threshold, the asynchronous training process is triggered. The asynchronous training process is used to instruct the load prediction model to be retrained based on the latest acquired multiple historical load evaluation indicator data within a preset time period based on the current moment.
7. The method according to any one of claims 1 to 6, characterized in that After determining the single-value load index corresponding to the i-th slave node based on the historical load evaluation index tensor corresponding to the i-th slave node and the predicted load evaluation index tensor, the method further includes: selecting a target survival time according to the single-value load index corresponding to each of the slave nodes; The target lifetime is fed back to the client.
8. A target address information acquisition device, characterized in that: include: The acquisition module is used to obtain the address query request sent by the client; a processing module, configured to process the plurality of historical load evaluation indicator data fed back by each slave node in advance after the acquisition module acquires the address query request, and acquire a historical load evaluation tensor corresponding to each slave node; A prediction module is configured to input a historical load evaluation index tensor corresponding to the i-th slave node into a pre-built load prediction model to obtain a predicted load evaluation index tensor corresponding to the i-th slave node in a preset future time period, where i is a positive integer; A determination module, configured to determine a single-value load index corresponding to the i-th slave node based on the historical load evaluation index tensor corresponding to the i-th slave node and the predicted load evaluation index tensor; a selection module, configured to select a preset number of target slave nodes from the plurality of slave nodes according to the single-value load index corresponding to each of the slave nodes, and obtain target address information corresponding to each of the target slave nodes; The sending module is used to feed back the target address information to the client.
9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method for obtaining target address information according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method for obtaining target address information according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Intelligent Kubernetes node maintenance method and device
CN120896908A