High-concurrency management method oriented to computing power host access
Through distributed token bucket network and intelligent prediction caching strategy, the traffic management problem of IoT platform under high concurrent requests is solved, and traffic control, reduced latency and resource utilization are improved.
Patent Information
- Application Number
- CN202510698279.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to effectively manage the traffic of the IoT platform under high concurrent requests, resulting in request blocking and resource waste. The caching strategy lacks the ability to predict future requests and cannot cope with complex traffic fluctuations.
The distributed token bucket network and intelligent prediction cache strategy are adopted to predict future traffic trends through machine learning, dynamically adjust the token generation rate and cache strategy, and combine the synchronization mechanism of the distributed token bucket network and intelligent prediction cache to optimize system resource allocation.
It realizes precise control of network traffic, reduces access latency, reduces server load, improves resource utilization, and avoids network congestion and system crashes.
Smart Images

Figure CN120378372A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to cloud computing / edge computing technology, artificial intelligence technology, and management operation support field, and particularly relates to a high-concurrency management method for accessing computing power hosts. Background Art
[0002] With the development of the Internet of Things, the number of devices has shown a geometric growth, and the concurrent request volume of the IoT platform has continued to rise. Handling a large number of concurrent requests has become a difficult problem. The existing token bucket algorithm will cause request blocking or resource waste under high concurrency. The caching strategy is based on a simple algorithm and lacks the ability to predict future requests, and cannot cope with complex traffic fluctuations. Therefore, combining efficient concurrent management with intelligent prediction caching is an important direction to improve the performance of the IoT platform. Summary of the Invention
[0003] To solve the above technical problems, the present invention provides a high-concurrency management method for accessing computing power hosts, which optimizes the system's concurrent processing ability, including improving the system response speed, enhancing resource utilization rate, etc.
[0004] The technical solution of the present invention is as follows:
[0005] A high-concurrency management method for accessing computing power hosts, including: constructing a distributed token bucket network; predicting future traffic trends through a machine learning algorithm based on historical traffic data and real-time request patterns; dynamically adjusting the token generation rate and caching strategy to optimize system resource allocation.
[0006] Further,
[0007] The synchronization mechanism of the distributed token bucket network includes: adopting the gossip protocol to regularly exchange token status information between adjacent nodes; when the local token consumption rate exceeds a preset threshold, requesting token reinforcement from adjacent nodes; ensuring the consistency and integrity of token information during transmission through a fault tolerance algorithm.
[0008] The detailed steps are as follows:
[0009] Step 1: Token Generation and Allocation
[0010] 1) Token Initialization
[0011] The token bucket of each node is initialized when the system starts, and the token generation rate r and the token bucket capacity b are set;
[0012] 2) Token Generation
[0013] Each node continuously generates tokens at the set token generation rate and puts the newly generated tokens into the token bucket; when the token bucket is full, the excess tokens will be discarded;
[0014] 3) Token Request Allocation
[0015] When a node receives a request, it attempts to obtain a token from the local token bucket; if there are available tokens in the token bucket, the request is allowed to pass and be processed; when the number of tokens in the node's local token bucket is insufficient to handle new requests and the local load exceeds a preset threshold, the node requests tokens from other nodes with lighter loads.
[0016] Furthermore,
[0017] The intelligent prediction caching strategy includes: analyzing behavior patterns,
[0018] identifying high-frequency accessed data items and access time windows; preloading data with high probability of access to the edge node cache based on the prediction model; implementing a multi-level cache architecture, including memory cache, disk cache, and distributed cache cluster.
[0019] Intelligent prediction caching includes:
[0020] 1) Data Collection and Preprocessing
[0021] Data collection: The system collects historical authentication request data of computing power hosts, including request timestamps, request types, request parameters, and response data;
[0022] Data preprocessing: Clean the collected data to remove noise data and abnormal data; perform normalization on the data to convert data in different dimensions to the same scale range for subsequent analysis;
[0023] 2) Prediction Model Construction and Training
[0024] Algorithm selection: Adopt machine learning algorithms, such as long short-term memory network and time series prediction algorithm, to construct a prediction model;
[0025] Model training: Divide the preprocessed historical request data into a training set and a test set; use the training set to train the model, and by adjusting the model parameters, enable the model to accurately learn the patterns and rules in the historical request data;
[0026] Model evaluation: Use the test set to evaluate the trained model, calculate the prediction accuracy and mean square error, and determine whether the performance of the model meets the requirements. If not, adjust the model parameters or select other algorithms to retrain;
[0027] 3) Cache Content Prediction and Update
[0028] Request prediction: Utilize the trained model to predict the possible request types and data requirements in the future period according to the current time and historical request patterns;
[0029] Cache Update: According to the prediction results, relevant data is pre-loaded into the cache in advance. When the system detects an actual request, it preferentially retrieves data from the cache for response. Meanwhile, according to the real-time request situation and prediction results, the cache content is dynamically adjusted.
[0030] It further includes: building a distributed monitoring system to collect the token usage and cache hit rate of each node in real time; generating a performance report based on the monitoring data for evaluating the overall system efficiency and optimizing the prediction model; implementing an adaptive feedback mechanism to adjust the prediction parameters and control strategies according to the actual operation effect. It includes:
[0031] Token Bucket Adjustment
[0032] According to the real-time traffic load situation, dynamically adjust the generation rate and distribution strategy of the token bucket; when the system load is below the preset value, reduce the token generation rate; when the system load is higher than the preset value, increase the token generation rate and adjust the token allocation strategy among nodes to ensure that the traffic can be reasonably controlled.
[0033] Cache Policy Adjustment
[0034] Combined with intelligent predictive caching, when there is a sudden increase in traffic, preferentially load the requested data into the cache; meanwhile, dynamically adjust the cache update frequency and replacement policy according to the traffic changes.
[0035] The beneficial effects of the present invention are
[0036] Precise Traffic Control
[0037] It can precisely control the network traffic to ensure that each node or service can obtain resources at a preset rate, avoiding network congestion or system crashes caused by excessive traffic.
[0038] Reduce Access Latency
[0039] By predicting the requests of the computing power host in advance and caching the relevant data, when the host issues a request, it can directly obtain data from the cache, greatly shortening the data access time and improving the system response speed.
[0040] Reduce Server Load
[0041] When a large number of data requests from computing power hosts can be obtained from the cache, the server only needs to process a small number of requests that cannot be hit in the cache, reducing the server pressure and enabling it to better handle other key tasks.
[0042] Improve Resource Utilization
[0043] The cache space is utilized reasonably, and the data most likely to be accessed is stored in the cache according to the prediction results, improving the cache hit rate and resource utilization. The intelligent prediction cache method analyzes and predicts based on factors such as data access frequency and time, retains important data in the cache, avoids waste of cache space, and enables the cache to play its maximum role. Brief Description of the Drawings
[0044] Figure 1 It is a schematic diagram of the workflow of the present invention. Detailed Implementation Manner
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] The present invention proposes a high-concurrency management method for accessing computing power hosts, which combines a distributed token bucket mechanism and an intelligent prediction cache strategy based on machine learning. Through this technical solution, the generation rate and distribution strategy of the token bucket can be dynamically adjusted according to the real-time traffic load situation. At the same time, combined with the intelligent prediction cache, the system will preferentially load possible connection authentication data when the device request connection surges, and dynamically adjust the cache content to avoid performance bottlenecks caused by unreasonable caching during high concurrency.
[0047] The present invention mainly includes the following steps:
[0048] 1) Token generation and distribution
[0049] 2) Token status synchronization
[0050] 3) Intelligent prediction cache
[0051] 4) Token bucket and cache dynamic adjustment
[0052] The detailed implementation is as follows:
[0053] The first step: Token generation and distribution
[0054] 1. Token initialization
[0055] The token bucket of each node is initialized when the system starts, and the token generation rate r (unit: number / second) and the token bucket capacity b (unit: number) are set. For example, if the token generation rate r = 10 and the token bucket capacity b = 100, it means that this node generates 10 tokens per second, and the token bucket can hold at most 100 tokens.
[0056] 2. Token Generation
[0057] Each node continuously generates tokens at the set token generation rate r and puts the newly generated tokens into the token bucket. When the token bucket is full, the excess tokens will be discarded.
[0058] 3. Token Request Allocation
[0059] When a node receives a request, it tries to obtain a token from its local token bucket. If there are available tokens in the token bucket, the request is allowed to pass and be processed; when the number of tokens in the node's local token bucket is insufficient to handle new requests and the local load exceeds a certain threshold, the node will request tokens from other nodes with lower load. Suppose node A requests n tokens from node B. If node B has enough remaining tokens in its own token bucket, it will transfer n tokens to node A and update the quantity in its own token bucket. At the same time, after receiving the tokens, node A updates the quantity in its local token bucket and continues to process the request.
[0060] Step 2: Token Status Synchronization
[0061] 1. Leader Election
[0062] When the system starts or the leader fails, a new leader is elected through an election. When a node does not receive the leader's heartbeat for a certain period of time, it will become a candidate and initiate an election. Other nodes vote as followers, and the candidate who receives more than half of the votes becomes the new leader.
[0063] 2. Token Status Synchronization
[0064] The leader is responsible for collecting the token bucket status information of each node and synchronizing the updated token allocation policy and status to other follower nodes. For example, when the status of a certain node's token bucket changes (such as changes in the number of tokens, request processing status, etc.), it will send the relevant information to the leader. The leader integrates the information and broadcasts it to all followers to ensure the consistency of the token status of each node.
[0065] Step 3: Intelligent Prediction Caching
[0066] 1. Data Collection and Preprocessing
[0067] Data Collection: The system collects the historical authentication request data of the computing power host, including information such as the request timestamp, request type, request parameters, and response data. For example, the records of the computing power host requesting platform authentication information at different time periods.
[0068] Data Preprocessing: Clean the collected data to remove noise data and abnormal data; perform normalization processing on the data to convert data in different dimensions to the same scale range for subsequent analysis.
[0069] 2. Prediction Model Construction and Training
[0070] Algorithm Selection: Machine learning algorithms such as Long Short-Term Memory Network (LSTM), Time Series Prediction Algorithm (such as ARIMA), etc. are used to construct the prediction model.
[0071] Model Training: The preprocessed historical request data is divided into a training set and a test set. The training set is used to train the model, and by adjusting the model parameters, the model can accurately learn the patterns and rules in the historical request data. For example, the Mean Squared Error (MSE) is used as the loss function, and the model parameters are updated through the backpropagation algorithm to continuously optimize the model performance.
[0072] Model Evaluation: The trained model is evaluated using the test set, and metrics such as prediction accuracy and mean squared error are calculated to determine whether the performance of the model meets the requirements. If not, adjust the model parameters or select other algorithms to retrain.
[0073] 3. Cache Content Prediction and Update
[0074] Request Prediction: Using the trained model, based on factors such as the current time and historical request patterns, predict the possible request types and data requirements in the future period. For example, predict that in a certain time period, the computing power hosts may restart on a large scale and initiate device connection authentication requests to the platform.
[0075] Cache Update: According to the prediction results, load the relevant data into the cache in advance. When the system detects an actual request, preferentially obtain the data from the cache for response. At the same time, dynamically adjust the cache content according to the real-time request situation and prediction results. For example, when it is found that some predicted requests do not occur while the frequency of other requests increases, replace the data in the cache in a timely manner to ensure the effectiveness and pertinence of the cache data.
[0076] Step 4: Token Bucket and Cache Dynamic Adjustment
[0077] 1. Token Bucket Adjustment
[0078] According to the real-time traffic load situation, dynamically adjust the generation rate r and distribution strategy of the token bucket. When the system load is low, appropriately reduce the token generation rate; when the system load is high, increase the token generation rate and adjust the token allocation strategy among nodes to ensure that the traffic can be reasonably controlled.
[0079] 2. Cache Policy Adjustment
[0080] Combined with intelligent predictive caching, when there is a sudden increase in traffic, the possible requested data is preferentially loaded into the cache. At the same time, the cache update frequency and replacement policy are dynamically adjusted according to traffic changes. For example, during peak traffic periods, increase the cache update frequency and adopt a replacement policy that combines the least recently used (LRU) with prediction to ensure that valid data is always stored in the cache.
[0081] The above are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are all included in the protection scope of the present invention.
Claims
1. A high-concurrency management method for the access of computing power hosts It is characterized in that it includes: constructing a distributed token bucket network; predicting future traffic trends through machine learning algorithms based on historical traffic data and real-time request patterns; dynamically adjusting the token generation rate and cache policy to optimize system resource allocation.
2. The method according to claim 1, characterized in that the synchronization mechanism of the distributed token bucket network includes: using the gossip protocol to regularly exchange token status information between adjacent nodes; when the local token consumption rate exceeds a preset threshold, requesting token reinforcement from adjacent nodes.
3. The method according to claim 2, characterized in that the detailed steps are as follows: The first step: Token generation and allocation 1) Token initialization The token bucket of each node is initialized when the system starts, setting the token generation rate and the token bucket capacity; 2) Token generation Each node continuously generates tokens at the set token generation rate and puts the newly generated tokens into the token bucket; When the token bucket is full, the excess tokens will be discarded; 3) Token request and allocation When a node receives a request, it tries to obtain a token from the local token bucket; If there are available tokens in the token bucket, the request is allowed to pass and be processed; When the number of tokens in the local token bucket of a node is insufficient to process new requests and the local load exceeds a preset threshold, the node will request tokens from other lightly loaded nodes.
4. The method according to claim 1, characterized in that the intelligent prediction cache policy includes: analyzing behavior patterns, identifying high-frequency access data items and access time windows; preloading data with a high probability of access to the edge node cache based on the prediction model; implementing a multi-level cache architecture, including memory cache, disk cache, and distributed cache clusters.
5. The method according to claim 4, characterized in that intelligent prediction caching includes: 1) Data collection and preprocessing Data collection: Collect historical authentication request data of the computing power host, including request timestamps, request types, request parameters, and response data; Data preprocessing: Clean the collected data to remove noise data and abnormal data; perform normalization processing on the data to convert data of different dimensions to the same scale range for subsequent analysis; 2) Prediction model construction and training Algorithm selection: Use machine learning algorithms to construct a prediction model; Model training: Divide the preprocessed historical request data into a training set and a test set; use the training set to train the model, and by adjusting the model parameters, make the model accurately learn the patterns and rules in the historical request data; Model evaluation: Use the test set to evaluate the trained model, calculate the prediction accuracy and mean square error, and determine whether the performance of the model meets the requirements. If not, adjust the model parameters or select other algorithms to retrain; 3) Cache content prediction and update Request prediction: Use the trained model to predict the possible request types and data requirements in the future period according to the current time and historical request patterns; Cache update: According to the prediction results, relevant data is loaded into the cache in advance. When an actual request is detected, data is preferentially retrieved from the cache for response. At the same time, according to the real-time request situation and prediction results, the cache content is dynamically adjusted.
6. The method according to claim 3, wherein it further includes: constructing a distributed monitoring system to collect the token usage and cache hit rate of each node in real time; generating a performance report based on the monitoring data for evaluating the overall system efficiency and optimizing the prediction model; implementing an adaptive feedback mechanism to adjust the prediction parameters and control strategy according to the actual operation effect.
7. The method according to claim 6, wherein Token bucket adjustment Dynamically adjust the generation rate and distribution strategy of the token bucket according to the real-time traffic load situation; when the system load is below the preset value, reduce the token generation rate; when the system load is higher than the preset value, increase the token generation rate and adjust the token allocation strategy among nodes to ensure that the traffic can be reasonably controlled.
8. The method according to claim 6, wherein Cache policy adjustment Combined with intelligent predictive caching, when the traffic surges, the requested data is preferentially loaded into the cache; at the same time, the cache update frequency and replacement policy are dynamically adjusted according to the traffic changes.