Server resource dynamic scheduling method for dealing with video stream high concurrent access

By building an access hotspot prediction model and a resource status perception model, combined with deep learning and graph neural networks, we can dynamically identify access hotspot areas and make refined allocations of server resources. This solves the response delay and lag issues of video servers during high-concurrency access, and achieves efficient resource scheduling and improved user experience.

CN120751207AActive Publication Date: 2025-10-03SBAIDA INTERNET OF THINGS TECH (BEIJING) CO LTD +1

Patent Information

Application Number
CN202511253962.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-10-03
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

When faced with high concurrent access to video streams, the existing video server scheduling system is unable to perceive and respond to dynamic changes in peak access areas in real time, and lacks the ability to model user behavior characteristics, resulting in problems such as response delays, video freezes, and service interruptions.

Method used

By building an access hotspot prediction model and a resource status perception model, combined with deep learning and graph neural networks, we can dynamically identify access hotspot areas and make refined allocations of server resources. We use a feedback-driven self-learning mechanism to optimize the scheduling strategy and achieve intelligent allocation and rapid response of server resources.

Benefits of technology

It significantly reduces system latency, improves access success rate and video playback experience, and has good generalization capabilities and deployment flexibility. It is suitable for high-concurrency video service scenarios such as live streaming platforms, short video applications, and online education.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751207A_ABST
    Figure CN120751207A_ABST
Patent Text Reader

Abstract

The invention discloses a server resource dynamic scheduling method for dealing with video stream high concurrent access, and particularly relates to the technical field of computer network and intelligent scheduling. A user access behavior data set is constructed; a deep learning model is adopted to train an access hot spot prediction model based on the time sequence features to predict a future access hot spot area and a peak trend; acquiring response delay, CPU / GPU occupancy rate and bandwidth load information of the heterogeneous server group, and generating a resource state multi-dimensional parameter set; carrying out joint modeling on the model and the parameter set, constructing a resource scheduling priority model by adopting a graph neural network, and generating an optimal scheduling path graph based on an A * improved algorithm; task transfer, instance elastic expansion, cache preheating and other scheduling operations are executed according to the model; according to the method, the resource utilization efficiency and the service quality of the video system in a high-concurrency scene can be improved, and the method has the advantages of being high in real-time performance, intelligent in scheduling and high in self-learning capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer networks and intelligent scheduling, and in particular to a method for dynamically scheduling server resources to cope with high-concurrency access to video streams. Background Art

[0002] With the widespread deployment of video platforms, live streaming services, and online conferencing systems, server systems face significant challenges in maintaining stability and responsiveness when faced with high levels of concurrent video stream access. This is especially true during hot events or emergencies, where large numbers of users access the same video resource within a short period of time. This can easily lead to an imbalance in server resource allocation, resulting in response delays, video freezes, service interruptions, and other issues that severely impact the user experience.

[0003] Existing video server scheduling systems often use static or preset load balancing strategies, which are unable to perceive and respond to dynamic changes in peak access areas in real time and lack the ability to elastically adapt to sudden concurrent requests. Furthermore, their ability to model video stream distribution patterns and user behavior characteristics is limited, making it difficult to predict access pressure concentration areas at the source and optimize resource scheduling in advance.

[0004] Therefore, it is urgent to propose a dynamic scheduling method for server resources for high-concurrency video streaming scenarios, which can establish a prediction model based on real-time access data, user behavior patterns and resource usage status, and dynamically identify access hotspots and weak nodes, thereby realizing refined and predictive allocation of server resources to improve system response efficiency and stability. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for dynamically scheduling server resources to cope with high concurrent access to video streams, so as to solve the shortcomings in the background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for dynamically scheduling server resources to cope with high concurrent access to video streams, comprising: S100, obtaining video access behavior data, video type information and server resource usage status within a fixed time period, and constructing a user access behavior dataset D; S200, based on the time series characteristics of user access behavior, using a deep learning model to train the dataset D, and establish an access hotspot prediction model S for predicting access hotspot areas and access peak trends of various video resources in the future time period; S300, obtaining the current service response delay, CPU / GPU occupancy rate, and network bandwidth load of the heterogeneous server group to form a multi-dimensional parameter set R of server resource status; S400, performing joint feature analysis on the access hotspot prediction model S and the resource status parameter set R, and constructing a resource scheduling priority model W for dynamically allocating resource weights and forming a scheduling path diagram; S500: Execute a dynamic scheduling strategy for server resources according to the model W; S600: Collect service response delay changes and user experience scores based on resource scheduling execution results to form a multi-dimensional feedback data set F; S700 : Based on the feedback data set F, adjust the access hotspot prediction model S and the scheduling priority model W in real time.

[0007] Preferably, the S200 includes: S201, extracting the time series features of access frequency, number of active users, and access bounce rate at each time granularity from the user access behavior dataset, and constructing a unified sequence feature matrix; S202. Using a bidirectional long short-term memory network with an attention mechanism to train the sequence feature matrix to enhance the model's ability to identify historical access fluctuation trends and future high-concurrency nodes; S203. Identify key time windows and influencing factors through the attention weight distribution output by the model, and generate a video resource access hotspot distribution map and access peak time prediction value based on the prediction results.

[0008] Preferably, the S300 includes: S301, classifying the configuration characteristics of each node in the server group, and dividing them into CPU-type, GPU-type, and hybrid heterogeneous server clusters according to computing architecture; S302: Using a lightweight edge monitoring agent to collect service response delay, CPU / GPU utilization, and network bandwidth load data of each node in real time, and send it to the central scheduling controller at a fixed period; S303: Through normalization and weighted aggregation processing, the service response delay, CPU / GPU utilization and network bandwidth load data are integrated to build a unified server resource state vector, and the multi-dimensional parameter set R is generated in combination with the node type label.

[0009] Preferably, the S400 includes: S401: Jointly encode the hotspot distribution map output by the access hotspot prediction model S and the resource status parameter set R to construct a scheduling input tensor that integrates access intensity and server load status; S402. Introduce a scheduling modeling method based on a graph neural network, using server nodes as graph vertices and inter-server links as edges, train a resource scheduling priority model W, and output the resource allocation weight of each node; S403: Calculate the scheduling path costs between different servers based on the output results of model W, and dynamically construct a scheduling path graph between servers using an improved minimum cost path algorithm.

[0010] Preferably, the improved minimum cost path algorithm is adopted, including: S431, taking the server node associated with the access hotspot area as the path starting point, initializing the path costs and priority indicators of all server nodes; S432. Introducing a dynamic bandwidth prediction function and a node load penalty factor into the path search process to improve the heuristic function in the traditional A* algorithm; S433. During each round of path expansion, the path costs of candidate nodes are dynamically updated, and the path with the lowest comprehensive cost is preferentially selected for search and recommendation; S434. Finally, one or more low-cost scheduling paths from the access pressure center node to the target resource node are output to guide task transfer or traffic guidance between servers.

[0011] Preferably, the S600 includes: S601. Record the average response delay of each server node before and after each resource scheduling. Set the delay change value ΔL as the difference between the average delay in the time window after scheduling and the delay before scheduling. The calculation expression is: ;in, Indicates the average response time of requests collected within 5 minutes after scheduling. It represents the average response time in the same window before scheduling, and ΔL is used to dynamically reflect the impact trend of scheduling strategy on node response performance.

[0012] Preferably, the S600 further includes: S602: Extracting the initial delay of video playback based on the user request processing log , lag rate and average bit rate The expression for calculating the user experience score U based on these three indicators is: ; Among them, α, β, and γ are normalized weight parameters, satisfying α+β+γ=1, and are set according to business needs.

[0013] Preferably, the S700 includes: S701: Perform distribution analysis on the service response delay changes and user experience scores collected in the feedback dataset F, and construct a scheduling effect evaluation label set to identify the quality of the model output results. S702: Fine-tune the access hotspot prediction model S using an incremental training strategy, and perform local retraining only on the time series samples corresponding to the high-error areas. S703: Based on the deviation between the historical scheduling path of the scheduling priority model W and the actual resource utilization, a reinforcement learning method is used to adjust the node weight update strategy in the model to optimize the resource allocation decision logic; S704. Periodically evaluate the performance changes of the adjusted model on the validation set and determine whether to push the updated model into the production environment.

[0014] In the above technical solution, the technical effects and advantages provided by the present invention are: 1. This invention achieves accurate prediction, intelligent allocation, and rapid response of server resources in high-concurrency video access scenarios by constructing a multi-stage video server resource scheduling process centered on access behavior modeling, hotspot prediction, resource status awareness, graph neural network scheduling, dynamic path construction, and containerized resource elasticity management. Compared to traditional static load balancing or rule-based scheduling methods, this invention dynamically generates scheduling paths and task migration plans based on real-time changes in user access trends and server operating status, significantly reducing system latency and improving access success rates and video playback experience.

[0015] 2. This invention introduces a feedback-driven self-learning mechanism. By collecting service response performance and user experience scores, it continuously optimizes the hotspot prediction model and scheduling strategy, achieving closed-loop self-adaptation and iterative model updates for the scheduling system. This method possesses excellent generalization and deployment flexibility, and can be widely applied to various high-concurrency video service scenarios, such as live streaming platforms, short video applications, online education, and large-scale conference systems, with significant engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments described in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0017] Figure 1 This is a mind map of the method of the present invention. DETAILED DESCRIPTION

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0019] For examples, see Figure 1As shown, the method for dynamically scheduling server resources to cope with high concurrent access to video streams described in this embodiment includes: S100, obtaining video access behavior data, video type information and server resource usage status within a fixed time period, and constructing a user access behavior dataset D; S200, based on the time series characteristics of user access behavior, using a deep learning model to train the dataset D, and establish an access hotspot prediction model S for predicting access hotspot areas and access peak trends of various video resources in the future time period; S300, obtaining the current service response delay, CPU / GPU occupancy rate, and network bandwidth load of the heterogeneous server group to form a multi-dimensional parameter set R of server resource status; S400, performing joint feature analysis on the access hotspot prediction model S and the resource status parameter set R, and constructing a resource scheduling priority model W for dynamically allocating resource weights and forming a scheduling path diagram; S500: Execute a dynamic scheduling strategy for server resources according to the model W; S600: Collect service response delay changes and user experience scores based on resource scheduling execution results to form a multi-dimensional feedback data set F; S700 : Based on the feedback data set F, adjust the access hotspot prediction model S and the scheduling priority model W in real time.

[0020] In the present invention, the specific implementation of step S100 includes the following: First, the server system collects raw access behavior data from the edge nodes and main server logs within a specified monitoring period (e.g., 7 consecutive days), including user ID, access timestamp, requested video ID, access duration, video type (on-demand or live), and other information.

[0021] S101: To extract multi-timescale features from access behavior, the system partitions the raw data into time windows of varying granularity (e.g., day, hour, minute). For example, access records are grouped into sliding windows of 24 hours, 60 minutes, and 1 minute, respectively, generating three subsets of access behavior: a daily, hourly, and minute-by-minute granularity subset. This multi-scale structure can reveal both cyclical trends in user access patterns (e.g., a concentrated visit every evening) and sudden events (e.g., a sudden surge in concurrent users upon the start of a live broadcast).

[0022] S102: For each access subset at each time granularity, the system clusters access behavior data using unsupervised clustering algorithms such as K-Means++ or DBSCAN. Clustering features may include access frequency, average duration, and bounce rate. The clustering results can automatically identify highly active user groups, typical access paths, and potential hot video resources, providing structured access patterns for subsequent hotspot prediction models.

[0023] S103: The system further incorporates video types (such as on-demand and live videos) as tag features, modeling access behavior hierarchically by content type. Live videos emphasize concurrency peaks and latency sensitivity, while on-demand videos prioritize access time distribution and cache hit rates. Ultimately, the system integrates multi-granular access features, access pattern clustering results, and video type information to construct a three-dimensional structured user access behavior dataset D encompassing time, behavioral characteristics, and content type dimensions.

[0024] This dataset D serves as an important input for the subsequent training of the access hotspot prediction model. It significantly improves the prediction model's accuracy and timeliness in identifying access hotspot areas and traffic trends, and is superior to existing static methods based only on access counts or historical average loads.

[0025] In the present invention, the specific implementation of step S200 includes the following: S201: The system first preprocesses the user access behavior dataset D constructed in step S100, extracting key time series features at multiple time scales, including metrics such as the number of video access requests per unit time (access frequency), the number of active users after deduplication, the bounce rate, and the average viewing duration. These features are arranged into multidimensional time series based on the time granularity (day, hour, minute) and uniformly constructed into a sequence feature matrix. Each row of the matrix represents multiple access metrics for a specific time window, and each column represents the time series of a specific feature.

[0026] S202: To learn and predict complex time series behavior, the system uses a bidirectional long short-term memory network with an attention mechanism for modeling and training. The model structure consists of an input layer, a bidirectional LSTM encoding layer, an attention weight calculation module, and an output layer. The bidirectional LSTM can simultaneously capture temporal dependencies between the past and the future, while the attention mechanism assigns higher weights to key time windows or feature dimensions, thereby improving the accuracy of predicting hotspot peaks. The model is trained using video access behavior data from several consecutive days, and the loss function uses a weighted combination of mean squared error and peak access prediction deviation.

[0027] S203: After model training is complete, the system uses the attention weights it outputs to visualize and analyze the prediction process, identifying key time windows and behavioral characteristics that are crucial for predicting peak traffic. For example, attention distribution can identify a potential concurrent node where live video traffic surges 30 minutes before an event. Ultimately, the model outputs a hotspot distribution map of video resource access (categorized by video ID) and predicted peak time periods, which serve as input for the subsequent scheduling model (step S400) to pre-prepare server resources or optimize load balancing strategies.

[0028] This method has stronger timeliness and prediction accuracy than traditional sliding average prediction or static load threshold identification methods, especially in significantly enhancing the scheduling response capability in burst access scenarios.

[0029] In the present invention, the specific implementation of step S300 includes the following: S301: The server system first categorizes the computing resources deployed in data centers and edge nodes in each region. Based on the node's hardware configuration characteristics (such as GPU availability, number of CPU cores, and memory capacity), the server group is divided into three categories: CPU-based servers (suitable for general request processing), GPU-based servers (suitable for high-concurrency image decoding and AI computing), and CPU+GPU hybrid servers (suitable for workloads requiring image preprocessing and deep learning model inference). This classification result serves as a resource status label for subsequent scheduling decisions.

[0030] S302: A lightweight edge monitoring agent component is deployed on each server node to regularly collect key performance indicators, including service response latency (average response time per request), CPU utilization, GPU utilization, and current inbound and outbound bandwidth load. The monitoring frequency is typically set to every 30 seconds to 1 minute. The collected data is sent to the central dispatch controller via an intranet channel to ensure data synchronization and low latency.

[0031] S303: After receiving performance data from each server node, the central dispatch controller first normalizes the different indicators to eliminate differences in numerical scales. Subsequently, it assigns weighted values ​​to the different indicators based on the type of server node (e.g., GPU-based servers prioritize AI tasks). For example, GPU occupancy is given a higher weight in GPU-based servers. Ultimately, the system generates a unified format of server resource state vectors. Each vector contains indicators such as response latency, CPU utilization, GPU utilization, and network bandwidth load, and is labeled with the server type, forming a structured multidimensional parameter set R.

[0032] This parameter set, R, fully describes the current operating status and scheduling capacity of each server and serves as the input for subsequent resource scheduling strategy calculations. This approach enables the system to achieve real-time perception and accurate modeling of the resource status of heterogeneous servers, significantly improving the accuracy and efficiency of resource allocation.

[0033] In the present invention, the specific implementation of step S400 includes the following: S401: The system first integrates the output of the access hotspot prediction model S in step S200—the predicted video access hotspot areas and traffic peak maps—with the multidimensional parameter set R of server resource status generated in step S300. Specifically, the system maps each predicted hotspot area with its potentially associated server node and generates a set of joint features for each server node, including the access intensity of the associated video (e.g., the expected number of concurrent users) and the current load status of the node (e.g., CPU / GPU utilization, response latency). These joint features are encoded as a scheduling input tensor, which serves as the input basis for the subsequent scheduling model.

[0034] S402: Based on the aforementioned input tensor, the system uses a graph neural network (GNN) to model resource scheduling. This model uses each server node in the server cluster as a graph vertex. Edges are created where network connections exist between servers. Edge weights can be set based on physical distance, network latency, or historical load transfer efficiency. The GNN model aggregates state information from neighboring nodes through a graph-structured propagation mechanism and, combined with local features, learns a scheduling priority weight for each server node. This priority value represents the server's scheduling suitability and distribution capabilities under the current access pressure. The training process uses historical scheduling efficiency and system response indicators as supervisory signals.

[0035] S403: Based on the scheduling priorities of each server node output by the GNN model, the system constructs a scheduling path diagram for the entire server graph. Using a modified minimum-cost path algorithm (such as a Dijkstra or A* variant), starting from the access pressure center, the system finds the optimal path for distributing tasks to other servers. The path cost comprehensively considers scheduling latency, link bandwidth, and node load. The resulting scheduling path diagram not only indicates how tasks should be migrated or replicated across multiple nodes but also provides a routing reference for subsequent resource allocation and caching strategies.

[0036] This method breaks through the limitations of traditional static scheduling or single-node priority sorting methods, and uses graph neural networks and graph path optimization algorithms to achieve deep coupling and automated reasoning of scheduling priority modeling and dynamic generation of scheduling paths, greatly improving scheduling intelligence and system response flexibility.

[0037] In the present invention, after the scheduling priority model W is generated in step S400, an improved minimum cost path algorithm (e.g., based on a variant of the A* algorithm) is further used to dynamically generate a scheduling path graph between servers. The specific implementation is as follows: S431: The system first uses the core server node associated with the hotspot area output by the access hotspot prediction model S as the path starting point, i.e., the access pressure center. The scheduling target is set to a server node cluster with available resources. The total path cost f(n) for all nodes is initialized to infinity, and f(start) for the starting node is initialized to 0. At the same time, a path heuristic evaluation value h(n) is defined for each node, initially estimated as the weighted predicted value of the network delay from that node to the target server.

[0038] S432: The system introduces two improvements based on the A* algorithm to improve the adaptability of scheduling paths: Dynamic Bandwidth Prediction Function: This function uses historical bandwidth usage data and a sliding window averaging method to predict the future available bandwidth of inter-server links, reflecting the stability of path transmission. This value is reflected in the edge weight and is used to adjust path selection preferences.

[0039] Node Load Penalty Factor: For server nodes with high CPU / GPU load or near-threshold response latency, a penalty coefficient, P(n), is introduced. This factor is weighted in the path cost g(n): g(n) = cumulative path transmission cost + P(n). P(n) is dynamically determined by node load, ensuring that scheduling paths avoid overloaded nodes.

[0040] The heuristic function f(n)=g(n)+h(n) is therefore updated to: f(n)=current cost+node load penalty+dynamic delay prediction value+inverse bandwidth prediction term.

[0041] S433: During path expansion, the system uses a priority queue to maintain a set of nodes to be expanded. Each time, the node with the smallest f(n) value is selected for search and recommendation. During each round of expansion, the corresponding g(n) and h(n) are recalculated based on the current node's adjacent edges (i.e., reachable servers). Whether the cost is better than the existing cost is determined by whether the path record is updated. This strategy supports real-time response to node status changes, ensuring that the path search process consistently expands towards low-load, highly available nodes.

[0042] S434: The search terminates when the path successfully reaches any node in the target server set or reaches the set maximum search depth limit. The system outputs one or more scheduling paths with the lowest overall cost, starting from the access hotspot area. These paths can be used for subsequent video task migration, traffic redirection, or content preloading. For multiple feasible paths, the system can further select a path based on the scheduling objective (e.g., latency-sensitive or throughput-optimized).

[0043] Compared with traditional static routing or simple shortest path methods, this path generation method has dynamic adaptation, load avoidance and prediction capabilities, and can significantly improve the real-time and stability of video server cluster resource scheduling in high-concurrency access scenarios.

[0044] In the present invention, the specific implementation of step S500 includes the following: S501: After calculating the resource scheduling priority model W and generating a scheduling path graph, the system automatically generates a set of scheduling instructions for task distribution and migration based on the resource weight allocation results and the optimal scheduling path for each server node. This instruction set includes the target node for hot video task distribution, cache deployment location, service instance scaling strategy, and other content. It is sent to the scheduling control center or Kubernetes controller in a YAML or JSON structure format for parsing and execution.

[0045] S502: For access to popular video resources, the system adopts a content slicing caching strategy. First, popular videos are segmented by time slice or bitrate (e.g., 10-second segments). Core segments are prioritized and deployed to high-priority nodes close to the user. These segments are preloaded before the user's first request, ensuring low-latency playback on the first access. Simultaneously, the system utilizes a user access trajectory prediction model to pre-determine the potential user path and pre-deploy likely video segments to nodes along that path.

[0046] S503: To improve its ability to cope with sudden concurrency, the system dynamically adjusts the number of service instances based on containerized deployment platforms (such as Kubernetes and Docker Swarm). For server nodes approaching overload, replica service instances are quickly launched on adjacent, less-loaded nodes, and traffic is automatically routed to the new nodes via the service mesh. Furthermore, task instances can be migrated across nodes in seconds using a container image migration mechanism. During this process, the system selects the target node with the lowest response latency based on the scheduling path map, enabling intelligent resource avoidance.

[0047] S504: During task execution, the system continuously monitors the service response status and instance processing efficiency of each node. If it detects a significant increase in response latency or a backlog in the processing queue, it immediately feeds the node status back to the central scheduling system, triggering a new round of retraining of model W and updating the path map. This mechanism forms a real-time feedback loop, enabling continuous optimization and rapid convergence of resource scheduling strategies in high-concurrency scenarios, and significantly reducing the risk of service bottlenecks caused by static scheduling lags.

[0048] This embodiment achieves efficient dynamic scheduling of server resources through four core means: resource scheduling weight guidance, hotspot preheating cache, elastic container expansion, and feedback closed-loop scheduling. It has significant stability and responsiveness advantages in complex application scenarios such as sudden access peaks and concentrated requests for hot videos.

[0049] In the present invention, step S600 is used to build a scheduling feedback closed loop, and its specific implementation is as follows: This step aims to establish a multidimensional feedback dataset F by collecting and analyzing the operational effects of resource scheduling policies after execution. This dataset provides a basis for subsequent model updates and scheduling optimization. This dataset covers two core feedback dimensions: changes in service response performance and user experience, reflecting the actual impact of scheduling policies on system operation and end users, respectively.

[0050] After completing the scheduling and execution of server resources, the system continuously samples the service response time of each server node. The main monitoring indicator is the average response delay, that is, the time taken from the user request to the server returning the first byte of data per unit time.

[0051] The specific implementation is as follows: Data collection cycle definition: Set the 5 minutes before and after the scheduling action as the performance evaluation window, and record the average response time of each node within the window. Calculation expression: ;in: Indicates the average response delay within 5 minutes before the schedule is executed; It represents the average response delay within 5 minutes after the scheduling is executed. ΔL represents the performance change brought about by the scheduling. If ΔL<0, it means that the response performance is improved.

[0052] Server nodes periodically record and process logs and timestamp data through edge probes or local agents. A central controller aggregates and calculates the difference. The delay change value serves as a key indicator for scheduling effectiveness and is used in the next round of scheduling weight adjustments.

[0053] In order to measure the pros and cons of resource scheduling strategies from the perspective of end users, this system introduces the User Experience Score (U), which integrates multiple user perception level indicators for unified quantification.

[0054] Specifically, the following indicators are included: : The initial loading time of the video, that is, the delay from the user clicking play to the start of the video; : Jam rate, which is the ratio of the total duration of video playback interruptions to the total video duration; : Average playback bitrate, reflecting the video quality level.

[0055] The formula for calculating the user experience score is: ; Among them: α, β, γ are weight coefficients set by experience, which are used to adjust the impact of each indicator on the overall score. Generally, α=0.3, β=0.4, and γ=0.3; All input indicators are normalized in the range of [0,1] to ensure comparability of scores; The higher the score U, the better the viewing experience of the user in that time period.

[0056] The system continuously collects these experience parameters while users watch videos, combines them with metadata such as video type (on-demand / live), access device type, etc., to form a user experience score log, and aggregates them to form an experience data subset.

[0057] After collecting data from the above two dimensions, the system stores the feedback results of each scheduling cycle in a structured manner as a multi-dimensional feedback dataset F. Each record contains the following fields: Scheduling timestamp; node ID and resource status snapshot; service response delay variation; user experience score; scheduling path and target node information; The feedback dataset F will serve as input for the next step of model training and parameter correction, providing historical scheduling effect data support for the access hotspot prediction model S and resource priority model W, thereby forming a self-learning and self-optimization closed-loop mechanism for the scheduling system.

[0058] In the present invention, step S700 is used to optimize and update the access hotspot prediction model S and the scheduling priority model W in real time after collecting feedback data, thereby building an adaptive and evolvable intelligent scheduling system.

[0059] This step builds a feedback-driven model update mechanism to enable the system to continuously learn and correct its own judgments during operation, thereby responding to dynamic factors such as changes in access behavior, fluctuations in server status, and updates to hot content.

[0060] S701: Construct a scheduling effect evaluation label set; the system first performs structured processing on the feedback data set F formed in step S600, and extracts the service response delay change value (ΔL) and user experience score (U) as the main evaluation indicators.

[0061] Discrete label generation: After normalizing ΔL and U, the results are divided into three levels: high efficiency (positive improvement), neutral (no significant change), and low efficiency (deterioration) using a threshold.

[0062] Example: If ΔL < -100ms and U increases by >10%, mark it as "Scheduling is good +"; if ΔL > +150ms or U decreases by more than 15%, mark it as "Scheduling is bad -".

[0063] Label set structure: Each label record is bound to the comparison result between the model output and actual feedback within a scheduling cycle, forming the "scoring label" of the training sample.

[0064] This step provides a real-world "supervision signal" for subsequent model training, guiding the model to learn and correct scheduling deviations.

[0065] S702: Incremental fine-tuning of the access hotspot prediction model S. The access hotspot prediction model S uses a time series deep neural network based on the Bi-LSTM with Attention architecture. To avoid resource consumption caused by frequent full training, the system adopts an incremental learning strategy: Error localization: Utilize the deviation between previous and subsequent predictions (such as the difference between predicted hotspot areas and actual access distribution) to filter out high-error samples; combined with the label set, only video resources and time periods with significant error rates are selected as "high-value training samples."

[0066] Local retraining: High-error samples are constructed into short-period time series micro-sets; the original model structure is retained, and only the weights of the later layers are fine-tuned to update the prediction preference without destroying the global stability of the model; the training cycle is greatly shortened, and fast updates can be completed within the scheduling period (such as every 15 minutes).

[0067] The technical advantages are: achieving rapid iteration of the prediction model, improving the accuracy and real-time performance of hotspot identification, and not affecting other stable prediction paths.

[0068] S703: Strategy optimization (reinforcement learning) of the scheduling priority model W. The scheduling priority model W uses a resource scheduling structure based on a graph neural network (GNN) to output the resource scheduling weights of the server nodes. The system further introduces a reinforcement learning mechanism (RL) to optimize the scheduling strategy: State-Action Modeling: Status: Consists of the current scheduling graph structure, node load status, predicted hotspot intensity, etc. Action: represents the scheduling path selection and resource migration amount decision.

[0069] Reward function design: The reward function combines ΔL and U in the feedback dataset; If the scheduling result improves (ΔL decreases, U increases), a positive reward is given, otherwise a penalty is imposed.

[0070] Policy Update: Optimize the scheduling path selection logic using strategies such as Q-learning or Actor-Critic. The model dynamically adjusts the weight update rate and priority determination rules for each node to avoid path congestion or resource misallocation. This step enables the scheduling model to have adaptive learning capabilities, improving its resource allocation efficiency under different system states.

[0071] S704: Model evaluation and deployment switching mechanism; To ensure stable operation of the updated model without causing system fluctuations, the system introduces a grayscale evaluation mechanism: Model comparison test: The new model and the previous version are run simultaneously on the offline validation set and some production traffic. Comparison indicators include hotspot prediction accuracy, average path latency, and changes in user experience scores.

[0072] Switching threshold setting: Set the switching conditions. For example, if the prediction accuracy is improved by >5% or the response delay is reduced by >10%, the old model will be automatically replaced. If the new model does not meet the set improvement standards, the current version will be retained and the switching will be delayed.

[0073] Deployment strategy: Use rolling updates and blue-green deployment mechanisms to minimize update risks and ensure system continuity.

[0074] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.

Claims

1. A method for dynamically scheduling server resources to cope with high concurrent access to video streams, characterized by: include: S100, obtaining video access behavior data, video type information and server resource usage status within a fixed time period, and constructing a user access behavior dataset D; S200, based on the time series characteristics of user access behavior, using a deep learning model to train the dataset D, and establish an access hotspot prediction model S for predicting access hotspot areas and access peak trends of various video resources in the future time period; S300, obtaining the current service response delay, CPU / GPU occupancy rate, and network bandwidth load of the heterogeneous server group to form a multi-dimensional parameter set R of server resource status; S400, performing joint feature analysis on the access hotspot prediction model S and the resource status parameter set R, and constructing a resource scheduling priority model W for dynamically allocating resource weights and forming a scheduling path diagram; S500: Execute a dynamic scheduling strategy for server resources according to the model W; S600: Collect service response delay changes and user experience scores based on resource scheduling execution results to form a multi-dimensional feedback data set F; S700 : Based on the feedback data set F, adjust the access hotspot prediction model S and the scheduling priority model W in real time.

2. A method for dynamically scheduling server resources to cope with high concurrent access of video streams according to claim 1, characterized in that: The S200 includes: S201, extracting the time series features of access frequency, number of active users, and access bounce rate at each time granularity from the user access behavior dataset, and constructing a unified sequence feature matrix; S202. Using a bidirectional long short-term memory network with an attention mechanism to train the sequence feature matrix to enhance the model's ability to identify historical access fluctuation trends and future high-concurrency nodes; S203. Identify key time windows and influencing factors through the attention weight distribution output by the model, and generate a video resource access hotspot distribution map and access peak time prediction value based on the prediction results.

3. The method for dynamically scheduling server resources for coping with high concurrent access of video streams according to claim 1, wherein S300 comprises: S301, classifying the configuration characteristics of each node in the server group, and dividing them into CPU-type, GPU-type, and hybrid heterogeneous server clusters according to computing architecture; S302: Using a lightweight edge monitoring agent to collect service response delay, CPU / GPU utilization, and network bandwidth load data of each node in real time, and send it to the central scheduling controller at a fixed period; S303: Through normalization and weighted aggregation processing, the service response delay, CPU / GPU utilization and network bandwidth load data are integrated to build a unified server resource state vector, and the multi-dimensional parameter set R is generated in combination with the node type label.

4. The method for dynamically scheduling server resources for coping with high concurrent access of video streams according to claim 1, wherein S400 comprises: S401: Jointly encode the hotspot distribution map output by the access hotspot prediction model S and the resource status parameter set R to construct a scheduling input tensor that integrates access intensity and server load status; S402. Introduce a scheduling modeling method based on a graph neural network, using server nodes as graph vertices and inter-server links as edges, train a resource scheduling priority model W, and output the resource allocation weight of each node; S403: Calculate the scheduling path costs between different servers based on the output results of model W, and dynamically construct a scheduling path graph between servers using an improved minimum cost path algorithm.

5. The method for dynamically scheduling server resources to cope with high concurrent access to video streams according to claim 4, wherein the improved minimum cost path algorithm is used, comprising: S431, taking the server node associated with the access hotspot area as the path starting point, initializing the path costs and priority indicators of all server nodes; S432. Introducing a dynamic bandwidth prediction function and a node load penalty factor into the path search process to improve the heuristic function in the traditional A* algorithm; S433. During each round of path expansion, the path costs of candidate nodes are dynamically updated, and the path with the lowest comprehensive cost is preferentially selected for search and recommendation; S434. Finally, one or more low-cost scheduling paths from the access pressure center node to the target resource node are output to guide task transfer or traffic guidance between servers.

6. The method for dynamically scheduling server resources for coping with high concurrent access of video streams according to claim 1, wherein S600 comprises: S601. Record the average response delay of each server node before and after each resource scheduling. Set the delay change value ΔL as the difference between the average delay in the time window after scheduling and the delay before scheduling. The calculation expression is: ;in, Indicates the average response time of requests collected within 5 minutes after scheduling. It represents the average response time in the same window before scheduling, and ΔL is used to dynamically reflect the impact trend of scheduling strategy on node response performance.

7. The method for dynamically scheduling server resources for coping with high concurrent access of video streams according to claim 6, wherein S600 further comprises: S602: Extracting the initial delay of video playback based on the user request processing log , lag rate and average bit rate The expression for calculating the user experience score U based on these three indicators is: ; Among them, α, β, and γ are normalized weight parameters, satisfying α+β+γ=1, and are set according to business needs.

8. The method for dynamically scheduling server resources for coping with high concurrent access of video streams according to claim 1, wherein S700 comprises: S701: Perform distribution analysis on the service response delay changes and user experience scores collected in the feedback dataset F, and construct a scheduling effect evaluation label set to identify the quality of the model output results. S702: Fine-tune the access hotspot prediction model S using an incremental training strategy, and perform local retraining only on the time series samples corresponding to the high-error areas. S703: Based on the deviation between the historical scheduling path of the scheduling priority model W and the actual resource utilization, a reinforcement learning method is used to adjust the node weight update strategy in the model to optimize the resource allocation decision logic; S704. Periodically evaluate the performance changes of the adjusted model on the validation set and determine whether to push the updated model into the production environment.

Citation Information

Patent Citations

  • Resource distribution back-to-source live video acceleration system and method

    CN117834956A

  • Visual scheduling method and system for scientific and technological achievement evaluation system

    CN120104850A

  • Operation management method and system of server cluster

    CN120416256A

  • Platform for orchestrating a scalable, privacy-enabled network of collaborative and negotiating agents utilizing modular hybrid computing architecture

    US20250259044A1

Cited By

  • One-key release and link generation method and system for virtual space application

    CN121037305A

  • Power supply safety real-time protection system based on power transmission line construction

    CN121235242A

  • E-commerce order efficient checking system based on hierarchical time wheel and intelligent prediction model

    CN121352932A

  • 4G three-card single standby method and system applied to monitoring camera

    CN121357525A

  • Cluster CPU cooperative frequency regulation method and system, computer equipment and medium

    CN122019189A