Server load balancing method and system based on edge computing

By analyzing multi-dimensional parameters in the edge computing environment and dynamically adjusting node grouping and task levels, the resource imbalance problem of traditional server load balancing methods in high-fluctuation scenarios is solved, achieving more efficient resource utilization and scheduling response.

CN121301008AActive Publication Date: 2026-01-09百信信息技术有限公司
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511464196.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-01-09
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

Traditional server load balancing methods cannot promptly identify node availability in scenarios with large network fluctuations, heterogeneous nodes, and frequent communication switching, leading to unbalanced resource distribution, frequent idle or overloaded nodes, reduced task distribution continuity, and fixed node groups, making it difficult to adapt to the needs of edge scenarios with high fluctuations and high dynamic loads.

Method used

The server load balancing method based on edge computing analyzes multi-dimensional parameters of heterogeneous server nodes, dynamically adjusts node grouping criteria, identifies bottlenecks, optimizes task level adaptation logic, realizes dynamic aggregation of node groups and resource allocation, and improves scheduling sensitivity and response continuity.

Benefits of technology

In edge deployment environments with high-frequency changes and uneven resources, task flow and node distribution tend to be more balanced, resource utilization and grouping structure are more flexible, and the resource utilization and scheduling response capability of server groups are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301008A_ABST
    Figure CN121301008A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of servers, in particular to a server load balancing method and system based on edge computing, and the method comprises the following steps: monitoring the state data of heterogeneous server nodes, identifying parameter fluctuation abnormal nodes, analyzing task scheduling and resource change trends, and screening and grouping the nodes based on an edge computing scene; and evaluating continuous online and communication active performance of the nodes, and classifying the nodes according to a grading standard to obtain an availability distribution structure. According to the method, the coupling rule of load fluctuation, resource collaboration and bandwidth flow dynamic change among the nodes is mined in real time, hierarchical dynamic collection of node availability is achieved, a resource grouping result is continuously optimized according to node task bearing capacity and multi-time-window communication activeness, node scheduling allocation fully adapts to a network state and a heterogeneous computing power scene, and the resource allocation efficiency is improved. Task circulation and node distribution tend to be balanced, overall resource utilization and a grouping structure of the cluster are more flexible, and the method can adapt to the problems of high-frequency change and uneven resources in an edge deployment environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of server technology, and in particular to a server load balancing method and system based on edge computing. Background Technology

[0002] The server domain involves the management and scheduling of computing resources, including server deployment, task allocation, data processing, and multi-server collaboration. This technical field encompasses server hardware architecture, software resource management, network communication, task distribution, fault tolerance mechanisms, and load balancing. Traditional server load balancing refers to distributing user requests across a server cluster using hardware load balancing devices or software algorithms to achieve a relative balance in workload among servers. Common methods for allocating and managing task requests include round-robin scheduling, least-connection allocation, allocation based on server response time, and session persistence-based scheduling strategies.

[0003] Traditional solutions lack multi-dimensional parameter linkage and dynamic hierarchical management. The scheduling method relies on a single load or static indicator. In scenarios with large network fluctuations, heterogeneous nodes, and frequent communication switching, it is impossible to identify node availability in a timely manner. This leads to unbalanced resource distribution, frequent idle or overload of some nodes, reduced task distribution continuity, and fixed node groups, resulting in some nodes being unable to participate in scheduling for a long time. The overall ability of the server group to adapt to complex environments is limited, making it difficult to meet the needs of edge scenarios with high fluctuations and high dynamic loads. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a server load balancing method and system based on edge computing.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a server load balancing method based on edge computing, comprising the following steps:

[0006] S1: Based on heterogeneous server nodes in edge computing scenarios, compare the fluctuations of each node at the same time, analyze the differences between parameters, filter the node data where the fluctuation difference between parameters exceeds the average change range, integrate various parameter interaction information, and obtain a global load feature set.

[0007] S2: Based on the global load feature set, determine the historical trend of GPU and CPU utilization, analyze the scheduling response frequency, screen nodes with abnormal synchronization of memory remaining and data traffic, calculate the bandwidth change trend, dynamically adjust the node grouping standard, and obtain the bottleneck layer identification quantity.

[0008] S3: Based on the bottleneck layer identifier, filter out non-abnormal nodes, analyze the current changes and historical extreme fluctuations of bandwidth, determine the memory change range, adjust the task level adaptation logic, and obtain the task node adaptation rate.

[0009] S4: Based on the task node adaptability, analyze the online duration and heartbeat response records of schedulable nodes within a preset time window, determine the communication status within minute, hour and day time windows, compare the integrity of periodic communication records, calculate the activity level of multi-time-window nodes, and obtain the periodic activity performance quantity.

[0010] S5: Based on the periodic active performance, determine the node hierarchical affiliation, analyze the distribution trend of periodic integral data, optimize the node grouping method, identify priority, schedulable and restricted allocation node groups, compare the affiliation relationship of each group, and obtain the availability distribution structure.

[0011] The present invention improves upon this invention by including the following features: the global load feature set includes running status features, fluctuation coupling features, and interaction distribution features; the bottleneck hierarchical identification quantity includes computing resource identification, storage resource identification, and network resource identification; the task node adaptability includes task matching indicators, resource response indicators, and allocation coordination indicators; the periodic activity performance quantity includes online activity indicators, communication integrity indicators, and time window performance indicators; and the availability distribution structure includes priority groups, scheduling groups, and restriction groups.

[0012] The present invention is improved in that the step of obtaining the global load feature set is specifically as follows:

[0013] S111: Based on heterogeneous server nodes in edge computing scenarios, analyze the GPU utilization, CPU utilization, remaining memory, network bandwidth utilization and data traffic parameters of each node, compare the changes of each parameter of each node at the same time point, determine whether there are differences in the synchronous change trend, identify node data with change amplitude exceeding the average range, and obtain abnormal parameter fluctuation data.

[0014] S112: Based on the abnormal parameter fluctuation data, compare the changing directions of multiple parameters of the same node in the continuous time series, analyze the synergy of the changes of similar parameters between nodes, determine the data segments with multiple parameters synchronously offset, and combine the associated parameters into a group to obtain the fluctuation difference judgment parameters;

[0015] S113: Based on the fluctuation difference determination parameters, calculate the data signals of GPU utilization, CPU utilization, memory remaining, network bandwidth utilization and data traffic of each node, analyze the coordinated change pattern of multiple state parameters at the same point in time, screen the combinations of parameter changes with linkage, and obtain the global load feature set.

[0016] The present invention is improved in that the step of obtaining the bottleneck layer identifier is specifically as follows:

[0017] S211: Based on the global load feature set, analyze the GPU utilization and CPU utilization of each node, compare the time-series change patterns of utilization data at multiple sampling times, determine the consistency of the change trends of similar parameters among nodes, identify utilization data that exhibit a unidirectional change direction within a preset time window, and obtain a utilization time-series trend sequence.

[0018] S212: Based on the utilization time-series trend sequence, determine the task scheduling records of each node, analyze the time-series relationship between the number of task receptions and the status response of each node in the scheduling cycle, identify the node parameters whose change rates of remaining memory and data traffic both exceed the preset change threshold in the same scheduling cycle, and aggregate them into the data combination of the corresponding node to obtain the synchronization abnormal parameter set.

[0019] S213: Based on the set of synchronization anomaly parameters, calculate the network bandwidth change sequence of each node, analyze the directional change of bandwidth change trend within the continuous scheduling period, determine the node group whose standard deviation of bandwidth change exceeds the preset stability threshold, adjust the node grouping standard, and obtain the bottleneck hierarchical identification quantity.

[0020] The present invention is improved in that the step of obtaining the task node fitness rate is specifically as follows:

[0021] S311: Based on the bottleneck hierarchical identification quantity, filter out nodes that are not marked as abnormal, analyze the correlation between the current change data of node bandwidth and the historical maximum fluctuation, compare the change trend of bandwidth data in each period, determine whether the change of node bandwidth is a continuous dynamic difference, and obtain the bandwidth dynamic response sequence.

[0022] S312: Based on the bandwidth dynamic response sequence, determine the change in the remaining memory of each node in the scheduling interval, analyze the change range of the current remaining memory interval of each node, and compare the node resource allocation capabilities by combining the node hierarchical identifier and the task scheduling priority order to obtain the task resource adaptation vector group.

[0023] S313: Based on the task resource adaptation vector group, adjust the adaptation logic between scheduling nodes and task levels, analyze the matching relationship between resource status and task requirements when each node allocates tasks, aggregate the adaptation performance of nodes for each task level, and obtain the task node adaptation rate.

[0024] The present invention is improved in that the step of obtaining the periodic active performance quantity is specifically as follows:

[0025] S411: Based on the task node adaptability, analyze the communication status of schedulable nodes in each time window, determine whether the heartbeat response and connection status of each node remain stable, identify nodes with interruptions or reconnections in the communication records, optimize the online communication data collection process, and obtain the communication integrity index.

[0026] S412: Based on the communication integrity index, determine the continuous communication performance of each node within multiple time windows, calculate the statistical changes in the number of task responses and communication intervals of the nodes, identify nodes with fluctuating periodic communication capabilities, and obtain multi-time-window communication continuity characteristics.

[0027] S413: Based on the multi-time-window communication continuity characteristics, calculate the difference between minute-level and hour-level task response frequencies, adjust the online performance of nodes within the day-level time window, optimize the communication performance and task response matching under multiple time windows, and obtain the periodic active performance quantity.

[0028] The present invention is improved in that the step of obtaining the availability distribution structure is specifically as follows:

[0029] S511: Based on the periodic active performance data, analyze the periodic integral data of each node, determine the category of the node under the classification standard, compare the distribution characteristics of the integral data of the node in the same period, filter the node data with priority of belonging, schedulable and restricted allocation, and obtain the node category determination sequence.

[0030] S512: Based on the node category determination sequence, optimize the node grouping mechanism, analyze the node periodic integral distribution trend, determine the difference in the affiliation boundary inside and outside the node group, split the affiliation relationship between nodes of the same type, identify the hierarchical structure of node affiliation, and obtain the node affiliation grouping structure.

[0031] S513: Based on the node affiliation grouping structure, compare the affiliation performance of each type of node among the different groups, aggregate the affiliation distribution patterns among node groups, sort out the hierarchical relationship of priority, schedulable and restricted allocation groups, and obtain the availability distribution structure.

[0032] The present invention is improved in that each node refers to every server node in the edge computing network, which is the basic object of load balancing scheduling; the historical trend refers to the change curve or pattern of server node parameters at multiple sampling time points; and the non-abnormal node refers to the server node that was not identified as a resource bottleneck, abnormal behavior or other special state during the anomaly screening process.

[0033] A server load balancing system based on edge computing, the system comprising:

[0034] The load feature extraction module is based on heterogeneous server nodes in edge computing scenarios. It compares the fluctuations of each node at the same time, analyzes the differences between parameters, filters fluctuating nodes, integrates various parameter interaction information, and obtains a global load feature set.

[0035] Based on the global load feature set, the bottleneck layering identification module determines the historical trend of GPU and CPU utilization, analyzes the scheduling response frequency, filters out synchronization abnormal nodes, calculates the bandwidth change trend, dynamically adjusts the node grouping standard, and obtains the bottleneck layering identification quantity.

[0036] Based on the bottleneck hierarchical identification quantity, the scheduling adaptation analysis module filters out non-abnormal nodes, analyzes the current changes and historical extreme fluctuations of bandwidth, determines the memory change range, adjusts the task level adaptation logic, and obtains the task node adaptation rate.

[0037] The activity assessment module analyzes the continuous online performance of nodes based on the task node adaptability, judges the communication status within minute, hour and day time windows, compares the integrity of periodic communication records, calculates the activity level of nodes in multiple time windows, and obtains the periodic activity performance quantity.

[0038] Based on the periodic active performance data, the availability aggregation module determines the hierarchical affiliation of nodes, analyzes the distribution trend of integral data, optimizes the node grouping method, identifies priority, schedulable and restricted allocation of node groups, compares the affiliation relationship of each group, and obtains the availability distribution structure.

[0039] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0040] In this invention, a multi-parameter synchronous sensing and hierarchical discrimination mechanism is adopted. By mining the coupling rules of load fluctuation, resource coordination and bandwidth traffic dynamic changes between nodes in real time, the hierarchical dynamic aggregation of node availability is realized. The resource grouping results are continuously optimized based on the node task carrying capacity and multi-time window communication activity. The node scheduling and allocation are fully adapted to the network status and heterogeneous computing power scenario. The task flow and node distribution tend to be balanced. The scheduling sensitivity and response continuity are further improved. The overall resource utilization and grouping structure of the cluster are more flexible and can adapt to the high-frequency changes and uneven resource distribution in the edge deployment environment. Attached Figure Description

[0041] Figure 1 This is a flowchart of the main steps of the present invention;

[0042] Figure 2 This is a flowchart illustrating the process of obtaining the global load feature set in this invention.

[0043] Figure 3 This is a flowchart illustrating the process of obtaining the bottleneck layer identifier in this invention.

[0044] Figure 4This is a flowchart illustrating the process of obtaining the task node fitness rate in this invention.

[0045] Figure 5 This is a flowchart illustrating the process of obtaining periodic active performance data in this invention.

[0046] Figure 6 This is a flowchart illustrating the process of obtaining the availability distribution structure in this invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0048] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0049] Example

[0050] Please see Figure 1 This invention provides a technical solution: a server load balancing method based on edge computing, comprising the following steps:

[0051] S1: Based on heterogeneous server nodes in edge computing scenarios, monitor real-time data of server status parameters, compare the synchronous fluctuation of each node at the same time point, determine the dynamic differences between each parameter, filter node data where the fluctuation difference between parameters exceeds the average change range, and integrate the interaction performance between multiple data to obtain a global load feature set.

[0052] S2: Based on the global load feature set, determine the historical trend of GPU utilization and CPU utilization at multiple sampling times, analyze the response frequency in the task scheduling process, screen nodes with abnormal fluctuations in memory remaining and data traffic synchronization, calculate the trend of bandwidth changes in continuous periods, adjust the node grouping criteria, and obtain the bottleneck layer identification quantity.

[0053] S3: Based on the bottleneck hierarchical identification quantity, filter nodes that are not marked as abnormal, analyze the dynamic relationship between the current change of node bandwidth and historical extreme fluctuations, determine the change range of remaining memory within the scheduling interval, optimize the matching mechanism when assigning tasks to nodes, and adjust the adaptation logic between scheduling nodes and task levels to obtain the task node adaptability rate.

[0054] S4: Based on the task node adaptability, analyze the online duration and heartbeat response records of schedulable nodes within a preset time window, determine the communication status of each node within the minute, hour and day time windows, compare whether the periodic communication records are complete, and calculate the node activity performance under multiple time windows to obtain the periodic activity performance quantity.

[0055] S5: Based on the periodic active performance, determine the category of the node in the classification standard, analyze the distribution trend of periodic integral data among the nodes, optimize the node grouping mechanism, identify the node groups with priority of affiliation, schedulable and restricted allocation, and compare the affiliation relationship between the node groups to obtain the availability distribution structure.

[0056] The global load feature set includes running status features, fluctuation coupling features, and interaction distribution features. The bottleneck hierarchical identification includes computing resource identification, storage resource identification, and network resource identification. The task node adaptability includes task matching indicators, resource response indicators, and allocation coordination indicators. The periodic activity performance includes online activity indicators, communication integrity indicators, and time window performance indicators. The availability distribution structure includes priority groups, scheduling groups, and restricted groups.

[0057] In S1, heterogeneous server nodes in edge computing scenarios refer to multiple server units distributed within an edge computing architecture that differ in hardware configuration or system architecture, typically deployed in geographical locations closer to the data source; server status parameters refer to real-time data items that reflect the operating status of each server, typically including GPU utilization, CPU utilization, remaining memory, network bandwidth utilization, and data traffic per unit time; each node refers to every server node in the edge computing network, which is the basic object of load balancing scheduling; synchronous fluctuation refers to the consistency or difference of status parameter changes across all server nodes at the same point in time, aiming to capture the collective changes in real-time load across the entire network; the relationships between parameters The dynamic difference refers to the relative numerical change of various server status parameters (such as CPU and GPU utilization) over time between the same node or different nodes; the average range of change refers to the average value of the change of similar parameters (such as CPU utilization of each node) within the statistical period, which serves as a reference benchmark for judging whether there is abnormal fluctuation; node data refers to the dataset composed of all real-time parameters collected from each server node, which is the basic data for subsequent judgment and analysis; the interaction performance between multiple data refers to the mutual influence characteristics of different parameters (such as bandwidth, CPU, GPU utilization, etc.) at the same time, for example, whether the bandwidth utilization increases synchronously when the CPU utilization increases.

[0058] In S2, historical trend refers to the curves or patterns of changes in server node parameters (such as GPU or CPU utilization) at multiple sampling time points, reflecting the load change pattern over time; task scheduling process refers to the process of allocating or adjusting computing tasks according to node status within the server cluster; response frequency refers to the response speed or number of times a server node responds to scheduling instructions or load changes, used to measure its activity level in task processing; synchronously abnormal nodes refer to server nodes whose remaining memory and data traffic fluctuate synchronously (such as simultaneous sudden drops or increases) within a certain period of time, indicating the existence of load bottlenecks; the trend in continuous period refers to the overall direction and characteristics of changes in parameters (such as bandwidth usage) within several consecutive time slices; node grouping criteria refer to the rules or thresholds for classifying nodes into different categories (such as high availability, schedulable, restricted allocation, etc.) based on indicators such as server load, response capability, and stability.

[0059] In S3, nodes not marked as abnormal refer to server nodes that were not identified as resource bottlenecks, abnormal behaviors, or other special states during the anomaly screening process, and are eligible to continue participating in scheduling. Historical extreme fluctuations refer to the maximum or minimum changes in node parameters (such as bandwidth) in historical sampling records, often used to determine the stability of the node. Dynamic relationships refer to the comparison between the current state of a node's bandwidth and its historical extreme fluctuations, used to reflect the dynamic applicability of its bandwidth resources. The change range within the scheduling interval refers to the actual change in the amount of memory remaining during task scheduling and allocation. The matching mechanism refers to the allocation rules established based on the node state parameters and task requirements, realizing the dynamic matching of task and node capabilities. The scheduling node refers to the server node that actually undertakes and executes the scheduling and allocation tasks. The adaptation logic refers to the judgment logic that optimizes and adjusts the task distribution and node allocation relationship based on the current state and resource constraints during the task and node allocation process.

[0060] In S4, a schedulable node refers to a server node that still meets the conditions for participating in task allocation after load balancing and resource filtering; the communication status within a time window refers to the communication results such as heartbeat packets and connection status between the server node and the management system within a window with time units such as minutes, hours, and days; periodic communication records refer to the continuous recording of all communication events of a node within a set period (minutes, hours, days), including online, offline, and reconnection statuses; node activity performance refers to the overall online status and task processing count of a node under multiple time windows, which is often used to reflect the stability and availability of a node.

[0061] In S5, the attribution category in the grading criteria refers to the specific category to which each node should be assigned based on the system's preset node classification rules (such as priority, schedulable, and restricted allocation). Periodic integral data refers to the numerical value or score reflecting node stability accumulated through weighting and statistical methods based on communication and online performance within each time window. Distribution trend refers to the distribution pattern of periodic integral data among all nodes, used to evaluate the overall availability structure. Grouping mechanism refers to the rules and procedures for implementing specific grouping of nodes based on indicators such as periodic integrals. Priority, schedulable, and restricted allocation refer to three different usage strategies for nodes: priority task assignment, general schedulable assignment, or limited task allocation. Attribution relationship refers to the hierarchical, subordinate, and cross-relationships between different node groups, which is an important basis for load balancing decisions.

[0062] Please see Figure 2 The specific steps for obtaining the global load feature set are as follows:

[0063] S111: Based on heterogeneous server nodes in edge computing scenarios, analyze the GPU utilization, CPU utilization, remaining memory, network bandwidth utilization and data traffic parameters of each node, compare the changes of each parameter of each node at the same time point, determine whether there are differences in the synchronous change trend, identify node data with change amplitude exceeding the average range, and obtain abnormal parameter fluctuation data.

[0064] A node identification mechanism is used to identify various server devices deployed on the edge side, and each node is assigned a unique number. Under a time synchronization mechanism, the GPU utilization, CPU utilization, remaining memory, network bandwidth utilization, and data traffic of all nodes are periodically collected every 10 seconds. The collected data records the actual value of each parameter at the current moment. At 10:00:00, server node A has a GPU utilization of 92%, a CPU utilization of 85%, remaining memory of 3.1GB, bandwidth utilization of 76%, and data traffic of 120MB / s. After collecting the parameters of all nodes at the same time, each parameter is summarized to form a data set of the same parameters for the entire network. The parameter value of each node at the current moment is compared with the network average value of that parameter. The deviation of each node from the parameter is calculated, and it is determined whether the deviation exceeds the baseline variation range. For example, the average bandwidth utilization rate is 58%, and the preset average variation reference range is set to fluctuate by 15%, that is, the critical value range is 49.3% to 66.7%. If the bandwidth utilization rate of a certain node is 76%, it is determined that the bandwidth utilization rate of that node exceeds the average range, and it is determined that the parameter has abnormal fluctuation. The node and the corresponding abnormal parameter are recorded as abnormal data items. At the same time, the judgment logic is executed for each parameter. All node parameters whose fluctuation range exceeds the reference range are combined to form parameter fluctuation abnormal data, and organized and stored according to the node number.

[0065] S112: Based on abnormal parameter fluctuation data, compare the changing directions of multiple parameters of the same node in a continuous time series, analyze the synergy of the changes of similar parameters between nodes, determine the data segments with multiple parameters synchronously offset, and combine the associated parameters into a group to obtain the fluctuation difference judgment parameters;

[0066] Extract the changing trends of multiple parameters for each node within a continuous time series. Set the analysis window to 60 seconds and record 5 sets of parameter data every 10 seconds within this time range. Determine the rising or falling state of the parameters by the changing direction of adjacent sampled values. Record the changing direction of different parameters for the same node separately, and establish the changing trajectory of each parameter for each time period. For example, the GPU utilization of node A shows a sequential increasing trend at 5 time points, the CPU utilization increases and then declines after the first 3 increases, the remaining memory continuously decreases, the data traffic fluctuates and increases, and the bandwidth remains stable. By comparing the changing direction of the parameters at each time point, determine whether there are multiple parameters. If upward or downward fluctuations occur simultaneously at the same time point, and at least three parameters show changes in the same direction at adjacent sampling times, it is determined that there are data segments with multiple parameters synchronously offset. Further statistical analysis is conducted on the consistency of the change direction of parameters across multiple nodes. For example, it is detected whether multiple nodes exhibit a coordinated behavior of increased GPU utilization, decreased memory, and increased data traffic within the same time period. If this change pattern is found to occur repeatedly across multiple nodes, the parameters are combined into a coordinated parameter group. This coordinated group is recorded as a fluctuation difference judgment parameter, reflecting the abnormal behavior pattern of multiple parameters coordinated by the node within that time period.

[0067] S113: Based on the fluctuation difference judgment parameters, calculate the data signals of GPU utilization, CPU utilization, memory remaining, network bandwidth utilization and data traffic of each node, analyze the coordinated change pattern of multiple state parameters at the same time point, screen the combination of parameter changes with linkage, and obtain the global load feature set.

[0068] Further, all state parameter data of the corresponding nodes within the time period are extracted and assembled into a state parameter set. Time-series data of GPU utilization, CPU utilization, remaining memory, bandwidth utilization, and data traffic are read separately. The numerical changes of the parameters at the same time point are compared and analyzed to determine whether there is a fixed coordinated change pattern among the parameters. For example, if the GPU, CPU, and bandwidth of node B increase simultaneously at 10:10:00, while memory decreases and data traffic increases, and similar change combinations appear in nodes C and D at the same time period, it indicates that this change combination has linkage among multiple nodes. Such linkage combinations are identified, and parameter combinations with stable repeatability across nodes are selected. The judgment criterion is that the combination changes simultaneously in the same direction 3 times or more in at least 3 nodes. Based on this repeatability pattern, a set of coordinated change parameter pairs is constructed. All parameter combinations that meet the conditions are output as a global load feature set, indicating that the resource load status of the current system has identifiable synchronous fluctuation characteristics among multiple nodes, providing a basis for subsequent node status determination and task allocation.

[0069] Please see Figure 3The specific steps for obtaining the bottleneck layer identifier are as follows:

[0070] S211: Based on the global load feature set, analyze the GPU utilization and CPU utilization of each node, compare the time-series change patterns of utilization data at multiple sampling times, determine the consistency of the change trends of similar parameters among nodes, identify utilization data that show a unidirectional change direction within a preset time window, and obtain the utilization time-series trend sequence.

[0071] The raw data sequences of GPU and CPU utilization for each server node at multiple consecutive sampling times are extracted. The sampling period is set to 10 seconds, and the sampling duration to 10 minutes, totaling 60 time points. Time-series data is constructed for the GPU and CPU utilization of each node in chronological order, recording the specific values ​​at each time point. For example, the GPU utilization of node A gradually increases from 65% to 89% within 10 minutes, while the CPU utilization changes from 72% to 91%. Then, for the GPU utilization sequence of the same node, a point-by-point directional judgment is performed to determine whether the data change direction is consistent between adjacent time points. If an upward trend occurs more than three times consecutively, it is recorded as an upward segment; otherwise, it is recorded as a downward segment or a fluctuating segment. The same judgment logic is applied to CPU utilization. Finally, the GPU and CPU utilization of each node are compared at their respective... Whether the trend of change direction in the time series is consistent: If the GPU utilization of node B continues to rise within 10 minutes and the CPU utilization also shows a synchronous rise, and there is no fluctuation or reversal trend in the middle, it is determined that the change trend between the GPU and CPU of this node is consistent. At the same time, by comparing the trend types of all nodes, it is statistically determined whether there are similar trends among multiple nodes in the same time period. For example, if at least 10 nodes show a stable upward trend in GPU, and 8 of them show a similar upward trend in CPU, it is determined that there is global trend consistency in this time period. Further, nodes that meet the trend consistency are selected from the original time series, and the start and end time points of the trend segment are marked. All utilization change sequences that show a continuous upward, downward or stable fluctuation trend are compiled and summarized into a utilization time series trend sequence.

[0072] S212: Based on the utilization time-series trend sequence, determine the task scheduling records of each node, analyze the time-series relationship between the number of task receptions and the status response of each node in the scheduling cycle, identify the node parameters whose change rates of remaining memory and data flow both exceed the preset change threshold in the same scheduling cycle, and aggregate them into the data combination of the corresponding node to obtain the set of synchronization abnormal parameters.

[0073] First, retrieve the task reception records for each node from the task scheduling log. Record the time and number of times each node was scheduled and assigned tasks within the time period corresponding to the trend sequence. For example, if node C was scheduled 3 times within 10 minutes, receiving tasks at the 2nd, 4th, and 8th minutes respectively, and based on the node status log associated with each scheduling record, extract the remaining memory and data traffic data for that node in the 30 seconds before, during, and after scheduling. Construct memory change sequences and traffic change sequences respectively, and perform trend judgment. If a node's remaining memory was 2.8GB before scheduling, decreased to 2.1GB during scheduling, and recovered to 2.6GB after scheduling, and the corresponding data traffic was 90MB / s and 150MB / s in the corresponding time periods, then... If the memory usage is 100MB / s, the memory usage decreases and the traffic increases, respectively. Then, by comparing whether the changes in the two parameters are synchronized, if both show a decreasing or increasing trend within 30 seconds after the scheduling starts, and the fluctuation amplitude exceeds the preset baseline range (set to ±20% of the average historical change amplitude), it is determined that the memory and traffic changes of the node are synchronized in this scheduling cycle. The scheduling cycle data that meets this condition is recorded as an abnormal cycle. The GPU utilization, CPU utilization, remaining memory, and data traffic of the node in the abnormal cycle are aggregated into a data set to form a set of synchronization abnormal parameters for multiple nodes, which is used to reflect the characteristics of nodes that exhibit load synchronization abnormality during the scheduling process.

[0074] S213: Based on the set of synchronization anomaly parameters, calculate the network bandwidth change sequence of each node, analyze the directional change of bandwidth change trend within the continuous scheduling period, determine the node group whose standard deviation of bandwidth change exceeds the preset stability threshold, adjust the node grouping standard, and obtain the bottleneck hierarchical identification quantity.

[0075] Network bandwidth sampling data of each node in the set is extracted to construct a node bandwidth change sequence. The direction and fluctuation amplitude of bandwidth utilization change of each node in continuous scheduling cycle are statistically analyzed, and the change amplitude of adjacent sampling points is calculated. For example, in the first scheduling, the bandwidth of node D increases from 48% to 71%, decreases to 53% in the second scheduling, and increases to 69% in the third scheduling. In continuous cycles, the bandwidth direction shows alternating increases and decreases, and the fluctuation amplitude is greater than 15%, indicating that its bandwidth change is unstable. Then, it is compared with other nodes to identify the set of nodes whose bandwidth direction changes frequently or fluctuates drastically in multiple continuous scheduling cycles. Based on the abnormal frequency and fluctuation intensity of each node in the set in multiple resource dimensions, hierarchical statistics are performed. Each node is scored according to the total number of fluctuations, the average fluctuation amplitude, and the number of abnormal overlaps in resource dimensions. The bottleneck hierarchical reference threshold is set to 15% above the average score of all nodes. When the score of a node exceeds the threshold, the node is classified as a bottleneck node. Otherwise, it enters the normal or low priority group. Based on the linkage abnormal performance and fluctuation frequency of multiple indicators such as bandwidth, memory, and traffic, the node's grouping identifier is updated, and the bottleneck hierarchical identifier of each node is output.

[0076] Please see Figure 4 The specific steps for obtaining the task node fitness rate are as follows:

[0077] S311: Based on the bottleneck hierarchical identification quantity, filter out nodes that are not marked as abnormal, analyze the correlation between the current change data of node bandwidth and the historical maximum fluctuation, compare the change trend of bandwidth data in each period, determine whether the change of node bandwidth shows continuous dynamic difference, and obtain the bandwidth dynamic response sequence.

[0078] The process involves retrieving the node identification data of all nodes that have completed the hierarchical determination, identifying the set of nodes not classified as abnormal, excluding nodes with "resource bottleneck," "communication anomaly," or "scheduling restriction" tags, and retaining only the node numbers under the schedulable and priority categories. Then, for each selected node, the process extracts its network bandwidth utilization sampling data within the current scheduling cycle, while simultaneously retrieving the node's historical bandwidth fluctuation records for the past 10 cycles. The maximum increase, maximum decrease, and average change range of bandwidth within each cycle are calculated, and the peak value of the maximum fluctuation in each historical cycle is used as the historical extreme value benchmark. For example, if node A's maximum increase is 26% and its maximum decrease is 19% in a historical cycle, the maximum historical extreme value is taken as the comparison benchmark. Finally, the bandwidth change sequence within the current scheduling cycle is retrieved and compared with the historical maximum fluctuation value in terms of direction and amplitude. The dual comparison is used. If the current bandwidth change is in the opposite direction to the historical trend or the change exceeds 75% of the historical maximum fluctuation range, the current period bandwidth is marked as unstable. The periodic trend comparison between the current and historical bandwidth changes is continued. That is, it is determined whether the change direction in each period is consistent. If the bandwidth utilization rate of node B increases by 12%, decreases by 15%, and then increases by 9% in three periods, it is determined that its bandwidth change direction has switched. Based on this fluctuation trend and the relationship between the change magnitude and the historical extreme value, it is confirmed whether it belongs to a continuous dynamic difference state. All the bandwidth change records of nodes that meet the condition of continuous change in bandwidth direction and change magnitude reaching more than the set threshold are compiled into a bandwidth dynamic response sequence. This threshold can be set as twice the standard deviation of the node bandwidth utilization rate and set as a difference range of not less than 12% through pre-evaluation.

[0079] S312: Based on the bandwidth dynamic response sequence, determine the change in the remaining memory of each node in the scheduling interval, analyze the change range of the current remaining memory interval of each node, and compare the node resource allocation capabilities by combining the node hierarchical identifier and the task scheduling priority order to obtain the task resource adaptation vector group.

[0080] Based on the bandwidth dynamic response sequence, the memory remaining amount sequence within the scheduling time period matching the bandwidth change is extracted for each node. The scheduling interval is set to 5 minutes per cycle. Six sets of memory remaining amount data are obtained within each cycle. The difference between the maximum and minimum values ​​in this sequence is calculated to determine the memory change amplitude for each node within the current scheduling cycle. For example, if the memory remaining amount of node C decreases from 6.8GB to 5.2GB, the current change amplitude is 1.6GB. This amplitude is then compared with the node's average memory fluctuation amplitude over the past 10 cycles. If the current fluctuation amplitude is greater than 1.5 times the historical average, it is marked as a resource fluctuation amplification state. Subsequently, the current layer identifier of each node is retrieved to determine whether it is in a bottleneck, restricted, or schedulable state, and this is combined with the tasks undertaken by the node in the task scheduling record. The task priority order is determined. For example, if a node is in the "priority group" and the proportion of high-load computing tasks in the task priority record exceeds 60%, it indicates that the node actually participates in high-priority task scheduling more frequently. Then, the bandwidth dynamic response intensity, memory remaining fluctuation intensity, hierarchical identifier, and task priority are combined to construct a four-dimensional data combination. These four parameters are assigned values ​​by setting weight rules. For example, the task priority adaptation value is set to 4, the hierarchical level adaptation value is 3, the memory fluctuation adaptation value is 2, and the bandwidth fluctuation suppression value is 1. After merging and calculating, the resource adaptation intensity of each node in the scheduling period is obtained. The multi-dimensional resource response capabilities of each node in the scheduling period are grouped into a task resource adaptation vector group, which represents the scheduling adaptation capability level of each node in response to the dynamic characteristics of resource load in the current period.

[0081] S313: Based on the task resource adaptation vector group, adjust the adaptation logic between scheduling nodes and task levels, analyze the matching relationship between resource status and task requirements when each node allocates tasks, aggregate the adaptation performance of nodes for each task level, and obtain the task node adaptation rate.

[0082] First, the resource status of each node under multiple task levels is matched and recorded item by item. The changes in bandwidth utilization, remaining memory, and CPU / GPU utilization of the node before and after being assigned tasks at different levels (high, medium, and low) are statistically analyzed. This data is then compared with the resource requirement standards corresponding to the task level. For example, for high-level tasks, bandwidth must be stable above 70%, remaining memory must not be less than 4GB, and CPU utilization must be between 60% and 85%. If node D has a CPU utilization of 91% and remaining memory of 2.9GB after being assigned a high-level task, it is marked as a resource mismatch, and the number of failed matches is recorded as 2. Then, its hierarchical identifier is used to determine whether the node should undertake this type of task. If the node belongs to the "schedulable group" rather than the "priority group," its scheduling adaptation is reduced in the matching logic. The scheduling engine dynamically adjusts the correspondence between task levels and node groups based on matching results. For example, the task allocation strategy for high-level tasks is changed from the original priority group + high-fitness nodes to only selecting nodes with the top 30% of fit vector values ​​in the priority group for scheduling. Then, the matching performance of each node for multiple task levels in a continuous scheduling cycle is classified and statistically analyzed. For example, if node E has parameter mismatch in 2 out of 3 high-level tasks and matches successfully in all 4 medium-level tasks, then the node's fitness rate for medium-level tasks is recorded as 100%, and for high-level tasks as 33%. Finally, based on the performance of each node in the task allocation process of different levels, its fitness rate is summarized and the task node fitness rate is output as a reference for the scheduling engine's priority options in the next round of task distribution.

[0083] Please see Figure 5 The specific steps for obtaining periodic active performance data are as follows:

[0084] S411: Based on the task node adaptability, analyze the communication status of schedulable nodes in each time window, determine whether the heartbeat response and connection status of each node remain stable, identify nodes with interruptions or reconnections in the communication records, optimize the online communication data collection process, and obtain the communication integrity index.

[0085] First, all nodes with fitness rates not lower than the set scheduling baseline are extracted to form a list of schedulable nodes. Based on this, the communication records of each node in different past time windows are retrieved. The time window granularity is set to three levels: minute, hour, and day. The heartbeat response records and connection status logs of each node within the corresponding time window are read. For example, if node A sends 360 heartbeats in one hour, with an expected interval of once every 10 seconds, the actual number of heartbeats is counted and the interruption point is recorded. If there are records of no response for more than 30 consecutive seconds, it is determined as a communication interruption. If the heartbeat recovery time is greater than 5 seconds and the number of state transitions exceeds 2, it is recorded as a reconnection event. Then, the total interruption duration and the number of reconnection events for each node across multiple time windows are summarized. The continuity and stability of each node's communication records are compared to identify nodes with unstable communication during specific time periods. The node numbers and corresponding interruption / reconnection information are archived. Then, an aggregation operation is performed on the full communication logs. The online status, connection interruption frequency, and continuous online duration of each node under each time window are merged and calculated, and organized into standard structured data entries. If a node is found to have a total interruption duration of more than 15 minutes or more than 6 reconnections in a day, it is classified as a node with incomplete communication records. Further, all communication status fields are aggregated and optimized. Information such as connection status, response latency, and connection breakpoints distributed in different log files are uniformly organized into a communication structure record table. Each record accurately corresponds to the node number and time window identifier. The completeness calculation of the aggregated communication logs of all nodes under all time windows is performed. The communication integrity index of each node is output according to the ratio of connection duration to the total expected online duration.

[0086] S412: Based on the communication integrity index, judge the continuous communication performance of each node in multiple time windows, calculate the statistical changes of the number of task responses and communication intervals of the nodes, identify nodes with fluctuating periodic communication capabilities, and obtain multi-time window communication continuity characteristics.

[0087] The continuous communication evaluation period is set to 24 hours per day, divided into hourly time windows. The communication integrity results of each node across all time windows are extracted, and a node communication continuity sequence is constructed chronologically. For each node, the number of time periods with communication integrity above 95% within the 24-hour window is counted. If node B has a communication integrity of 99% for 19 hours within 24 hours and below 80% for the remaining 5 hours, its communication stability is considered to fluctuate. Next, the task response records and communication behavior of each node are matched and analyzed. The start and completion times of each task allocation in the task scheduling log are retrieved, the corresponding node numbers are counted, and the communication status for the 30 seconds before and after scheduling begins is obtained. If a node... If heartbeat packet loss or connection interruption occurs, it is recorded as a communication anomaly during the task. The number of abnormal responses of each node in all scheduling cycles is accumulated and the ratio is calculated with the total number of task responses. If the ratio exceeds 20%, the node is judged as having poor communication response stability. At the same time, the time distribution of node communication intervals is analyzed. Within a continuous time window, the interval between two connections is calculated to see if it remains within a specified threshold range. For example, an interval fluctuation of more than 15 seconds is considered discontinuous communication. All nodes with abnormal intervals are marked and included in the communication fluctuation node list. The connection persistence, communication status during task response, and time interval fluctuation information of each node in multiple time windows are combined as the basis for identifying fluctuations in periodic communication capabilities and organized into multi-time window communication continuity characteristics.

[0088] S413: Based on the continuity characteristics of multi-time-window communication, calculate the difference in task response frequency between minutes and hours, and adjust the online performance of nodes within a day-level time window using the following formula:

[0089] ;

[0090] Optimize communication performance and task response matching across multiple time windows to obtain periodic active performance metrics. ,in, Representative node Task response frequency within a minute-level time window Representative node Task response frequency within an hourly time window Representative node The proportion of online time slots within a daily time window Represents the total number of nodes;

[0091] Periodic activity performance refers to the overall activity level of a node in an edge computing server load balancing scenario, based on the communication and task scheduling response characteristics of the node within different time windows. Its calculation integrates multiple parameters such as the difference in task response frequency at the minute and hour level, and the proportion of online time periods within a day-level time window, reflecting the communication continuity and task response capability of each node over multiple periods. It can be used to quantify the overall online activity of nodes at different time scales, serving as the basis for decisions such as node grouping, scheduling priority determination, and resource allocation.

[0092] Extract the minute-level task response frequency of each node. Hourly task response frequency and the proportion of online time periods within the daily time window The raw data comes from the node response logs recorded in the scheduling system and the communication activity information summarized by the online monitoring unit. The task response frequency is measured by the number of response events per unit time. For example, the raw data for nodes 1 to 5 are as follows:

[0093] ;

[0094] ;

[0095] ;

[0096] right and The maximum-minimum normalization method was used for processing, and the normalized frequency values ​​are as follows:

[0097] ;

[0098] ;

[0099] Since these are proportional parameters, no normalization is needed. Substitute the parameters and execute the formula:

[0100] ;

[0101] The node-by-node calculation is as follows:

[0102] Node 1:

[0103] ;

[0104] ;

[0105] The product is ;

[0106] Node 2:

[0107] ;

[0108] ;

[0109] The product is ;

[0110] Node 3:

[0111] ;

[0112] ;

[0113] The product is ;

[0114] Node 4:

[0115] ;

[0116] ;

[0117] The product is ;

[0118] Node 5:

[0119] ;

[0120] ;

[0121] The product is ;

[0122] Summing the results gives ;

[0123] Substituting into the formula, we get:

[0124] ;

[0125] Periodic active performance The overall interval division is as follows:

[0126] when When the node is in a highly coordinated state, it is determined to be a highly coordinated segment with excellent node status continuity and strong adaptability.

[0127] when At that time, it was determined to be a moderately coordinated stage, with slight differences in scheduling response among nodes;

[0128] when At this point, it is determined to be a weak coordination period, and frequency differences begin to affect the stability of multi-time-window activity;

[0129] when When this occurs, it is determined to be an uncoordinated segment, where the node response exhibits periodic faults or unstable fluctuations.

[0130] result As it belongs to the highly coordinated segment, the generated periodic active performance data can be directly used as an important basis for grouping nodes in subsequent steps. It is used to identify that the task response frequency and online behavior of nodes at different time scales have a high degree of consistency and synchronization, which is manifested as stable task scheduling response, no significant fluctuations in communication links, and stable online status of nodes within the daily scheduling cycle, without short-term frequent fluctuations or response disconnection.

[0131] Please see Figure 6 The specific steps for obtaining the availability distribution structure are as follows:

[0132] S511: Based on the periodic active performance, analyze the periodic integral data of each node, determine the category of the node under the classification standard, compare the distribution characteristics of the integral data of the node in the same period, filter the node data with priority of belonging, schedulable and restricted allocation, and obtain the node category determination sequence.

[0133] The system retrieves raw metrics from each node, including online activity duration, communication integrity, and task response frequency, collected at the minute, hour, and day levels. These metrics are then normalized within their respective time windows, and periodic integral data is calculated. For example, if node A has a total online time of 23 hours and 45 minutes, a heartbeat packet loss rate of 0.8%, and a total of 138 task responses within 24 hours, its integral value is calculated as 93.7 points. A baseline classification rule is then established, with a maximum score of 100. Nodes scoring 80 points or higher are marked as priority groups, those between 60 and 79 points as schedulable groups, and those below 59 points as restricted groups. All nodes are then classified according to this standard and assigned to their respective groups. The system first categorizes nodes and then compares the point data of all nodes horizontally on an hourly basis. It then counts the number and percentage of node points distributed across the three category intervals within each period. If the percentage of restricted group nodes exceeds 40% in a certain period, that period is recorded as a period of poor load stability. The system then analyzes the change trajectory of each node's category in consecutive periods in chronological order to determine whether its affiliation remains stable. For example, if node B enters the priority group three times and the schedulable group once in four consecutive periods, it is determined that the node's periodic affiliation tends to be priority. Finally, the hierarchical affiliation of each node in all periods is recorded in sequence to form a periodic category determination sequence indexed by the node number, which is used to identify its usage priority in each time period.

[0134] S512: Based on the node category determination sequence, optimize the node grouping mechanism, analyze the node periodic integral distribution trend, determine the difference in the affiliation boundary inside and outside the node group, split the affiliation relationship between nodes of the same type, identify the hierarchical structure of node affiliation, and obtain the node affiliation grouping structure.

[0135] First, a grouping status table is constructed for all nodes in the latest time period. Each node is initially grouped into three categories: priority, schedulable, and restricted, based on its classification identifier. Then, a horizontal comparison is made based on the change trajectory of its integral value in past consecutive periods. For example, node C currently has an integral of 78.6 points and is classified into the schedulable group. However, its integrals in the previous two periods were 83.1 points and 85.4 points respectively, showing a downward trend. It is determined to be a boundary node of the priority group and recorded as a boundary offset point. The proportion of all such nodes in each group is counted. If the proportion of boundary offset nodes in a group exceeds 25%, the grouping mechanism is reconstructed. Nodes with similar integral change trends in the current group are split into new subgroups. At the same time, the integral values ​​of nodes with similar values ​​in different groups are compared. For example, if node D has a score of 79.9 and is assigned to the schedulable group, while node E has a score of 80.1 and is assigned to the priority group, the difference is only 0.2 points. In this case, the assignment judgment has boundary ambiguity. Such node pairs are collected into candidate pairs and assigned to a unified group for correction. The correction process is based on the average score over three consecutive periods. For example, if node D's average score over three periods is 81.2 points, it should be upgraded to the priority group. Furthermore, the assignment evolution structure of each node in different time periods is constructed, and it is recorded whether it has jumped from the restricted group to the priority group or from the priority group to the restricted group. Its assignment fluctuation pattern is identified, and a grouping result structure map is formed. The group to which the node belongs in the current period and its assignment hierarchy with surrounding nodes are output to obtain the node assignment grouping structure.

[0136] S513: Based on the node affiliation grouping structure, compare the affiliation performance of each type of node among the different groups, aggregate the affiliation distribution patterns among node groups, sort out the hierarchical relationship of priority, schedulable and restricted allocation groups, and obtain the availability distribution structure.

[0137] First, the resource utilization rate, scheduling response count, and task level matching results of each type of node in the most recent period are statistically analyzed, and a comparison relationship is established between the indicators and their current group. If the resource matching failure rate of a priority group node exceeds 50% in the last three high-priority task scheduling, its actual adaptability is marked as deviating from its group level. Simultaneously, the total number of nodes of each type in its group, actual scheduling frequency, and average score change are aggregated to analyze whether its performance in the current group is better than, equal to, or lower than the group average. For example, if node F is in the restricted group but its performance in the last two periods is better than the average level of the schedulable group, it is recorded as a potential upgrade node. Then, node migration records between all groups are analyzed. The current grouping structure is cross-analyzed to construct a migration path graph from the restricted group to the priority group. The historical affiliation performance of each type of node in other groups and the task response quality score in the corresponding group are marked. The core nodes with continuous priority performance in the priority group, the stable mid-level nodes in the schedulable group, and the edge nodes in the restricted group that have not yet reached the migration threshold are summarized. The nodes are assigned level coefficients according to the average resource response, scheduling participation success rate and integral fluctuation range. Then, the overall distribution of nodes in each group is output as a multi-level availability map. The resource allocation adaptation status and response capability hierarchy of the priority group, schedulable group and restricted group in the current period are clearly organized and output as the availability distribution structure of the current scheduling system.

[0138] A server load balancing system based on edge computing, comprising:

[0139] The load feature extraction module is based on heterogeneous server nodes in edge computing scenarios. It compares the fluctuations of each node at the same time, analyzes the differences between parameters, filters fluctuating nodes, integrates various parameter interaction information, and obtains a global load feature set.

[0140] The bottleneck hierarchical identification module is based on a global load feature set to determine the historical trends of GPU and CPU utilization, analyze scheduling response frequency, screen out synchronization abnormal nodes, calculate bandwidth change trends, dynamically adjust node grouping criteria, and obtain the bottleneck hierarchical identification quantity.

[0141] The scheduling adaptation analysis module filters out non-abnormal nodes based on the bottleneck hierarchical identifier, analyzes the current changes and historical extreme fluctuations in bandwidth, determines the memory change range, adjusts the task level adaptation logic, and obtains the task node adaptation rate.

[0142] The activity assessment module analyzes the continuous online performance of nodes based on the task node adaptability, judges the communication status within minute, hour and day time windows, compares the integrity of periodic communication records, calculates the activity level of nodes in multiple time windows, and obtains the periodic activity performance quantity.

[0143] The availability aggregation module determines the hierarchical affiliation of nodes based on periodic active performance, analyzes the distribution trend of integral data, optimizes the node grouping method, identifies priority, schedulable and restricted allocation of node groups, compares the affiliation relationship of each group, and obtains the availability distribution structure.

[0144] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A server load balancing method based on edge computing, characterized in that, Includes the following steps: S1: Based on heterogeneous server nodes in edge computing scenarios, compare the fluctuations of each node at the same time, analyze the differences between parameters, filter the node data where the fluctuation difference between parameters exceeds the average change range, integrate various parameter interaction information, and obtain a global load feature set. S2: Based on the global load feature set, determine the historical trend of GPU and CPU utilization, analyze the scheduling response frequency, screen nodes with abnormal synchronization of memory remaining and data traffic, calculate the bandwidth change trend, dynamically adjust the node grouping standard, and obtain the bottleneck layer identification quantity. S3: Based on the bottleneck layer identifier, filter out non-abnormal nodes, analyze the current changes and historical extreme fluctuations of bandwidth, determine the memory change range, adjust the task level adaptation logic, and obtain the task node adaptation rate. S4: Based on the task node adaptability, analyze the online duration and heartbeat response records of schedulable nodes within a preset time window, determine the communication status within minute, hour and day time windows, compare the integrity of periodic communication records, calculate the activity level of multi-time-window nodes, and obtain the periodic activity performance quantity. S5: Based on the periodic active performance, determine the node hierarchical affiliation, analyze the distribution trend of periodic integral data, optimize the node grouping method, identify priority, schedulable and restricted allocation node groups, compare the affiliation relationship of each group, and obtain the availability distribution structure.

2. The server load balancing method based on edge computing according to claim 1, characterized in that, The global load feature set includes running status features, fluctuation coupling features, and interaction distribution features. The bottleneck hierarchical identification quantity includes computing resource identification, storage resource identification, and network resource identification. The task node adaptability includes task matching indicators, resource response indicators, and allocation coordination indicators. The periodic activity performance quantity includes online activity indicators, communication integrity indicators, and time window performance indicators. The availability distribution structure includes priority groups, scheduling groups, and restriction groups.

3. The server load balancing method based on edge computing according to claim 1, characterized in that, The specific steps for obtaining the global load feature set are as follows: S111: Based on heterogeneous server nodes in edge computing scenarios, analyze the GPU utilization, CPU utilization, remaining memory, network bandwidth utilization and data traffic parameters of each node, compare the changes of each parameter of each node at the same time point, determine whether there are differences in the synchronous change trend, identify node data with change amplitude exceeding the average range, and obtain abnormal parameter fluctuation data. S112: Based on the abnormal parameter fluctuation data, compare the changing directions of multiple parameters of the same node in the continuous time series, analyze the synergy of the changes of similar parameters between nodes, determine the data segments with multiple parameters synchronously offset, and combine the associated parameters into a group to obtain the fluctuation difference judgment parameters; S113: Based on the fluctuation difference determination parameters, calculate the data signals of GPU utilization, CPU utilization, memory remaining, network bandwidth utilization and data traffic of each node, analyze the coordinated change pattern of multiple state parameters at the same point in time, screen the combinations of parameter changes with linkage, and obtain the global load feature set.

4. The server load balancing method based on edge computing according to claim 1, characterized in that, The specific steps for obtaining the bottleneck hierarchical identifier are as follows: S211: Based on the global load feature set, analyze the GPU utilization and CPU utilization of each node, compare the time-series change patterns of utilization data at multiple sampling times, determine the consistency of the change trends of similar parameters among nodes, identify utilization data that exhibit a unidirectional change direction within a preset time window, and obtain a utilization time-series trend sequence. S212: Based on the utilization time-series trend sequence, determine the task scheduling records of each node, analyze the time-series relationship between the number of task receptions and the status response of each node in the scheduling cycle, identify the node parameters whose change rates of remaining memory and data traffic both exceed the preset change threshold in the same scheduling cycle, and aggregate them into the data combination of the corresponding node to obtain the synchronization abnormal parameter set. S213: Based on the set of synchronization anomaly parameters, calculate the network bandwidth change sequence of each node, analyze the directional change of bandwidth change trend within the continuous scheduling period, determine the node group whose standard deviation of bandwidth change exceeds the preset stability threshold, adjust the node grouping standard, and obtain the bottleneck hierarchical identification quantity.

5. The server load balancing method based on edge computing according to claim 1, characterized in that, The steps for obtaining the fitness rate of the task nodes are as follows: S311: Based on the bottleneck hierarchical identification quantity, filter out nodes that are not marked as abnormal, analyze the correlation between the current change data of node bandwidth and the historical maximum fluctuation, compare the change trend of bandwidth data in each period, determine whether the change of node bandwidth is a continuous dynamic difference, and obtain the bandwidth dynamic response sequence. S312: Based on the bandwidth dynamic response sequence, determine the change in the remaining memory of each node in the scheduling interval, analyze the change range of the current remaining memory interval of each node, and compare the node resource allocation capabilities by combining the node hierarchical identifier and the task scheduling priority order to obtain the task resource adaptation vector group. S313: Based on the task resource adaptation vector group, adjust the adaptation logic between scheduling nodes and task levels, analyze the matching relationship between resource status and task requirements when each node allocates tasks, aggregate the adaptation performance of nodes for each task level, and obtain the task node adaptation rate.

6. The server load balancing method based on edge computing according to claim 1, characterized in that, The specific steps for obtaining the periodic active performance data are as follows: S411: Based on the task node adaptability, analyze the communication status of schedulable nodes in each time window, determine whether the heartbeat response and connection status of each node remain stable, identify nodes with interruptions or reconnections in the communication records, optimize the online communication data collection process, and obtain the communication integrity index. S412: Based on the communication integrity index, determine the continuous communication performance of each node within multiple time windows, calculate the statistical changes in the number of task responses and communication intervals of the nodes, identify nodes with fluctuating periodic communication capabilities, and obtain multi-time-window communication continuity characteristics. S413: Based on the multi-time-window communication continuity characteristics, calculate the difference between minute-level and hour-level task response frequencies, adjust the online performance of nodes within the day-level time window, optimize the communication performance and task response matching under multiple time windows, and obtain the periodic active performance quantity.

7. The server load balancing method based on edge computing according to claim 1, characterized in that, The specific steps for obtaining the availability distribution structure are as follows: S511: Based on the periodic active performance data, analyze the periodic integral data of each node, determine the category of the node under the classification standard, compare the distribution characteristics of the integral data of the node in the same period, filter the node data with priority of belonging, schedulable and restricted allocation, and obtain the node category determination sequence. S512: Based on the node category determination sequence, optimize the node grouping mechanism, analyze the node periodic integral distribution trend, determine the difference in the affiliation boundary inside and outside the node group, split the affiliation relationship between nodes of the same type, identify the hierarchical structure of node affiliation, and obtain the node affiliation grouping structure. S513: Based on the node affiliation grouping structure, compare the affiliation performance of each type of node among the different groups, aggregate the affiliation distribution patterns among node groups, sort out the hierarchical relationship of priority, schedulable and restricted allocation groups, and obtain the availability distribution structure.

8. The server load balancing method based on edge computing according to claim 1, characterized in that, Each node refers to every server node in the edge computing network, which is the basic object of load balancing scheduling. The historical trend refers to the change curve or pattern of server node parameters at multiple sampling time points. The non-abnormal node refers to the server node that was not identified as a resource bottleneck, abnormal behavior or other special state during the anomaly screening process.

9. A server load balancing system based on edge computing, characterized in that, The system is used to implement the server load balancing method based on edge computing as described in any one of claims 1-8, and the system comprises: The load feature extraction module is based on heterogeneous server nodes in edge computing scenarios. It compares the fluctuations of each node at the same time, analyzes the differences between parameters, filters fluctuating nodes, integrates various parameter interaction information, and obtains a global load feature set. Based on the global load feature set, the bottleneck layering identification module determines the historical trend of GPU and CPU utilization, analyzes the scheduling response frequency, filters out synchronization abnormal nodes, calculates the bandwidth change trend, dynamically adjusts the node grouping standard, and obtains the bottleneck layering identification quantity. Based on the bottleneck hierarchical identification quantity, the scheduling adaptation analysis module filters out non-abnormal nodes, analyzes the current changes and historical extreme fluctuations of bandwidth, determines the memory change range, adjusts the task level adaptation logic, and obtains the task node adaptation rate. The activity assessment module analyzes the continuous online performance of nodes based on the task node adaptability, judges the communication status within minute, hour and day time windows, compares the integrity of periodic communication records, calculates the activity level of nodes in multiple time windows, and obtains the periodic activity performance quantity. Based on the periodic active performance data, the availability aggregation module determines the hierarchical affiliation of nodes, analyzes the distribution trend of integral data, optimizes the node grouping method, identifies priority, schedulable and restricted allocation of node groups, compares the affiliation relationship of each group, and obtains the availability distribution structure.

Citation Information

Patent Citations

  • Server resource scheduling system integrating AI and edge computing

    CN119718682A

  • Multi-domain computing resource aggregation method and system based on virtualized user network

    CN120281776A

  • Intelligent equipment remote cooperative control method and system based on Internet of Things

    CN120343067A

  • Visual operation and maintenance data analysis system

    CN120407350A

  • Edge computing server information processing system based on deep reinforcement learning

    CN120492179A