Server load balancing method and system based on edge computing

By analyzing multi-dimensional parameters in edge computing scenarios and dynamically adjusting node grouping and task levels, the problem of resource distribution imbalance under high dynamic load in traditional server load balancing methods is solved, achieving flexible adaptation of node scheduling and efficient utilization of resources.

CN121301008BActive Publication Date: 2026-04-10百信信息技术有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
百信信息技术有限公司
Filing Date
2025-10-14
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Traditional server load balancing methods cannot promptly identify node availability in scenarios with large network fluctuations, heterogeneous nodes, and frequent communication switching, leading to unbalanced resource distribution, frequent idle or overloaded nodes, reduced task distribution continuity, and fixed node groups, making it difficult to adapt to the needs of edge scenarios with high fluctuations and high dynamic loads.

Method used

The server load balancing method based on edge computing analyzes multi-dimensional parameters of heterogeneous server nodes, dynamically adjusts node grouping criteria, optimizes task level adaptation logic, identifies priority, schedulable and restricted allocation of node groups, realizes hierarchical dynamic aggregation of node availability, and continuously optimizes resource grouping results based on node task carrying capacity and multi-time window communication activity.

Benefits of technology

It achieves node scheduling and allocation that fully adapts to network conditions, makes task flow and node distribution more balanced, improves scheduling sensitivity and response continuity, and makes resource utilization and grouping structure more flexible, adapting to high-frequency changes and uneven resource distribution in edge deployment environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301008B_ABST
    Figure CN121301008B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of servers, in particular to a server load balancing method and system based on edge computing, which comprises the following steps: monitoring state data of heterogeneous server nodes based on an edge computing scenario, identifying parameter fluctuation abnormal nodes, analyzing task scheduling and resource change trends, screening and grouping nodes, evaluating node continuous online and communication active performance, classifying nodes according to grading standards, and obtaining an availability distribution structure. The application realizes hierarchical dynamic collection of node availability by real-time mining of the coupling rules of load fluctuation, resource cooperation and bandwidth flow dynamic change among nodes, the resource grouping result is continuously optimized according to the node task bearing capacity and multi-time window communication activity, the node scheduling and distribution fully adapt to the network state and heterogeneous computing power scene, the task flow and node distribution tend to be balanced, the resource utilization and grouping structure of the cluster as a whole are more flexible, and the problems of high-frequency change and resource unevenness under the edge deployment environment can be adapted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of servers, in particular to a server load balancing method and system based on edge computing. BACKGROUND

[0002] The field of servers involves the management and scheduling of computing resources, including the deployment of servers, task allocation, data processing, and the cooperative work of multiple servers. This technical field covers server hardware architecture, software resource management, network communication, task distribution, fault tolerance mechanisms, and load balancing. Traditional server load balancing refers to the distribution of user requests among a server cluster through hardware balancing devices or software algorithms to achieve relative balance of task loads between servers. Common methods include round-robin scheduling, least connection allocation, allocation based on server response time, and scheduling strategies based on session preservation.

[0003] Traditional solutions lack multi-dimensional parameter linkage and dynamic grading, and the scheduling method relies on a single load or static indicators. In scenarios with large network fluctuations, heterogeneous nodes, and frequent communication switching, it is difficult to identify node availability in a timely manner, resulting in unbalanced resource distribution, frequent idle or overload of some nodes, reduced task distribution continuity, and fixed node grouping, which leads to some nodes being unable to participate in scheduling for a long time. The overall adaptability of the server group to complex environments is limited, making it difficult to meet the needs of actual high-volatility and high-dynamic-load edge scenarios. SUMMARY

[0004] The purpose of the present application is to solve the shortcomings in the prior art and to provide a server load balancing method and system based on edge computing.

[0005] To achieve the above purpose, the present application adopts the following technical solution: a server load balancing method based on edge computing, comprising the following steps:

[0006] S1: Based on the heterogeneous server nodes in the edge computing scenario, compare the fluctuations of each node at the same time, analyze the differences between parameters, filter node data with fluctuation differences exceeding the average change range, integrate various parameter interaction information, and obtain a global load feature set;

[0007] S2: Based on the global load feature set, judge the historical trend of GPU and CPU utilization, analyze the scheduling response frequency, filter nodes with simultaneous abnormalities in memory remaining capacity and data traffic, calculate the bandwidth change trend, dynamically adjust the node grouping standard, and obtain a bottleneck layer identification quantity;

[0008] S3: filtering non-abnormal nodes based on the bottleneck hierarchical identification quantity, analyzing current bandwidth change and historical extreme fluctuation, judging memory change interval, adjusting task level adaptation logic, and obtaining task node adaptation rate;

[0009] S4: based on the task node adaptation rate, analyzing the online duration of the schedulable node within the preset time window and the heartbeat packet response record, judging the communication state within the minute, hour and day level time window, comparing the periodic communication record integrity, calculating the multi-time window node activity level, and obtaining the periodic activity performance quantity;

[0010] S5: based on the periodic activity performance quantity, judging node hierarchical attribution, analyzing periodic integral data distribution trend, optimizing node grouping method, identifying priority, schedulable and restricted allocation node groups, comparing the attribution relationship of each group, and obtaining the availability distribution structure.

[0011] The application improves that the global load feature set includes running state features, fluctuation coupling features and interaction distribution features, the bottleneck hierarchical identification quantity includes computing resource identification, storage resource identification and network resource identification, the task node adaptation rate includes task matching indicators, resource response indicators and allocation coordination indicators, the periodic activity performance quantity includes online activity indicators, communication integrity indicators and time window performance indicators, and the availability distribution structure includes priority groups, scheduling groups and restricted groups.

[0012] The application improves that the acquisition step of the global load feature set is specifically:

[0013] S111: based on the heterogeneous server nodes in the edge computing scene, analyzing the GPU utilization, CPU utilization, memory remaining amount, network bandwidth utilization and data flow parameters of each node, comparing the parameter changes of each node at the same time point, judging whether there is a difference in the synchronous change trend, identifying node data with a change amplitude exceeding the average interval, and obtaining parameter fluctuation abnormal data;

[0014] S112: based on the parameter fluctuation abnormal data, comparing the change direction of multiple parameters of the same node in a continuous time sequence, analyzing the cooperativity of the same type of parameter changes between nodes, judging the data segment with multiple parameters synchronously deviated, and combining the associated parameters into a group, and obtaining the fluctuation difference determination parameter;

[0015] S113: based on the fluctuation difference determination parameter, calculating the data signals of the GPU utilization, CPU utilization, memory remaining amount, network bandwidth utilization and data flow of each node, analyzing the cooperative change mode of multiple state parameters at the same time point, screening the combination with linkage between parameter changes, and obtaining the global load feature set.

[0016] The application improves that the bottleneck hierarchical identification quantity acquisition step is specifically:

[0017] S211: Based on the global load feature set, analyze the GPU utilization and CPU utilization of each node, compare the time sequence change rule of the utilization data at multiple sampling moments, judge the consistency of the change trend of the same parameters between nodes, identify the utilization data showing a one-way change direction in a preset time window, and obtain a utilization time sequence trend sequence;

[0018] S212: Based on the utilization time sequence trend sequence, judge the task scheduling record of each node, analyze the time sequence relationship between the task receiving frequency of each node and the scheduling period state response, identify the node parameters whose memory remaining amount and data flow change rate both exceed the preset change threshold in the same scheduling period, and aggregate them into the data combination of the corresponding node, to obtain a synchronous abnormal parameter set;

[0019] S213: Based on the synchronous abnormal parameter set, calculate the network bandwidth change sequence of each node, analyze the direction change of the bandwidth change trend in the continuous scheduling period, judge the node group whose bandwidth change standard deviation exceeds the preset stability threshold, adjust the grouping standard of the node, and obtain the bottleneck hierarchical identification quantity.

[0020] The application improves that the task node adaptation rate acquisition step is specifically:

[0021] S311: Based on the bottleneck hierarchical identification quantity, filter the nodes not marked as abnormal, analyze the relevance of the current change data and the historical maximum fluctuation of the node bandwidth, compare the change trend of the bandwidth data in each period, judge whether the node bandwidth change shows continuous dynamic difference, obtain a bandwidth dynamic response sequence;

[0022] S312: According to the bandwidth dynamic response sequence, judge the memory remaining amount change of each node in the scheduling interval, analyze the change amplitude of the current memory remaining interval of each node, combine the node hierarchical identification and the task scheduling priority, compare the ability performance of the node resource allocation, and obtain a task resource adaptation vector group;

[0023] S313: Based on the task resource adaptation vector group, adjust the adaptation logic between the scheduling node and the task level, analyze the matching relationship between the resource state and the task demand when each node allocates tasks, aggregate the adaptation performance of the node for each task level, and obtain a task node adaptation rate.

[0024] The application improves that the periodic active performance quantity acquisition step is specifically:

[0025] S411: Based on the task node adaptability, analyze the communication state of the schedulable node in each time window, judge whether the node heartbeat packet response and connection state remain stable, identify the nodes with interruption or reconnection in the communication record, optimize the collection process of online communication data, and obtain the communication integrity index;

[0026] S412: Based on the communication integrity index, judge the continuous communication performance of each node in multiple time windows, calculate the statistical change of the task response times and communication intervals of the node, identify the nodes with periodic communication ability fluctuation, and obtain the multi-time window communication continuity feature;

[0027] S413: Based on the multi-time window communication continuity feature, calculate the difference between the minute-level and hour-level task response frequency, adjust the online performance of the node in the day-level time window, optimize the communication performance and task response matching in multiple time windows, and obtain the periodic active performance quantity.

[0028] The application improves that the acquisition step of the availability distribution structure is specifically:

[0029] S511: Based on the periodic active performance quantity, analyze the periodic integral data of each node, judge the belonging category of the node under the grading standard, compare the distribution characteristics of the integral data of the node in the same period, filter the node data of priority belonging, schedulable and limited allocation, and obtain the node category judgment sequence;

[0030] S512: Based on the node category judgment sequence, optimize the node grouping division mechanism, analyze the node periodic integral distribution trend, judge the belonging boundary difference between the inside and outside of the node group, split the belonging relationship between the same type of nodes, identify the hierarchical structure of node belonging, and obtain the node belonging grouping structure;

[0031] S513: Based on the node belonging grouping structure, compare the belonging performance of nodes of each category between different groups, aggregate the belonging distribution mode between node groups, sort out the hierarchical relationship of priority, schedulable and limited allocation groups, and obtain the availability distribution structure.

[0032] The application improves that each node is a basic object of load balancing scheduling in the edge computing network, the historical trend refers to the change curve or law of the server node parameters at multiple sampling time points, and the non-abnormal node refers to a server node that is not identified as a resource bottleneck, abnormal behavior or other special state in the abnormality screening process.

[0033] The server load balancing system based on edge computing comprises:

[0034] The load feature extraction module is based on a heterogeneous server node in an edge computing scene, compares fluctuations of each node at the same time, analyzes differences between parameters, screens fluctuating nodes, integrates various parameter interaction information, and obtains a global load feature set;

[0035] The bottleneck layer identification module is based on the global load feature set, judges historical trends of GPU and CPU utilization, analyzes scheduling response frequency, screens synchronous abnormal nodes, calculates bandwidth change trends, dynamically adjusts node grouping standards, and obtains bottleneck layer identification quantities;

[0036] The scheduling adaptation analysis module is based on the bottleneck layer identification quantities, screens non-abnormal nodes, analyzes current changes and historical extreme fluctuations of bandwidth, judges memory change intervals, adjusts task level adaptation logic, and obtains task node adaptation rates;

[0037] The activity evaluation module is based on the task node adaptation rates, analyzes continuous online performance of nodes, judges communication states in minute, hour and day level time windows, compares completeness of periodic communication records, calculates multi-time window node activity levels, and obtains periodic activity performance quantities;

[0038] The availability aggregation module is based on the periodic activity performance quantities, judges node classification attribution, analyzes integral data distribution trends, optimizes node grouping methods, identifies node groups of priority, schedulable and limited allocation, compares attribution relationships of each group, and obtains availability distribution structures.

[0039] Compared with the prior art, the application has the advantages and positive effects that:

[0040] In the application, a multi-parameter synchronous sensing and hierarchical discrimination mechanism is adopted, the coupling rules of load fluctuations between nodes, resource cooperation and dynamic changes of bandwidth flow are mined in real time, hierarchical dynamic aggregation of node availability is realized, resource grouping results are continuously optimized according to node task bearing capacity and multi-time window communication activity, node scheduling and distribution fully adapt to network state and heterogeneous computing power scene, task flow and node distribution tend to be balanced, scheduling sensitivity and response continuity are further improved, resource utilization and grouping structure of the whole cluster are more flexible, and the problems of high frequency change and resource unevenness in the edge deployment environment can be adapted. BRIEF DESCRIPTION OF DRAWINGS

[0041] Figure 1 The main step flowchart of the application is shown in the figure;

[0042] Figure 2 The acquisition flowchart of the global load feature set in the application is shown in the figure;

[0043] Figure 3 The acquisition flowchart of the bottleneck layer identification quantities in the application is shown in the figure;

[0044] Figure 4Flow chart for obtaining task node adaptability in the application;

[0045] Figure 5 Flow chart for obtaining periodic activity performance in the application;

[0046] Figure 6 Flow chart for obtaining availability distribution structure in the application. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.

[0048] In the description of the present application, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, in the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise explicitly and specifically limited.

[0049] EMBODIMENT

[0050] Please refer to Figure 1 The present application provides a technical solution: a server load balancing method based on edge computing, comprising the following steps:

[0051] S1: Based on the heterogeneous server nodes in the edge computing scenario, monitor the real-time data of the server state parameters, compare the synchronous fluctuation of each node at the same time point, judge the dynamic difference between each parameter, filter the node data whose fluctuation difference between parameters exceeds the average change range, and integrate the interaction between multiple data to obtain a global load feature set;

[0052] S2: Based on the global load feature set, judge the historical variation trend of GPU utilization rate and CPU utilization rate at multiple sampling time points, analyze the response frequency in the task scheduling process, filter the nodes whose memory remaining amount fluctuation and data flow appear synchronous abnormality, calculate the trend of bandwidth change in continuous period, adjust the node grouping standard, and obtain the bottleneck layer identification quantity;

[0053] S3: Based on the bottleneck hierarchical identification quantity, screening the nodes not marked as abnormal, analyzing the dynamic relationship between the current change and the historical extreme fluctuation of the node bandwidth, judging the change interval of the memory remaining amount in the scheduling interval, optimizing the matching mechanism of the node in task allocation, adjusting the adaptation logic between the scheduling node and the task level, and obtaining the task node adaptation rate;

[0054] S4: Based on the task node adaptation rate, analyzing the online duration and heartbeat packet response record of the schedulable node within the preset time window, judging the communication state of each node within the minute, hour and day time window, comparing whether the periodic communication record is complete, and calculating the node active performance under multiple time windows, obtaining the periodic active performance quantity;

[0055] S5: Based on the periodic active performance quantity, judging the belonging category of the node in the grading standard, analyzing the distribution trend of the periodic integral data among the nodes, optimizing the grouping mechanism of the nodes, identifying the node groups of belonging priority, schedulable and limited allocation, and comparing the belonging relationship between each node group, obtaining the availability distribution structure.

[0056] The global load feature set includes running state features, fluctuation coupling features, and interaction distribution features. The bottleneck hierarchical identification quantity includes computing resource identification, storage resource identification, and network resource identification. The task node adaptation rate includes task matching indicators, resource response indicators, and allocation coordination indicators. The periodic active performance quantity includes online active indicators, communication completeness indicators, and time window performance indicators. The availability distribution structure includes priority groups, scheduling groups, and limited groups.

[0057] In S1, the heterogeneous server nodes in the edge computing scenario refer to multiple server units with different hardware configurations or system architectures distributed in the edge computing architecture, which are usually deployed in a geographic location closer to the data source; the server state parameters refer to real-time data items that can reflect the running state of each server, typically including GPU utilization, CPU utilization, memory remaining amount, network bandwidth utilization, data traffic per unit time, etc.; each node refers to each server node in the edge computing network, which is the basic object of load balancing scheduling; the synchronous fluctuation refers to the consistency or difference of the state parameter changes of all server nodes at the same time point, aiming to capture the collective changes of the real-time load of the whole network; the dynamic difference between parameters refers to the relative numerical change amplitude of each server state parameter (such as CPU and GPU utilization, etc.) generated by the real-time changes over time in the same node or different nodes; the average change range refers to the average value of the change amplitude of the same type of parameters (such as the CPU utilization of each node) within the statistical period, which is used as a reference benchmark to judge whether it is an abnormal fluctuation; the node data refers to the data set composed of all real-time parameters collected from each server node, which is the basic data for subsequent judgment and analysis; the interaction performance between multiple data refers to the mutual influence characteristics between different parameters (such as bandwidth, CPU, GPU utilization, etc.) within the same time, for example, whether the bandwidth usage rate is synchronously improved when the CPU utilization rate rises.

[0058] In S2, the historical variation trend refers to the change curve or law of the server node parameters (such as GPU or CPU utilization) at multiple sampling time points, reflecting the change mode of the load over time; the task scheduling process refers to the process of distributing or adjusting computing tasks according to the node state within the server cluster; the response frequency refers to the response speed or response times of the server node to the scheduling instructions or load changes, which is used to measure its activity level in task processing; the synchronously abnormal node refers to those server nodes whose memory remaining amount and data traffic parameters have synchronous fluctuations (such as simultaneous sharp drop or sharp rise) within a certain period of time, indicating a load bottleneck; the trend in consecutive periods refers to the overall direction and characteristics of the parameter (such as bandwidth usage) changes within several consecutive time slices; the node grouping standard refers to the rules or thresholds for dividing nodes into different categories (such as high availability, schedulable, limited allocation, etc.) according to server load, response ability, stability, etc.

[0059] In S3, the node not marked as abnormal refers to the server node not identified as a resource bottleneck, abnormal behavior or other special state in the abnormality screening process, which is the object that can continue to participate in scheduling; the historical extreme fluctuation refers to the maximum or minimum change of the node parameter (such as bandwidth) in the historical sampling record, which is often used to determine the stability of the node; the dynamic relationship refers to the comparison relationship between the current state of the bandwidth of a certain node and its historical extreme fluctuation, which is used to reflect the dynamic applicability of its bandwidth resources; the change interval in the scheduling interval refers to the actual change range of the memory remaining amount in the task scheduling and distribution process; the matching mechanism refers to the allocation rule established based on the node state parameter and the task demand characteristics, which realizes the dynamic matching of the task and the node capability; the scheduling node refers to the server node that actually undertakes and executes the scheduling and distribution task; the adaptation logic refers to the judgment logic for optimizing and adjusting the task distribution and node allocation relationship according to the current state and resource constraints in the task and node allocation process.

[0060] In S4, the schedulable node refers to the server node that still meets the conditions for participating in task allocation after load balancing and resource screening; the communication state within the time window refers to the communication results such as heartbeat packet and connection state between the server node and the management system within a time window with units of minutes, hours, days, etc.; the periodic communication record refers to the continuous record of all communication events of the node within a set period (minutes, hours, days), including online, offline, reconnection and other states; the node active performance refers to the comprehensive online state, task processing times and other performances of the node under multiple time windows, which is often used to reflect the stability and availability of the node.

[0061] In S5, the belonging category in the grading standard refers to the specific category to which each node should belong according to the system preset node classification rules (such as priority, schedulable, limited allocation, etc.); the periodic integral data refers to the value or score reflecting the stability of the node accumulated by weighting, statistical and other methods based on the communication and online performance within each time window; the distribution trend refers to the distribution pattern of the periodic integral data among all nodes, which is used to evaluate the overall availability structure; the grouping division mechanism refers to the rules and processes for implementing specific grouping of nodes according to periodic integral and other indicators; the belonging priority, schedulable and limited allocation refer to the three different use strategies of the node being classified by the system as priority to undertake tasks, ordinary schedulable or only limitedly participating in task allocation; the belonging relationship refers to the hierarchical, subordinate, cross and other logical relationships between different node groups, which is an important basis for load balancing decision.

[0062] Please refer to Figure 2 The acquisition step of the global load feature set is specifically:

[0063] S111: Based on the heterogeneous server nodes in the edge computing scenario, analyze the GPU utilization, CPU utilization, memory remaining amount, network bandwidth utilization and data flow parameters of each node, compare the changes of each parameter of each node at the same time point, judge whether there is difference in the synchronous change trend, identify the node data whose change amplitude exceeds the average interval, and obtain the parameter fluctuation abnormal data;

[0064] Through the node identification mechanism, confirm various server devices deployed on the edge side, and assign an independent number to each node. Under the time synchronization mechanism, periodically collect the GPU utilization, CPU utilization, memory remaining amount, network bandwidth utilization and data flow of all nodes, with a sampling period of every 10 seconds. The collected content records the actual value of each parameter at the current time. The GPU utilization of server node A at 10:00:00 is 92%, the CPU utilization is 85%, the memory remaining amount is 3.1GB, the bandwidth utilization is 76%, and the data flow is 120MB / s. After collecting the parameters of all nodes at the same time point, each parameter is summarized to form a data set of the same parameters in the whole network. The parameter value of each node at the current time is compared with the average value of the whole network for that parameter, and the deviation of each node on that parameter is calculated. It is judged whether the deviation value exceeds the reference change interval. For example, the average value of the bandwidth utilization is 58%, and the preset average change reference interval is set to float up and down by 15%, i.e. the critical value interval is 49.3% to 66.7%. If the bandwidth utilization of a node is 76%, it is judged that the bandwidth utilization of the node exceeds the average interval, and it is determined that the parameter has abnormal fluctuation. The node and the corresponding abnormal parameter are recorded as abnormal data items. The same judgment logic is performed for each parameter, and all node parameters whose fluctuation amplitude exceeds the reference interval are combined to form parameter fluctuation abnormal data, which is organized and stored according to the node number.

[0065] S112: Based on the parameter fluctuation abnormal data, compare the change direction of multiple parameters of the same node in the continuous time sequence, analyze the coordination of the same type of parameter changes between nodes, judge the data segment with multiple parameters synchronously deviating, and combine the related parameters into a group to obtain the fluctuation difference judgment parameter;

[0066] Extract the trend of multiple parameter changes of each node in a continuous time series, set the analysis window to 60 seconds, record 5 groups of parameter data every 10 seconds in this time range, judge the rising or falling state of the parameters by the change direction of adjacent sampling values, record the change direction of different parameters of the same node, and establish the change trajectory of each parameter by time period, for example, the GPU utilization of node A shows a trend of increasing in turn in 5 time points, the CPU utilization decreases after increasing for the first 3 times, the memory remaining amount continuously decreases, the data flow fluctuates and increases, and the bandwidth remains stable, by comparing the change direction of the parameters at each time point, it is judged whether multiple parameters simultaneously fluctuate upward or downward at the same time, if at least 3 parameters show the same direction change at adjacent sampling time, it is determined that there is a data segment with multiple parameters synchronous offset, further, the consistency of the change direction of the parameters among multiple nodes is counted, for example, whether the GPU utilization increases, the memory decreases, and the data flow increases are detected in multiple nodes in the same time period, if such change mode repeatedly appears in multiple nodes, the parameters are combined into a linkage parameter group, and the linkage combination is recorded as a fluctuation difference judgment parameter, reflecting the behavior pattern of multiple parameter coordinated abnormality of the node in the time period.

[0067] S113: Based on the fluctuation difference judgment parameter, calculate the data signals of GPU utilization, CPU utilization, memory remaining amount, network bandwidth utilization and data flow of each node, analyze the coordinated change mode of multiple state parameters at the same time point, select the combination with linkage change among the parameters, and obtain the global load feature set;

[0068] Further extract all state parameter data of the corresponding node in the time period, assemble into a state parameter set, read the time series data of GPU utilization, CPU utilization, memory remaining amount, bandwidth utilization and data flow respectively, compare and analyze the numerical change of the parameters at the same time point, judge whether there is a fixed coordinated change mode, for example, at 10:10:00, the GPU, CPU and bandwidth of node B increase at the same time, the memory decreases, and the data flow increases, similar change combination also appears in node C and node D in the same time period, which shows that this change combination has linkage among multiple nodes, identify such linkage combination and select the parameter combination with cross-node stability and repeatability from it, the judgment standard is that the combination changes in the same direction at the same time in at least 3 nodes for 3 times or more, according to this repeatability rule, a coordinated change parameter pair set is constructed, and all parameter combinations meeting the conditions are output as a global load feature set, indicating that there is an identifiable synchronous fluctuation feature between multiple nodes in the current system resource load state, which provides a basis for subsequent node state judgment and task allocation.

[0069] Please refer to Figure 3The step of obtaining the bottleneck hierarchical identification quantity is specifically as follows:

[0070] S211: Based on the global load feature set, analyze the GPU utilization and CPU utilization of each node, compare the time sequence variation law of the utilization data at multiple sampling time points, judge the consistency of the variation trend of the same parameters between nodes, identify the utilization data showing a one-way variation direction within a preset time window, and obtain the utilization time sequence trend sequence;

[0071] Extract the original data sequence of the GPU utilization and CPU utilization of each server node at multiple continuous sampling time points, set the sampling period to 10 seconds, the sampling duration to 10 minutes, and the total number of time points to 60. For the GPU and CPU utilization of each node, construct the time sequence data in the order of the time axis, record the specific value of each time point, for example, the GPU utilization of node A increases from 65% to 89% within 10 minutes, and the CPU utilization changes from 72% to 91%. Then, for the GPU utilization sequence of the same node, perform point-by-point direction judgment to determine whether the data variation direction of each adjacent two time points is consistent, such as an upward trend for more than 3 times in a row, which is recorded as an upward segment, otherwise as a downward segment or a fluctuation segment. The same judgment logic is performed on the CPU utilization. Then, compare whether the variation direction trends of the GPU and CPU of each node in their respective time sequence sequences are consistent. If the GPU utilization of node B continuously increases within 10 minutes and the CPU utilization also shows a synchronous increase, and there is no fluctuation reversal trend in the middle, it is determined that the variation trend between the GPU and CPU of the node is consistent. By comparing the trend types of all nodes, it is determined whether there are similar trends in multiple nodes within the same time period, for example, the GPU of at least 10 nodes shows a stable upward trend, and the CPU of 8 of them also shows a similar upward trend. It is determined that there is a global trend consistency in this time period. Further, filter the nodes that meet the trend consistency from the original time sequence sequence, mark the start and end time points of the trend segment, and arrange and summarize all utilization rate variation sequences that show continuous increase, decrease or stable fluctuation trend into the utilization rate time sequence trend sequence.

[0072] S212: Based on the utilization time sequence trend sequence, judge the task scheduling record of each node, analyze the time sequence relationship between the task receiving times and the scheduling period state response of each node, identify the node parameters whose memory remaining amount and data flow change rate exceed the preset change threshold within the same scheduling period, and aggregate them into the data combination of the corresponding node to obtain the synchronous abnormal parameter set;

[0073] First, the task receiving records of each node are retrieved from the task scheduling log, recording the time points and times at which each node is scheduled and allocated tasks within the time period corresponding to the trend sequence, for example, node C is scheduled 3 times within 10 minutes, starting to receive tasks at the 2nd, 4th and 8th minutes, according to each scheduling record, the node state log is associated, the remaining memory and data flow data of the node within 30 seconds before, during and after the scheduling are extracted, the memory change sequence and the flow change sequence are constructed respectively, the change trend is judged, if the memory remaining amount of a node before scheduling is 2.8GB, it decreases to 2.1GB during scheduling, and it recovers to 2.6GB after scheduling, and the data flow within the corresponding time period is 90MB / s, 150MB / s and 100MB / s respectively, then the memory reduction and flow increase are recorded respectively, and then the change directions of the two parameters are compared, if both of them show downward or upward trend within 30 seconds after the start of scheduling, and the fluctuation amplitude exceeds the preset reference interval, the reference interval is set to ±20% of the average of the historical fluctuation amplitude, then it is judged that the memory and flow change of the node within the scheduling period exist synchronization, the scheduling period data that meets the condition is recorded as an abnormal period, and the GPU utilization, CPU utilization, memory remaining amount and data flow of the node within the abnormal period are aggregated into a data combination, forming a synchronous abnormal parameter set corresponding to multiple nodes, reflecting the node characteristics showing load synchronization abnormal state in the scheduling process.

[0074] S213: Based on the synchronous abnormal parameter set, the network bandwidth change sequence of each node is calculated, the direction change of the bandwidth change trend within the continuous scheduling period is analyzed, the node group whose standard deviation of bandwidth change exceeds the preset stable threshold is judged, the grouping standard of the node is adjusted, and the bottleneck hierarchical identification quantity is obtained;

[0075] The network bandwidth sampling data of each node in the collection is extracted, the node bandwidth change sequence is constructed, the bandwidth usage change direction and fluctuation amplitude of each node in the continuous scheduling period are counted, and the change amplitude of adjacent sampling points is calculated. For example, the bandwidth of node D rises from 48% to 71% in the first scheduling, and drops to 53% in the second scheduling, and rises to 69% in the third scheduling. In the continuous period, the bandwidth direction shows alternating rise and fall, and the fluctuation amplitude is greater than 15%. It is judged that the bandwidth change is unstable, and then compared with other nodes, a node set with frequent direction change or severe fluctuation in bandwidth in multiple continuous scheduling periods is identified. The abnormal frequency and fluctuation intensity of each node in the set in multiple resource dimensions are hierarchically counted. According to the total fluctuation number, average fluctuation amplitude and abnormal coincidence number of resource dimensions of the node, each node is scored. The bottleneck hierarchical reference threshold is set to be 15% higher than the average score of all nodes. When the score of a node exceeds the threshold, the node is divided into a bottleneck node. Otherwise, it enters the ordinary or low priority group. According to the linkage abnormal performance and fluctuation frequency of the node in the bandwidth, memory and traffic multiple indexes, the grouping identifier of the node is updated, and the bottleneck hierarchical identifier quantity of each node is output.

[0076] Please refer to Figure 4 The task node adaptation rate acquisition step is specifically:

[0077] S311: Based on the bottleneck hierarchical identifier quantity, the nodes not marked as abnormal are screened, the correlation between the current change data and the historical maximum fluctuation of the node bandwidth is analyzed, the fluctuation trend of the bandwidth data in each period is compared, whether the node bandwidth change shows continuous dynamic difference is judged, and the bandwidth dynamic response sequence is obtained.

[0078] The node identification data of all completed hierarchical determinations are called, a node set in which is not classified as an abnormal state is identified, nodes with "resource bottleneck", "communication anomaly" or "scheduling limit" identification are excluded, only node numbers in a schedulable and priority category are retained, then for each node screened, network bandwidth usage sampling data of the node in the current scheduling period is extracted, historical bandwidth fluctuation records of the node in the past 10 periods are called, the maximum rising amplitude, the maximum falling amplitude and the average change interval of the bandwidth in each period are counted, the peak value of the maximum fluctuation amplitude in each historical period is calculated as a historical extreme value benchmark, for example, the maximum rise of node A in the historical period is 26%, the maximum fall is 19%, the historical extreme value takes the maximum as the comparison benchmark, the bandwidth change sequence in the current scheduling period is called again, and the direction and amplitude are compared with the historical maximum fluctuation value, if the current bandwidth change shows the opposite direction or the change amount exceeds 75% of the historical maximum fluctuation, the current period bandwidth performance is marked as unstable, the periodic trend comparison of the current and historical bandwidth changes is continued, that is, whether the change direction in each period is consistent is judged, if the bandwidth usage rate of node B increases by 12%, decreases by 15% and then increases by 9% in three periods, it is judged that the bandwidth change direction changes, according to the fluctuation trend and the change amplitude relationship between it and the historical extreme value, whether it belongs to the continuous dynamic difference state is confirmed, the bandwidth change records of all nodes that meet the bandwidth direction continuous change and the change amplitude above the set threshold are arranged into a bandwidth dynamic response sequence, the threshold can be set to twice the standard deviation of the node bandwidth usage rate, and is set to a difference interval of not less than 12% through prior evaluation.

[0079] S312: According to the bandwidth dynamic response sequence, the memory remaining amount of each node in the scheduling interval is judged, the change amplitude of the current memory remaining interval of each node is analyzed, the node hierarchical identification and the task scheduling priority order are combined, the ability performance of node resource allocation is compared, and a task resource adaptation vector group is obtained.

[0080] According to the bandwidth dynamic response sequence, the memory remaining amount sequence in the scheduling time period matching the bandwidth change is extracted node by node, the scheduling interval is set to be a period of every 5 minutes, 6 groups of memory remaining amount data are obtained in each period, the maximum value and the minimum value in the sequence are subjected to difference operation, the memory change amplitude of each node in the current scheduling period is obtained, for example, the memory remaining amount of node C decreases from 6.8 GB to 5.2 GB, then the current change amplitude is 1.6 GB, then the amplitude is compared with the average memory fluctuation amplitude in the historical 10 periods of the node, if the current fluctuation amplitude is greater than 1.5 times of the historical average, then it is marked as resource fluctuation amplification state, then the hierarchical identifier of each node is called, whether it is in the bottleneck, limit or schedulable state is judged, and the grade order of the task undertaken by the node in the task scheduling record is combined, for example, a node is a "priority group", there is a high load calculation task in the task grade record, and the proportion of the high load calculation task is more than 60%, then it is indicated that the node actually participates in the high grade task scheduling frequency is higher, then the bandwidth dynamic response strength, the memory remaining fluctuation strength, the hierarchical identifier and the task grade are combined into a four-dimensional data group, the four types of parameters are valued through setting the weight rule, for example, the task grade adaptation value is set to be 4, the hierarchical grade adaptation value is 3, the memory fluctuation adaptation value is 2, and the bandwidth fluctuation suppression value is 1, after the combination calculation, the resource adaptation strength of each node in the scheduling period is obtained, the multi-dimensional resource response ability of each node in the scheduling period is merged into a task resource adaptation vector group, which represents the scheduling adaptation ability level of each node to the resource load dynamic characteristics in the current period.

[0081] S313: Based on the task resource adaptation vector group, the adaptation logic between the scheduling node and the task grade is adjusted, the matching relationship between the resource state and the task demand when each node allocates tasks is analyzed, the adaptation performance of the node to each task grade is aggregated, and a task node adaptation rate is obtained.

[0082] First, the resource state of each node at multiple task levels is recorded item by item, and the bandwidth utilization, memory remaining amount and CPU / GPU utilization change data of the node before and after the allocation of high, medium and low level tasks are counted. The data is compared with the resource demand standard corresponding to the task level to determine whether the resource is matched. For example, in the high level task standard, the bandwidth needs to be stable at more than 70%, the memory remaining amount should not be less than 4GB, and the CPU utilization needs to be between 60%-85%. If the CPU utilization of node D reaches 91% and the memory remaining amount is 2.9GB after the allocation of high level tasks, it is marked as resource mismatch, and the number of matching failure items is recorded as 2. Then, whether the node should undertake such tasks is determined according to its hierarchical identification. If the node belongs to the "schedulable group" but not the "priority group", its scheduling adaptation level is reduced in the matching logic. According to the matching result, the corresponding relationship between the task level and the node group is dynamically adjusted. For example, the task allocation strategy of high level tasks is adjusted from the original priority group + high adaptation nodes to only selecting nodes with adaptation vector value within the top 30% in the priority group for scheduling. Then, the matching performance of each node to multiple level tasks in the continuous scheduling period is classified and counted. For example, node E has 2 parameter mismatches in 3 high level tasks and all 4 medium level tasks are matched successfully. Therefore, the adaptation rate of node E to medium level tasks is 100% and to high level tasks is 33%. Finally, according to the performance of each node in the process of task allocation at different levels, the adaptation capacity proportion of each node is summarized, and the task node adaptation rate is output as a reference item for the next round of task distribution of the scheduling engine.

[0083] Please refer to Figure 5 The acquisition step of the periodic active performance quantity is specifically:

[0084] S411: Based on the task node adaptation rate, analyze the communication state of the schedulable node in each time window, determine whether the heartbeat packet response and connection state of each node remain stable, identify the nodes with interrupted or reconnected communication records, optimize the collection process of online communication data, and obtain the communication integrity index.

[0085] First, all nodes with an adaptation rate not lower than the set scheduling reference are extracted to form a schedulable node list, on the basis of which the communication records of each node in different time windows in the past are called, and the time window granularity is set to three levels of minute, hour and day. The heartbeat packet response records and connection state logs of each node in the corresponding time window are read, for example, node A sends a total of 360 heartbeat packets in 1 hour, with a pre-interval of 10 seconds. The actual number of times is counted and the breakpoint position is recorded. If there is a continuous non-response record for more than 30 seconds, it is determined that the communication is interrupted. If the heartbeat recovery time is greater than 5 seconds and the state switching times exceed 2 times, it is recorded as a reconnection event. Then the total interruption time and the number of reconnection events of each node in multiple time windows are summarized. Whether the communication records of each node are continuous and stable is compared, the nodes with unstable communication in a specific period are identified, and the node number and corresponding interruption / reconnection information are archived. After that, the aggregation operation is performed on the full communication log. The online state, connection interruption frequency and continuous online time of each node in each time window are merged and calculated, and are arranged into a standard structured data entry. If it is found that the cumulative interruption time of a node in a day exceeds 15 minutes or the number of reconnections exceeds 6 times, it is classified as a node with incomplete communication records. Further, all communication state fields are collected and optimized. The connection state, response delay, connection breakpoint and other information distributed in different log files are uniformly arranged in the communication structure record table. Each record accurately corresponds to the node number and time window identifier. The collection results of all nodes in all time windows are calculated for integrity. According to the ratio of connection duration to total online expected duration, the communication integrity index of each node is output.

[0086] S412: Based on the communication integrity index, the continuous communication performance of each node in multiple time windows is judged, the statistical changes of the number of task responses and the communication interval of the node are calculated, the nodes with periodic communication ability fluctuations are identified, and the multi-time window communication continuity feature is obtained.

[0087] The continuous communication evaluation period is set to 24 hours per day, divided into hourly time windows. The communication integrity results of each node across all time windows are extracted, and a node communication continuity sequence is constructed chronologically. For each node, the number of time periods with communication integrity above 95% within the 24-hour window is counted. If node B has a communication integrity of 99% for 19 hours within 24 hours and below 80% for the remaining 5 hours, its communication stability is considered to fluctuate. Next, the task response records and communication behavior of each node are matched and analyzed. The start and completion times of each task allocation in the task scheduling log are retrieved, the corresponding node numbers are counted, and the communication status for the 30 seconds before and after scheduling begins is obtained. If a node... If heartbeat packet loss or connection interruption occurs, it is recorded as a communication anomaly during the task. The number of abnormal responses of each node in all scheduling cycles is accumulated and the ratio is calculated with the total number of task responses. If the ratio exceeds 20%, the node is judged as having poor communication response stability. At the same time, the time distribution of node communication intervals is analyzed. Within a continuous time window, the interval between two connections is calculated to see if it remains within a specified threshold range. For example, an interval fluctuation of more than 15 seconds is considered discontinuous communication. All nodes with abnormal intervals are marked and included in the communication fluctuation node list. The connection persistence, communication status during task response, and time interval fluctuation information of each node in multiple time windows are combined as the basis for identifying fluctuations in periodic communication capabilities and organized into multi-time window communication continuity characteristics.

[0088] S413: Based on the continuity characteristics of multi-time-window communication, calculate the difference in task response frequency between minutes and hours, and adjust the online performance of nodes within a day-level time window using the following formula:

[0089] ;

[0090] Optimize communication performance and task response matching across multiple time windows to obtain periodic active performance metrics. ,in, Representative node Task response frequency within a minute-level time window Representative node Task response frequency within an hourly time window Representative node The proportion of online time slots within a daily time window Represents the total number of nodes;

[0091] Periodic activity performance quantity refers to the node activity degree characterized by the communication and task scheduling response characteristics of the node in different time windows in the edge computing server load balancing scene. It calculates the differences in minute-level and hour-level task response frequency, the proportion of online period in day-level time window and other parameters, reflecting the communication continuity and task response ability of each node in multiple periods, which can be used to quantify the comprehensive online activity of the node in different time scales, as the basis data for node grouping attribution, scheduling priority determination, resource allocation and other decisions;

[0092] The minute-level task response frequency of each node is extracted , the hour-level task response frequency , and the proportion of online period in day-level time window . The original data comes from the node response log recorded in the scheduling system and the communication activity information collected by the online monitoring unit. The task response frequency is measured by the number of response events in unit time. For example, the original data of node 1 to node 5 is as follows:

[0093] ;

[0094] ;

[0095] ;

[0096] and are processed by the maximum and minimum value normalization method. The normalized frequency values are as follows:

[0097] ;

[0098] ;

[0099] is a proportional parameter and does not need to be normalized. After substituting the parameters, the formula is executed as follows:

[0100] ;

[0101] The following is calculated node by node:

[0102] Node 1:

[0103] ;

[0104] ;

[0105] The product is ;

[0106] Node 2: ​

[0107] ;

[0108] ;

[0109] The product is ;

[0110] Node 3:

[0111] ;

[0112] ;

[0113] The product is ;

[0114] Node 4:

[0115] ;

[0116] ;

[0117] The product is ;

[0118] Node 5:

[0119] ;

[0120] ;

[0121] The product is ;

[0122] Summing the results gives ;

[0123] Substituting into the formula, we get:

[0124] ;

[0125] Periodic active performance The overall interval division is as follows:

[0126] when When the node is in a highly coordinated state, it is determined to be a highly coordinated segment with excellent node status continuity and strong adaptability.

[0127] when At that time, it was determined to be a moderately coordinated stage, with slight differences in scheduling response among nodes;

[0128] when At this point, it is determined to be a weak coordination period, and frequency differences begin to affect the stability of multi-time-window activity;

[0129] when When the result is determined as a loss of coordination section, the node is determined to exist a periodic fault or fluctuation instability.

[0130] Results The result is determined as a highly coordinated section, and the generated periodic activity performance quantity can be directly used as an important basis for determining the grouping of the nodes in the subsequent step, for identifying that the nodes have high consistency and synchronization between the task response frequency and online behavior at different time scales, which is manifested as stable task scheduling response, no significant fluctuations in communication link, and stable online of the nodes within the day-level scheduling period, without short-term frequent fluctuations or response disconnection.

[0131] Referring to Figure 6 The acquisition step of the availability distribution structure is specifically as follows:

[0132] S511: Based on the periodic activity performance quantity, analyze the periodic integral data of each node, determine the belonging category of the node under the hierarchical standard, compare the distribution characteristics of the integral data of the node within the same period, filter the node data belonging to the priority, schedulable and restricted allocation, and obtain the node category determination sequence;

[0133] Call the online active duration, communication integrity, task response frequency and other original indexes of each node obtained in the three dimensions of minute, hour and day, calculate the periodic integral data after normalizing each index in the corresponding time window, for example, node A has a total online time of 23 hours and 45 minutes within 24 hours, a heartbeat packet loss rate of 0.8%, and a total task response frequency of 138 times, then after proportional assignment, the integral value is 93.7 points, then set the reference division rule of the belonging hierarchy, take the full score of 100 points as the upper limit, mark the nodes with 80 points and above as priority grouping, mark the nodes between 60 points and 79 points as schedulable grouping, and mark the nodes below 59 points as restricted grouping, then classify all nodes according to the standard, and record them in the corresponding category, then compare the integral data of all nodes horizontally in hours, and count the number and proportion of nodes in the three category intervals within each period, if it is found that the proportion of the number of nodes in the restricted grouping in a period is higher than 40%, the period is recorded as a period with poor load stability, further analyze the change trajectory of the category of each node in the continuous period, and determine whether the belonging remains stable, for example, node B enters the priority group three times and the schedulable group once in four consecutive periods, then it is determined that the node tends to belong to the priority, finally, record the hierarchical belonging of each node in all periods in order to form a periodic category determination sequence indexed by the node number, which is used to identify the use priority of the node in each period.

[0134] S512: Based on the node category decision sequence, optimize the node grouping mechanism, analyze the node cycle integral distribution trend, judge the difference of the belonging boundary inside and outside the node group, split the belonging relationship between the same type of nodes, identify the hierarchical structure of node belonging, and get the node belonging grouping structure;

[0135] First, build the grouping state table of all nodes in the latest time period, divide each node into three types of nodes: priority, schedulable and limit initial grouping according to its classification identifier, and then compare the integral value change trajectory in the past continuous period, for example, node C current integral is 78.6 points, which is classified into schedulable group, but its integral in the previous two periods is 83.1 points and 85.4 points respectively, showing a downward trend, which is judged as the boundary node of priority group, recorded as boundary offset point, and the proportion of all such nodes in each group is calculated, if the proportion of boundary offset nodes in a group exceeds 25%, the grouping mechanism reconstruction will be started, the nodes with similar integral value change trend in the current group are split into new subgroups, and the belonging of integral value close nodes between different groups is compared, for example, node D integral is 79.9 points, which is classified into schedulable group, and node E integral is 80.1 points, which is classified into priority group, only 0.2 points difference, so the belonging judgment has boundary ambiguity, collect such nodes into candidate pairs, and uniformly judge the belonging correction, the average integral value of three consecutive periods is the main judgment basis in the correction process, for example, the average integral value of node D in three periods is 81.2 points, which should be upgraded to priority group, further build the belonging evolution structure of each node in different time period, and record whether it has ever jumped from limit group to priority group or from priority group to limit group, identify its belonging fluctuation mode, form the grouping result structure diagram, output the node belonging group in the current period and its belonging hierarchical structure with the surrounding nodes, and get the node belonging grouping structure.

[0136] S513: Based on the node belonging grouping structure, compare the belonging performance of nodes in different groups, aggregate the belonging distribution mode between node groups, sort out the hierarchical relationship of priority, schedulable and limit allocation groups, and get the availability distribution structure;

[0137] First, the resource utilization, scheduling response times and task level matching results of each type of node in the last cycle are counted, and the indicators are compared with their current assigned group. If the resource matching failure rate of a priority group node in the last three high priority level task scheduling exceeds 50%, it is marked that its actual adaptability deviates from its group level. At the same time, the total number of each type of node in its assigned group, the actual scheduling frequency, the average integral change and other indicators are aggregated to analyze whether its performance in the current group is better than, equal to or lower than the average level in the group. For example, node F is a restricted group, but its performance in the last two cycles is better than the average level of the schedulable group, so it is recorded as a potential promotion node. Then, cross analysis is performed on the node migration records between all groups and the current group structure to construct a migratable path graph from the restricted group to the priority group, label the historical performance of each type of node in other groups and the task response quality score in the corresponding group, and summarize the core nodes with sustained priority performance in the priority group, stable middle-level nodes in the schedulable group, and edge nodes in the restricted group that have not reached the migration threshold. The average resource response value, scheduling participation success rate and integral fluctuation amplitude of the nodes are assigned with level coefficients respectively. Then, the overall distribution of nodes in each group is output as a multi-level availability map. The resource allocation adaptability and response capability level of the priority group, the schedulable group and the restricted group in the current cycle are clearly sorted and output as the availability distribution structure of the current scheduling system.

[0138] The server load balancing system based on edge computing includes:

[0139] The load feature extraction module is based on heterogeneous server nodes in an edge computing scenario, compares the fluctuations of each node at the same time, analyzes the differences between parameters, filters fluctuating nodes, integrates various parameter interaction information, and obtains a global load feature set;

[0140] The bottleneck layer identification module is based on the global load feature set, judges the historical trends of GPU and CPU utilization, analyzes the scheduling response frequency, filters synchronous abnormal nodes, calculates the bandwidth change trend, dynamically adjusts the node grouping standard, and obtains a bottleneck layer identification quantity;

[0141] The scheduling adaptability analysis module is based on the bottleneck layer identification quantity, filters non-abnormal nodes, analyzes the current change and historical extreme fluctuation of the bandwidth, judges the memory change interval, adjusts the task level adaptation logic, and obtains a task node adaptability rate;

[0142] The activity evaluation module is based on the task node adaptability rate, analyzes the continuous online performance of the node, judges the communication state in the minute, hour and day level time window, compares the completeness of the period communication record, calculates the node activity level in multiple time windows, and obtains a period activity performance quantity;

[0143] The availability aggregation module judges node hierarchical attribution based on periodic active performance, analyzes integral data distribution trend, optimizes node grouping mode, identifies priority, schedulable and limited distribution node groups, compares attribution relationship of each group, and obtains availability distribution structure.

[0144] The above merely describes the preferred embodiments of the present application, but not other forms of the present application, any skilled person in the art can use the disclosed technical content to make changes or modifications into equivalent embodiments applied to other fields, but any simple modification, equivalent change and modification made to the above embodiments without departing from the technical solution content of the present application, according to the technical essence of the present application, still belongs to the protection scope of the technical solution of the present application.

Claims

1. A method for server load balancing based on edge computing, characterized in that, The method comprises the following steps: S1: based on the edge computing scene of the heterogeneous server node, comparing the fluctuations of each node at the same time, analyzing the differences between parameters, screening the node data whose fluctuation difference between parameters exceeds the average change range, integrating various parameter interaction information, and obtaining a global load feature set; S2: based on the global load feature set, judging the historical trend of GPU and CPU utilization, analyzing the scheduling response frequency, screening the nodes whose memory remaining amount and data traffic appear synchronous abnormalities, calculating the bandwidth change trend, dynamically adjusting the node grouping standard, and obtaining a bottleneck layer identification quantity; S3: based on the bottleneck layer identification quantity, screening the non-abnormal nodes, analyzing the current change and historical extreme fluctuation of bandwidth, judging the memory change interval, adjusting the task level adaptation logic, and obtaining a task node adaptation rate; The task node adaptation rate acquisition step is specifically: S311: based on the bottleneck layer identification quantity, screening the nodes not marked as abnormal, analyzing the relevance of the current change data and the historical maximum fluctuation of the node bandwidth, comparing the change trend of the bandwidth data in each cycle, judging whether the node bandwidth change shows continuous dynamic difference, and obtaining a bandwidth dynamic response sequence; S312: according to the bandwidth dynamic response sequence, judging the memory remaining amount change of each node in the scheduling interval, analyzing the change amplitude of the current memory remaining interval of each node, combining the node layer identification and the task scheduling priority, comparing the ability performance of node resource allocation, and obtaining a task resource adaptation vector group; S313: based on the task resource adaptation vector group, adjusting the adaptation logic between the scheduling node and the task level, analyzing the matching relationship between the resource state and the task demand when each node allocates tasks, aggregating the adaptation performance of the node for each task level, and obtaining a task node adaptation rate; S4: based on the task node adaptation rate, analyzing the online duration and heartbeat packet response record of the schedulable node in the preset time window, judging the communication state in the minute, hour and day level time window, comparing the completeness of the cycle communication record, calculating the node activity level in multiple time windows, and obtaining a cycle activity performance quantity; The cycle activity performance quantity acquisition step is specifically: S411: based on the task node adaptation rate, analyzing the communication state of the schedulable node in each time window, judging whether the heartbeat packet response and connection state of each node remain stable, identifying the nodes with interruptions or reconnections in the communication record, optimizing the collection process of online communication data, and obtaining a communication integrity index; S412: based on the communication integrity index, judging the continuous communication performance of each node in multiple time windows, calculating the statistical change of the task response times and communication intervals of the node, identifying the nodes with fluctuations in periodic communication ability, and obtaining a multi-time window communication continuity feature; S413: based on the multi-time window communication continuity feature, calculating the difference between the minute level and hour level task response frequencies, adjusting the online performance of the node in the day level time window, optimizing the communication performance and task response matching in multiple time windows, and obtaining a cycle activity performance quantity; S5: judging node hierarchical attribution based on the periodic active performance quantity, analyzing periodic integral data distribution trend, optimizing node grouping mode, identifying priority, schedulable and limited distribution node groups, comparing each group attribution relationship, and obtaining availability distribution structure; The global load feature set includes running state features, fluctuation coupling features and interaction distribution features, the bottleneck hierarchical identification quantity includes computing resource identification, storage resource identification and network resource identification, the task node adaptation rate includes task matching indicators, resource response indicators and distribution coordination indicators, the periodic active performance quantity includes online active indicators, communication integrity indicators and time window performance indicators, and the availability distribution structure includes priority groups, scheduling groups and limited groups. 2.The edge computing based server load balancing method according to claim 1, characterized in that, The obtaining step of the global load feature set is specifically: S111: based on the heterogeneous server nodes in the edge computing scene, analyzing GPU utilization, CPU utilization, memory remaining amount, network bandwidth utilization and data flow parameters of each node, comparing parameter changes of each node at the same time point, judging whether there is a difference in synchronous change trend, identifying node data with change amplitude exceeding the average interval, and obtaining parameter fluctuation abnormal data; S112: based on the parameter fluctuation abnormal data, comparing the change direction of multiple parameters of the same node in continuous time sequence, analyzing the cooperativity of the same type of parameter changes between nodes, judging the data segment with multiple parameters synchronously offset, and combining the associated parameters into a group, and obtaining fluctuation difference judgment parameters; S113: based on the fluctuation difference judgment parameters, calculating data signals of GPU utilization, CPU utilization, memory remaining amount, network bandwidth utilization and data flow of each node, analyzing the cooperative change mode of multiple state parameters at the same time point, screening combinations with change linkage between parameters, and obtaining a global load feature set. 3.The edge computing based server load balancing method of claim 1, wherein, The obtaining step of the bottleneck hierarchical identification quantity is specifically: S211: based on the global load feature set, analyzing GPU utilization and CPU utilization of each node, comparing time sequence change law of utilization data at multiple sampling times, judging consistency of change trend of the same type of parameters between nodes, identifying utilization data showing one-way change direction within a preset time window, and obtaining utilization time sequence trend sequence; S212: based on the utilization time sequence trend sequence, judging task scheduling records of each node, analyzing time sequence relationship between task receiving times and scheduling period state response of each node, identifying node parameters whose memory remaining amount and data flow change rate exceed a preset change threshold in the same scheduling period, and aggregating them into a data combination of the corresponding node, and obtaining a synchronous abnormal parameter set; S213: based on the synchronous abnormal parameter set, calculating network bandwidth change sequence of each node, analyzing direction variation of bandwidth change trend in continuous scheduling period, judging node group whose bandwidth change standard deviation exceeds a preset stability threshold, adjusting grouping standard of the node, and obtaining a bottleneck hierarchical identification quantity. 4.The edge computing based server load balancing method of claim 1, wherein, The obtaining step of the availability distribution structure is specifically: S511: Based on the periodic active performance quantity, analyze the periodic integral data of each node, judge the belonging category of the node under the grading standard, compare the distribution characteristics of the integral data of the node in the same period, screen the node data of priority, schedulable and limited allocation, and obtain the node category judgment sequence; S512: Based on the node category judgment sequence, optimize the node grouping division mechanism, analyze the distribution trend of the node periodic integral, judge the belonging boundary difference between the inside and outside of the node group, split the belonging relationship between the same type of nodes, identify the hierarchical structure of the node belonging, and obtain the node belonging grouping structure; S513: Based on the node belonging grouping structure, compare the belonging performance of nodes of each category between different groups, aggregate the belonging distribution mode between node groups, sort out the hierarchical relationship of priority, schedulable and limited allocation groups, and obtain the availability distribution structure.

5. The edge-computing-based server load balancing method according to claim 1, wherein, Each node refers to each server node in the edge computing network, which is the basic object of load balancing scheduling, the historical trend refers to the change curve or law of the server node parameters at multiple sampling time points, and the non-abnormal node refers to the server node which is not identified as a resource bottleneck or an abnormal behavior state in the abnormality screening process.

6. A server load balancing system based on edge computing, characterized in that, The system is used to implement the server load balancing method based on edge computing according to any one of claims 1-5, and the system comprises: The load feature extraction module compares the fluctuations of each node at the same time based on the heterogeneous server nodes in the edge computing scene, analyzes the differences between parameters, screens the fluctuation nodes, integrates the interaction information of various parameters, and obtains a global load feature set; The bottleneck hierarchical identification module judges the historical trend of GPU and CPU utilization based on the global load feature set, analyzes the scheduling response frequency, screens the synchronous abnormal nodes, calculates the bandwidth change trend, dynamically adjusts the node grouping standard, and obtains a bottleneck hierarchical identification quantity; The scheduling adaptation analysis module screens the non-abnormal nodes based on the bottleneck hierarchical identification quantity, analyzes the current change and historical extreme fluctuation of the bandwidth, judges the memory change interval, adjusts the task level adaptation logic, and obtains a task node adaptation rate; The activity evaluation module analyzes the continuous online performance of the node based on the task node adaptation rate, judges the communication state in the minute, hour and day level time window, compares the completeness of the periodic communication record, calculates the node activity level in multiple time windows, and obtains a periodic active performance quantity; The availability aggregation module judges the node grading belonging based on the periodic active performance quantity, analyzes the integral data distribution trend, optimizes the node grouping method, identifies the node grouping of priority, schedulable and limited allocation, compares the belonging relationship of each group, and obtains an availability distribution structure; The global load feature set comprises running state features, fluctuation coupling features and interaction distribution features, the bottleneck hierarchical identification quantity comprises computing resource identification, storage resource identification and network resource identification, the task node adaptation rate comprises task matching indicators, resource response indicators and allocation coordination indicators, the periodic active performance quantity comprises online activity indicators, communication completeness indicators and time window performance indicators, and the availability distribution structure comprises priority groups, scheduling groups and limited groups.

Citation Information

Patent Citations

  • Intelligent equipment remote cooperative control method and system based on Internet of Things

    CN120343067A

  • Visual operation and maintenance data analysis system

    CN120407350A

  • Edge computing server information processing system based on deep reinforcement learning

    CN120492179A