Method and System for Constructing a System Performance Optimization Platform Based on Big Data
By analyzing the distributed computing cluster of cloud service systems, identifying and optimizing the logical link and load balancing of virtual machines, and building a performance optimization solution, the problem of inefficient system performance optimization in the existing technology is solved, and more efficient data processing and resource utilization is achieved.
Patent Information
- Application Number
- CN202411510086.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-10-28
AI Technical Summary
The existing system performance optimization platform construction methods are inefficient when processing massive data, especially when reusing computing resources, resulting in reduced system performance optimization efficiency.
By obtaining the distributed computing cluster of the cloud service system, identifying the functions and configuration parameters of the service virtual machine, analyzing the logical links of the interactive virtual machine, calculating the link transmission capacity and load balancing deviations, and combining the cluster operation data and log information, a performance optimization solution is built.
It improves the performance optimization efficiency of big data systems, ensures the rapidity and stability of data processing, and optimizes resource allocation and load balancing.
Smart Images

Figure CN119025289B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and particularly relates to a method and system for constructing a system performance optimization platform based on big data. Background Art
[0002] The performance optimization of a big data system refers to, for a big data processing system, on the premise of ensuring the data processing quality and efficiency, improving the system performance through various technical means and optimization strategies. It can help the system process massive data faster and provide higher concurrency, lower latency, better scalability, and higher throughput.
[0003] The existing method for constructing a system performance optimization platform adopts a batch processing method, which is suitable for scenarios of static data mining of massive data. The mode is to store first and then calculate. The data may be reused or processed repeatedly, increasing the reuse of computing resources and thus reducing the efficiency of system performance optimization. Summary of the Invention
[0004] The present invention provides a method and system for constructing a system performance optimization platform based on big data, and its main purpose is to improve the efficiency of system performance optimization of big data.
[0005] To achieve the above object, a method for constructing a system performance optimization platform based on big data provided by the present invention includes:
[0006] Obtain a cloud service system to be optimized, extract the distributed computing cluster of the cloud service system, identify the service virtual machines in the distributed computing cluster, query the service functions and configuration parameters corresponding to the service virtual machines, analyze the interactive virtual machines in the service virtual machines according to the service functions, and determine the logical links corresponding to the interactive virtual machines;
[0007] Calculate the link transmission capacity corresponding to the logical link according to the configuration parameters, and calculate the cluster transmission capacity corresponding to the distributed computing cluster according to the link transmission capacity;
[0008] Calculate the load balancing deviation between the logical links according to the configuration parameters, perform normalization processing on the load balancing deviation to obtain the target balance degree, collect the cluster operation data corresponding to the distributed computing cluster, and analyze the cluster state corresponding to the distributed computing cluster according to the cluster operation data;
[0009] Schedule the cluster logs corresponding to the distributed computing cluster, calculate the cluster center load corresponding to the distributed computing cluster according to the cluster logs, set the scheduling priority corresponding to the distributed computing cluster in combination with the cluster transmission capacity, the target balance degree, and the cluster center load, and construct a performance optimization plan for the distributed computing cluster according to the scheduling priority and the cluster status.
[0010] Optionally, the analyzing the interactive virtual machines in the service virtual machine according to the service function includes:
[0011] Identify the function information corresponding to the service function, and determine the function modules and module components corresponding to the service virtual machine according to the function information;
[0012] Schedule the module data flow corresponding to the function module, and analyze the data endpoints corresponding to the module data flow;
[0013] Access the component code corresponding to the module component, and analyze the dependency relationship between the component codes;
[0014] Analyze the interactive virtual machines in the service virtual machine according to the data endpoints and the dependency relationship.
[0015] Optionally, the determining the logical link corresponding to the interactive virtual machine includes:
[0016] Obtain the virtual network configuration information corresponding to the interactive virtual machine, and query the communication devices and device configuration information corresponding to the interactive virtual machine;
[0017] Analyze the device topology structure corresponding to the communication device according to the device configuration information;
[0018] Combine the device topology structure and the network configuration information to determine the network topology structure corresponding to the interactive virtual machine;
[0019] Determine the logical link corresponding to the interactive virtual machine according to the network topology structure.
[0020] Optionally, the calculating the link transmission capacity corresponding to the logical link according to the configuration parameters includes:
[0021] Calculate the idle bandwidth corresponding to the logical link according to the configuration parameters;
[0022] Perform path tracing on each link in the logical link to obtain a tracing path;
[0023] Count the number of passing times corresponding to each link in the logical link according to the tracing path;
[0024] Calculate the link transmission capacity corresponding to the logical link according to the number of passages and the idle bandwidth by using the following formula:
[0025]
[0026] Wherein, A represents the link transmission capacity corresponding to the logical link, represents the idle bandwidth of the i-th logical link, represents the number of passages of the i-th logical link, and i represents the serial number corresponding to the logical link
[0027] Optionally, calculating the idle bandwidth corresponding to the logical link according to the configuration parameters includes:
[0028] Extract the network bandwidth corresponding to the interactive virtual machine from the configuration parameters;
[0029] Calculate the link bandwidth corresponding to the logical link according to the network bandwidth;
[0030] Count the number of upstream and downstream virtual machines corresponding to the logical link to obtain the first virtual machine number and the second virtual machine number;
[0031] Calculate the global bandwidth corresponding to the logical link according to the first virtual machine number and the second virtual machine number by using the following formula:
[0032]
[0033] Wherein, D represents the global bandwidth corresponding to the logical link, F1 and F2 respectively represent the first virtual machine number and the second virtual machine number of the F-th logical link, represents the total bandwidth of the a-th distributed computing cluster, a represents the serial number corresponding to the distributed computing cluster, and R represents the number of distributed computing clusters;
[0034] Calculate the idle bandwidth corresponding to the logical link according to the global bandwidth and the link bandwidth.
[0035] Optionally, calculating the cluster transmission capacity corresponding to the distributed computing cluster according to the link transmission capacity includes:
[0036] Count the cluster failures corresponding to the distributed computing cluster, and calculate the cluster failure rate corresponding to the distributed computing cluster according to the cluster failures;
[0037] Allocate the cluster weight corresponding to the distributed computing cluster according to the cluster failure rate;
[0038] Combined with the cluster weight and the link transmission capacity, the cluster transmission capacity corresponding to the distributed computing cluster is calculated through the following formula:
[0039]
[0040] Among them, G represents the cluster transmission capacity corresponding to the distributed computing cluster, represents the cluster weight, a is the serial number of the distributed computing cluster, represents the b-th link transmission capacity, b is the serial number corresponding to the link transmission capacity, and q is the total number corresponding to the link transmission capacity.
[0041] Optionally, calculating the load balancing deviation between the logical links according to the configuration parameters includes:
[0042] Monitoring the link traffic value corresponding to the logical link and recording the monitoring period corresponding to the link traffic value;
[0043] Performing abnormal elimination on the link traffic value to obtain a target traffic value;
[0044] Combined with the target traffic value and the monitoring period, calculating the average traffic corresponding to the logical link;
[0045] Combined with the configuration parameters and the average traffic, calculating the load balancing deviation corresponding to the logical link.
[0046] Optionally, the calculating the load balancing deviation corresponding to the logical link by combining the configuration parameters and the average traffic includes:
[0047] Extracting the bandwidth parameter and the port parameter in the configuration parameters;
[0048] Calculating the traffic extreme value corresponding to the logical link according to the bandwidth parameter and the port parameter;
[0049] Calculating the load balancing deviation corresponding to the logical link according to the traffic extreme value and the average traffic;
[0050] Calculating the coefficient of variation corresponding to the load balancing deviation according to the load balancing deviation;
[0051] Combined with the load balancing deviation and the coefficient of variation, calculating the load balancing deviation corresponding to the logical link.
[0052] Optionally, calculating the cluster center load corresponding to the distributed computing cluster according to the cluster log includes:
[0053] Searching for the cluster center node corresponding to the distributed computing cluster according to the cluster log;
[0054] Analyze the node performance metrics corresponding to the cluster central node by using a preset performance analysis tool;
[0055] Extract the metric log information corresponding to the node performance metrics from the cluster log;
[0056] Calculate the metric load corresponding to the node performance metrics according to the metric log information;
[0057] Calculate the cluster central load corresponding to the distributed computing cluster according to the metric load.
[0058] A system for constructing a system performance optimization platform based on big data, characterized in that the system includes:
[0059] A logical link determination module, configured to obtain a cloud service system to be optimized, extract the distributed computing cluster of the cloud service system, identify service virtual machines in the distributed computing cluster, query the service functions and configuration parameters corresponding to the service virtual machines, and according to the service functions, analyze the interactive virtual machines in the service virtual machines, and determine the logical links corresponding to the interactive virtual machines;
[0060] A transmission capacity calculation module, configured to calculate the link transmission capacity corresponding to the logical link according to the configuration parameters, and calculate the cluster transmission capacity corresponding to the distributed computing cluster according to the link transmission capacity;
[0061] A status analysis module, configured to calculate the load balancing deviation between the logical links according to the configuration parameters, perform normalization processing on the load balancing deviation to obtain a target balance degree, collect the cluster operation data corresponding to the distributed computing cluster, and analyze the cluster status corresponding to the distributed computing cluster according to the cluster operation data;
[0062] A performance optimization module, configured to schedule the cluster log corresponding to the distributed computing cluster, calculate the cluster central load corresponding to the distributed computing cluster according to the cluster log, combine the cluster transmission capacity, the target balance degree and the cluster central load, set the scheduling priority corresponding to the distributed computing cluster, and construct a performance optimization plan for the distributed computing cluster according to the scheduling priority and the cluster status.
[0063] The present invention queries the service functions and configuration parameters corresponding to the service virtual machines, thereby understanding the technical services and corresponding attribute information of the service virtual machines. According to the service functions, the interactive virtual machines in the service virtual machines are analyzed to obtain the virtual machines that interact with each other in the service virtual machines. The present invention calculates the link transmission capacity corresponding to the logical link according to the configuration parameters, and can understand the allocable bandwidth value corresponding to the logical link, thereby facilitating the subsequent calculation and processing of the cluster transmission capacity. The present invention calculates the load balancing deviation between the logical links according to the configuration parameters, and can understand the load balancing degree between the logical links, thereby analyzing the stability between the logical links. The present invention calculates the cluster center load corresponding to the distributed computing cluster according to the cluster log, and can understand the resource usage of the central nodes in the distributed computing cluster, thereby facilitating the evaluation of the overall performance of the distributed computing cluster and providing a basis for setting the subsequent scheduling priorities, and facilitating the improvement of the performance optimization efficiency of the distributed computing cluster. Therefore, a method and system for constructing a system performance optimization platform based on big data provided by the embodiments of the present invention can improve the system performance optimization efficiency of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a schematic flowchart of a method for constructing a system performance optimization platform based on big data provided by an embodiment of the present invention;
[0065] Figure 2 It is a functional block diagram of a system for constructing a system performance optimization platform based on big data provided by an embodiment of the present invention.
[0066] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0068] The embodiments of the present application provide a method for constructing a system performance optimization platform based on big data. In the embodiments of the present application, the execution subject of the method for constructing a system performance optimization platform based on big data includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiments of the present application. In other words, the method for constructing a system performance optimization platform based on big data can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0069] Referring to Figure 1 As shown, it is a schematic flowchart of a method for constructing a system performance optimization platform based on big data provided by an embodiment of the present invention. In this embodiment, the method for constructing a system performance optimization platform based on big data includes steps S1 - S4.
[0070] S1. Obtain the cloud service system to be optimized, extract the distributed computing cluster of the cloud service system, identify the service virtual machines in the distributed computing cluster, query the service functions and configuration parameters corresponding to the service virtual machines, analyze the interactive virtual machines in the service virtual machines according to the service functions, and determine the logical links corresponding to the interactive virtual machines.
[0071] The present invention queries the service functions and configuration parameters corresponding to the service virtual machine, thereby understanding the technical services and corresponding attribute information of the service virtual machine. According to the service functions, the interactive virtual machines in the service virtual machine are analyzed, so as to obtain the virtual machines that interact with each other in the service virtual machine. Among them, the cloud service system is a system that provides services such as computing, storage, network, and applications through cloud computing technology. The distributed computing cluster is a cluster of computing resources composed of multiple servers in the cloud service system. The service virtual machine is a virtual computing environment in a physical computer in the distributed computing cluster. The service functions and the configuration parameters are respectively the technical services and related attribute information corresponding to the service virtual machine. The interactive virtual machine is a virtual machine that interacts with each other in the service virtual machine. Optionally, the service virtual machines in the distributed computing cluster can be obtained by identifying the IP addresses or host names in the distributed computing cluster; the service functions and configuration parameters corresponding to the service virtual machine can be obtained by querying the official documents of the service machine, and the official documents contain the introduction and description of the service functions and configuration parameters.
[0072] As an embodiment of the present invention, the analyzing the interactive virtual machines in the service virtual machine according to the service functions includes: identifying the function information corresponding to the service functions, determining the function modules and module components corresponding to the service virtual machine according to the function information, scheduling the module data flow corresponding to the function modules, analyzing the data endpoints corresponding to the module data flow, accessing the component code corresponding to the module components, analyzing the dependency relationship between the component codes, and analyzing the interactive virtual machines in the service virtual machine according to the data endpoints and the dependency relationship.
[0073] Among them, the function information is the description information in the service functions. The function module is the business logic component of the service virtual machine. It implements specific functions or services and provides external interfaces for other components or systems to call. The module component is the component in the service virtual machine used to manage and organize function modules. It includes functions such as module loading, initialization, runtime management, and communication. The module data flow is the description of data input, processing, and output in the function module. The data endpoint is the data source and data destination corresponding to the module data flow. The component code is the constituent code of the module component, such as the code corresponding to the constituent functions of the module component. The dependency relationship indicates the relationship that the component codes need to depend on the functions, data, or interfaces of other modules when performing tasks.
[0074] Optionally, the function information corresponding to the service function can be obtained through OCR recognition technology; the function modules and module components corresponding to the service virtual machine can be determined according to the information semantics corresponding to the function information; the module data flow corresponding to the function module can be implemented through a scheduling algorithm, such as the highest response ratio first scheduling algorithm; the data endpoints corresponding to the module data flow can be obtained by analyzing the data format and type of the module data flow, and the data format and type obtained after different modules process the data will be different; the component code corresponding to the module component can be obtained by accessing through an IDE tool; the dependency relationship between the component codes can be realized by analyzing the reference relationship between the component codes, such as calling the functions of other components; the interactive virtual machine in the service virtual machine can determine the module where the data source is located and the module where the data is output according to the input and output in the data endpoints, and combine the dependency relationship to analyze the virtual machines with interactions.
[0075] By determining the logical link corresponding to the interactive virtual machine, the data transmission path between the interactive virtual machines can be obtained, providing a basis for the subsequent calculation of the link transmission capacity. Among them, the logical link is the data transmission path between the interactive virtual machines.
[0076] As an embodiment of the present invention, the determination of the logical link corresponding to the interactive virtual machine includes: obtaining the virtual network configuration information corresponding to the interactive virtual machine, querying the communication devices and device configuration information corresponding to the interactive virtual machine, analyzing the device topology structure corresponding to the communication devices according to the device configuration information, combining the device topology structure and the network configuration information to determine the network topology structure corresponding to the interactive virtual machine, and determining the logical link corresponding to the interactive virtual machine according to the network topology structure.
[0077] Among them, the virtual network configuration information is the IP address and subnet information corresponding to the interactive virtual machine, the communication devices are the network devices for communication between the interactive virtual machines, such as virtual interactive machines and physical interactive machines, the device configuration information is the information about the port connection and VLAN configuration corresponding to the communication devices, the device topology structure is the connection method and relationship between the communication devices, and the network topology structure is the network communication connection method and relationship between the interactive virtual machines.
[0078] Optionally, the virtual network configuration information corresponding to the interactive virtual machine can be obtained by using a network management tool, and the network management tool includes an SNMP (Simple Network Management Protocol) management tool; the communication device and device configuration information corresponding to the interactive virtual machine can be queried through a virtualization management platform, such as VMware vSphere; the device topology corresponding to the communication device can be analyzed based on the device configuration information. In the device configuration information, it usually includes information such as the IP address, subnet mask, gateway, and MAC address of each device. These information can help determine the connection method between devices, such as connecting through a router, switch, or direct connection. In addition, the physical connection between devices can be understood by viewing the port information in the device configuration information. For example, in the switch configuration, the connection object of each port will be specified, such as which server, virtual machine, or other switch it is connected to; the network topology corresponding to the interactive virtual machine can be obtained by combining the device topology and the network configuration information. Through the device topology and the network configuration information, the subnet where the interactive virtual machine is located and the network connection method with other virtual machines or physical devices can be determined; the logical link corresponding to the interactive virtual machine can be obtained according to the connection method in the network topology.
[0079] S2. Calculate the link transmission capacity corresponding to the logical link according to the configuration parameters, and calculate the cluster transmission capacity corresponding to the distributed computing cluster according to the link transmission capacity.
[0080] According to the configuration parameters of the present invention, calculating the link transmission capacity corresponding to the logical link can understand the allocable bandwidth value corresponding to the logical link, and further facilitate the calculation and processing of the subsequent cluster transmission capacity. Among them, the link transmission capacity represents the data transmission rate corresponding to the logical link.
[0081] As an embodiment of the present invention, calculating the link transmission capacity corresponding to the logical link according to the configuration parameters includes: calculating the idle bandwidth corresponding to the logical link according to the configuration parameters, performing path tracing on each link in the logical link to obtain a tracing path, and according to the tracing path, counting the passing times corresponding to each link in the logical link. According to the passing times and the idle bandwidth, use the following formula to calculate the link transmission capacity corresponding to the logical link:
[0082]
[0083] Among them, A represents the link transmission capacity corresponding to the logical link, represents the idle bandwidth of the i-th logical link, It represents the number of passes of the i-th logical link, and i represents the serial number corresponding to the logical link.
[0084] Among them, the idle bandwidth represents the remaining transmission volume corresponding to the logical link. The path tracing is the path tracking of the logical link during the processing of the distributed computing cluster, and the number of passes is the number of times each link in the logical link is passed through during the processing of the distributed computing cluster.
[0085] Optionally, the path tracing of each link in the logical link can be implemented through a path tracking algorithm, such as depth-first search; the number of passes corresponding to each link in the logical link can be obtained by statistically analyzing the traced path.
[0086] Further, as an optional embodiment of the present invention, calculating the idle bandwidth corresponding to the logical link according to the configuration parameters includes: extracting the network bandwidth corresponding to the interactive virtual machine from the configuration parameters, calculating the link bandwidth corresponding to the logical link according to the network bandwidth, counting the number of upstream and downstream virtual machines corresponding to the logical link to obtain the first virtual machine number and the second virtual machine number, and calculating the global bandwidth corresponding to the logical link through the following formula according to the first virtual machine number and the second virtual machine number:
[0087]
[0088] Among them, D represents the global bandwidth corresponding to the logical link, and F1 and F2 respectively represent the first virtual machine number and the second virtual machine number of the F-th logical link. represents the total bandwidth corresponding to the a-th distributed computing cluster, a represents the serial number corresponding to the distributed computing cluster, and R represents the number of distributed computing clusters;
[0089] Calculate the idle bandwidth corresponding to the logical link according to the global bandwidth and the link bandwidth.
[0090] Among them, the network bandwidth is the data transmission volume corresponding to the interactive virtual machine, the link bandwidth represents the total transmission volume corresponding to the logical link, the first virtual machine number and the second virtual machine number respectively represent the number of virtual machines placed in the front and back positions corresponding to the logical link, and the global bandwidth represents the minimum guaranteed bandwidth reserved for the distributed computing cluster in advance to ensure the normal operation of tasks.
[0091] Optionally, the link bandwidth corresponding to the logical link can be obtained by summing the network bandwidths corresponding to the logical link; the number of upstream and downstream virtual machines corresponding to the logical link can be obtained by counting according to the corresponding link topology diagram; the idle bandwidth corresponding to the logical link can be obtained by subtracting the global bandwidth from the link bandwidth.
[0092] According to the link transmission capacity, the present invention calculates the cluster transmission capacity corresponding to the distributed computing cluster, and then the allocable bandwidth value corresponding to the distributed computing cluster can be obtained, providing a basis for subsequent performance optimization processing, where the cluster transmission capacity represents the allocable bandwidth value corresponding to the distributed computing cluster.
[0093] As an embodiment of the present invention, calculating the cluster transmission capacity corresponding to the distributed computing cluster according to the link transmission capacity includes: counting the cluster failures corresponding to the distributed computing cluster, calculating the cluster failure rate corresponding to the distributed computing cluster according to the cluster failures, allocating the cluster weight corresponding to the distributed computing cluster according to the cluster failure rate, and combining the cluster weight and the link transmission capacity, and calculating the cluster transmission capacity corresponding to the distributed computing cluster through the following formula:
[0094]
[0095] where G represents the cluster transmission capacity corresponding to the distributed computing cluster, represents the cluster weight, a is the serial number of the distributed computing cluster, represents the b-th link transmission capacity, b is the serial number corresponding to the link transmission capacity, and q is the total number of link transmission capacities.
[0096] Among them, the cluster failures are all failures that occur in the distributed computing cluster, such as node failures, network failures, storage failures, etc., the cluster failure rate represents the probability of failure of the distributed computing cluster within a specific time, and the cluster weight represents the importance assigned to the distributed computing cluster when processing tasks.
[0097] Optionally, the cluster failures corresponding to the distributed computing cluster can be obtained by counting historical failure data; the cluster failure rate corresponding to the distributed computing cluster can be obtained by calculating the proportion of the frequency corresponding to the cluster failures in the total number of failures; the cluster weight corresponding to the distributed computing cluster can be set according to the value of the cluster failure rate. For example, if the failure rate is 10, the cluster weight can be set to 0.9.
[0098] S3. Calculate the load balancing deviation between the logical links according to the configuration parameters, perform normalization processing on the load balancing deviation to obtain the target balance degree, collect the cluster operation data corresponding to the distributed computing cluster, and analyze the cluster state corresponding to the distributed computing cluster according to the cluster operation data.
[0099] According to the configuration parameters of the present invention, the load balancing deviation between the logical links is calculated, so as to understand the load balancing degree between the logical links, and further analyze the stability between the logical links. Among them, the load balancing deviation represents the degree of load distribution deviation between the logical links.
[0100] As an embodiment of the present invention, calculating the load balancing deviation between the logical links according to the configuration parameters includes: monitoring the link traffic value corresponding to the logical link, recording the monitoring period corresponding to the link traffic value, removing anomalies from the link traffic value to obtain the target traffic value, combining the target traffic value and the monitoring period to calculate the average traffic corresponding to the logical link, and combining the configuration parameters and the average traffic to calculate the load balancing deviation corresponding to the logical link.
[0101] Among them, the link traffic value is the data transmission volume corresponding to the logical link within the monitoring period, and the target traffic value is the traffic value obtained after removing the anomalies in the link traffic value. Optionally, the link traffic value corresponding to the logical link can be monitored by a monitoring system, such as the Nagios system; the average traffic corresponding to the logical link can be obtained by calculating the ratio of the link traffic value and the monitoring period.
[0102] Further, as an optional embodiment of the present invention, combining the configuration parameters and the average traffic to calculate the load balancing deviation corresponding to the logical link includes: extracting the bandwidth parameter and the port parameter in the configuration parameters, calculating the traffic extreme value corresponding to the logical link according to the bandwidth parameter and the port parameter, calculating the load balancing deviation corresponding to the logical link according to the traffic extreme value and the average traffic, calculating the coefficient of variation corresponding to the load balancing deviation according to the load balancing deviation, and combining the load balancing deviation and the coefficient of variation to calculate the load balancing deviation corresponding to the logical link.
[0103] Among them, the traffic extreme value represents the maximum transmission traffic and the minimum transmission traffic corresponding to the logical link, the minimum transmission traffic is set to 0, the load balancing deviation represents the load distribution balance degree corresponding to the logical link, and the coefficient of variation represents the distribution volatility corresponding to the load balancing deviation.
[0104] Optionally, the bandwidth parameter and the port parameter in the configuration parameters can be obtained by an extraction function compiled by a scripting language, such as the JAVA language; the calculation method for the traffic extreme value corresponding to the logical link is: maximum traffic = bandwidth in the bandwidth parameter * port rate in the port parameter; the calculation formula for the load balancing deviation corresponding to the logical link is: load balancing deviation = (maximum traffic - minimum traffic) / average traffic; the coefficient of variation corresponding to the load balancing deviation can be obtained by calculating the standard deviation of the load balancing deviation; the load balancing deviation corresponding to the logical link can be obtained by calculating the coefficient of variation / the load balancing deviation.
[0105] By normalizing the load balancing deviation, the present invention can eliminate the influence between the load balancing deviations, facilitate the comparison between the load balancing deviations, analyze the cluster state corresponding to the distributed computing cluster according to the cluster operation data, and further understand the health state, load distribution, etc. corresponding to the distributed computing cluster, which is convenient for fault self-checking or determining the load to be optimized. Among them, the target balance degree is the balance degree obtained after unifying the load balancing deviations to the same standard, the cluster operation data is the data generated during the operation of the distributed computing cluster, such as component status, performance indicators, and event record data, and the cluster state is a description of the operation state of the distributed computing cluster. Optionally, the normalization process of the load balancing deviation can be realized by the maximum-minimum normalization method; the cluster operation data corresponding to the distributed computing cluster can be realized by sensors, such as energy consumption sensors and load sensors; the cluster state corresponding to the distributed computing cluster can be analyzed according to the cluster operation data. For example, whether there is abnormal data in the operation data, and the abnormal state of the cluster can be judged according to the abnormal data. The completion state corresponding to the operation data can be used to analyze whether the corresponding cluster is idle.
[0106] S4. Schedule the cluster log corresponding to the distributed computing cluster, calculate the cluster center load corresponding to the distributed computing cluster according to the cluster log, combine the cluster transmission capacity, the target balance degree, and the cluster center load, set the scheduling priority corresponding to the distributed computing cluster, and construct a performance optimization plan for the distributed computing cluster according to the scheduling priority and the cluster state.
[0107] Based on the cluster log, the present invention calculates the cluster center load corresponding to the distributed computing cluster, which can understand the resource usage of the central nodes in the distributed computing cluster, and thus facilitate the evaluation of the overall performance of the distributed computing cluster, providing a basis for setting subsequent scheduling priorities and facilitating the improvement of the performance optimization efficiency of the distributed computing cluster. Among them, the cluster log is the information generated and recorded by each node in the distributed computing cluster, and the cluster center load refers to the load situation of the central nodes in the distributed computing cluster.
[0108] As an embodiment of the present invention, the calculating the cluster center load corresponding to the distributed computing cluster according to the cluster log includes: according to the cluster log, finding the cluster center node corresponding to the distributed computing cluster, using a preset performance analysis tool to analyze the node performance indicators corresponding to the cluster center node, extracting the index log information corresponding to the node performance indicators from the cluster log, calculating the index load corresponding to the node performance indicators according to the index log information, and calculating the cluster center load corresponding to the distributed computing cluster according to the index load.
[0109] Among them, the cluster center node is the control node or management node corresponding to the distributed computing cluster, responsible for coordinating and scheduling tasks and resource allocation in the cluster. The performance analysis tool is a tool used to analyze each index in the cluster center node to determine the corresponding node performance indicators, such as the Perf tool. The node performance indicators are the standards for evaluating the performance and efficiency corresponding to the cluster center node, such as CPU utilization rate, response time, memory utilization rate, and throughput, etc.
[0110] Optionally, the cluster center node corresponding to the distributed computing cluster can be found through the cluster configuration file; the index load corresponding to the node performance indicators can be calculated according to the index log information. For example, according to the extracted index log information, calculate the average value of the performance indicators of each node within a given time period. For the CPU utilization rate, the average value of the CPU utilization rates of all nodes can be calculated, and the index load can be obtained according to the average value; the cluster center load corresponding to the distributed computing cluster can be obtained by summing the index loads through a load balancing algorithm.
[0111] The present invention sets the scheduling priority corresponding to the distributed computing cluster by combining the cluster transmission capacity, the target balance degree, and the cluster center load, so as to optimize the allocation of resources and enable the cluster to obtain a faster response. Among them, the scheduling priority represents the priority degree of resource scheduling corresponding to the distributed computing cluster. Optionally, the scheduling priority corresponding to the distributed computing cluster can be set by combining the values corresponding to the cluster transmission capacity, the target balance degree, and the cluster center load. For example, if the cluster transmission capacity is high, it indicates that transmission resources such as network bandwidth and storage bandwidth are sufficient, and data transmission can be completed faster, and the data transmission between tasks will not have a great impact on the entire cluster. If the target balance degree is high, then the scheduling priority should first consider allocating tasks to nodes with more idle resources to balance the load of each node in the cluster. If the cluster center load is high, that is, the number of tasks waiting to be executed in the task queue is large, these tasks may need to be completed more urgently. In this case, the scheduling priority can be set according to factors such as the importance, urgency, and data transmission requirements of the tasks, and these high-priority tasks can be scheduled first; the performance optimization scheme of the distributed computing cluster can be constructed according to the scheduling priority and the cluster state. For example, if the cluster state is normal and in an idle state, and the scheduling priority is high, then this cluster can be scheduled as a load optimization cluster.
[0112] The present invention queries the service functions and configuration parameters corresponding to the service virtual machine, and then understands the technical services and corresponding attribute information of the service virtual machine. According to the service functions, the interactive virtual machines in the service virtual machine are analyzed to obtain the virtual machines that interact with each other in the service virtual machine. The present invention calculates the link transmission capacity corresponding to the logical link according to the configuration parameters, and can understand the allocable bandwidth value corresponding to the logical link, which is convenient for subsequent calculation and processing of the cluster transmission capacity. The present invention calculates the load balance deviation between the logical links according to the configuration parameters, and can understand the load balance degree between the logical links, and then analyze the stability between the logical links. The present invention calculates the cluster center load corresponding to the distributed computing cluster according to the cluster log, and can understand the resource usage of the central node in the distributed computing cluster, which is convenient for evaluating the overall performance of the distributed computing cluster and provides a basis for setting the subsequent scheduling priority, and is convenient for improving the performance optimization efficiency of the distributed computing cluster. Therefore, a method for constructing a system performance optimization platform based on big data provided by the embodiments of the present invention can improve the system performance optimization efficiency of data.
[0113] As Figure 2 shown, it is a functional module diagram of a system for constructing a system performance optimization platform based on big data provided by an embodiment of the present invention.
[0114] The system 100 for building a system performance optimization platform based on big data according to the present invention can be installed in an electronic device. According to the functions achieved, the system 100 for building a system performance optimization platform based on big data can include a logical link determination module 101, a transmission capacity calculation module 102, a status analysis module 103, and a performance optimization module 104. The modules in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0115] In this embodiment, the functions of each module / unit are as follows:
[0116] The logical link determination module 101 is used to obtain the cloud service system to be optimized, extract the distributed computing cluster of the cloud service system, identify the service virtual machines in the distributed computing cluster, query the service functions and configuration parameters corresponding to the service virtual machines, analyze the interactive virtual machines in the service virtual machines according to the service functions, and determine the logical links corresponding to the interactive virtual machines;
[0117] The transmission capacity calculation module 102 is used to calculate the link transmission capacity corresponding to the logical link according to the configuration parameters, and calculate the cluster transmission capacity corresponding to the distributed computing cluster according to the link transmission capacity;
[0118] The status analysis module 103 is used to calculate the load balancing deviation between the logical links according to the configuration parameters, perform normalization processing on the load balancing deviation to obtain the target balance degree, collect the cluster operation data corresponding to the distributed computing cluster, and analyze the cluster status corresponding to the distributed computing cluster according to the cluster operation data;
[0119] The performance optimization module 104 is used to schedule the cluster logs corresponding to the distributed computing cluster, calculate the cluster center load corresponding to the distributed computing cluster according to the cluster logs, combine the cluster transmission capacity, the target balance degree, and the cluster center load, set the scheduling priority corresponding to the distributed computing cluster, and construct a performance optimization plan for the distributed computing cluster according to the scheduling priority and the cluster status.
[0120] Specifically, each module in the system 100 for building a system performance optimization platform based on big data in the embodiment of the present application uses the same technical means as those in the Figure 1 method for building a system performance optimization platform based on big data described above, and can produce the same technical effects, which will not be elaborated here.
[0121] In several embodiments provided by the present invention, it should be understood that the provided methods and systems can be implemented in other ways. For example, the method embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for constructing a system performance optimization platform based on big data, characterized in that, The method includes: Obtain a cloud service system to be optimized, extract the distributed computing cluster of the cloud service system, identify the service virtual machines in the distributed computing cluster, query the service functions and configuration parameters corresponding to the service virtual machines, analyze the interactive virtual machines in the service virtual machines according to the service functions, and determine the logical links corresponding to the interactive virtual machines; Calculate the link transmission capacity corresponding to the logical link according to the configuration parameters, and calculate the cluster transmission capacity corresponding to the distributed computing cluster according to the link transmission capacity. Among them, calculating the link transmission capacity corresponding to the logical link according to the configuration parameters includes: Calculate the idle bandwidth corresponding to the logical link according to the configuration parameters; Perform path tracing on each link in the logical link to obtain a traced path; Statistically count the passing times corresponding to each link in the logical link according to the traced path; Calculate the link transmission capacity corresponding to the logical link by using the following formula according to the passing times and the idle bandwidth: Among them, A represents the link transmission capacity corresponding to the logical link, B i represents the idle bandwidth of the i-th logical link, C i represents the number of times the i-th logical link has passed through, and i represents the serial number corresponding to the logical link; Among them, calculating the idle bandwidth corresponding to the logical link according to the configuration parameters includes: Extract the network bandwidth corresponding to the interactive virtual machine from the configuration parameters; Calculate the link bandwidth corresponding to the logical link according to the network bandwidth; Statistically count the number of upstream and downstream virtual machines corresponding to the logical link to obtain the first virtual machine number and the second virtual machine number; Calculate the global bandwidth corresponding to the logical link by using the following formula according to the first virtual machine number and the second virtual machine number: D = min(F1, F2) * E a Among them, D represents the global bandwidth corresponding to the logical link, F1 and F2 respectively represent the number one of virtual machines and the number two of virtual machines of the F-th logical link, and E a represents the total bandwidth corresponding to the a-th distributed computing cluster, and a represents the serial number corresponding to the distributed computing cluster; Calculate the idle bandwidth corresponding to the logical link according to the global bandwidth and the link bandwidth; Among them, calculating the cluster transmission capacity corresponding to the distributed computing cluster according to the link transmission capacity includes: Statistically count the cluster failures corresponding to the distributed computing cluster, and calculate the cluster failure rate corresponding to the distributed computing cluster according to the cluster failures; Allocate the cluster weight corresponding to the distributed computing cluster according to the cluster failure rate; Calculate the cluster transmission capacity corresponding to the distributed computing cluster by using the following formula in combination with the cluster weight and the link transmission capacity: Among them, G represents the cluster transmission capacity corresponding to the distributed computing cluster, ω a represents the cluster weight corresponding to the a-th distributed computing cluster, A b represents the b-th link transmission capacity, b is the serial number corresponding to the link transmission capacity, and q is the total number corresponding to the link transmission capacity; Calculate the load balancing deviation between the logical links according to the configuration parameters, perform normalization processing on the load balancing deviation to obtain a target balance degree, collect the cluster operation data corresponding to the distributed computing cluster, and analyze the cluster state corresponding to the distributed computing cluster according to the cluster operation data; Schedule the cluster logs corresponding to the distributed computing cluster, calculate the cluster central load corresponding to the distributed computing cluster according to the cluster logs, set the scheduling priority corresponding to the distributed computing cluster in combination with the cluster transmission capacity, the target balance degree and the cluster central load, and construct a performance optimization plan for the distributed computing cluster according to the scheduling priority and the cluster state.
2. The method for constructing a system performance optimization platform based on big data according to claim 1, characterized in that Analyzing the interactive virtual machines in the service virtual machines according to the service functions includes: Identify the function information corresponding to the service function, and determine the function module and module components corresponding to the service virtual machine according to the function information; Schedule the module data stream corresponding to the function module, and analyze the data endpoints corresponding to the module data stream; Access the component code corresponding to the module component, and analyze the dependency relationship between the component codes; Analyze the interactive virtual machine in the service virtual machine according to the data endpoints and the dependency relationship; 3. A method for constructing a system performance optimization platform based on big data according to claim 1, characterized in that The determination of the logical link corresponding to the interactive virtual machine includes: Obtain the virtual network configuration information corresponding to the interactive virtual machine, and query the communication device and device configuration information corresponding to the interactive virtual machine; Analyze the device topology structure corresponding to the communication device according to the device configuration information; Combine the device topology structure and the network configuration information to determine the network topology structure corresponding to the interactive virtual machine; Determine the logical link corresponding to the interactive virtual machine according to the network topology structure; 4. A method for constructing a system performance optimization platform based on big data according to claim 1, characterized in that, The calculation of the load balancing deviation between the logical links according to the configuration parameters includes: Monitor the link traffic value corresponding to the logical link, and record the monitoring period corresponding to the link traffic value; Perform anomaly rejection on the link traffic value to obtain the target traffic value; Combine the target traffic value and the monitoring period to calculate the average traffic corresponding to the logical link; Combine the configuration parameters and the average traffic to calculate the load balancing deviation corresponding to the logical link; 5. A method for constructing a system performance optimization platform based on big data as claimed in claim 4, wherein, The combination of the configuration parameters and the average traffic to calculate the load balancing deviation corresponding to the logical link includes: Extract the bandwidth parameter and port parameter in the configuration parameters; Calculate the traffic extreme value corresponding to the logical link according to the bandwidth parameter and the port parameter; Calculate the load balancing deviation corresponding to the logical link according to the traffic extreme value and the average traffic; Calculate the coefficient of variation corresponding to the load balancing deviation according to the load balancing deviation; Combine the load balancing deviation and the coefficient of variation to calculate the load balancing deviation corresponding to the logical link; 6. A method for constructing a system performance optimization platform based on big data according to claim 1, characterized in that The calculation of the cluster center load corresponding to the distributed computing cluster according to the cluster log includes: Find the cluster center node corresponding to the distributed computing cluster according to the cluster log; Use a preset performance analysis tool to analyze the node performance indicators corresponding to the cluster center node; Extract the index log information corresponding to the node performance indicators from the cluster log; Calculate the index load corresponding to the node performance indicators according to the index log information; Calculate the cluster center load corresponding to the distributed computing cluster according to the index load; 7. A system for constructing a system performance optimization platform based on big data, characterized in that, The system includes: A logical link determination module, configured to obtain a cloud service system to be optimized, extract the distributed computing cluster of the cloud service system, identify the service virtual machines in the distributed computing cluster, query the service functions and configuration parameters corresponding to the service virtual machines, and analyze the interactive virtual machines in the service virtual machines according to the service functions, and determine the logical links corresponding to the interactive virtual machines; A transmission capacity calculation module, configured to calculate the link transmission capacity corresponding to the logical link according to the configuration parameters, and calculate the cluster transmission capacity corresponding to the distributed computing cluster according to the link transmission capacity. Wherein, calculating the link transmission capacity corresponding to the logical link according to the configuration parameters includes: Calculating the idle bandwidth corresponding to the logical link according to the configuration parameters; Performing path tracing on each link in the logical link to obtain a traced path; Counting the number of passes corresponding to each link in the logical link according to the traced path; Calculating the link transmission capacity corresponding to the logical link according to the number of passes and the idle bandwidth by using the following formula: Among them, A represents the link transmission capacity corresponding to the logical link, and B i represents the idle bandwidth of the i-th logical link, and C i represents the number of times the i-th logical link has passed, and i represents the serial number corresponding to the logical link; Wherein, calculating the idle bandwidth corresponding to the logical link according to the configuration parameters includes: Extracting the network bandwidth corresponding to the interactive virtual machine from the configuration parameters; Calculating the link bandwidth corresponding to the logical link according to the network bandwidth; Counting the number of upstream and downstream virtual machines corresponding to the logical link to obtain the first number of virtual machines and the second number of virtual machines; Calculating the global bandwidth corresponding to the logical link according to the first number of virtual machines and the second number of virtual machines by using the following formula: D = min(F1, F2) * E a Wherein, D represents the global bandwidth corresponding to the logical link, F1 and F2 respectively represent the number one of virtual machines and the number two of virtual machines of the F-th logical link, and E a represents the total bandwidth corresponding to the a-th distributed computing cluster, and a represents the serial number corresponding to the distributed computing cluster; Calculating the idle bandwidth corresponding to the logical link according to the global bandwidth and the link bandwidth; Wherein, calculating the cluster transmission capacity corresponding to the distributed computing cluster according to the link transmission capacity includes: Counting the cluster failures corresponding to the distributed computing cluster, and calculating the cluster failure rate corresponding to the distributed computing cluster according to the cluster failures; Allocating the cluster weight corresponding to the distributed computing cluster according to the cluster failure rate; Combining the cluster weight and the link transmission capacity, and calculating the cluster transmission capacity corresponding to the distributed computing cluster by using the following formula: Among them, G represents the cluster transmission capacity corresponding to the distributed computing cluster, ω a represents the cluster weight corresponding to the a-th distributed computing cluster, A b represents the b-th link transmission capacity, b is the serial number corresponding to the link transmission capacity, and q is the total number corresponding to the link transmission capacity; A status analysis module, configured to calculate the load balancing deviation between the logical links according to the configuration parameters, perform normalization processing on the load balancing deviation to obtain a target balance degree, collect the cluster operation data corresponding to the distributed computing cluster, and analyze the cluster status corresponding to the distributed computing cluster according to the cluster operation data; A performance optimization module, configured to schedule the cluster logs corresponding to the distributed computing cluster, calculate the cluster central load corresponding to the distributed computing cluster according to the cluster logs, combine the cluster transmission capacity, the target balance degree and the cluster central load, set the scheduling priority corresponding to the distributed computing cluster, and construct a performance optimization scheme for the distributed computing cluster according to the scheduling priority and the cluster status.
Citation Information
Patent Citations
Cluster resource load balancing method and device, electronic equipment and medium
CN114443284A
Service node scheduling method and device, computer equipment, storage medium and program product
CN118784727A