A cloud computing-based big data cloud platform management system and method
By intelligently collecting multi-source data and analyzing multi-dimensional data, combined with a dynamic equilibrium model for cloud platforms, the problem of low resource utilization efficiency in traditional cloud platform management has been solved, achieving dynamic resource scheduling and efficient platform response capabilities.
Patent Information
- Application Number
- CN202511486826.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-10-17
AI Technical Summary
Traditional big data cloud platform management technologies struggle to detect dynamic changes in business fluctuations, resource load, and service efficiency in real time, resulting in low resource utilization efficiency and an inability to meet the demands for high availability and high elasticity.
Through multi-source data intelligent collection, data cleaning and standardization, multi-dimensional data analysis and intelligent decision optimization modules, the business fluctuation coefficient, resource load coefficient and service energy efficiency coefficient are calculated and input into the cloud platform dynamic balancing model to achieve dynamic resource scheduling.
It enables multi-dimensional quantitative evaluation of the cloud platform's operational status, improves resource utilization efficiency and platform responsiveness, and enhances the model's flexibility and accuracy.
Smart Images

Figure CN121000720B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud computing technology, and in particular to a cloud computing-based big data cloud platform management system and method. Background Technology
[0002] With the rapid development of information technology, the integrated application of cloud computing and big data technologies has become a core driving force for the digital transformation of various industries. More and more enterprises and institutions rely on cloud platforms to deploy business and process data. Massive amounts of structured and unstructured data are continuously generated by IoT devices, business systems, and cloud logs, posing unprecedented challenges to the resource management, data processing, and service scheduling capabilities of cloud platforms.
[0003] Currently, traditional big data cloud platform management technologies mostly adopt static resource allocation models, relying on preset rules for load scheduling, making it difficult to perceive dynamic changes in business fluctuations, resource load, and service efficiency in real time. Although some systems have introduced basic load balancing mechanisms, they lack the ability to collaboratively analyze multi-dimensional data when facing complex scenarios such as sudden changes in user behavior, device performance degradation, or abnormal network traffic, thus failing to achieve intelligent dynamic optimization of resources.
[0004] Existing technologies have revealed significant shortcomings in practical applications: On the one hand, static resource scheduling strategies lead to low resource utilization efficiency, often resulting in some devices being overloaded while others are idle, which seriously affects the platform's response speed; on the other hand, traditional models lack a dynamic weight adjustment mechanism for assessing the platform's health status, making it difficult to adapt to the differences in the importance of business fluctuations, resource load, and service efficiency under different business scenarios, resulting in insufficient accuracy of resource scheduling strategies and failing to meet the high availability and high elasticity requirements of modern cloud platforms.
[0005] Therefore, it is essential to invent a cloud platform management system and method based on cloud computing to solve the above problems. Summary of the Invention
[0006] The purpose of this invention is to provide a big data cloud platform management system and method based on cloud computing to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a cloud computing-based big data platform management system, comprising the following modules:
[0008] The multi-source data intelligent acquisition module is used to collect structured and unstructured data from IoT devices, business system interfaces, and cloud logs through a distributed sensor network, forming a cloud real-time data resource pool; the cloud real-time data resource pool includes user behavior feature datasets, device status monitoring datasets, and network traffic feature datasets;
[0009] The data cleaning and standardization module is used to perform outlier removal, missing value imputation and dimension normalization on the cloud real-time data resource pool to generate a standardized data resource pool.
[0010] The multi-dimensional data analysis module is used to perform time-series decomposition, cluster analysis, and association rule mining on the standardized data resource pool to obtain the business fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient. The business fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient are then input into the cloud platform dynamic equilibrium model to output the platform health index.
[0011] The business fluctuation coefficient is specifically as follows:
[0012] ,
[0013] in, , Where α1 is the delay-sensitive term, α2 is the behavior discrete term, and ω 1i ω 2i S is the weighting factor, n is the total number of users, and S is the weighting factor. i S is the service call delay time for the i-th user. max L is the service latency threshold. i Let i be the login frequency of the i-th user. Let σ be the mean login frequency of all users. L Let be the standard deviation of login frequency, π be the value of pi, sin be the sine function, and tanh be the hyperbolic tangent function.
[0014] The resource load factor is specifically:
[0015] ,
[0016] in, , Where β1 is the multi-resource coupling overload term, β2 is the equipment pressure balancing term, k1 is the adjustment coefficient, m is the total number of equipment, and λ j M is the baseline resource threshold for the j-th device. j Let D be the memory usage of the j-th device. j G represents the disk I / O throughput of the j-th device. j Let C be the GPU memory usage of the j-th device. j Let P be the CPU core temperature of the j-th device. j Let P be the packet loss rate of the network where the j-th device is located. max Let ln be the maximum tolerable packet loss rate, where ln is the logarithm to the base e, and e is the natural constant.
[0017] The service energy efficiency coefficient is specifically:
[0018] ,
[0019] in, , , Where γ1 is the bandwidth utilization term, γ2 is the request stability term, γ3 is the retransmission penalty term, and B max B is the system's maximum theoretical bandwidth. p (t) represents the peak network bandwidth at time t, P l (t) represents the packet loss rate at time t, H r (t) represents the HTTP request-response ratio at time t. ref For the ideal HTTP request-response ratio, R c (t) represents the number of TCP retransmissions at time t, k2 is the attenuation coefficient, max is the maximum value function, and e is the natural constant;
[0020] The cloud platform dynamic balancing model is as follows:
[0021] ,
[0022] Where α is the business fluctuation coefficient, β is the resource load coefficient, γ is the service energy efficiency coefficient, δ1, δ2 and δ3 are dynamic weight coefficients, δ1+δ2+δ3=1 and δ1, δ2 and δ3∈[0,1];
[0023] The intelligent decision optimization module is used to trigger resource scheduling strategies based on the platform's health index;
[0024] The dynamic feedback optimization module optimizes and standardizes the dynamic weighting coefficients based on historical platform health index data collected within a preset time window and their corresponding business fluctuation coefficients, resource load coefficients, and service efficiency coefficients.
[0025] Preferably, the user behavior feature dataset includes the user's login frequency and the user's service call latency; the device status monitoring dataset includes CPU core temperature, memory usage, disk I / O throughput, and GPU memory usage; and the network traffic feature dataset includes network bandwidth peak, packet loss rate, TCP retransmission count, and HTTP request-response ratio.
[0026] Preferably, the adjustment coefficient k1 is dynamically set according to the total number of devices m:
[0027] When m < 100, then k1 ∈ [0.8, 1.2];
[0028] When m≥100, then k1∈[0.3,0.7].
[0029] Preferably, the initial values of the dynamic weight coefficients δ1, δ2, and δ3 are set according to the business scenario:
[0030] If it is the peak business period, the initial values of δ1, δ2, and δ3 are set to δ1 = 0.5, δ2 = 0.3, and δ3 = 0.2;
[0031] If it is the resource - shortage period, the initial values of δ1, δ2, and δ3 are set to δ1 = 0.2, δ2 = 0.5, and δ3 = 0.3;
[0032] If it is the service - sensitive period, the initial values of δ1, δ2, and δ3 are set to δ1 = 0.2, δ2 = 0.3, and δ3 = 0.5;
[0033] If it is the normal period, the initial values of δ1, δ2, and δ3 are set to δ1 = 0.4, δ2 = 0.3, and δ3 = 0.3.
[0034] Preferably, the resource scheduling strategy is as follows:
[0035] If the platform health index 0 ≤ E < E1, trigger the first - level scheduling strategy and start the emergency downgrading strategy;
[0036] If the platform health index E1 ≤ E < E2, trigger the second - level scheduling strategy and start the load - balancing strategy;
[0037] If the platform health index E2 ≤ E < E3, trigger the third - level scheduling strategy and start the elastic expansion strategy;
[0038] If the platform health index E ≥ E3, trigger the fourth - level scheduling strategy and only monitor without scheduling.
[0039] Preferably, the execution process of the dynamic feedback optimization module is as follows:
[0040] A1. Obtain the business fluctuation coefficient α, t resource load coefficient β, t and service energy - efficiency coefficient γ t at the current time t; <000A2. Calculate the co - entropy value of the three - dimensional coefficients:
[0041] , where b is a very small constant to prevent the denominator from being zero;
[0042] A3. Update the dynamic weight coefficients according to the entropy value state:
[0043] , where τ is the entropy - sensitive adjustment factor, is the old weight, is the old weight, θ k ∈{α, β, γ}, k = 1, 2, 3;
[0044] A4. Constraint normalization of the weight coefficients:
[0045] ,
[0046] in, , For temporary weights, This is the final weight.
[0047] A cloud computing-based big data platform management method specifically includes the following steps:
[0048] S1. The multi-source data intelligent acquisition module collects structured and unstructured data from IoT devices, business system interfaces, and cloud logs through a distributed sensor network, forming a cloud real-time data resource pool; the cloud real-time data resource pool includes user behavior feature datasets, device status monitoring datasets, and network traffic feature datasets;
[0049] S2. The data cleaning and standardization module performs outlier removal, missing value imputation and dimension normalization on the cloud real-time data resource pool to generate a standardized data resource pool.
[0050] S3. The multi-dimensional data analysis module performs time-series decomposition, cluster analysis, and association rule mining on the standardized data resource pool to obtain the business fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient. The business fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient are then input into the cloud platform dynamic equilibrium model to output the platform health index.
[0051] The business fluctuation coefficient is specifically as follows:
[0052] ,
[0053] in, , Where α1 is the delay-sensitive term, α2 is the behavior discrete term, and ω 1i ω 2i S is the weighting factor, n is the total number of users, and S is the weighting factor. i S is the service call delay time for the i-th user. max L is the service latency threshold. i Let i be the login frequency of the i-th user. Let σ be the mean login frequency of all users. L Let be the standard deviation of login frequency, π be the value of pi, sin be the sine function, and tanh be the hyperbolic tangent function.
[0054] The resource load factor is specifically:
[0055] ,
[0056] in, , Where β1 is the multi-resource coupling overload term, β2 is the equipment pressure balancing term, k1 is the adjustment coefficient, m is the total number of equipment, and λ j M is the baseline resource threshold for the j-th device. j Let D be the memory usage of the j-th device. j G represents the disk I / O throughput of the j-th device. j Let C be the GPU memory usage of the j-th device. j Let P be the CPU core temperature of the j-th device. j Let P be the packet loss rate of the network where the j-th device is located. max Let ln be the maximum tolerable packet loss rate, where ln is the logarithm to the base e, and e is the natural constant.
[0057] The service energy efficiency coefficient is specifically:
[0058] ,
[0059] in, , , Where γ1 is the bandwidth utilization term, γ2 is the request stability term, γ3 is the retransmission penalty term, and B max B is the system's maximum theoretical bandwidth. p (t) represents the peak network bandwidth at time t, P l (t) represents the packet loss rate at time t, H r (t) represents the HTTP request-response ratio at time t. ref For the ideal HTTP request-response ratio, R c (t) represents the number of TCP retransmissions at time t, k2 is the attenuation coefficient, max is the maximum value function, and e is the natural constant;
[0060] The cloud platform dynamic balancing model is as follows:
[0061] ,
[0062] Where α is the business fluctuation coefficient, β is the resource load coefficient, γ is the service energy efficiency coefficient, δ1, δ2 and δ3 are dynamic weight coefficients, δ1+δ2+δ3=1 and δ1, δ2 and δ3∈[0,1];
[0063] S4. The intelligent decision-making optimization module triggers resource scheduling strategies based on the platform's health index.
[0064] S5. The dynamic feedback optimization module optimizes and standardizes the dynamic weighting coefficient based on the historical data of the platform health index collected within a preset time window and its corresponding business fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient.
[0065] The technical effects and advantages of this invention are as follows:
[0066] This invention utilizes a multi-dimensional data analysis module to perform time-series decomposition, cluster analysis, and association rule mining on a standardized data resource pool. It calculates business fluctuation coefficients, resource load coefficients, and service efficiency coefficients, and inputs these into a cloud platform dynamic balancing model to output a platform health index. This achieves a multi-dimensional quantitative assessment of the cloud platform's operational status, accurately reflecting the platform's health condition.
[0067] This invention, through an intelligent decision optimization module, triggers different levels of resource scheduling strategies based on the platform's health index, including emergency degradation, load balancing, and elastic scaling, thereby achieving dynamic intelligent scheduling of cloud platform resources and improving resource utilization efficiency and platform responsiveness.
[0068] This invention optimizes and standardizes dynamic weight coefficients based on historical data and corresponding coefficients of the platform health index within a preset time window through a dynamic feedback optimization module. This enables the weight configuration of the cloud platform dynamic balancing model to adapt to changes in business scenarios, thereby improving the model's flexibility and accuracy. Attached Figure Description
[0069] Figure 1 This is a schematic diagram of the system module connections of the present invention.
[0070] Figure 2 This is a schematic diagram of the execution process of the dynamic feedback optimization module of the present invention.
[0071] Figure 3 This is a schematic diagram of the method steps of the present invention. Detailed Implementation
[0072] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0073] This invention provides, for example Figure 1 The cloud platform management system based on cloud computing shown includes the following modules:
[0074] The multi-source data intelligent acquisition module is used to collect structured and unstructured data from IoT devices, business system interfaces, and cloud logs through a distributed sensor network, forming a cloud real-time data resource pool; the cloud real-time data resource pool includes user behavior feature datasets, device status monitoring datasets, and network traffic feature datasets;
[0075] Furthermore, in the above technical solution, the user behavior feature dataset includes the user's login frequency and the user's service call latency; the device status monitoring dataset includes CPU core temperature, memory usage, disk I / O throughput, and GPU memory usage; and the network traffic feature dataset includes network bandwidth peak, packet loss rate, TCP retransmission count, and HTTP request-response ratio.
[0076] It should be noted that the login frequency of the user is obtained by obtaining the login operation records of the user account through the business system interface or cloud logs. Specifically, the login events of the user authentication system are monitored by the API interface, or the identity authentication logs, such as Apache or Nginx logs and user center database logs, are parsed to aggregate and count the number of logins per unit time by user ID.
[0077] The user's service call latency is collected from the response time monitoring of the business system interface. Probes, such as OpenTelemetry and Zipkin, are implanted in the service interface call chain to record the time difference between request sending and response receiving, or the latency is obtained by parsing the response time of HTTP or API requests in the logs.
[0078] The CPU core temperature is collected by hardware sensors of IoT devices, such as servers and edge nodes. The CPU temperature sensor data is read through hardware management interfaces, such as IPMI, WMI or system tools, such as Linux lm-sensors, so as to obtain the temperature of each core in real time.
[0079] The memory usage rate is collected from the memory status information of the device operating system. The ratio of the used capacity of physical memory and virtual memory to the total capacity is obtained through system APIs, such as Python's psutil library and Java's ManagementFactory. The memory usage rate is then calculated.
[0080] The disk I / O throughput is obtained by monitoring the I / O operations of the storage device. The throughput is obtained by using operating system tools, such as iostat in Linux, perfmon in Windows, or the storage device management interface, to count the amount of data read and written by the disk per unit time.
[0081] The GPU memory usage rate is collected from the memory usage status of the GPU hardware. It is obtained through the driver interface provided by the GPU manufacturer, such as NVIDIA's nvidia-smi, AMD's rocm-smi, or the GPU monitoring API of deep learning frameworks such as TensorFlow, to obtain the ratio of used memory capacity to total capacity.
[0082] The network bandwidth peak is collected through traffic monitoring of network devices, such as routers and switches, using SNMP protocol or traffic monitoring tools, such as Prometheus combined with Node Exporter, to collect real-time bandwidth data of network interfaces and count the peak traffic per unit time.
[0083] The packet loss rate is collected based on the data packet transmission status of network devices. The number of input and output data packets of the network interface is analyzed by SNMP or network packet capture tools such as Wireshark, and the ratio of the number of lost packets to the total number of transmitted packets is calculated to obtain the packet loss rate.
[0084] The number of TCP retransmissions is collected from the TCP connection state of the network protocol stack. The TCP retransmission flags, such as duplicate ACKs and timeout retransmissions, are captured through the network statistics interface of the operating system, such as netstat-s in Linux or network analysis tools, such as tcpdump, and the number of retransmissions is obtained.
[0085] The HTTP request-response ratio is collected through web server access logs or reverse proxy monitoring data. This involves parsing HTTP server logs, such as Nginx's access.log, and calculating the ratio of successfully responded requests to the total number of requests per unit time. Alternatively, it can be achieved by using APM tools, such as New Relic, to monitor the ratio of HTTP response status codes, such as 2xx and 3xx success codes to 4xx and 5xx error codes.
[0086] The data cleaning and standardization module is used to perform outlier removal, missing value imputation and dimension normalization on the cloud real-time data resource pool to generate a standardized data resource pool.
[0087] It should be noted that the execution process of the data cleaning and standardization module is as follows: For the login frequency of the user, firstly, the mean and standard deviation of the login frequency of all users are calculated, and values that exceed the mean ± 3 times the standard deviation are judged as outliers and removed. For the missing login frequency data in the log records, the mean of the user's historical login frequency is used for interpolation. Finally, the login frequency data is linearly transformed to the [0, 1] interval through the Max-Min standardization method to eliminate the influence of the unit and generate standardized user login frequency data.
[0088] For the service call delay time of the user, the box plot method is first used to identify and remove abnormal delay times that are more than 1.5 times the interquartile range outside the upper and lower quartiles. For the missing delay time data, linear interpolation is used to interpolate based on the data before and after the time series. Then, the delay time is transformed into standard normal distribution data with a mean of 0 and a standard deviation of 1 through the Z-score standardization method to achieve dimensional normalization.
[0089] For the CPU core temperature, first set a hardware safe temperature threshold, then determine the temperature value exceeding the threshold as an outlier and replace it with the threshold. For missing temperature data, use the average temperature of other cores in the same device for interpolation. Finally, linearly map the temperature data to the interval [0, 1], where 0 corresponds to the lowest safe temperature and 1 corresponds to the highest safe temperature, thus completing the dimensional normalization.
[0090] For the memory usage rate, outliers exceeding 100% or less than 0 are first truncated to 100% or 0. For missing memory usage rate data, the average memory usage rate of the same device in the same time period is used for imputation. Then, the memory usage rate is converted into standardized data in the range of [0, 1] by Max-Min standardization, which is convenient for subsequent analysis.
[0091] For the disk I / O throughput, firstly, the sliding window mid-value filtering method is used to remove sudden abnormally high or low throughput data. For missing throughput data, the forward padding method is used for imputation. Then, the data is logarithmically transformed and normalized to the [0, 1] interval by Max-Min to cope with the large range of throughput data.
[0092] For the GPU memory usage rate, outliers exceeding the range of 0 to 100% are first truncated to 0 or 100%. For missing memory usage rate data, the mean memory usage rate under similar tasks on the same device is used for interpolation. Finally, the data is scaled to the [0, 1] interval through linear transformation to achieve dimensional normalization.
[0093] For the network bandwidth peak, firstly, peak values exceeding the maximum theoretical bandwidth of the network device are identified as outliers and replaced with the maximum theoretical bandwidth value. For missing bandwidth peak data, the average bandwidth peak value at adjacent time points is used for interpolation. Then, the bandwidth peak data is divided by the maximum theoretical bandwidth to obtain standardized data in the [0, 1] interval.
[0094] For the packet loss rate, abnormal packet loss rates that are greater than 1 or less than 0 are first truncated to 1 or 0. For missing packet loss rate data, they are interpolated based on the average packet loss rate during similar network load periods. Since the packet loss rate itself is in the range of [0, 1], the data range is directly maintained.
[0095] For the TCP retransmission count, the upper quartile of the retransmission count is first calculated, and data exceeding 1.5 times the upper quartile are judged as outliers and removed. For missing retransmission count data, the mean of the historical retransmission count of the same network connection is used for interpolation. Finally, the retransmission count is converted into standardized data in the range [0, 1] by Max-Min standardization.
[0096] For the HTTP request-response ratio, abnormal response ratios that exceed the range of 0 to 1 are first truncated to 0 or 1. For missing response ratio data, the mean response ratio of the same type of service in the same time period is used for interpolation. Since the response ratio itself is a proportional value, it is directly standardized to ensure that the data has a uniform dimension in the range of [0, 1].
[0097] The multi-dimensional data analysis module is used to perform time-series decomposition, cluster analysis, and association rule mining on the standardized data resource pool to obtain the business fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient. The business fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient are then input into the cloud platform dynamic equilibrium model to output the platform health index.
[0098] The business fluctuation coefficient is specifically as follows:
[0099] ,
[0100] in, , Where α1 is the delay-sensitive term, α2 is the behavior discrete term, and ω 1i ω 2i S is the weighting factor, n is the total number of users, and S is the weighting factor. i S is the service call delay time for the i-th user. max L is the service latency threshold. i Let i be the login frequency of the i-th user. Let σ be the mean login frequency of all users. L Let be the standard deviation of login frequency, π be the mathematical constant pi, sin be the sine function, and tanh be the hyperbolic tangent function.
[0101] It is important to understand that the core design goal of the aforementioned business fluctuation coefficient α is to construct an intelligent indicator that quantifies system stability by comprehensively evaluating the suddenness of service latency and the discreteness of user behavior. Specifically, the latency sensitivity term α1 measures the service call latency S of user i. i With preset threshold S max The ratio is mapped to the nonlinear periodic interval of the sine function, when S i Approaching S maxWhen α1 approaches its peak value of 1, it amplifies the critical risk of delay, while at low latency the function value is flat, avoiding small fluctuations from interfering with the judgment and achieving precise focus on service anomalies; the behavioral discrete term α2 is obtained through... Calculate the login frequency Li of user i relative to the overall mean. The degree of deviation is determined, and the saturation property of the hyperbolic tangent function tanh is used to compress the discrete values to the range of (0, 1). This not only identifies abnormal user behavior, such as high-frequency abnormal logins, but also suppresses the excessive influence of extreme outliers on the results, ensuring the robustness of the group behavior analysis. The two calculation results are respectively weighted by configurable weighting factors ω. 1i ω 2i Adjusting contribution levels allows for differentiated weight allocation based on user value or business scenario, such as assigning higher ω to VIP users. 1i To strictly control latency, the formula ultimately integrates all user data using a normalization framework: the numerator is the sum of the weighted latency and discrete signal, and the denominator is the sum of the weighting factors, eliminating the scale effect of the total number of users n; then, the ratio is subtracted from 1, so that the output value α∈[0,1] intuitively maps the stability of the service, where the stability is optimal when α approaches 1.
[0102] The resource load factor is specifically:
[0103] ,
[0104] in, , Where β1 is the multi-resource coupling overload term, β2 is the equipment pressure balancing term, k1 is the adjustment coefficient, m is the total number of equipment, and λ j M is the baseline resource threshold for the j-th device. j Let D be the memory usage of the j-th device. j G represents the disk I / O throughput of the j-th device. j Let C be the GPU memory usage of the j-th device. j Let P be the CPU core temperature of the j-th device. j Let P be the packet loss rate of the network where the j-th device is located. max Let ln be the maximum tolerable packet loss rate, ln be the logarithm to the base e, and e be the natural constant.
[0105] Furthermore, in the above technical solution, the adjustment coefficient k1 is dynamically set according to the total number of devices m: if it is a small-to-medium scale system, i.e., the total number of devices is less than 100, then k1 can be set between 0.8 and 1.2; if it is a large-scale system, i.e., the total number of devices exceeds 100, then k1 can be set between 0.3 and 0.7.
[0106] It is important to understand that the core design goal of the resource load coefficient β is to construct a unified comprehensive index for quantifying system overload risk through a two-dimensional coupled analysis of hardware resource conflicts and network quality. The subtraction framework of its formula maps the output to the [0, 1] interval: when β approaches 0, it represents high load risk; when β approaches 1, it represents healthy load. The exponential decay structure of the multi-resource coupled overload term β1 reveals the synergistic amplification effect of hardware resources. The numerator, memory utilization M... j Disk I / O throughput D j GPU memory usage (G) j The multiplicative design deliberately intensifies the resource competition relationship, while the baseline resource threshold λ on the denominator side... j CPU core temperature C j The product of C constitutes a dynamic fault-tolerant mechanism: j Increasing the denominator when λ rises weakens the harmfulness of other resource anomalies. j It supports differential calibration of safety boundaries based on device performance baselines, inner layer To compress the order of magnitude of resource conflict values and avoid extreme value dominance, an outer exponential function is used. The aggregation weight of the conflict signals is controlled by adjusting the coefficient k1, ultimately outputting a smooth decay curve that conforms to the principle of entropy increase. The geometric mean structure of the equipment pressure balancing term β2 focuses on the network's bottleneck effect. Convert the packet loss rate of device j into a relative maximum tolerance threshold P. max The occupancy ratio, multiplied by the m-th root, makes β2 extremely sensitive to single-point failures. When all devices P j Equilibrium is lower than P max When β2 approaches 1, the intermediate value is output to guide and optimize nodes with high packet loss when the distribution is non-uniform. The product of the two terms, β1×β2, contains deep business logic: hardware overload risk and network risk are coupled through multiplication to form a synergistic amplification. β only approaches 0 significantly when both deteriorate simultaneously.
[0107] The service energy efficiency coefficient is specifically:
[0108] ,
[0109] in, , , Where γ1 is the bandwidth utilization term, γ2 is the request stability term, γ3 is the retransmission penalty term, and B max B is the system's maximum theoretical bandwidth. p (t) represents the peak network bandwidth at time t, P l (t) represents the packet loss rate at time t, H r (t) represents the HTTP request-response ratio at time t. ref For the ideal HTTP request-response ratio, R c(t) represents the number of TCP retransmissions at time t, k2 is the attenuation coefficient, max is the maximum value function, and e is the natural constant.
[0110] It should be noted that the attenuation coefficient k2 can be set as follows: when the packet loss rate is low (less than 0.1%), k2 can be set between 0.5 and 0.7 to quickly penalize retransmissions; when the packet loss rate is high (more than 1%), k2 can be set between 0.1 and 0.3 to mitigate the penalty and avoid misjudgment.
[0111] It is important to understand that the core design of the service energy efficiency coefficient γ lies in forcing the coordinated optimization of three dimensions—bandwidth utilization efficiency, request processing stability, and transmission reliability—through a multiplicative concatenation framework. Attenuation of any one of these sub-items leads to a collapse in overall energy efficiency. The structure of the bandwidth utilization term γ1 is constructed with a dual constraint mechanism: numerator B... p (t) measures the actual channel capacity, while 1-P l (t) Penalize P in the form of a packet loss rate compensation factor l (t) The data integrity loss caused by the denominator B max To achieve normalization, this design ensures that γ1 is only used when bandwidth is efficiently utilized, i.e., B. p (t) approaches B max And low packet loss, i.e., P l (t) approaches 1 as it approaches 0, and any idle resources or transmission distortion significantly reduce its value; the piecewise function of the stability term γ2 is requested through H r (t) and H ref Dynamic comparison achieves bidirectional risk suppression: when H r (t)≤H ref When γ2=1, fluctuations within the benchmark are allowed, while H r (t) > H ref At that time, penalty items The risk of response delay caused by overloaded requests is quantified as a decreasing function; in the exponential decay model of the retransmission penalty term γ3, R c (t) Directly maps the frequency of transport layer faults, with k2 controlling the penalty intensity: exponential function e -x The characteristics of this feature cause a severe penalty to be triggered on the first retransmission, while the marginal effect of subsequent retransmissions diminishes, accurately simulating the avalanche effect of network congestion; the final product It contains deep business logic: three dimensions constitute a series of rigid constraints. For example, when the bandwidth is fully loaded with no packet loss and zero retransmission, if the request response efficiency drops by 15%, the overall γ will still drop to 0.85. The bottleneck can be quickly located by the output value: the decline of γ1 indicates the failure of bandwidth planning, the decline of γ2 reveals the insufficient computing resources, and the collapse of γ3 exposes the network link failure.
[0112] The cloud platform dynamic balancing model is as follows:
[0113] ,
[0114] Among them, α is the business fluctuation coefficient, β is the resource load coefficient, γ is the service energy efficiency coefficient, δ1, δ2, and δ3 are dynamic weight coefficients, and δ1 + δ2 + δ3 = 1 and δ1, δ2, and δ3 ∈ [0, 1].
[0115] Furthermore, in the above technical solution, the initial values of the dynamic weight coefficients δ1, δ2, and δ3 are set according to the business scenario: If it is the peak business period, the initial values of δ1, δ2, and δ3 are set to δ1 = 0.5, δ2 = 0.3, and δ3 = 0.2; if it is the resource shortage period, the initial values of δ1, δ2, and δ3 are set to δ1 = 0.2, δ2 = 0.5, and δ3 = 0.3; if it is the service sensitive period, the initial values of δ1, δ2, and δ3 are set to δ1 = 0.2, δ2 = 0.3, and δ3 = 0.5; if it is the normal period, the initial values of δ1, δ2, and δ3 are set to δ(1 = 0.4, δ2 = 0.3, and δ3 = 0.3.
[0116] An intelligent decision-making optimization module, which is used to trigger a resource scheduling strategy according to the platform health index;
[0117] Furthermore, in the above technical solution, the resource scheduling strategy is:
[0118] If the platform health index 0 ≤ E < E1, trigger a first-level scheduling strategy and start an emergency downgrading strategy;
[0119] If the platform health index E1 ≤ E < E2, trigger a second-level scheduling strategy and start a load balancing strategy;
[0120] If the platform health index E2 ≤ E < E3, trigger a third-level scheduling strategy and start an elastic expansion strategy;
[0121] If the platform health index E ≥ E3, trigger a fourth-level scheduling strategy and only monitor without scheduling.
[0122] It should be noted that the initial values of E1, E2, and E3 are set to E1 = 0.5, E2 = 0.7, and E3 = 0.85, and are calibrated in real time according to historical operation data:
[0123] ,
[0124] Among them, , among which, is the new threshold adjusted at the current moment, is the old threshold at the previous moment, η is the learning rate, with a default value of 0.01, T is the time window, with a default value of 24 hours, I is the indicator function, which is 1 when the condition is satisfied and 0 otherwise, α t is the business fluctuation coefficient at time t, βt Let E be the resource load factor at time t. k Let be the initial threshold value, where k = 1, 2, 3.
[0125] It is important to know that the emergency degradation strategy includes core service circuit breaking: forcibly shutting down non-critical business functions, such as data analysis and log auditing, while retaining only core functional modules, such as user authentication and payment transactions, thereby reducing resource consumption through business degradation;
[0126] Traffic interception: Inject circuit breaker rules into the load balancing layer to reject new requests or return a preset static page, such as HTTP 503, to avoid a cascading failure effect;
[0127] Strong resource reclamation: Terminate low-priority processes, such as batch tasks, to release CPU and memory resources, and to protect high-temperature devices, i.e., C. j Devices exceeding the threshold will perform load shifting or hardware frequency limiting.
[0128] Dynamic weight freeze: Pause the weight adjustment of the dynamic feedback optimization module and lock δ. k To ensure safety and prevent strategy oscillation, k = 1, 2, 3;
[0129] The load balancing strategy includes multi-dimensional routing: combining the resource load coefficient β, requests are dynamically distributed: nodes with high memory utilization (i.e., memory utilization exceeding 80%) receive computationally intensive tasks; nodes with low disk I / O throughput (i.e., disk I / O throughput below 50MB / s) process streaming data acquisition; and based on the service energy efficiency coefficient γ, requests with an HTTP request-response ratio exceeding 2 are preferentially routed to network areas with a packet loss rate below 0.1%.
[0130] Resource rebalancing: Real-time migration of unbalanced loads in device status monitoring datasets, such as migrating tasks between devices with GPU memory usage differences exceeding 30%, and automatically identifying resource hotspot clusters using cluster analysis;
[0131] Elastic preheating: Pre-start backup containers, accounting for 5% of the total, but do not connect to traffic for the time being, to ensure a seamless switch to the expansion strategy when E2≤E;
[0132] The elastic scaling strategy includes horizontal scaling: when the business fluctuation coefficient α exceeds 0.7, the login authentication nodes are scaled up according to the user behavior feature dataset. Where k is the elasticity coefficient, with a default value of 1.2; when the resource load coefficient β is less than 0.4, the number of high-load devices is doubled, such as expanding the capacity of a device group with memory utilization exceeding 90% by 200%;
[0133] Vertical optimization: Network traffic feature dataset shows peak bandwidth exceeding B maxWhen the network efficiency reaches 80%, upgrade the network interface to a redundant link; when the service energy efficiency coefficient γ is below 0.6, add GPU resources to the computing nodes.
[0134] Cost constraint: If E≥E3 for three consecutive time windows after expansion, automatically shrink to the base size to avoid resource idleness;
[0135] The dynamic feedback optimization module optimizes and standardizes the dynamic weighting coefficients based on historical platform health index data collected within a preset time window and their corresponding business fluctuation coefficients, resource load coefficients, and service efficiency coefficients.
[0136] Furthermore, in the above technical solution, the execution process of the dynamic feedback optimization module is as follows: Figure 2 As shown:
[0137] A1. Obtain the business fluctuation coefficient α at the current time t. t Resource load factor β t Service energy efficiency coefficient γ t ;
[0138] A2. Calculate the co-entropy value of the three-dimensional coefficients: Where b is a minimal constant to prevent the denominator from being zero;
[0139] A3. Update the dynamic weight coefficients based on the entropy value state:
[0140] Where τ is an entropy-sensitive adjustment factor, For the old weight, For the old weights, θ k ∈{α, β, γ}, k=1, 2, 3;
[0141] A4. Constraint normalization of the weight coefficients:
[0142] ,
[0143] in, , For temporary weights, This is the final weight.
[0144] It's important to know that the default value for 'b' is 10. -6 ;
[0145] The default value of the entropy-sensitive adjustment factor τ is 0.05, but it can be set according to different states: if it is a stable state, τ can be set between 0.08 and 0.12; if it is a fluctuating state, τ can be set between 0.01 and 0.03.
[0146] This invention provides, for example Figure 3The method for managing a big data cloud platform based on cloud computing, as shown, specifically includes the following steps:
[0147] S1. The multi-source data intelligent acquisition module collects structured and unstructured data from IoT devices, business system interfaces, and cloud logs through a distributed sensor network, forming a cloud real-time data resource pool; the cloud real-time data resource pool includes user behavior feature datasets, device status monitoring datasets, and network traffic feature datasets;
[0148] S2. The data cleaning and standardization module performs outlier removal, missing value imputation and dimension normalization on the cloud real-time data resource pool to generate a standardized data resource pool.
[0149] S3. The multi-dimensional data analysis module performs time-series decomposition, cluster analysis, and association rule mining on the standardized data resource pool to obtain the business fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient. The business fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient are then input into the cloud platform dynamic equilibrium model to output the platform health index.
[0150] S4. The intelligent decision-making optimization module triggers resource scheduling strategies based on the platform's health index.
[0151] S5. The dynamic feedback optimization module optimizes and standardizes the dynamic weighting coefficient based on the historical data of the platform health index collected within a preset time window and its corresponding business fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient.
[0152] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A cloud computing-based big data platform management system, characterized in that, It includes the following modules: A multi-source data intelligent acquisition module, which is used to collect structured and unstructured data from Internet of Things devices, business system interfaces and cloud logs through a distributed sensor network to form a cloud real-time data resource pool; the cloud real-time data resource pool includes a user behavior feature data set, a device status monitoring data set and a network traffic feature data set; A data cleaning and standardization module, which is used to perform outlier removal, missing value imputation and dimension normalization on the cloud real-time data resource pool to generate a standardized data resource pool; A multi-dimensional data analysis module, which is used to perform time series decomposition, clustering analysis and association rule mining on the standardized data resource pool to obtain a business fluctuation coefficient, a resource load coefficient and a service energy efficiency coefficient, and input the business fluctuation coefficient, the resource load coefficient and the service energy efficiency coefficient into a cloud platform dynamic balance model to output a platform health index; The business fluctuation coefficient is specifically: , in, , Where α1 is the delay-sensitive term, α2 is the behavior discrete term, and ω 1i ω 2i S is the weighting factor, n is the total number of users, and S is the weighting factor. i S is the service call delay time for the i-th user. max L is the service latency threshold. i Let i be the login frequency of the i-th user. Let σ be the mean login frequency of all users. L Let be the standard deviation of login frequency, π be the value of pi, sin be the sine function, and tanh be the hyperbolic tangent function. The resource load coefficient is specifically: , in, , Where β1 is the multi-resource coupling overload term, β2 is the equipment pressure balancing term, k1 is the adjustment coefficient, m is the total number of equipment, and λ j M is the baseline resource threshold for the j-th device. j Let D be the memory usage of the j-th device. j G represents the disk I / O throughput of the j-th device. j Let C be the GPU memory usage of the j-th device. j P represents the CPU core temperature of the j-th device. j Let P be the packet loss rate of the network where the j-th device is located. max Let ln be the maximum tolerable packet loss rate, where ln is the logarithm to the base e, and e is the natural constant. The service energy efficiency coefficient is specifically: , in, , , Where γ1 is the bandwidth utilization term, γ2 is the request stability term, γ3 is the retransmission penalty term, and B max B is the system's maximum theoretical bandwidth. p (t) represents the peak network bandwidth at time t, P l (t) represents the packet loss rate at time t, H r (t) represents the HTTP request-response ratio at time t. ref For the ideal HTTP request-response ratio, R c (t) represents the number of TCP retransmissions at time t, k2 is the attenuation coefficient, max is the maximum value function, and e is the natural constant; The cloud platform dynamic balance model is specifically: , Where, α is the business fluctuation coefficient, β is the resource load coefficient, γ is the service energy efficiency coefficient, δ1, δ2 and δ3 are dynamic weight coefficients, δ1 + δ2 + δ3 = 1 and δ1, δ2 and δ3 ∈ [0, 1]; An intelligent decision-making optimization module, which is used to trigger a resource scheduling strategy according to the platform health index; A dynamic feedback optimization module, which optimizes and standardizes the dynamic weight coefficients based on the historical data of the platform health index collected within a preset time window and its corresponding business fluctuation coefficient, resource load coefficient and service energy efficiency coefficient.
2. The big data cloud platform management system based on cloud computing according to claim 1, characterized in that, The user behavior feature data set includes the login frequency of users and the service call delay time of users; the device status monitoring data set includes the CPU core temperature, memory usage rate, disk I / O throughput and GPU video memory occupancy rate; the network traffic feature data set includes the network bandwidth peak value, packet loss rate, TCP retransmission times and HTTP request response ratio.
3. The big data cloud platform management system based on cloud computing according to claim 1, characterized in that, The adjustment coefficient k1 is dynamically set according to the total number of devices m: When m < 100, then k1 ∈ [0.8, 1.2]; When m ≥ 100, then k1 ∈ [0.3, 0.7].
4. The big data cloud platform management system based on cloud computing according to claim 1, characterized in that, The initial values of the dynamic weight coefficients δ1, δ2 and δ3 are set according to the business scenario: If it is the business peak period, the initial values of δ1, δ2 and δ3 are set to δ1 = 0.5, δ2 = 0.3, δ3 = 0.2; If it is the resource shortage period, the initial values of δ1, δ2 and δ3 are set to δ1 = 0.2, δ2 = 0.5, δ3 = 0.3; If it is the service sensitive period, the initial values of δ1, δ2 and δ3 are set to δ1 = 0.2, δ2 = 0.3, δ3 = 0.5; If it is the normal period, the initial values of δ1, δ2 and δ3 are set to δ1 = 分 0.4, δ2 = 0.3, δ3 = 0.
3.
5. The big data cloud platform management system based on cloud computing according to claim 1, characterized in that, The resource scheduling strategy is: If the platform health index 0 ≤ E < E1, trigger a first-level scheduling strategy and start an emergency downgrading strategy; If the platform health index E1 ≤ E < E2, trigger a second-level scheduling strategy and start a load balancing strategy; If the platform health index E2 ≤ E < E3, trigger the three - level scheduling strategy and start the elastic expansion strategy; If the platform health index E ≥ E3, trigger the four - level scheduling strategy, only monitor without scheduling.
6. The big data cloud platform management system based on cloud computing according to claim 1, characterized in that, The execution process of the dynamic feedback optimization module is as follows: A1. Obtain the business fluctuation coefficient α at the current time t. t Resource load factor β t Service energy efficiency coefficient γ t ; A2. Calculate the co-entropy value of the three-dimensional coefficients: Where b is a minimal constant to prevent the denominator from being zero; A3. Update the dynamic weight coefficient according to the entropy value state: Where τ is an entropy-sensitive adjustment factor. For the old weight, For the old weights, θ k ∈{α, β, γ}, k=1, 2, 3; A4. Perform constraint normalization on the weight coefficient: , in, , For temporary weights, This is the final weight.
7. A big data cloud platform management method based on cloud computing specifically includes the following steps: S1. The multi - source data intelligent acquisition module collects structured and unstructured data through a distributed sensor network in Internet of Things devices, business system interfaces, and cloud logs to form a cloud real - time data resource pool; the cloud real - time data resource pool includes a user behavior feature data set, a device status monitoring data set, and a network traffic feature data set; S2. The data cleaning and standardization module performs outlier removal, missing value imputation, and dimension normalization on the cloud real - time data resource pool to generate a standardized data resource pool; S3. The multi - dimensional data analysis module performs time - series decomposition, clustering analysis, and association rule mining on the standardized data resource pool to obtain a business fluctuation coefficient, a resource load coefficient, and a service energy efficiency coefficient, and inputs the business fluctuation coefficient, the resource load coefficient, and the service energy efficiency coefficient into the cloud platform dynamic equilibrium model to output the platform health index; The business fluctuation coefficient is specifically: , in, , Where α1 is the delay-sensitive term, α2 is the behavior discrete term, and ω 1i ω 2i S is the weighting factor, n is the total number of users, and S is the weighting factor. i S is the service call delay time for the i-th user. max L is the service latency threshold. i Let i be the login frequency of the i-th user. Let σ be the mean login frequency of all users. L Let be the standard deviation of login frequency, π be the value of pi, sin be the sine function, and tanh be the hyperbolic tangent function. The resource load coefficient is specifically: , in, , Where β1 is the multi-resource coupling overload term, β2 is the equipment pressure balancing term, k1 is the adjustment coefficient, m is the total number of equipment, and λ j M is the baseline resource threshold for the j-th device. j Let D be the memory usage of the j-th device. j G represents the disk I / O throughput of the j-th device. j Let C be the GPU memory usage of the j-th device. j P represents the CPU core temperature of the j-th device. j Let P be the packet loss rate of the network where the j-th device is located. max Let ln be the maximum tolerable packet loss rate, where ln is the logarithm to the base e, and e is the natural constant. The service energy efficiency coefficient is specifically: , in, , , Where γ1 is the bandwidth utilization term, γ2 is the request stability term, γ3 is the retransmission penalty term, and B max B is the system's maximum theoretical bandwidth. p (t) represents the peak network bandwidth at time t, P l (t) represents the packet loss rate at time t, H r (t) represents the HTTP request-response ratio at time t. ref For the ideal HTTP request-response ratio, R c (t) represents the number of TCP retransmissions at time t, k2 is the attenuation coefficient, max is the maximum value function, and e is the natural constant; The cloud platform dynamic equilibrium model is specifically: , Among them, α is the business fluctuation coefficient, β is the resource load coefficient, γ is the service energy efficiency coefficient, δ1, δ2, and δ3 are dynamic weight coefficients, δ1 + δ2 + δ3 = 1 and δ1, δ2, and δ3 ∈ [0, 1]; S4. The intelligent decision - making optimization module triggers the resource scheduling strategy according to the platform health index; S5. The dynamic feedback optimization module optimizes and standardizes the dynamic weight coefficient based on the historical data of the platform health index collected within a preset time window and its corresponding business fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient.
Citation Information
Patent Citations
Method and device for acquiring health index of business system on cloud platform, and electronic equipment
CN115277474A
Cloud platform computing power resource performance monitoring and real-time scheduling optimization method
CN119883651A