Big data cloud platform management system and method based on cloud computing
By intelligently collecting multi-source data and analyzing multi-dimensional data, a platform health index is generated, which solves the problems of low resource utilization efficiency and insufficient scheduling strategies in traditional cloud platform management, and realizes dynamic balance and efficient resource scheduling of the cloud platform.
Patent Information
- Application Number
- CN202511486826.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-17
AI Technical Summary
Traditional big data cloud platform management technologies suffer from low resource utilization efficiency, insufficient accuracy of resource scheduling strategies, and difficulty in adapting to complex scenarios such as sudden changes in user behavior, equipment performance degradation, and abnormal network traffic, thus failing to meet the requirements of high availability and high elasticity.
By employing a multi-source data intelligent acquisition module, a data cleaning and standardization module, a multi-dimensional data analysis module, and an intelligent decision optimization module, the platform health index is generated by calculating the business fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient, thereby achieving dynamic balancing and resource scheduling of the cloud platform.
It enables multi-dimensional quantitative assessment of the cloud platform's operational status, improves resource utilization efficiency and platform responsiveness, enhances model flexibility and accuracy, and adapts to changes in different business scenarios.
Smart Images

Figure CN121000720A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of cloud computing, and particularly relates to a big data cloud platform management system and method based on cloud computing. BACKGROUND
[0002] With the rapid development of information technology, the application of cloud computing and big data technology has become the core driving force for promoting the digital transformation of various industries. More and more enterprises and institutions rely on cloud platforms to deploy business and process data. Massive Internet of Things devices, business systems and cloud logs continuously generate large amounts of structured and unstructured data, which poses unprecedented challenges to the resource management, data processing and service scheduling capabilities of cloud platforms.
[0003] At present, the traditional big data cloud platform management technology mostly adopts a static resource allocation mode, relies on preset rules for load scheduling, and is difficult to real-time perceive the dynamic changes of business fluctuations, resource loads and service energy efficiency. Although some systems introduce a basic load balancing mechanism, in the face of complex scenarios such as sudden changes in user behavior, device performance degradation or network traffic anomalies, they lack the ability to analyze multi-dimensional data in coordination and cannot achieve intelligent dynamic optimization of resources.
[0004] The prior art has exposed significant defects in actual application: on the one hand, the static resource scheduling strategy leads to low resource utilization efficiency, and often some devices are overloaded while others are idle, which seriously affects the platform response speed; on the other hand, the traditional model lacks a dynamic weight adjustment mechanism for evaluating the health status of the platform, and is difficult to adapt to the importance differences of business fluctuations, resource loads and service energy efficiency in different business scenarios, resulting in insufficient precision of resource scheduling strategy and inability to meet the needs of modern cloud platforms for high availability and high elasticity.
[0005] Therefore, it is necessary to invent a big data cloud platform management system and method based on cloud computing to solve the above problems. SUMMARY
[0006] The purpose of the present application is to provide a big data cloud platform management system and method based on cloud computing to solve the problems raised in the background.
[0007] To achieve the above purpose, the present application provides the following technical solution: a big data cloud platform management system based on cloud computing, comprising the following modules: A multi-source data intelligent acquisition module is used to acquire structured and unstructured data in Internet of Things devices, business system interfaces and cloud logs through a distributed sensor network to form a cloud real-time data resource pool; the cloud real-time data resource pool includes a user behavior feature data set, a device state monitoring data set and a network traffic feature data set; A data cleaning and standardization module is configured to perform outlier elimination, missing value interpolation, and dimension normalization on the cloud real-time data resource pool to generate a standardized data resource pool; A multi-dimensional data analysis module is configured to perform time series decomposition, clustering analysis, and association rule mining on the standardized data resource pool to obtain a business fluctuation coefficient, a resource load coefficient, and a service energy efficiency coefficient, and input the business fluctuation coefficient, the resource load coefficient, and the service energy efficiency coefficient into a cloud platform dynamic balancing model to output a platform health index; The business fluctuation coefficient is specifically: , wherein, , wherein, α1 is a delay sensitive term, α2 is a behavior dispersion term, ω 1i , ω 2i are weight factors, n is the total number of users, S i is the service call delay time of the i-th user, S max is a service delay threshold, L i is the login frequency of the i-th user, is the average value of the login frequencies of all users, σ L is the standard deviation of the login frequencies, π is a circular constant, sin is a sine function, and tanh is a hyperbolic tangent function. The resource load coefficient is specifically: , wherein, , wherein, β1 is a multi-resource coupling overload term, β2 is a device pressure balancing term, k1 is an adjustment coefficient, m is the total number of devices, λ j is the baseline resource threshold of the j-th device, M j is the memory usage rate of the j-th device, D j is the disk I / O throughput of the j-th device, G j is the GPU memory occupancy rate of the j-th device, C j is the CPU core temperature of the j-th device, P j is the packet loss rate of the network where the j-th device is located, P max is the maximum tolerable packet loss rate, ln is the logarithm with base e, and e is a natural constant. The service energy efficiency coefficient is specifically: , wherein, , , wherein, γ1 is a bandwidth utilization term, γ2 is a request stability term, γ3 is a retransmission penalty term, Bmax B is the system's maximum theoretical bandwidth. p (t) represents the peak network bandwidth at time t, P l (t) represents the packet loss rate at time t, H r (t) represents the HTTP request-response ratio at time t. ref For the ideal HTTP request-response ratio, R c (t) represents the number of TCP retransmissions at time t, k2 is the attenuation coefficient, max is the maximum value function, and e is the natural constant; The cloud platform dynamic balancing model is as follows: , Where α is the business fluctuation coefficient, β is the resource load coefficient, γ is the service energy efficiency coefficient, δ1, δ2 and δ3 are dynamic weight coefficients, δ1+δ2+δ3=1 and δ1, δ2 and δ3∈[0,1]; The intelligent decision optimization module is used to trigger resource scheduling strategies based on the platform's health index; The dynamic feedback optimization module optimizes and standardizes the dynamic weighting coefficients based on historical platform health index data collected within a preset time window and their corresponding business fluctuation coefficients, resource load coefficients, and service efficiency coefficients.
[0008] Preferably, the user behavior feature dataset includes the user's login frequency and the user's service call latency; the device status monitoring dataset includes CPU core temperature, memory usage, disk I / O throughput, and GPU memory usage; and the network traffic feature dataset includes network bandwidth peak, packet loss rate, TCP retransmission count, and HTTP request-response ratio.
[0009] Preferably, the adjustment coefficient k1 is dynamically set according to the total number of devices m: When m < 100, then k1 ∈ [0.8, 1.2]; When m≥100, then k1∈[0.3,0.7].
[0010] Preferably, the initial values of the dynamic weighting coefficients δ1, δ2, and δ3 are set according to the business scenario: During peak business periods, the initial values of δ1, δ2, and δ3 are set as follows: δ1 = 0.5, δ2 = 0.3, and δ3 = 0.2. During periods of resource scarcity, the initial values of δ1, δ2, and δ3 are set as δ1=0.2, δ2=0.5, and δ3=0.3. If it is a service-sensitive period, the initial values of δ1, δ2 and δ3 are set as δ1=0.2, δ2=0.3 and δ3=0.5; If it is a regular period, the initial values of δ1, δ2 and δ3 are set as δ1=0.4, δ2=0.3 and δ3=0.3.
[0011] Preferably, the resource scheduling strategy is: If the platform health index 0≤E<E1, a first-level scheduling strategy is triggered, and an emergency degradation strategy is started; If the platform health index E1≤E<E2, a second-level scheduling strategy is triggered, and a load balancing strategy is started; If the platform health index E2≤E<E3, a third-level scheduling strategy is triggered, and an elastic expansion strategy is started; If the platform health index E≥E3, a fourth-level scheduling strategy is triggered, and only monitoring is performed without scheduling.
[0012] Preferably, the execution process of the dynamic feedback optimization module is: A1, obtaining a service fluctuation coefficient α of a current time t t , a resource load coefficient β t and a service energy efficiency coefficient γ t ; A2, calculating a synergy entropy value of the three-dimensional coefficient: wherein b is a minimum constant for preventing the denominator from being zero; A3, updating the dynamic weight coefficient according to the entropy value state: wherein τ is an entropy sensitive adjustment factor, is an old weight, is an old weight, θ k ∈{α, β, γ}, k=1, 2, 3; A4, performing constraint normalization on the weight coefficient: , wherein, , is a temporary weight, is a maximum weight.
[0013] A big data cloud platform management method based on cloud computing, specifically comprising the following steps: S1, a multi-source data intelligent acquisition module acquires structured and unstructured data in Internet of Things devices, business system interfaces and cloud logs through a distributed sensor network, forming a cloud real-time data resource pool; the cloud real-time data resource pool includes a user behavior feature data set, a device state monitoring data set and a network traffic feature data set; S2, a data cleaning and standardization module performs outlier rejection, missing value interpolation and dimensionless normalization processing on the cloud real-time data resource pool, generating a standardized data resource pool; S3, the multi-dimensional data analysis module performs time series decomposition, cluster analysis and association rule mining on the standardized data resource pool, obtains business fluctuation coefficient, resource load coefficient and service energy efficiency coefficient, and inputs the business fluctuation coefficient, resource load coefficient and service energy efficiency coefficient into the cloud platform dynamic balance model, and outputs the platform health index; S4, the intelligent decision optimization module triggers resource scheduling strategy according to the platform health index; S5, the dynamic feedback optimization module optimizes and standardizes the dynamic weight coefficient based on the platform health index historical data and the corresponding business fluctuation coefficient, resource load coefficient and service energy efficiency coefficient collected in the preset time window.
[0014] The technical effects and advantages of the present application are: The multi-dimensional data analysis module of the present application performs time series decomposition, cluster analysis and association rule mining on the standardized data resource pool, calculates the business fluctuation coefficient, resource load coefficient and service energy efficiency coefficient, and inputs the platform health index output by the cloud platform dynamic balance model, realizes the multi-dimensional quantitative evaluation of the cloud platform running state, and accurately reflects the platform health status. The intelligent decision optimization module of the present application triggers different levels of resource scheduling strategy according to the platform health index, including emergency degradation, load balancing and elastic expansion, realizes the dynamic intelligent scheduling of the cloud platform resources, and improves the resource utilization efficiency and platform response ability. The dynamic feedback optimization module of the present application optimizes and standardizes the dynamic weight coefficient based on the platform health index historical data and the corresponding coefficient in the preset time window, so that the weight configuration of the cloud platform dynamic balance model can adapt to the business scene change, and the flexibility and accuracy of the model are improved. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 The system module connection diagram of the present application.
[0016] Figure 2 The execution process flow diagram of the dynamic feedback optimization module of the present application.
[0017] Figure 3 The method step flow diagram of the present application. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described clearly and completely in the embodiments of the present application combined with the drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0019] The application provides a cloud computing-based big data cloud platform management system as shown in the accompanying drawings. Figure 1 The cloud computing-based big data cloud platform management system comprises the following modules. A multi-source data intelligent collection module is used to collect structured and unstructured data in Internet of Things devices, business system interfaces and cloud logs through a distributed sensor network to form a cloud real-time data resource pool; the cloud real-time data resource pool comprises a user behavior feature data set, a device state monitoring data set and a network traffic feature data set. Further, in the above technical solution, the user behavior feature data set comprises a user login frequency and a user service call delay time; the device state monitoring data set comprises a CPU core temperature, a memory usage rate, a disk I / O throughput and a GPU video memory occupancy rate; and the network traffic feature data set comprises a network bandwidth peak value, a packet loss rate, a TCP retransmission number and an HTTP request response ratio.
[0020] It should be noted that the user login frequency is obtained from login operation records of a user account through a business system interface or a cloud log, specifically, a login event of a user authentication system is listened to by using an API interface, or identity authentication logs, such as Apache or Nginx logs, user center database logs, are parsed to aggregate and statistically obtain the login times within a unit time according to a user ID; The user service call delay time is collected from response time monitoring of a business system interface, a probe, such as a distributed tracking tool of OpenTelemetry, Zipkin, etc., is implanted in a service interface call link to record a time difference between request sending and response receiving, or a response time consumption of an HTTP or API request is parsed through a log to obtain the delay time; The CPU core temperature is collected by a hardware sensor of an Internet of Things device, such as a server or an edge node, and CPU temperature sensor data is read by means of a hardware management interface, such as IPMI, WMI, or a system tool, such as lm-sensors of Linux, to obtain the temperature of each core in real time; The memory usage rate is collected from memory state information of a device operating system, and a ratio of a used capacity to a total capacity of physical memory and virtual memory is obtained by means of a system API, such as a psutil library of Python or ManagementFactory of Java, to calculate the memory usage rate; The disk I / O throughput is obtained by means of I / O operation monitoring of a storage device, and a data amount of disk reading and writing within a unit time is counted by means of an operating system tool, such as iostat of Linux or perfmon of Windows, or a storage device management interface to obtain the throughput; The GPU memory occupancy rate is collected from the memory usage state of GPU hardware, and the ratio of used memory capacity to total capacity is obtained through the driving interface provided by the GPU manufacturer, such as nvidia-smi of NVIDIA, rocm-smi of AMD, or the GPU monitoring API of a deep learning framework such as TensorFlow; The network bandwidth peak is collected by monitoring the traffic of network devices such as routers and switches, and the real-time bandwidth data of network interfaces is collected by means of SNMP protocol or traffic monitoring tools such as Prometheus combined with Node Exporter, and the peak traffic in a unit of time is counted; The packet loss rate is collected based on the data packet transmission state of network devices, and the input and output packet numbers of network interfaces are analyzed by means of SNMP or network packet capture tools such as Wireshark to calculate the ratio of lost packets to total transmitted packets to obtain the packet loss rate; The TCP retransmission number is collected from the TCP connection state of the network protocol stack, and the retransmission number is counted by capturing TCP retransmission flags such as repeated ACK and timeout retransmission through the network statistics interface of the operating system such as netstat-s of Linux or network analysis tools such as tcpdump; The HTTP request response ratio is collected through the access log of the Web server or the monitoring data of the reverse proxy, the HTTP server log such as access.log of Nginx is parsed, the ratio of the number of successful responses to the total number of requests in a unit of time is counted, or the proportion of HTTP response status codes such as 2xx, 3xx success codes and 4xx, 5xx error codes is monitored through APM tools such as New Relic.
[0021] The data cleaning and standardization module is used for outlier removal, missing value interpolation and dimensionless normalization processing of the cloud real-time data resource pool to generate a standardized data resource pool; It should be noted that the execution process of the data cleaning and standardization module is as follows: for the login frequency of the user, first calculate the mean and standard deviation of all user login frequencies, and determine the values exceeding ±3 times the mean as outliers and remove them, for the missing login frequency data in the log record, use the mean of the historical login frequency of the user for interpolation, and finally linearly transform the login frequency data to the [0, 1] interval by the Max-Min standardization method to eliminate the influence of dimension and generate standardized user login frequency data; For the service call delay time of the user, the abnormal delay time outside 1.5 times of the interquartile range of the upper and lower quartiles is first identified and removed by the box plot method, for the missing delay time data, linear interpolation method is used to interpolate according to the data before and after the time sequence, and then the delay time is converted into standard normal distribution data with mean of 0 and standard deviation of 1 by Z-score standardization method, realizing dimensionless normalization; For the CPU core temperature, first set the hardware safety temperature threshold, and determine the temperature value exceeding the threshold as an abnormal value and replace it with the threshold, for the missing temperature data, use the mean of other core temperatures of the same device to interpolate, and finally linearly map the temperature data to the [0, 1] interval, where 0 corresponds to the lowest safety temperature and 1 corresponds to the highest safety temperature, completing dimensionless normalization; For the memory usage, first truncate the abnormal values exceeding 100% or less than 0 to 100% or 0, for the missing memory usage data, use the mean of the memory usage of the same device in the same time period to interpolate, and then convert the memory usage to standardized data in the [0, 1] interval by Max-Min standardization, which is convenient for subsequent analysis; For the disk I / O throughput, first remove the abnormal high or low throughput data in the sliding window median filter, for the missing throughput data, use forward filling method to interpolate, and then after logarithmic transformation of the data, Max-Min standardization to [0, 1] interval to cope with the characteristics of large span of throughput data; For the GPU memory occupancy rate, first truncate the abnormal values exceeding 0-100% range to 0 or 100%, for the missing memory occupancy rate data, use the mean of the memory occupancy rate under similar tasks of the same device to interpolate, and finally linearly transform the data to [0, 1] interval to realize dimensionless normalization; For the network bandwidth peak value, first determine the peak value exceeding the maximum theoretical bandwidth of the network device as an abnormal value and replace it with the maximum theoretical bandwidth value, for the missing bandwidth peak value data, use the mean of the bandwidth peak value at adjacent time points to interpolate, and then divide the bandwidth peak value data by the maximum theoretical bandwidth to get standardized data in the [0, 1] interval; For the packet loss rate, first truncate the abnormal packet loss rate exceeding 1 or less than 0 to 1 or 0, for the missing packet loss rate data, interpolate according to the mean of the packet loss rate in similar periods of network load, since the packet loss rate itself is in the [0, 1] interval, directly keep the data range; For the TCP retransmission number, first calculate the upper quartile of the retransmission number, and determine the data exceeding 1.5 times of the upper quartile as an abnormal value and eliminate it, for the missing retransmission number data, use the mean value of the same network connection history retransmission number to interpolate, and finally convert the retransmission number to standardized data in the interval [0, 1] through Max-Min standardization; For the HTTP request response ratio, first truncate the abnormal response ratio exceeding the range of 0 to 1 to 0 or 1, for the missing response ratio data, use the mean value of the response ratio of the same type of service in the same time period to interpolate, since the response ratio itself is a proportion value, directly standardize the processing to ensure that the data is in the interval [0, 1] uniform dimension.
[0022] The multi-dimensional data analysis module is used for time series decomposition, cluster analysis and association rule mining on the standardized data resource pool, to obtain business fluctuation coefficient, resource load coefficient and service energy efficiency coefficient, and input the business fluctuation coefficient, resource load coefficient and service energy efficiency coefficient into the cloud platform dynamic balance model to output the platform health index; The business fluctuation coefficient is specifically: , Wherein, , Wherein, α1 is a delay sensitive term, α2 is a behavior dispersion term, ω 1i , ω 2i is a weight factor, n is the total number of users, S i is the service call delay time of the i-th user, S max is the service delay threshold, L i is the login frequency of the i-th user, is the mean value of the login frequency of all users, σ L is the standard deviation of the login frequency, π is the circular constant, sin is the sine function, and tanh is the hyperbolic tangent function.
[0023] It should be noted that the core design goal of the business fluctuation coefficient α is to construct an intelligent index quantifying system stability by comprehensively evaluating the burstiness of service delay and the discreteness of user behavior, wherein the delay sensitive term α1 maps the ratio of the service call delay time S i of the user i to the preset threshold S max to the nonlinear periodic interval of the sine function, when S i approaches S max , α1 approaches the peak value 1, thereby amplifying the delay critical risk, and when the function value is flat at low delay, the small fluctuation is avoided to interfere with the judgment, and the service exception is accurately focused; the behavior dispersion term α2 calculates the login frequency Li of the user i relative to the overall mean value degree of deviation, and compresses the discrete values to the range of (0, 1) using the saturation characteristics of the hyperbolic tangent function tanh, which not only identifies user abnormal behaviors such as high-frequency abnormal login, but also suppresses the excessive influence of extreme outliers on the results, ensuring the robustness of the group behavior analysis. The results of the two calculations are respectively adjusted by configurable weight factors ω 1i 、ω 2i to adjust the contribution degree and support differentiated allocation of weights according to user values or business scenarios, such as assigning higher ω 1i to VIP users to strictly control the delay. The formula finally integrates all user data in a normalized framework: the numerator is the sum of the weighted delay and the discrete signal, and the denominator is the sum of the weight factors, which eliminates the scale effect of the total number of users n; then 1 is subtracted from the ratio, so that the output value α∈[0, 1] directly maps the business stability, where α tends to 1 for the best stability.
[0024] The resource load coefficient is specifically: , wherein, , wherein, β1 is a multi-resource coupling overload term, β2 is a device pressure balancing term, k1 is an adjustment coefficient, m is the total number of devices, λ j is the baseline resource threshold of the jth device, M j is the memory usage rate of the jth device, D j is the disk I / O throughput of the jth device, G j is the GPU memory occupancy rate of the jth device, C j is the CPU core temperature of the jth device, P j is the packet loss rate of the network where the jth device is located, P max is the maximum tolerable packet loss rate, ln is the logarithm with base e, and e is the natural constant.
[0025] Further, in the above technical solution, the adjustment coefficient k1 is dynamically set according to the total number of devices m: if it is a small or medium-sized system, i.e. the total number of devices is less than 100, then k1 can be set to between 0.8 and 1.2; if it is a large-scale system, i.e. the total number of devices exceeds 100, then k1 can be set to between 0.3 and 0.7; It should be noted that the core design goal of the resource load coefficient β is to construct a comprehensive index for quantifying the system overload risk through two-dimensional coupling analysis of hardware resource conflicts and network quality, and the subtraction framework of the formula maps the output to the interval [0, 1]: when β tends to 0, it represents a high load risk, and when β tends to 1, it represents a healthy load, wherein the exponential decay structure of the multi-resource coupling overload term β1 reveals the linkage amplification effect of hardware resources, and the numerator end memory usage rate M j , disk I / O throughput D jGPU memory occupancy rate G j The continuous multiplication design deliberately strengthens the resource competition relationship, and the product of the denominator end baseline resource threshold λ j and the CPU core temperature C j constitutes a dynamic fault tolerance mechanism: C j increases the denominator to weaken the harmfulness of other resource exceptions, and λ j supports differentiated calibration of the safety boundary according to the performance baseline difference of the device, and the inner layer compresses the order of magnitude of the resource conflict value to avoid extreme values from dominating, and the outer layer exponential function controls the aggregation weight of the conflict signal through the adjustment coefficient k1, and finally outputs a smooth decay curve conforming to the principle of entropy increase; The geometric mean structure of the device pressure balancing term β2 focuses on the network short board effect: Converts the packet loss rate of device j into the occupancy ratio of the relative maximum tolerance threshold P max The continuous multiplication and mth root structure makes β2 extremely sensitive to single-point faults, and when all devices P j are balanced below P max , β2 tends to 1, and non-uniform distribution outputs intermediate values to guide and optimize high packet loss nodes; The product of the two terms β1×β2 contains deep business logic: hardware overload risk and network risk are coupled through multiplication to form synergistic amplification, and β only tends to 0 when both are simultaneously deteriorated.
[0026] The service energy efficiency coefficient is specifically: , Wherein, , , , wherein γ1 is the bandwidth utilization term, γ2 is the request stability term, γ3 is the retransmission penalty term, B max is the maximum theoretical bandwidth of the system, B p (t) is the network bandwidth peak value at time t, P l (t) is the packet loss rate at time t, H r (t) is the HTTP request response ratio at time t, H ref is the ideal HTTP request response ratio, R c (t) is the number of TCP retransmissions at time t, k2 is the decay coefficient, max is the maximum value function, and e is the natural constant.
[0027] It should be noted that the value of the decay coefficient k2 can be set as follows: when the packet loss rate is low, i.e. the packet loss rate is less than 0.1%, in order to quickly punish retransmission, the value of k2 can be set to between 0.5 and 0.7; when the packet loss rate is high, i.e. the packet loss rate exceeds 1%, in order to slow down the punishment and avoid misjudgment, the value of k2 can be set to between 0.1 and 0.3; It is to be understood that the core design of the service energy efficiency coefficient γ is to force the three dimensions of bandwidth utilization efficiency, request processing stability and transmission reliability to be optimized through the multiplication series framework, and the decline of any sub-item will lead to the collapse of the overall energy efficiency. The structure of the bandwidth utilization item γ1 has a double constraint mechanism: the numerator B p (t) measures the actual channel capacity, and 1-P l (t) punishes P l (t) causes loss of data integrity, and the denominator B max The design realizes normalization processing, so that γ1 only approaches 1 when the bandwidth is efficiently utilized, that is, B p (t) approaches B max , and P l (t) approaches 0; the request stability item γ2 is a segmented function that realizes bidirectional risk suppression through the dynamic comparison of H r (t) and H ref : when H r (t)≤H ref , γ2=1 allows fluctuations within the benchmark, and when H r (t)>H ref , the punishment item quantifies the response delay risk caused by overloaded requests as a decreasing function; in the exponential decay model of the retransmission punishment item γ3, R c (t) directly maps the transmission layer failure frequency, and k2 controls the punishment intensity: the characteristics of the exponential function e -x cause the first retransmission to trigger a severe punishment, and the marginal effect of subsequent retransmissions decreases, accurately simulating the avalanche effect of network congestion; the final product contains deep business logic: the three dimensions form a series of rigid constraints, such as when the bandwidth is fully loaded without packet loss and zero retransmission, if the request response efficiency decreases by 15%, the overall γ will drop to 0.85, and the output value can quickly locate the bottleneck: the decrease of γ1 indicates a failure in bandwidth planning, the decrease of γ2 reveals insufficient computing resources, and the collapse of γ3 exposes network link faults.
[0028] The cloud platform dynamic balancing model is specifically: , wherein α is a business fluctuation coefficient, β is a resource load coefficient, γ is a service energy efficiency coefficient, δ1, δ2 and δ3 are dynamic weight coefficients, δ1+δ2+δ3=1 and δ1, δ2 and δ3∈[0, 1].
[0029] Further, in the above technical solution, the initial values of the dynamic weight coefficients δ1, δ2 and δ3 are set according to the business scenario: if it is a business peak period, the initial values of δ1, δ2 and δ3 are set as δ1=0.5, δ2=0.3 and δ3=0.2; if it is a resource shortage period, the initial values of δ1, δ2 and δ3 are set as δ1=0.2, δ2=0.5 and δ3=0.3; if it is a service sensitive period, the initial values of δ1, δ2 and δ3 are set as δ1=0.2, δ2=0.3 and δ3=0.5; and if it is a regular period, the initial values of δ1, δ2 and δ3 are set as δ1=0.4, δ2=0.3 and δ3=0.3.
[0030] an intelligent decision optimization module configured to trigger a resource scheduling strategy according to the platform health index; Further, in the above technical solution, the resource scheduling strategy is: if the platform health index 0≤E<E1, a first-level scheduling strategy is triggered to start an emergency degradation strategy; if the platform health index E1≤E<E2, a second-level scheduling strategy is triggered to start a load balancing strategy; if the platform health index E2≤E<E3, a third-level scheduling strategy is triggered to start an elastic expansion strategy; if the platform health index E≥E3, a fourth-level scheduling strategy is triggered to only monitor without scheduling.
[0031] It should be noted that the initial values of E1, E2 and E3 are set as E1=0.5, E2=0.7 and E3=0.85, and are calibrated in real time according to historical operation data: , wherein, wherein, is a new threshold value adjusted at the current moment, is an old threshold value at the last moment, η is a learning rate, the default value is 0.01, T is a time window, the default value is 24 hours, I is an indication function, which is 1 when the condition is met, otherwise it is 0, and α t is a business fluctuation coefficient at moment t, β t is a resource load coefficient at moment t, E k is a threshold initial value, wherein k=1, 2, 3.
[0032] It should be noted that the emergency degradation strategy includes core service fusing: forcibly closing non-critical services such as data analysis and log auditing, and only retaining core function modules such as user authentication and payment transactions, to reduce resource consumption through business degradation; traffic interception: injecting a fusing rule at the load balancing layer to reject new requests or return a preset static page such as HTTP503, to avoid the snowball effect; Resource intensive recovery: terminate low-priority processes such as batch tasks, release CPU and memory resources, and perform hardware frequency limiting on high-temperature devices, i.e., C j devices perform load migration or hardware frequency limiting when the threshold is exceeded; Dynamic weight freezing: suspend weight adjustment of the dynamic feedback optimization module, lock δ k For safety value, prevent policy shock, where k = 1, 2, 3; The load balancing strategy includes multi-dimensional routing: combine resource load coefficient β to dynamically distribute requests: nodes with high memory usage, i.e., memory usage exceeding 80%, receive compute-intensive tasks; nodes with low disk I / O throughput, i.e., disk I / O throughput below 50MB / s, handle streaming data collection; based on service energy efficiency coefficient γ, preferentially route requests with HTTP request response ratio exceeding 2 to network areas with packet loss rate below 0.1%; Resource rebalancing: migrate unbalanced loads in device state monitoring datasets in real time, such as tasks between devices with GPU memory usage difference exceeding 30%, and automatically identify resource hotspot clusters using cluster analysis; Elastic preheating: pre-start standby containers, accounting for 5% of the total, but do not connect to traffic, ensuring seamless switching to expansion strategy when E2 ≤ E; The elastic expansion strategy includes horizontal expansion: when the service fluctuation coefficient α exceeds 0.7, expand login authentication nodes based on user behavior feature dataset: , where k is the elasticity coefficient, default 1.2; when resource load coefficient β is below 0.4, double the high-load devices, such as devices with memory usage exceeding 90% by 200%; Vertical optimization: when network traffic feature dataset shows that bandwidth peak exceeds 80% of B max , upgrade network interface to redundant link; when service energy efficiency coefficient γ is below 0.6, add GPU resources to compute nodes; Cost constraint: if E ≥ E3 for 3 consecutive time windows after expansion, automatically scale down to baseline size to avoid resource idling; Dynamic feedback optimization module: based on platform health index historical data collected within the preset time window and its corresponding service fluctuation coefficient, resource load coefficient, and service energy efficiency coefficient, optimize and standardize dynamic weight coefficient.
[0033] Further, in the above technical solution, the execution process of the dynamic feedback optimization module is as shown in Figure 2 : A1, obtain service fluctuation coefficient α t , resource load coefficient β t , and service energy efficiency coefficient γ t at current time t; A2, calculate the collaborative entropy value of the three-dimensional coefficient: Wherein, b is a minimum constant to prevent the denominator from being zero; A3, update the dynamic weight coefficient according to the entropy value state: Wherein, τ is an entropy sensitive adjustment factor, is the old weight, is the old weight, θ k ∈{α,β,γ},k=1,2,3; A4, constraint normalization of weight coefficient: , Wherein, , is a temporary weight, is the maximum weight.
[0034] It should be known that the default value of b is 10 -6 ; The default value of the entropy sensitive adjustment factor τ is 0.05, which can be set according to different states: if it is a stable state, τ can be set to between 0.08 and 0.12; if it is a fluctuation state, τ can be set to between 0.01 and 0.03.
[0035] The present application provides a kind of big data cloud platform management method based on cloud computing as Figure 3 Shown in the figure, specifically comprising the following steps: S1, multi-source data intelligent acquisition module is collected structured and unstructured data in Internet of Things equipment, business system interface and cloud log through distributed sensor network, forms cloud real-time data resource pool;The cloud real-time data resource pool includes user behavior characteristic data set, equipment state monitoring data set and network traffic characteristic data set; S2, data cleaning and standardization module carries out outlier rejection, missing value interpolation and dimensionless normalization processing to the cloud real-time data resource pool, generates standardization data resource pool; S3, multi-dimensional data analysis module carries out time series decomposition, clustering analysis and association rule mining to the standardization data resource pool, obtains business fluctuation coefficient, resource load coefficient and service energy efficiency coefficient, and inputs business fluctuation coefficient, resource load coefficient and service energy efficiency coefficient into cloud platform dynamic balance model, exports platform health index; S4, intelligent decision optimization module triggers resource scheduling strategy according to the platform health index; S5, dynamic feedback optimization module optimizes and standardizes dynamic weight coefficient based on platform health index historical data and its corresponding business fluctuation coefficient, resource load coefficient, service energy efficiency coefficient collected in preset time window.
[0036] It should be pointed out finally that the above only describes the preferred embodiments of the present application and is not intended to limit the present application. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent replacements to some of the technical features, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A cloud computing-based big data cloud platform management system, characterized in that, Comprise the following modules: Multi-source data intelligent acquisition module, for through distributed sensor network in the Internet of Things device, business system interface and cloud log collection get structured and unstructured data, form cloud real-time data resource pool; The cloud real-time data resource pool includes user behavior characteristic data set, equipment state monitoring data set and network traffic characteristic data set; Data cleaning and standardization module, for the cloud real-time data resource pool is excluded, missing value interpolation and dimensionless processing, generate standardization data resource pool; Multi-dimensional data analysis module, for the standardization data resource pool is time decomposition, cluster analysis and association rule mining, get business fluctuation coefficient, resource load coefficient and service energy efficiency coefficient, and business fluctuation coefficient, resource load coefficient and service energy efficiency coefficient input cloud platform dynamic balance model, output platform health index; The business fluctuation coefficient, specifically: , wherein, , wherein, a1 is the delay sensitive term, a2 is the behavior dispersion term, ω 1i , ω 2i is the weight factor, n is the total number of users, S i is the service invocation delay time of the i-th user, S max is the service delay threshold, L i is the login frequency of the i-th user, is the mean of all users' login frequency, σ L is the standard deviation of login frequency, π is the circular constant, sin is the sine function, and tanh is the hyperbolic tangent function. The resource load coefficient, specifically: , wherein, , wherein, β1 is a multi-resource coupling overload term, β2 is a device pressure balancing term, k1 is an adjustment coefficient, m is the total number of devices, λ j is the baseline resource threshold of the jth device, M j is the memory usage rate of the jth device, D j is the disk I / O throughput of the jth device, G j is the GPU memory occupancy rate of the jth device, C j is the CPU core temperature of the jth device, P j is the packet loss rate of the network where the jth device is located, P max is the maximum tolerable packet loss rate, ln is the logarithm with base e, and e is the natural constant; The service energy efficiency coefficient, specifically: , wherein, , , wherein, γ1 is a bandwidth utilization term, γ2 is a request stability term, γ3 is a retransmission penalty term, B max is the maximum theoretical bandwidth of the system, B p (t) is the peak network bandwidth at time t, P l (t) is the packet loss rate at time t, H r (t) is the HTTP request response ratio at time t, H ref is the ideal HTTP request response ratio, R c (t) is the number of TCP retransmissions at time t, k2 is a decay coefficient, max is a maximum function, and e is the natural constant; The cloud platform dynamic balance model, specifically: , Wherein, α is business fluctuation coefficient, β is resource load coefficient, γ is service energy efficiency coefficient, δ1, δ2 and δ3 are dynamic weight coefficients, δ1+δ2+δ3=1 and δ1, δ2 and δ3∈[0, 1]; Intelligent decision optimization module, for triggering resource scheduling strategy according to the platform health index; Dynamic feedback optimization module, based on the platform health index historical data and its corresponding business fluctuation coefficient, resource load coefficient, service energy efficiency coefficient collected in the preset time window, optimize and standardize the dynamic weight coefficient. 2.The big data cloud platform management system based on cloud computing according to claim 1, wherein, The user behavior characteristic data set includes user login frequency and user service call delay time;The equipment state monitoring data set includes CPU core temperature, memory usage, disk I / O throughput and GPU video memory occupancy;The network traffic characteristic data set includes network bandwidth peak, packet loss rate, TCP retransmission times and HTTP request response ratio. 3.The cloud computing-based big data cloud platform management system of claim 1, wherein, The adjustment coefficient k1 is dynamically set according to the total number of devices m: When m<100, then k1∈[0.8, 1.2]; When m≥100, then k1∈[0.3, 0.7].
4. The big data cloud platform management system based on cloud computing according to claim 1, characterized in that, The initial value of the dynamic weight coefficient δ1, δ2 and δ3 is set according to the business scenario: If it is business peak period, the initial value of δ1, δ2 and δ3 is set as δ1=0.5, δ2=0.3, δ3=0.2; If it is resource shortage period, the initial value of δ1, δ2 and δ3 is set as δ1=0.2, δ2=0.5, δ3=0.3; If it is service sensitive period, the initial value of δ1, δ2 and δ3 is set as δ1=0.2, δ2=0.3, δ3=0.5; If it is a regular period, the initial value of δ1, δ2 and δ3 is set as δ1=0.4, δ2=0.3, δ3=0.
3.
5. The big data cloud platform management system based on cloud computing according to claim 1, characterized in that, The resource scheduling strategy is: If the platform health index 0≤E<E1, trigger level one scheduling strategy, start emergency degradation strategy; If the platform health index E1≤E<E2, trigger level two scheduling strategy, start load balancing strategy; If the platform health index E2≤E<E3, a third-level scheduling strategy is triggered, and an elastic expansion strategy is started; If the platform health index E≥E3, a fourth-level scheduling strategy is triggered, and only monitoring is performed without scheduling. 6.The big data cloud platform management system based on cloud computing according to claim 1, wherein, The execution process of the dynamic feedback optimization module is as follows: A1, obtain service fluctuation coefficient a of current time t t , resource load coefficient β t and service energy efficiency coefficient γ t ; A2, the collaborative entropy value of the three-dimensional coefficient is calculated: wherein b is a minimum constant to prevent the denominator from being zero; A3, updating the dynamic weight coefficient according to the entropy value state; where τ is an entropy sensitivity adjustment factor, is the old weight, is the old weight, θ k ∈ {α, β, γ}, k = 1, 2, 3; A4, constraint normalization of the weight coefficient; , wherein, , is a temporary weight, is a final weight.
7. A cloud computing-based big data cloud platform management method, specifically comprising the following steps: S1, a multi-source data intelligent acquisition module acquires structured and unstructured data in an Internet of Things device, a business system interface and a cloud log through a distributed sensor network to form a cloud real-time data resource pool; the cloud real-time data resource pool includes a user behavior feature data set, a device state monitoring data set and a network traffic feature data set; S2, a data cleaning and standardization module performs outlier rejection, missing value interpolation and dimension normalization processing on the cloud real-time data resource pool to generate a standardized data resource pool; S3, a multi-dimensional data analysis module performs time series decomposition, clustering analysis and association rule mining on the standardized data resource pool to obtain a business fluctuation coefficient, a resource load coefficient and a service energy efficiency coefficient, and inputs the business fluctuation coefficient, the resource load coefficient and the service energy efficiency coefficient into a cloud platform dynamic balancing model to output a platform health index; S4, an intelligent decision optimization module triggers a resource scheduling strategy according to the platform health index; S5, a dynamic feedback optimization module optimizes and standardizes a dynamic weight coefficient based on platform health index historical data and corresponding business fluctuation coefficients, resource load coefficients and service energy efficiency coefficients collected in a preset time window.
Citation Information
Patent Citations
Method for automatically balancing cloud platform resources
CN110543355A
Method and device for acquiring health index of business system on cloud platform, and electronic equipment
CN115277474A
Cloud platform computing power resource performance monitoring and real-time scheduling optimization method
CN119883651A
Automatic load balancing system for cloud servers
DE202024105913U1