Data processing resource dynamic allocation method for data-in-data station
Through the dynamic weight resource allocation algorithm, combined with Kafka and Hadoop systems to process real-time and offline data of the data middle platform, predict future data volume and dynamically adjust resources, solving the problem of rigid resource allocation in Taichung in data and improving resource utilization.
Patent Information
- Application Number
- CN202510872884.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-27
AI Technical Summary
In the prior art, the data middle platform lacks dynamic resource allocation during real-time and offline data processing, resulting in rigid resource allocation and inability to meet future data demands, affecting the system's resource utilization rate.
The dynamic weight resource allocation algorithm is used to process real-time data through the Kafka system, the Hadoop system processes offline data, and predicts future data volume based on the timing analysis model, and dynamically adjusts resource weights to match future resource requirements.
It realizes dynamic differentiated resource allocation for real-time and offline data processing modules, avoids excessive or insufficient resources, improves resource utilization, and solves the resource rigidity problem in traditional static allocation strategies.
Smart Images

Figure CN120371544A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method for dynamically allocating data processing resources for a data middle platform. Background Art
[0002] The data middle platform is the precipitation of the business and data of existing / newly built information systems, and is an intermediate and supporting platform for realizing data empowerment for new businesses and new applications. With the in-depth promotion of the construction of smart power plants, thermal power generation enterprises rely on the data middle platform to integrate multi-source heterogeneous data (including DCS real-time and historical flow data, equipment ledger maintenance and repair data, equipment vibration monitoring, environmental protection CEMS data, etc.) to support the real-time calculation and operation of key algorithm modules such as combustion optimization, equipment fault prediction, and economic analysis. Such modules need to access massive multi-source heterogeneous data sets frequently, and have high real-time performance and strong burstiness. For example, resource competition between different real-time intelligent prediction services, and resource allocation conflicts between existing services and new service model training, all pose severe challenges to the global allocation ability of system resources. Especially in the context of continuous expansion of business scale and continuous improvement of model complexity, the demand for efficient and reasonable allocation of system resources has become increasingly urgent, and the pressure of resource contention has increased significantly.
[0003] Publication No. CN113836235B discloses a data processing method based on a data middle platform and related devices. Real-time data is processed through a Kafka system to obtain real-time data calculation results, and offline data is processed through a Hadoop system. The Kafka system and the Hadoop system are integrated to give full play to their respective advantages while making up for their own disadvantages through other systems.
[0004] However, the above application still has the following problems: Real-time data processing and offline data processing are two core business scenarios. When performing real-time data processing and offline data processing, the above application lacks the allocation of computing resources, while traditional data processing systems usually adopt a static resource allocation strategy, pre-dividing a fixed proportion of computing resources for real-time and offline modules. For example, the queue quota of Hadoop YARN in the Hadoop system, or the resource request limit of Kubernetes. Pre-dividing a fixed proportion of computing resources for real-time and offline modules has the problems of rigid resource allocation and inability to perform forward-looking resource scheduling based on future data volume prediction. Summary of the Invention
[0005] To solve the technical problems in the background art, the present invention proposes a method for dynamically allocating data processing resources for a data middle platform.
[0006] A method for dynamically allocating data processing resources for a data middle platform proposed by the present invention is characterized by including: Step S1: Process the real-time data of the data middle platform through the Kafka system and generate the real-time data processing result; Step S2: Process the offline data of the data middle platform through the Hadoop system and generate the offline data processing result; Step S3: Based on the time series analysis model, predict the data volumes of the real-time data processing result and the offline data processing result in the future unit time, and output the real-time data volume prediction value and the offline data volume prediction value; Step S4: Use the dynamic weight resource allocation algorithm to allocate the computing resource amount for the real-time data volume prediction value and the offline data volume prediction value.
[0007] Further, the process of outputting the real-time data volume prediction value and the offline data volume prediction value includes: Judge the stationarity of the real-time data processing result and the offline data processing result by using the time series diagram, autocorrelation diagram and statistical test method for the real-time data processing result and the offline data processing result; If it is stationary, the ARIMA model is used for the real-time data processing result or the offline data processing result to generate the corresponding real-time data volume prediction value in the future unit time and the offline data volume prediction value in the offline data processing module; If it is non-stationary, the real-time data processing result or the offline data processing result is subjected to differencing processing, and then the ARIMA model is used for fitting. The ARIMA model is used to generate the corresponding real-time data volume prediction value in the future unit time and the offline data volume prediction value in the offline data processing module.
[0008] Further, the process of the dynamic weight resource allocation algorithm includes: Set the data complexity, the current system load and the response speed requirement corresponding to the real-time data volume prediction value or the offline data volume prediction value as the influencing factors for resource allocation, and obtain the resource weight through calculation; allocate the computing resource amount according to the resource weight.
[0009] Further, the process of obtaining the resource weight includes: Let f1 be the real-time data volume prediction value or the offline data volume prediction value; obtain the data type of f1, marked as {image; video; text; numerical value; BP neural network; LSTM neural network; MSET structure; office file; other file}; and obtain the corresponding value of the data type; Let f2 be the data complexity, which is used to record the data call and resource occupancy at each moment; Let f3 be the current system load; Let f4 be the response speed requirement, and f4 is the preset response time threshold; Perform normalization processing on f1, f2, f3 and f4 to obtain and , and ; Among them, the normalization process of f1 includes: Set the weight corresponding to the data type of f1, marked as the data type weight, and obtain the weighted sum of the values of the corresponding data type of f1 according to the data type weight ; Calculate the resource weight RW: The calculation formula of RW is: ; Among them, α, β, γ, and δ are respectively , , and corresponding weights.
[0010] Furthermore, the process of the data complexity f2 includes: The value range of f2 is [0,1], and the calculation formula of f2 is: ; Among them, C ty is the data type diversity index, , dt is the total number of data types in f1, and tdt is the preset maximum total number of data types; C st is the data structure complexity index, , nl is the total number of nested layers of the data structures of each data type in f1, fc is the total number of data fields of each data type in f1, mnl is the preset maximum number of nested layers, and mfc is the preset maximum number of fields; C re is the autocorrelation degree of f1, and the value range is [0,1]; C vo is the data volatility index, , is the standard deviation of the values in f1, is the average value of the values in f1; , , and are respectively the weights of C ty , C st , C re and C vo , .
[0011] Furthermore, the process of resource allocation includes: Substitute the predicted value of the real-time data volume into the RW calculation formula to calculate the resource weight of the predicted value of the real-time data volume ; Substitute the predicted value of the offline data volume obtained from the data prediction module into the resource weight calculation unit to calculate the resource weight of the predicted value of the offline data volume. ; Calculate the resource weight of the real-time data processing module. , ; Calculate the resource weight of the offline data processing module. , ; Suppose the total calculated resource amount is R, and the calculated resource amount of the real-time data processing module is ; The calculated resource amount of the offline data processing module is .
[0012] Furthermore, the security enhancement module: is used to encrypt and store the real-time data processing result and the offline data processing result in the unified data storage module, and distinguish the processing permissions of internal and external data requests through the access control policy.
[0013] Furthermore, it is characterized by including the following modules: Real-time data processing module: processes the real-time data of the data middle platform through the Kafka system and generates the real-time data processing result; Offline data processing module: processes the offline data of the data middle platform through the Hadoop system and generates the offline data processing result; Unified data storage module: is used to store the real-time data processing result in the real-time data processing module and the offline data processing result in the offline data processing module; Data prediction module: obtains the real-time data processing result and the offline data processing result in the unified data storage module, predicts the data volume of the real-time data processing module and the offline data processing module within the future unit time based on the time series analysis model, and outputs the predicted value of the real-time data volume and the predicted value of the offline data volume; Intelligent scheduling module: obtains the predicted value of the real-time data volume and the predicted value of the offline data volume of the data prediction module, and allocates the calculated resource amount to the real-time data processing module and the offline data processing module within the future unit time by using the dynamic weight resource allocation algorithm.
[0014] In the present invention, a method for dynamically allocating data processing resources for a data middle platform has the following beneficial technical effects: The influencing factors of resource allocation, such as the predicted value of real-time data volume or offline data volume, data complexity, the current load of the system, and the response speed requirement, are introduced, and the resource weights can be dynamically adjusted according to the predicted value of real-time data volume or offline data volume. This method can more accurately match the resource requirements of the real-time data processing module and the offline data processing module in the future unit time, realize the dynamic differential resource allocation of the real-time data processing module and the offline data processing module, avoid over-allocation or under-allocation of resources, solve the problem of rigid resource allocation in the traditional static resource allocation strategy, and greatly improve the resource utilization rate. Description of the Drawings
[0015] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments
[0016] During the implementation and deployment of the intelligent power plant system of a thermal power generation enterprise, the following technical problems are found through actual operation monitoring: With the continuous increase in the number of algorithm modules such as equipment safety warning and economic analysis, the overall load of the system has increased significantly, and the response delay of key business data processing has increased significantly.
[0017] After in-depth analysis and stress testing, the root causes of the problems are determined as follows: (1) The warning modules generally adopt technical solutions such as neural networks and multi-state estimation. Their reasoning processes need to frequently access and fuse multi-source heterogeneous real-time data from different systems. The linear growth of the number of modules leads to an exponential increase in the concurrent occupation of computing resources (CPU), memory (Memory), and I / O bandwidth, becoming the primary factor for the increase in system load.
[0018] (2) When deploying new business modules, it is necessary to call a large amount of historical offline data to train and optimize the model parameters. This training task has the characteristics of being computationally intensive and competes fiercely with the existing online warning services for cluster resources (especially GPU / CPU computing power and storage I / O), resulting in further deterioration of the real-time service response delay and even triggering task timeout alarms.
[0019] The above problems seriously restrict the scalability and reliability of the intelligent power plant system. To solve the core contradiction of the conflict between real-time service delay and offline training resources, it is necessary to use a dynamic data processing resource allocation method for the data middle platform to optimize the resource scheduling efficiency, ensure the real-time performance of key services, and improve the overall availability of the system.
[0020] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference signs denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.
[0021] Hadoop YARN is a core component of the Hadoop system in the prior art. It is a distributed resource management and scheduling framework, mainly used for efficiently allocating and managing computing resources in a cluster, and supporting multiple distributed computing frameworks to run on the same cluster; Kubernetes is an open-source container orchestration platform in the prior art, used for automating the deployment, scaling, and management of containerized applications.
[0022] As Figure 1 shown, a method for dynamically allocating data processing resources for a data middle platform includes: Step S1: Process the real-time data of the data middle platform through the Kafka system and generate real-time data processing results; Step S2: Process the offline data of the data middle platform through the Hadoop system and generate offline data processing results; Step S3: Based on a time series analysis model, predict the data volumes of the real-time data processing results and the offline data processing results in the future unit time, and output the real-time data volume prediction value and the offline data volume prediction value; Step S4: Use a dynamic weight resource allocation algorithm to allocate computing resources for the real-time data volume prediction value and the offline data volume prediction value.
[0023] In an optional embodiment, the process of outputting the real-time data volume prediction value and the offline data volume prediction value includes: Judge the stationarity of the real-time data processing results and the offline data processing results by using a time series diagram, an autocorrelation diagram, and a statistical test method for the real-time data processing results and the offline data processing results; If it is stationary, the ARIMA model is used for the real-time data processing results or the offline data processing results to generate the corresponding real-time data volume prediction value and the offline data volume prediction value in the offline data processing module in the future unit time; If it is non-stationary, the real-time data processing results or the offline data processing results are subjected to differencing processing, and then the ARIMA model is used for fitting. The ARIMA model is used to generate the corresponding real-time data volume prediction value and the offline data volume prediction value in the offline data processing module in the future unit time.
[0024] Generate ARIMA model parameters through residual analysis.
[0025] The ARIMA model is an autoregressive integrated moving average model, which is a statistical model used in the prior art to analyze and predict time series data; In an optional embodiment, the dynamic weight resource allocation algorithm includes the following steps: Obtain the predicted value of the real-time data volume and the predicted value of the offline data volume of the data prediction module, set the predicted value of the real-time data volume or the predicted value of the offline data volume, data complexity, the current system load, and the response speed requirement as the influencing factors for resource allocation, and calculate the resource weight through calculation; Substitute the predicted value of the real-time data volume and the predicted value of the offline data volume obtained in the data prediction module into the resource weight calculation step, calculate the resource weight of the real-time data processing module and the resource weight of the offline data processing module, and allocate the computing resource amount for the real-time data processing module and the offline data processing module.
[0026] The intelligent scheduling module uses the dynamic weight resource allocation algorithm to allocate the computing resource amount for the real-time data processing module and the offline data processing module in the future unit time. Different from the traditional static resource allocation method, this algorithm introduces the predicted value of the real-time data volume or the predicted value of the offline data volume, data complexity, the current system load, and the response speed requirement as the influencing factors for resource allocation, and can dynamically adjust the resource weight according to the predicted value of the real-time data volume or the predicted value of the offline data volume. This method can more accurately match the resource requirements of the system in the future unit time, avoid over-allocation or under-allocation of resources, solve the problem of rigid resource allocation in the traditional static resource allocation strategy, and greatly improve the resource utilization rate.
[0027] In an optional embodiment, in the resource weight calculation step: Let f1 be the predicted value of the real-time data volume or the predicted value of the offline data volume; Let f2 be the data complexity, which is used to record the data call and resource occupancy at each moment; obtain the data type of f1, marked as {image; video; text; numerical value; BP neural network; LSTM neural network; MSET structure; office class file; other files}; and obtain the corresponding numerical value of the data type; Let f3 be the current system load; Let f4 be the response speed requirement, and f4 is the preset response time threshold; Perform normalization processing on f1, f2, f3, and f4 to obtain 、 、 and ; Among them, the normalization processing process of f1 includes: Set the weight corresponding to the data type of f1, marked as the data type weight, and obtain the weighted sum of the values of the corresponding data type of f1 according to the data type weight ; It should be further noted that the meaning of f1 is [total size of image files in Mb; total size of video files in Mb, total size of text files in Mb, number of values, total number of neurons in the BP neural network; total number of neurons in the LSTM; total dimension of the MSET structure, total size of OFFICE, total size of other files]; For the data after normalization of the above F1 real-time and F1 offline data sets, the calculation method is: weighted sum of each item in F1, with weights [b1; b2; b3; ……; bi], that is, [0.5; 2; 1; ……; bi], then f1 = ∑ai * bi; where i represents the number of the data type corresponding to f1, and the value is a positive integer; ai represents the value of the data type corresponding to i; bi represents the weight of the data type corresponding to i; Calculate the computing resource weight RW: The calculation formula of RW is: ; where α, β, γ, and δ are respectively , , and the corresponding weights.
[0028] In an optional embodiment, in the resource allocation step, the calculation of the data complexity f2 is as follows: Preferably, in the resource allocation step, the calculation of the data complexity f2 is as follows: The value range of f2 is [0, 1], and the calculation formula of f2 is: ; where C ty is the data type diversity index, , dt is the total number of data types in f1, and tdt is the preset maximum total number of data types; C st is the data structure complexity index, , nl is the total number of nested layers of the data structures of each data type in f1, fc is the total number of data fields of each data type in f1, mnl is the preset maximum number of nested layers, and mfc is the preset maximum number of fields; C re is the autocorrelation degree of f1, and the value range is [0, 1]; C vo is the data volatility index, , is the standard deviation of the values in f1, is the average value of the numerical values in f1; , , and are the weights of C ty , C st , C re and C vo respectively, .
[0029] In an alternative embodiment, , , , The numerical values are obtained by normalizing the data C ty , C st , C re and C vo , and then calculated by the entropy weight method to obtain the numerical values of , , and respectively.
[0030] In an alternative embodiment, in the resource allocation step: Substitute the predicted value of the real-time data volume obtained in the data prediction module into the resource weight calculation unit to calculate the resource weight of the predicted value of the real-time data volume ; Substitute the predicted value of the offline data volume obtained in the data prediction module into the resource weight calculation unit to calculate the resource weight of the predicted value of the offline data volume ; Calculate the resource weight of the real-time data processing module , ; Calculate the resource weight of the offline data processing module , ; Suppose the total resource amount is R, and the computing resource amount of the real-time data processing module is ; The computing resource amount of the offline data processing module is .
[0031] The present invention also provides a system for a data processing resource dynamic allocation method for a data middle platform, including the following modules: Real-time data processing module: processes the real-time data of the data middle platform through the Kafka system and generates real-time data processing results; Offline data processing module: processes the offline data of the data middle platform through the Hadoop system and generates offline data processing results; Unified Data Storage Module: It is used to store the real-time data processing results in the real-time data processing module and the offline data processing results in the offline data processing module; Data Prediction Module: It obtains the real-time data processing results and offline data processing results in the unified data storage module, predicts the data volumes of the real-time data processing module and the offline data processing module within the next unit time based on the time series analysis model, and outputs the predicted real-time data volume value and the predicted offline data volume value; Intelligent Scheduling Module: It obtains the predicted real-time data volume value and the predicted offline data volume value of the data prediction module, and uses the dynamic weight resource allocation algorithm to allocate the computing resource amounts for the real-time data processing module and the offline data processing module within the next unit time.
[0032] Meanwhile, the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
[0033] In the embodiments provided by the present invention, it should be understood that the disclosed system or method can be implemented in other ways. For example, the above-described invention embodiments are merely illustrative. For example, the division of modules is only a logical function division, and there may be other division methods in actual implementation.
[0034] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0035] In addition, the functional modules in each embodiment of the present invention can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.
[0036] For those skilled in the art of operation and maintenance, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the basic characteristics of the present invention.
[0037] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.
Claims
1. A method for dynamically allocating data processing resources for a data middle platform, characterized in that, Including: Step S1: Process the real-time data of the data middle platform through the Kafka system and generate the real-time data processing result; Step S2: Process the offline data of the data middle platform through the Hadoop system and generate the offline data processing result; Step S3: Predict the data volumes of the real-time data processing result and the offline data processing result within the next unit time, and output the real-time data volume prediction value and the offline data volume prediction value; Step S4: Use the dynamic weight resource allocation algorithm to allocate the computing resource volume for the real-time data volume prediction value and the offline data volume prediction value.
2. The data processing resource dynamic allocation method for a data middle platform according to claim 1, wherein The process of outputting the real-time data volume prediction value and the offline data volume prediction value includes: Judge the stationarity of the real-time data processing result and the offline data processing result by using the time series diagram, autocorrelation diagram and statistical test method for the real-time data processing result and the offline data processing result; If it is stationary, the ARIMA model is used for the real-time data processing result or the offline data processing result to generate the corresponding real-time data volume prediction value within the next unit time and the offline data volume prediction value in the offline data processing module; If it is non-stationary, the real-time data processing result or the offline data processing result is subjected to differencing processing, and then the ARIMA model is used for fitting. The ARIMA model is used to generate the corresponding real-time data volume prediction value within the next unit time and the offline data volume prediction value in the offline data processing module.
3. A method for dynamically allocating data processing resources for a data middle platform according to claim 1, characterized in that The process of the dynamic weight resource allocation algorithm includes: Set the data complexity, the current system load and the response speed requirement corresponding to the real-time data volume prediction value or the offline data volume prediction value as the influencing factors for resource allocation, and calculate the resource weight; allocate the computing resource volume according to the resource weight.
4. A data processing resource dynamic allocation method for a data middle platform according to claim 2, characterized in that The process of obtaining the resource weight includes: Let f1 be the real-time data volume prediction value or the offline data volume prediction value; obtain the data type of f1, marked as {image; video; text; numerical value; BP neural network; LSTM neural network; MSET structure; office type file; other files}; and obtain the corresponding value of the data type; Let f2 be the data complexity, which is used to record the data call and resource occupancy at each moment; Let f3 be the current system load; Let f4 be the response speed requirement, and f4 is the preset response time threshold; Normalize f1, f2, f3, and f4 to obtain , , and ; Among them, the normalization process of f1 includes: Set the weight corresponding to the data type of f1, marked as the data type weight, and obtain the weighted sum of the values of the data type corresponding to f1 according to the data type weight ; Calculate the resource weight RW: The RW calculation formula is as follows: ; where α, β, γ, and δ are the weights corresponding to , , and respectively.
5. A method for dynamically allocating data processing resources for a data middle platform according to claim 4, characterized in that The process of the data complexity f2 includes: The value range of f2 is [0, 1], and the calculation formula of f2 is: ; Among them, C ty is the data type diversity index, , dt is the total number of data types in f1, and tdt is the preset maximum total number of data types; C st is the data structure complexity index, , nl is the total number of nested layers of the data structures of each data type in f1, fc is the total number of data fields of each data type in f1, mnl is the preset maximum number of nested layers, and mfc is the preset maximum number of fields; C re is the autocorrelation correlation degree of f1, and the value range is [0, 1]; C vo is the data volatility index, which is the standard deviation of the values in f1, and is the average value of the values in f1; , , and are the weights of C ty , C st , C re and C vo respectively, .
6. A method for dynamically allocating data processing resources for a data middle platform according to claim 3, characterized in that, The process of resource volume allocation includes: Substitute the predicted value of the real-time data volume into the RW calculation formula to calculate the resource weight of the predicted value of the real-time data volume ; Substitute the predicted value of the offline data volume obtained in the data prediction module into the resource weight calculation unit to calculate the resource weight of the predicted value of the offline data volume ; Calculate the resource weights of the real-time data processing module , ; Calculate the resource weights of the offline data processing module , ; Assume that the total computing resource is R, and the computing resource of the real-time data processing module is ; The computing resource amount of the offline data processing module is .
7. A method for dynamically allocating data processing resources for a data middle platform according to claim 1, characterized in that Security enhancement module: used to encrypt and store the real-time data processing result and the offline data processing result in the unified data storage module, and distinguish the processing permissions of internal and external data requests through the access control policy.
8. A method for dynamically allocating data processing resources for a data middle platform according to any one of claims 1-7, characterized in that Including the following modules: Real-time data processing module: Process the real-time data of the data middle platform through the Kafka system and generate the real-time data processing result; Offline data processing module: Process the offline data of the data middle platform through the Hadoop system and generate the offline data processing result; Unified data storage module: used to store the real-time data processing result in the real-time data processing module and the offline data processing result in the offline data processing module; Data Prediction Module: Obtain the real-time data processing results and offline data processing results from the unified data storage module, predict the data volumes of the real-time data processing module and the offline data processing module in the future unit time based on the time series analysis model, and output the predicted value of the real-time data volume and the predicted value of the offline data volume; Intelligent Scheduling Module: Obtain the predicted value of the real-time data volume and the predicted value of the offline data volume from the data prediction module, and use the dynamic weight resource allocation algorithm to allocate the computing resource volumes for the real-time data processing module and the offline data processing module in the future unit time.
Citation Information
Patent Citations
Data processing method based on data middle platform and related equipment
CN113836235B
Dynamic allocation method for container resources in cluster
CN111124689A
Data processing method based on data medium table and related equipment thereof
CN113836235A
Method for automatically configuring resources of data center
CN119094335A
Super-fusion computing power scheduling method and system based on prediction model
CN119336516A