Real-time data integration and analysis system based on cloud computing
By providing real-time data integration and analysis systems in the cloud computing environment, the problem of insufficient real-time processing capabilities of heterogeneous data in the cloud computing environment is solved, and efficient and accurate data analysis and decision-making support is achieved. Especially in a multi-tenant environment, resource allocation and task scheduling are optimized.
Patent Information
- Application Number
- CN202510115746.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The lack of a unified data integration and analysis framework in the cloud computing environment of the existing technology has led to insufficient real-time processing capabilities of cross-source and heterogeneous data. Especially in a multi-tenant environment, it may lead to delays, loss of data processing or deviations in analysis results, making it difficult to meet the needs of high-precision real-time decision-making.
Provide a real-time data integration and analysis system based on cloud computing, including real-time data acquisition module, real-time data analysis module, tenant resource scheduling module, task scheduling adjustment module and analysis result extraction module. The system combines resource scheduling and task scheduling to achieve efficient data integration and analysis through real-time data acquisition, field verification, format conversion, windowed calculation and joint aggregation operations.
Through a unified data integration framework, heterogeneous data format inconsistency, structural differences and inefficient processing are solved, the accuracy and response speed of real-time data processing are improved, and high-precision data analysis and decision-making support are ensured. Especially in a multi-tenant environment, resource allocation and task scheduling are optimized to avoid resource overload and data loss.
Smart Images

Figure CN120045321A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and more specifically, to a real-time data integration and analysis system based on cloud computing. Background Art
[0002] In the prior art, real-time data processing in a cloud computing environment mainly relies on a single functional module to implement specific tasks (such as streaming data processing or batch data analysis); due to the lack of a unified data integration and analysis framework, the real-time processing ability of cross-source and heterogeneous data is insufficient. Especially in a multi-tenant environment, dynamic resource allocation may lead to data processing delays, losses, or biases in analysis results, making it difficult to meet the requirements for high-precision real-time decision-making. Summary of the Invention
[0003] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a real-time data integration and analysis system based on cloud computing to solve the problems raised in the above background art.
[0004] To achieve the above object, the present invention provides the following technical solutions:
[0005] A real-time data integration and analysis system based on cloud computing, comprising:
[0006] A real-time data acquisition module, configured to obtain real-time data from multiple heterogeneous sources, and perform field verification processing and format conversion processing on the real-time data to generate preprocessed real-time data and transmit it to the real-time data analysis module;
[0007] A real-time data analysis module, configured to perform windowed calculation and joint aggregation operation on the preprocessed real-time data based on a distributed computing algorithm, generate intermediate analysis results and transmit them to the tenant resource scheduling module and the analysis result extraction module;
[0008] A tenant resource scheduling module, configured to analyze the load time series of different resource types based on the intermediate analysis results, and evaluate the heterogeneous resource allocation balance state of each computing node;
[0009] A task scheduling adjustment module, when the heterogeneous resource allocation balance state of any computing node exceeds a preset range, dynamically adjusts the task scheduling strategy, and generates an optimized task allocation plan and transmits it to the real-time data analysis module to optimize the data processing process;
[0010] An analysis result extraction module, configured to perform further processing on the intermediate analysis results to complete the integration and analysis of real-time data.
[0011] In a preferred embodiment, obtaining real-time data from multiple heterogeneous sources, and performing field verification processing and format conversion processing on the real-time data to generate preprocessed real-time data, specifically:
[0012] Obtain real-time data from multiple heterogeneous data sources, including sensors, devices, databases, and cloud services. The real-time data is continuously changing in each data source and has a time series.
[0013] Perform field verification processing on the received real-time data to ensure that each data field conforms to the predetermined rules.
[0014] After the field verification is qualified, perform data format conversion processing to unify the different formats of each heterogeneous data source into a standard format.
[0015] Generate preprocessed real-time data in a unified standard format based on the results of field verification and format conversion processing.
[0016] In a preferred embodiment, perform windowed calculation and joint aggregation operation on the preprocessed real-time data based on a distributed computing algorithm to generate an intermediate analysis result. Specifically:
[0017] Divide the preprocessed real-time data into multiple time windows, and the data within each time window is regarded as a processing unit. The size of the window is dynamically adjusted according to the predetermined rules, and the window division is based on timestamps.
[0018] Within each time window, use the distributed computing algorithm to perform calculation processing on the preprocessed real-time data, specifically including performing mathematical statistical operations on the preprocessed real-time data within each window.
[0019] Perform joint aggregation operation on the calculation results of different time windows, and perform global data analysis by merging the results of multiple windows to generate an intermediate analysis result.
[0020] Generate an intermediate analysis result after the windowed calculation and joint aggregation operation are completed. The intermediate analysis result is the intermediate data after distributed computing processing.
[0021] In a preferred embodiment, analyze the load time series of different resource types based on the intermediate analysis result to evaluate the heterogeneous resource allocation balance state of each computing node. Specifically:
[0022] Extract the load time series of each computing node from the intermediate analysis result. The load time series reflects the resource consumption situation of each computing node at different time points.
[0023] By taking the load time series of each resource as input and using the Transformer architecture to model the time dependence to capture the temporal characteristics of the load changes of different resource types.
[0024] Apply the discrete-time Lyapunov entropy to each resource load time series, and analyze the rate and trend of load changes to generate a balance state indicator.
[0025] In a preferred embodiment, the Lyapunov entropy calculation formula is: Where, H i represents the Lyapunov entropy of the i-th resource, δx k represents the state difference at the k-th iteration, b is the maximum number of iterations for calculation, and k is the numbering of the iteration times.
[0026] In a preferred embodiment, the resource pressure balance entropy calculation process is as follows:
[0027] Fuse the entropy values of each dimension from different resource types to calculate the resource pressure balance entropy: Where, H balance is the resource pressure balance entropy, w i is the weight of each resource type, H i is the Lyapunov entropy value of the corresponding resource type, N is the total number of resource types, and i is the index of the resource type.
[0028] In a preferred embodiment, when the heterogeneous resource allocation equilibrium state of any computing node exceeds the preset range, dynamically adjust the task scheduling strategy and generate an optimized task allocation plan to optimize the data processing process, specifically:
[0029] Compare the resource pressure balance entropy of the computing node with the preset resource allocation equilibrium range to determine whether it exceeds the resource allocation equilibrium range;
[0030] When the resource pressure balance entropy of a certain computing node exceeds the resource allocation equilibrium range, confirm that the resource allocation state of this node is unbalanced;
[0031] Dynamically generate an optimized task scheduling strategy based on the calculation result of the resource pressure balance entropy, and adjust the task allocation to improve the resource load;
[0032] Generate a new task allocation plan based on the optimized task scheduling strategy, and reallocate the resources to each computing node.
[0033] In a preferred embodiment, perform further processing on the intermediate analysis results to complete the integration and analysis of real-time data, specifically:
[0034] By integrating the intermediate analysis results from different computing nodes, unify the data format and standardize the data scale;
[0035] Apply the corresponding analysis algorithms to the integrated intermediate analysis results to deeply analyze the data. The analysis process includes trend identification, anomaly detection, and pattern mining of the data;
[0036] Generate a real-time data analysis report through the results of the in-depth analysis. The report content includes key analysis results, trend predictions, and suggestions.
[0037] The technical effects and advantages of the real-time data integration and analysis system based on cloud computing of the present invention:
[0038] 1. Efficiently obtain real-time data from multiple heterogeneous sources, and generate standardized preprocessed data through field verification and format conversion processing. Compared with the prior art, the present invention solves the problems of inconsistent data formats, structural differences, and low processing efficiency from different data sources through a unified data integration framework. The system not only supports the flexible integration of heterogeneous data but also ensures the high efficiency of data transmission and processing, further improving the accuracy and response speed of real-time data processing. Through a standardized data processing process, errors caused by inconsistent data formats or processing delays are avoided, ensuring high-precision data analysis and decision support.
[0039] 2. By introducing distributed computing algorithms and multi-dimensional analysis methods, perform windowed calculations and joint aggregation operations on the preprocessed real-time data, effectively improving the ability to process large-scale real-time data. Analyze the resource allocation balance state of each computing node in combination with the intermediate analysis results, and dynamically adjust the task scheduling strategy, which helps to optimize the allocation of computing resources. Especially in a multi-tenant environment, the system can flexibly adjust the task allocation scheme according to changes in resource pressure, avoiding resource overload and data loss. The system can respond in real-time to dynamically changing computing requirements, ensuring the stability and accuracy of the data processing process, thereby providing efficient and accurate real-time data analysis support for decision-makers. Brief Description of the Drawings
[0040] Figure 1 It is a schematic structural diagram of the real-time data integration and analysis system based on cloud computing of the present invention. Detailed Embodiments
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0042] Embodiment: Figure 1The structure diagram of the real-time data integration and analysis system based on cloud computing according to the present invention is given. The real-time data integration and analysis system based on cloud computing includes:
[0043] A real-time data acquisition module, which is used to obtain real-time data from multiple heterogeneous sources, and perform field verification processing and format conversion processing on the real-time data to generate preprocessed real-time data and transmit it to the real-time data analysis module;
[0044] A real-time data analysis module, which is used to perform windowed calculation and joint aggregation operation on the preprocessed real-time data based on distributed computing algorithms, generate intermediate analysis results and transmit them to the tenant resource scheduling module and the analysis result extraction module;
[0045] The tenant resource scheduling module analyzes the load time series of different resource types based on the intermediate analysis results, and evaluates the heterogeneous resource allocation balance state of each computing node;
[0046] The task scheduling adjustment module, when the heterogeneous resource allocation balance state of any computing node exceeds the preset range, dynamically adjusts the task scheduling strategy, and generates an optimized task allocation plan and transmits it to the real-time data analysis module to optimize the data processing process;
[0047] The analysis result extraction module performs further processing on the intermediate analysis results to complete the integration and analysis of real-time data.
[0048] Obtain real-time data from multiple heterogeneous sources, and perform field verification processing and format conversion processing on the real-time data to generate preprocessed real-time data. Specifically:
[0049] Obtain real-time data from multiple heterogeneous data sources. The data sources include sensors, devices, databases, and cloud services. The real-time data is real-time data that continuously changes in each data source and has a time series:
[0050] Obtain real-time data from multiple heterogeneous data sources. The data sources include sensors, devices, databases, and cloud services, etc. These data sources can provide real-time data with different data formats and structures. Specifically, the real-time data collected by sensors may represent physical quantities measured by sensors (such as temperature, humidity, pressure, etc.) in the form of a time series; the real-time data that devices may provide includes status monitoring data (such as current, voltage, power, etc.); the real-time data in the database can be query or updated data records; and the real-time data provided by cloud services may include API requests or real-time data streams. During this process, the real-time data continuously flows into the system from each data source with its inherent change characteristics, and usually presents a continuously changing time series form.
[0051] Perform field verification processing on the received real-time data to ensure that each data field conforms to the predetermined rules:
[0052] Verify the fields of each piece of data to ensure that they conform to the preset rules. The field verification process specifically includes: verifying whether the data type of each data field conforms to the expected type (such as integer, floating decimal, character, etc.); checking whether the field value is within the valid range (for example, whether the temperature data is within a reasonable temperature range and whether the voltage data exceeds the safety value); and verifying the existence of missing or abnormal data (such as missing value filling, null value handling, etc.). The purpose of field verification is to ensure that each data field is valid and to avoid deviations in subsequent analysis results caused by invalid data.
[0053] After the field verification is qualified, perform data format conversion processing to unify the different formats of each heterogeneous data source into a standard format:
[0054] Perform format conversion processing on real-time data. This processing process is to unify the real-time data from different heterogeneous data sources into a standard format. Since different data sources (such as sensors, devices, databases, and cloud services) may adopt different data formats (such as JSON format, XML format, CSV format, binary stream, etc.), it is necessary to unify them into a standardized format (such as unified into JSON or other standard formats) through format conversion to ensure that subsequent processing modules can correctly parse and analyze the data. The purpose of format conversion is to improve the compatibility of data processing and ensure the smooth progress of subsequent analysis links.
[0055] Generate preprocessed real-time data in a unified standard format based on the results of field verification and format conversion processing:
[0056] Generate preprocessed real-time data in a unified standard format. These processed data have removed invalid fields, the format has been standardized, and can be used as input data for subsequent analysis. At this stage, all data sources have been unified into a processable standard format, enabling the system to process subsequent tasks more efficiently without worrying about the processing difficulties caused by data source heterogeneity.
[0057] Perform windowed calculation and joint aggregation operations on the preprocessed real-time data based on distributed computing algorithms to generate intermediate analysis results, specifically:
[0058] Divide the preprocessed real-time data into multiple time windows, and the data within each time window is regarded as a processing unit; the size of the window is dynamically adjusted according to the predetermined rules, and the window division is based on timestamps:
[0059] The preprocessed real-time data is divided into multiple time windows according to timestamps. The size of each time window is dynamically adjusted by predetermined rules, usually determined according to the changes in the data stream or calculation requirements.
[0060] For example, the length of the time window can be in time units such as seconds, minutes, hours, etc., or be adapted according to the change frequency of the real-time data to ensure that a certain number of data points are included within each time window. The basis for dividing the time window is the timestamp, which is the time identifier for each data record and can ensure that the data is evenly distributed within consecutive time periods.
[0061] Within each time window, a distributed computing algorithm is used to perform computational processing on the preprocessed real-time data, specifically including performing mathematical statistical operations on the preprocessed real-time data within each window:
[0062] Specifically, the system performs mathematical statistical operations on the data within each time window. These operations can include but are not limited to summation, mean calculation, maximum value calculation, minimum value calculation, standard deviation calculation, etc. The selection of each statistical operation depends on the goal of data processing, aiming to extract the key features or trends of the data within the time window and provide valuable information for subsequent analysis.
[0063] For example, the mean and standard deviation can be used to describe the data distribution, while the maximum and minimum values help identify extreme cases of the data.
[0064] Perform a joint aggregation operation on the calculation results of different time windows. By combining the results of multiple windows, perform global data analysis to generate intermediate analysis results:
[0065] Perform a joint aggregation operation on the calculation results of different time windows. The purpose of the joint aggregation operation is to perform global data integration on the processing results of each time window, thereby generating an overall analysis result. This aggregation operation can be in ways such as weighted aggregation, maximum value merging, minimum value merging, etc., or other merging methods. The specific method is determined according to the task or goal to be analyzed.
[0066] For example, through weighted aggregation, the results of different time windows can be weighted and merged according to their importance to ensure that the data of time windows with greater influence occupies a larger proportion in the final result.
[0067] After the windowed calculation and joint aggregation operation are completed, intermediate analysis results are generated. The intermediate analysis results are the intermediate data after distributed computing processing:
[0068] The intermediate analysis results are the data after distributed computing processing, containing the data characteristics of different time windows and global analysis information. This intermediate result can reflect the overall trend, key change points, and possible abnormal situations of the preprocessed real-time data.
[0069] The intermediate analysis results are usually the basic data for further decision-making or optimization processes, providing reliable input for subsequent decision support systems or control mechanisms.
[0070] Analyze the load time series of different resource types based on the intermediate analysis results to evaluate the heterogeneous resource allocation balance state of each computing node, specifically as follows:
[0071] Extract the load time series of each computing node from the intermediate analysis results. The load time series reflects the resource consumption of each computing node at different time points:
[0072] Extract the load time series of each computing node from the intermediate analysis results processed by distributed computing. The load time series consists of resource consumption data at multiple time points, reflecting the changes in various resources such as computing resources, memory, bandwidth, and storage consumed by the computing node at each time point. These time series data are the basis for evaluating the resource load balance of the computing node.
[0073] The load time series can be regarded as a dynamic data set, where each data point represents the resource consumption state at a moment. Through these time series data, the load situation of each node within a given time can be analyzed, providing data support for subsequent load analysis.
[0074] By taking the load time series of each resource as input and using the Transformer architecture to model the time dependence, the temporal characteristics of the load changes of different resource types can be captured:
[0075] Use the Transformer architecture to model these time series to capture the time dependence. Transformer is a deep learning model based on the self-attention mechanism and is widely used in the analysis of time series data. This architecture can effectively capture the long-term dependence relationships in time series data and reveal the trends and periodic characteristics of time series under complex non-linear patterns.
[0076] By taking the load time series of each resource as input, the Transformer model can understand the mutual relationships between each time point through the self-attention mechanism, automatically identify and extract useful temporal characteristics.
[0077] The processing ability of Transformer makes it more advantageous than traditional time series analysis methods (such as ARIMA models and LSTM networks) in capturing the changes in different resource loads and can accurately model the temporal correlations in the load data.
[0078] Apply the discrete-time Lyapunov entropy to the load time series of each resource to analyze the rate and trend of load changes to generate balance state indicators:
[0079] Lyapunov entropy is a tool used to quantify the degree of chaos in a dynamic system. In the analysis of load time series, Lyapunov entropy can measure the rate of load change and the non-linear characteristics of the change trend.
[0080] Specifically, Lyapunov entropy analyzes the stability and complexity of time series by calculating the sensitivity of the system (i.e., how sensitive the system state is to changes in initial conditions). For the load time series of each resource, Lyapunov entropy can reveal its rate and trend of change over time. A higher Lyapunov entropy value usually indicates unstable load changes, while a lower value indicates relatively stable load changes. By calculating Lyapunov entropy, the load fluctuation characteristics of each resource can be evaluated, providing a more detailed basis for system resource scheduling and load balancing.
[0081] The formula for calculating Lyapunov entropy is: where H i represents the Lyapunov entropy of the i-th resource, δx k represents the state difference at the k-th iteration, n is the maximum number of iterations for calculation, and k is the iteration number.
[0082] The entropy values of each dimension are fused to calculate the resource pressure balance entropy, and the heterogeneous resource allocation equilibrium state of each computing node is evaluated through the resource pressure balance entropy:
[0083] After the calculation of Lyapunov entropy is completed, the entropy values of each dimension from different resource types are fused to calculate the resource pressure balance entropy. The resource pressure balance entropy is a comprehensive indicator used to measure whether the allocation of heterogeneous resources (such as computing, storage, network, etc.) is balanced and reflects the scheduling ability of the system under various resource pressures.
[0084] The resource pressure balance entropy can be calculated in the following way: where H balance is the resource pressure balance entropy, w i is the weight of each resource type, H i is the Lyapunov entropy value of the corresponding resource type, N is the total number of resource types, and i is the index of the resource type.
[0085] A comprehensive resource pressure balance entropy is calculated through the weighted sum of the weights and entropy values.
[0086] The heterogeneous resource allocation equilibrium state of each computing node is evaluated through the calculated resource pressure balance entropy. The lower the resource pressure balance entropy, the more balanced the resource load distribution of each computing node and the stronger the resource scheduling ability of the system; conversely, a higher balance entropy value indicates uneven resource allocation, and there may be overload or resource waste.
[0087] Based on the resource pressure balance entropy of each node, it is possible to further determine whether it is necessary to adjust the resource allocation strategy to avoid performance bottlenecks or failures caused by resource overload on certain nodes.
[0088] When the heterogeneous resource allocation equilibrium state of any computing node exceeds the preset range, dynamically adjust the task scheduling strategy and generate an optimized task allocation plan to optimize the data processing process. Specifically:
[0089] Compare the resource pressure balance entropy of the computing node with the preset resource allocation equilibrium range to determine whether it exceeds the resource allocation equilibrium range:
[0090] Compare the calculated resource pressure balance entropy with the preset resource allocation equilibrium range. The resource allocation equilibrium range is usually a range of upper and lower limits, which are dynamically set by the system according to historical data or the resource capabilities of the computing nodes.
[0091] When the resource pressure balance entropy of a certain computing node exceeds the resource allocation equilibrium range, confirm that the resource allocation state of this node is unbalanced:
[0092] If the resource pressure balance entropy value of a certain computing node exceeds the preset resource allocation equilibrium range, it is considered that the resource load state of this node is abnormal and does not meet the system's requirements for balanced resource allocation. At this time, the system will identify the abnormal node, providing a basis for subsequent adjustment of the task scheduling strategy.
[0093] After the comparison is performed, if the resource pressure balance entropy of a certain computing node exceeds the resource allocation equilibrium range, the system further confirms that the resource allocation state of this computing node is unbalanced. This confirmation process is to analyze the load time series of the computing node to determine the resource consumption situation of the node at different time points, and then obtain its overall resource consumption trend.
[0094] At this time, the resource usage situation reflected by the load time series shows that the node may be overloaded during certain time periods, resulting in the resource pressure balance entropy value of this node exceeding the resource allocation equilibrium range. Through the analysis of the time series, it is possible to accurately identify the time periods of abnormal node load and further confirm that its resource allocation state is unbalanced. This process provides data support for subsequent adjustment of the task scheduling strategy.
[0095] Dynamically generate an optimized task scheduling strategy based on the calculation result of the resource pressure balance entropy, and adjust the task allocation to improve the resource load:
[0096] Based on the calculation results of the resource pressure balance entropy of the node, an optimized task scheduling strategy is dynamically generated. There are mainly two factors for the generation of the task scheduling strategy: First, the load situation of the node; Second, the current resource requirements of the system and the resource consumption characteristics of the tasks.
[0097] The optimization of the task scheduling strategy usually includes the following aspects of adjustment:
[0098] Load balancing: Re-evaluate the load situation of each computing node. According to the current resource usage of the node, transfer the overloaded task or resource allocation from the node with higher load to the node with lower load to achieve global load balancing;
[0099] Task priority adjustment: Adjust the priority of the tasks according to the level of resource pressure of the node. Give priority to allocating tasks with higher resource requirements to nodes with lighter load to avoid overloading some nodes;
[0100] Dynamic resource allocation: Adjust the execution time of the tasks and the number of task allocations according to the real-time resource requirements of the computing node, so that the resource load of each node is within an acceptable range, thereby reducing the risk of resource overload.
[0101] Based on the optimized task scheduling strategy, a new task allocation plan is generated, and the resources are re-allocated to each computing node:
[0102] The generation process of the task allocation plan includes the following aspects: Dynamically migrate the tasks already allocated to overloaded nodes to nodes with lighter load. Adjust the resource allocation ratio between computing nodes, and re-adjust the resources available to each node according to the resource requirements of the tasks and the resource capabilities of the computing nodes. Apply the new task allocation plan to the system in real time to ensure that the task scheduling is carried out according to the new plan. The system will monitor the resource usage of the nodes to ensure balanced resource allocation and avoid the emergence of new load overload phenomena.
[0103] Generating an optimized task allocation plan and transmitting it to the real-time data analysis module can effectively improve the efficiency and accuracy of data processing. When the resource allocation of the computing nodes is unbalanced, the task scheduling adjustment module will adjust the task allocation strategy according to the real-time analysis results, avoiding the situation of resource overload or idleness. Through reasonable resource scheduling and task allocation, the system can ensure that the load of each computing node is in the best state, reducing delays or calculation errors caused by resource bottlenecks. This optimized task allocation plan helps to improve the stability of real-time data analysis, avoiding data processing delays, losses or inaccuracies caused by uneven resource allocation, thereby ensuring the high precision of real-time decision-making and the overall performance of the system. Especially in a multi-tenant environment, it ensures that the resource requirements of different users can be evenly met, improving the overall processing capacity and service quality of the system.
[0104] Further processing is performed on the intermediate analysis results to complete the integration and analysis of real-time data, specifically as follows:
[0105] By integrating the intermediate analysis results from different computing nodes, the data format is unified and the data scale is standardized:
[0106] Collect the intermediate analysis results generated by each computing node. Each computing node may generate data in different formats, which may include different measurement units, different data representation methods, etc.
[0107] In this step, data integration processing is carried out to ensure that all data can be unified into a standard data format. The integration process includes converting multiple data formats into the same data structure. For example, converting time-series data into a unified timestamp format, unifying the units of numerical data, and performing encoding conversion on character-type data, etc.
[0108] Standardizing the data scale means unifying the scale of the data generated by different computing nodes. For example, through normalization or standardization methods, the data from different sources can be compared and analyzed under the same dimension. Through this processing, it is ensured that all intermediate analysis results are in the same dimension and range, facilitating the processing of subsequent analysis algorithms.
[0109] Apply corresponding analysis algorithms to the integrated intermediate analysis results to deeply analyze the data. The analysis process includes trend identification, anomaly detection, and pattern mining of the data:
[0110] Extract valuable knowledge from the intermediate analysis results, and identify trends, anomalies, or potential patterns in the data. The selection of analysis algorithms should be based on the nature of the data and the analysis objectives. Time series analysis, regression analysis, clustering analysis, classification algorithms, or machine learning methods, etc. can be used.
[0111] The specific analysis process includes:
[0112] By analyzing the change trends in historical data, predict the future direction of the data. For example, use methods such as moving average method and exponential smoothing to identify the long-term trend of the data;
[0113] Compare the current data with historical data to detect data points that deviate from the normal range. Abnormal data usually needs to be marked and fed back as important information to subsequent processing links;
[0114] Search for potential laws and patterns in multi-dimensional data, such as finding patterns of uneven resource allocation or periodic characteristics of load fluctuations, etc.
[0115] Generate a real-time data analysis report through the results of in-depth analysis. The report content includes key analysis results, trend predictions, and suggestions:
[0116] Key analysis results: The most important conclusions drawn based on data analysis, such as the imbalance of system load, low resource utilization efficiency, etc.
[0117] Trend prediction: Based on the analysis results of historical data, predict future resource consumption, system load, etc., predict possible high-load periods, or predict system performance bottlenecks.
[0118] Optimization suggestions: Based on the analysis results, put forward optimization suggestions in aspects such as data processing, resource scheduling, and load balancing to improve system performance or optimize resource allocation.
[0119] The generated analysis report provides decision support for system administrators or decision-makers, helping them to further process data, optimize the system, or adjust resource allocation.
[0120] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulation to get a formula closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0121] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0122] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0123] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0124] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in an electrical, mechanical, or other form.
[0125] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0126] In addition, the functional modules in each embodiment of this application can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0127] If the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0128] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
[0129] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A real-time data integration and analysis system based on cloud computing, characterized by: include: A real-time data acquisition module is used to obtain real-time data from multiple heterogeneous sources, and perform field verification processing and format conversion processing on the real-time data to generate pre-processed real-time data to transmit to the real-time data analysis module; A real-time data analysis module is used to perform windowed calculations and joint aggregation operations on pre-processed real-time data based on a distributed computing algorithm, generate intermediate analysis results, and transmit them to the tenant resource scheduling module and the analysis result extraction module; The tenant resource scheduling module analyzes the load time series of different resource types based on the intermediate analysis results and evaluates the balanced state of heterogeneous resource allocation of each computing node; The task scheduling adjustment module dynamically adjusts the task scheduling strategy when the heterogeneous resource allocation balance state of any computing node exceeds the preset range, and generates an optimized task allocation plan and transmits it to the real-time data analysis module to optimize the data processing process; The analysis result extraction module performs further processing on the intermediate analysis results to complete the integration and analysis of real-time data.
2. The real-time data integration and analysis system based on cloud computing according to claim 1, characterized in that: Acquire real-time data from multiple heterogeneous sources and perform field validation and format conversion on the real-time data to generate pre-processed real-time data. Specifically: Acquire real-time data from multiple heterogeneous data sources, including sensors, devices, databases, and cloud services. Real-time data is real-time data that continuously changes in each data source and has a time series. Perform field verification on the received real-time data to ensure that each data field complies with the predetermined rules; After the field verification is qualified, data format conversion is performed to unify the different formats of various heterogeneous data sources into a standard format; Based on the field verification and format conversion processing results, pre-processed real-time data in a unified standard format is generated.
3. The real-time data integration and analysis system based on cloud computing according to claim 1, characterized in that: Based on the distributed computing algorithm, window computing and joint aggregation operations are performed on the pre-processed real-time data to generate intermediate analysis results, specifically: The pre-processed real-time data is divided into multiple time windows, and the data in each time window is regarded as a processing unit; the size of the window is dynamically adjusted according to the predetermined rules, and the window division is based on the timestamp; In each time window, a distributed computing algorithm is used to perform computational processing on the preprocessed real-time data, specifically including performing mathematical statistical operations on the preprocessed real-time data in each window; Perform joint aggregation operations on the calculation results of different time windows, combine the results of multiple windows, perform global data analysis, and generate intermediate analysis results; After the window calculation and joint aggregation operation are completed, the intermediate analysis results are generated. The intermediate analysis results are intermediate data processed by distributed computing.
4. The real-time data integration and analysis system based on cloud computing according to claim 1, characterized in that: Based on the intermediate analysis results, the load time series of different resource types are analyzed to evaluate the balanced state of heterogeneous resource allocation of each computing node. Specifically: Extract the load time series of each computing node from the intermediate analysis results. The load time series reflects the resource consumption of each computing node at different time points. By taking the load time series of each resource as input, the Transformer architecture is used to model the time dependency to capture the timing characteristics of load changes of different resource types. Discrete-time Lyapunov entropy is applied to each resource load time series to analyze the rate and trend of load change to generate a balance state indicator.
5. The real-time data integration and analysis system based on cloud computing according to claim 4, characterized in that: The Lyapunov entropy calculation formula is: Among them, H i represents the Lyapunov entropy of the i-th resource, δx k represents the state difference at the kth iteration, n is the maximum number of iterations calculated, and k is the number of iterations.
6. The real-time data integration and analysis system based on cloud computing according to claim 4, characterized in that: The calculation process of resource pressure balance entropy is: The entropy values of each dimension from different resource types are integrated to calculate the resource pressure balance entropy: Among them, H balance is the resource pressure balance entropy, w i is the weight of each resource type, H i is the Lyapunov entropy value of the corresponding resource type, N is the total number of resource types, and i is the index of the resource type.
7. The real-time data integration and analysis system based on cloud computing according to claim 1, characterized in that: When the balanced state of heterogeneous resource allocation of any computing node exceeds the preset range, the task scheduling strategy is dynamically adjusted and an optimized task allocation plan is generated to optimize the data processing process. Specifically: Compare the resource pressure balance entropy of the computing node with the preset resource allocation balance range to determine whether it exceeds the resource allocation balance range; When the resource pressure balance entropy of a computing node exceeds the resource allocation balance range, it is confirmed that the resource allocation state of the node is unbalanced; Dynamically generate optimized task scheduling strategies based on the calculation results of resource pressure balance entropy, and adjust task allocation to improve resource load; A new task allocation plan is generated based on the optimized task scheduling strategy to reallocate resources to each computing node.
8. The real-time data integration and analysis system based on cloud computing according to claim 1, characterized in that: Further processing is performed on the intermediate analysis results to complete the integration and analysis of real-time data, specifically: By integrating the intermediate analysis results from different computing nodes, the data format is unified and the data scale is standardized; Apply corresponding analysis algorithms to the integrated intermediate analysis results to conduct in-depth analysis of the data. The analysis process includes trend identification, anomaly detection and pattern mining of the data; Generate real-time data analysis reports based on the results of in-depth analysis, including key analysis results, trend forecasts and recommendations.
Citation Information
Cited By
Heterogeneous energy data processing method and device and storage medium
CN120821734A
Isomerization energy data processing method, device and storage medium
CN120821734B