Park management service method and equipment

Through multi-source heterogeneous data cleaning and microservice architecture, combined with an integrated learning framework and fault tolerance mechanism, the bottlenecks of data quality and system performance of park enterprises are solved, efficient energy optimization, traffic scheduling and risk warning are achieved, and park operation efficiency and security are improved.

CN120234531APending Publication Date: 2025-07-01HANGZHOU XIGU IND HOLDINGS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337977.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Park enterprises face the quality problems of massive multi-source heterogeneous data, complex system performance bottlenecks and single model algorithm accuracy problems, resulting in insufficient data analysis accuracy and system scalability, which cannot meet the needs of intelligent operations.

Method used

Multi-source heterogeneous data cleaning, microservice architecture and integrated learning framework are adopted to generate energy optimization, flow scheduling and risk warning strategies through data cleaning, missing value filling, outlier value processing, feature extraction and multi-model training, and combine fault tolerance mechanisms and heartbeat detection to ensure stable operation of the system.

Benefits of technology

It realizes intelligent analysis and decision-making optimization of multi-dimensional data in the park, improves operational efficiency and security, and ensures the real-time and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234531A_ABST
    Figure CN120234531A_ABST
Patent Text Reader

Abstract

The invention provides a park management service method and equipment, and relates to the technical field of security, and the method comprises the steps: obtaining multi-source heterogeneous data which comprises energy use data, network use data and security event data; performing data cleaning, data missing value filling and data abnormal value processing on the multi-source heterogeneous data to obtain standardized data; inputting the standardized data into a pre-constructed real-time processing system to obtain feature data; and inputting the feature data into a pre-constructed integrated learning framework. According to the method, a prediction result is generated through feature extraction, multi-model training and an integrated learning framework, and then comprehensive operation decisions such as energy optimization, flow scheduling and risk early warning are made. In the decision execution process, a fault-tolerant mechanism, heartbeat detection and other methods are adopted to ensure stable operation of the system. According to the system, intelligent analysis and decision optimization of park multi-dimensional data are realized, and the park operation efficiency and safety are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of security technologies, and in particular, to a park management service method and device. Background Art

[0002] During the daily operation of park enterprises, a large amount of energy consumption data, network usage data, and security incident data are generated. These data are crucial for the energy management, network security, and property security of enterprises. However, due to diverse data sources and inconsistent formats, the data quality is uneven, with a large number of missing values and outliers, seriously affecting the data availability and analysis accuracy. In addition, there are numerous enterprises in the park and complex business systems, posing high requirements for real-time data processing and analysis. The traditional monolithic architecture can no longer meet the performance and scalability requirements of the system, and it is urgent to introduce a microservices architecture to split complex business logics into multiple independent services, improving the system's elasticity and fault tolerance. At the same time, a single prediction model is difficult to adapt to complex and changing business scenarios, and the prediction accuracy cannot meet the actual needs. It is necessary to explore advanced machine learning algorithms such as ensemble learning, and by combining the prediction results of multiple basic models, significantly improve the overall prediction performance. The technical challenges faced by park enterprises can be summarized as: the quality problem of massive multi-source heterogeneous data, the performance bottleneck problem of complex systems, and the accuracy problem of single model algorithms. These problems are intertwined and closely linked. It is urgent for park managers and technical teams to work together to comprehensively apply advanced technologies such as big data processing, microservices architecture, and machine learning to escort the intelligent operation of park enterprises. Summary of the Invention

[0003] The present invention provides a park management service method, mainly including: Obtaining multi-source heterogeneous data, where the multi-source heterogeneous data includes energy consumption data, network usage data, and security incident data; performing data cleaning, filling missing data values, and processing outliers on the multi-source heterogeneous data to obtain standardized data; inputting the standardized data into a pre-constructed real-time processing system to obtain feature data; inputting the feature data into a pre-constructed ensemble learning framework, and performing prediction through multiple basic models in the ensemble learning framework to obtain an ensemble prediction result; generating a comprehensive operation decision based on the ensemble prediction result, including an energy optimization strategy, a traffic scheduling strategy, and a risk warning strategy; and feeding back the comprehensive operation decision to the park management system.

[0004] Further, the data cleaning, data missing value filling, and data outlier handling include: converting the multi-source heterogeneous data into a unified format, removing redundant fields, and formatting the timestamps to obtain standardized data with a unified format; for the standardized data, using an interpolation algorithm to fill in the missing values, and when the data missing rate exceeds a preset threshold, supplementing the missing values by the weighted average method of neighboring data points; performing outlier detection on the complete data set, and when a data point deviates from the mean by more than three standard deviations, marking it as an outlier and correcting the outlier by the median replacement method.

[0005] Further, the real-time processing system adopts a microservices architecture, splitting the data processing tasks into multiple independent services. When the data volume exceeds the processing capacity of a single node, the computing nodes are dynamically increased through an elastic scaling mechanism, and heartbeat detection and data consistency verification are adopted. Further, the construction of the integrated learning framework includes: for the feature data, training a linear regression model to predict energy consumption data, training a decision tree model to predict network usage data, and training a random forest model to predict security event data to obtain multiple basic prediction results; inputting the multiple basic prediction results into the integrated learning framework, and using the weighted average method to combine the basic models. When the prediction error of a certain model is lower than the preset threshold, its weight coefficient is increased.

[0006] Further, the energy optimization strategy is based on the energy consumption prediction results, combines the energy-saving characteristics of the equipment and the electricity price ladder, and formulates the equipment power consumption priority and the on / off time for each period through an integer programming algorithm to generate an energy-saving scheduling plan.

[0007] Further, the traffic scheduling strategy is based on the network usage prediction results, adopts the ant colony optimization algorithm to solve the multi-objective optimization problem, balance the network load, and reasonably allocate bandwidth resources to generate a traffic scheduling plan.

[0008] Further, the risk warning strategy is based on the security event prediction results, adopts the fuzzy comprehensive evaluation method to evaluate the possibility and impact degree of each risk event occurring, generates a risk level matrix, and formulates physical prevention, technical prevention, and human prevention measures for high-risk events.

[0009] Further, the generation of the comprehensive operation decision adopts the analytic hierarchy process, taking the energy optimization, traffic scheduling, and risk warning strategies as the criterion layer and the specific measures as the scheme layer, determining the weights through pairwise comparison, and running the analytic hierarchy process model to obtain the strategy combination with the highest comprehensive score.

[0010] A park management service device, based on the park management service method according to any one of claims 1-8, characterized in that: a data acquisition and preprocessing module, a real-time processing system module, an integrated learning framework module, and a comprehensive operation decision module; A data acquisition and preprocessing module, which is used to obtain multi-source heterogeneous data and perform data cleaning, filling of missing data values, and processing of data outliers on the multi-source heterogeneous data to obtain standardized data; A real-time processing system module, which is used to input the standardized data into a pre-constructed real-time processing system to obtain feature data; An integrated learning framework module, which is used to input the feature data into a pre-constructed integrated learning framework and perform predictions through multiple base models in the integrated learning framework to obtain an integrated prediction result; A comprehensive operation decision-making module, which is used to generate comprehensive operation decisions according to the integrated prediction result, including energy optimization strategies, traffic scheduling strategies, and risk warning strategies.

[0011] A computer-readable storage medium, on which a computer program is stored, characterized in that: when the computer program is executed by a processor, the steps of the park management service method described in any one of claims 1 to 8 are implemented.

[0012] The technical solution provided by the embodiment of the present invention may include the following beneficial effects: The present invention discloses an intelligent park service method. The system obtains energy consumption, network usage, and security event data from multi-source heterogeneous data, and processes it through technologies such as data cleaning, standardization, and filling of missing values to obtain a complete data set. An outlier detection algorithm based on statistical distribution is used to correct data outliers. The micro-service architecture and elastic expansion mechanism are used to achieve large-scale real-time data processing. Prediction results are generated through feature extraction, multi-model training, and an integrated learning framework, and then comprehensive operation decisions such as energy optimization, traffic scheduling, and risk warning are formulated. During the decision execution process, the present invention uses methods such as a fault tolerance mechanism and heartbeat detection to ensure the stable operation of the system. The system realizes intelligent analysis and decision optimization of multi-dimensional data in the park, and improves the operation efficiency and security of the park. Description of the Drawings

[0013] Figure 1 It is a flowchart of a park management service method of the present invention. Detailed Embodiments

[0014] In order to make the purpose, technical solution, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the drawings and specific embodiments.

[0015] As Figure 1 , a park management service method in this embodiment may specifically include: Step S101, obtain energy consumption data, network usage data, and security event data from multi-source heterogeneous data, and use data cleaning technology to perform standardization processing on the data, remove redundant fields, and unify the time stamp format to obtain standardized data.

[0016] According to the pre-established list of multi-source heterogeneous data sources, energy consumption data, network usage data and security event data are obtained through API interfaces or crawlers. For the obtained original data, firstly, data format conversion is performed, all data are converted into CSV format, and uniformly encoded into UTF-8 to obtain the first data with uniform format. For the first data, a field matching algorithm is used to determine whether there is redundancy in the fields of each data. If there is a redundant field, the redundant field is removed to obtain the second data after the redundant field is removed. According to the preset standardization rules, the second data is subjected to standardization processing such as data type conversion and data format unification. Among them, the numerical data is normalized, the categorical data is encoded, and the text data is segmented and vectorized to obtain the standardized third data. For the third data, the timestamp field in the data is obtained, and the timestamp is converted into a unified date and time format (for example, yyyy-MM-ddHH:mm:ss) through the timestamp formatting function to obtain the fourth data with unified timestamp. The local sensitive hashing algorithm is used to determine whether there is duplicate data in the fourth data. If there is duplicate data, the duplicate data is removed to obtain the fifth data after deduplication. The field information of the fifth data is obtained, and the field names of different data sources are mapped to unified standard field names through the field mapping table to obtain the sixth data with standardized field names. For the sixth data, the data quality detection rules are used to determine whether the key fields of each data are complete and whether the values ​​are legal. If they are incomplete or illegal, the data is removed to obtain the seventh data with qualified quality. The seventh data is segmented according to the data source type to obtain energy consumption data, network usage data and security event data respectively. For the segmented data, field splicing, data merging and other processing are performed according to the preset data integration rules to obtain the integrated energy consumption data, network usage data and security event data. The integrated data is stored in the corresponding data tables of the Hive data warehouse respectively, and the partitions and indexes of the data tables are established to facilitate subsequent data query and analysis. The data table adopts the Parquet column storage format and the compression algorithm adopts Snappy, which can significantly improve the storage and access efficiency of data. Finally, based on the Hive data warehouse, efficient management and analysis of standardized energy consumption, network and security event data can be achieved, providing data support for the operation management and security prevention and control of the smart park.

[0017] Exemplarily, multi-source heterogeneous data acquisition is the primary step in smart park data processing. Taking energy consumption data as an example, power consumption data can be collected in real time through the API interface of smart meters; network usage data can be extracted from the logs of park network devices; security event data can be obtained by connecting to access control systems and surveillance cameras. These raw data are in various formats and need to be uniformly converted to the CSV format and encoded with UTF-8 for subsequent processing. Data redundancy processing is crucial for improving data quality. For example, the energy consumption data may simultaneously contain "total electricity consumption" and "sum of sub-item electricity consumption". Redundant fields can be identified and removed through a field matching algorithm, retaining the most original and valuable data. Data standardization is a key step in ensuring data consistency. For the electricity quantity values in energy consumption data, the min-max normalization method can be used to map them to the 0-1 interval; for categorical data such as device types in network usage data, one-hot encoding can be used to convert them into numerical types; for text data such as security event descriptions, they can be converted into vectors through word segmentation and TF-IDF algorithms. Timestamp unification processing helps with the time series analysis of data. Different formats of time information (such as Unix timestamps, ISO8601 format, etc.) are uniformly converted to the format of "yyyy-MM-dd HH:mm:ss", facilitating time comparison and aggregation analysis across data sources. Data deduplication is an important means to ensure data uniqueness. Similar data can be quickly identified through the locality-sensitive hashing algorithm. For example, duplicate reported energy consumption data caused by network fluctuations can be effectively removed by this method. Standardization of field names is conducive to the unified management of data. Through a predefined field mapping table, the data fields from different sources are uniformly named. For example, "elec_consumption", "power_usage", etc. are uniformly mapped to "electricity_consumption". Data quality detection ensures the reliability of the data stored in the database. For example, the rationality of the electricity quantity values in energy consumption data is verified, and obvious abnormal data points, such as negative values or values exceeding the rated power of the device, are excluded. Data segmentation and integration are the basis for realizing multi-dimensional analysis. The processed data is stored classified according to energy consumption, network usage, and security events, and correlation analysis is carried out according to business requirements. For example, by combining energy consumption data with network usage data, the energy consumption of IT devices can be analyzed in depth. Adopting the Hive data warehouse and Parquet columnar storage format, combined with the Snappy compression algorithm, not only greatly improves the data storage efficiency but also significantly enhances the query performance. This storage method is particularly suitable for scenarios such as smart parks that require frequent multi-dimensional data analysis, providing strong data support for applications such as real-time monitoring, energy consumption optimization, and security management.

[0018] Step S102. For the missing problems in the standardized data, an interpolation algorithm is used to fill in the missing values. If the data missing rate exceeds the threshold of 0.1%, the missing values are supplemented by the weighted average method of adjacent data points to obtain a complete data set.

[0019] For the standardized data set, first, count the proportion of missing values in the data set and calculate the missing rate. Compare the missing rate with the preset threshold of 1%. According to the comparison result, different missing value filling strategies are adopted. If the missing rate does not exceed 1%, the spline interpolation method is used for missing value estimation. Taking the position of the missing value as the interpolation point, select multiple data points before and after the missing value as sample points, use the cubic spline interpolation function to construct an interpolation polynomial, and calculate the estimated value of the missing value according to the interpolation polynomial. Fill the estimated value into the position corresponding to the missing value to complete the repair of the missing value. If the missing rate exceeds 1%, the KNN algorithm is used for missing value estimation. For each missing value, select the K nearest data points in its neighborhood. The neighborhood range can be set according to the characteristics of the data distribution. Calculate the Euclidean distance between the missing value and each adjacent point, and calculate the weight coefficient based on the distance. The weight coefficient is inversely proportional to the distance. Perform weighted averaging on the values of the K adjacent points and the corresponding weight coefficients to obtain the estimated value of the missing value. Fill the estimated value into the position corresponding to the missing value to complete the repair of the missing value. Traverse all the missing values in the data set and repeat the above missing value repair steps until all the missing values are estimated and filled. After filling, output the repaired standardized data set as the input for subsequent data analysis and machine learning.

[0020] Exemplarily, the handling of missing values in the dataset is crucial for ensuring data quality. First of all, counting the proportion of missing values is a key step in evaluating data integrity. Taking the energy consumption data of a smart campus as an example, assume that in the electricity meter data of a certain building, due to communication failures, data for some time periods is missing. Through calculation, it is known that the missing rate is 0.8%, which does not exceed the preset threshold of 1%. In this case, it is more appropriate to use the spline interpolation method for missing value estimation. The advantage of spline interpolation is that it can maintain the continuity and smoothness of the data, and is especially suitable for time series data. For example, for the missing 15-minute electricity consumption data, the data for 2 hours before and after the missing point can be selected as sample points to construct a cubic spline interpolation function. This method can well capture the intraday change trend of electricity consumption, such as the increase in electricity consumption during working hours and the decrease at night. However, when dealing with network usage data, the missing rate may exceed 1%. Assume that due to equipment failures, 2.5% of the traffic data of a certain network segment in the campus is missing. At this time, the KNN algorithm becomes a better choice. The advantage of the KNN algorithm is that it can perform similarity matching using multi-dimensional features and is suitable for complex data structures. When applying the KNN algorithm, 10 time points before and after the missing data point can be selected as the reference range. For each missing traffic data, calculate its Euclidean distance from other data points in the reference range in terms of time and known features (such as IP address, port number, etc.). Select the 5 points with the closest distance, and perform weighted average calculation according to the reciprocal of the distance as the weight. This method not only considers the continuity in time but also incorporates the feature information of network usage, and can estimate the missing values more accurately. When dealing with security incident data, the handling of missing values requires more caution. Assume that 0.5% of the data in the access control system records is missing, which is lower than the 1% threshold. Although spline interpolation can be used, considering the discreteness and importance of security incidents, simple numerical interpolation may introduce incorrect information. In this case, spline interpolation can be combined with logical judgment. For example, for the missing access control records, in addition to considering the continuity of the time series, information such as the weekday pattern and employee shift schedule also needs to be combined for reasonable estimation. Through these missing value handling methods, not only are the data gaps filled, but the internal logic and time series characteristics of the data are also maintained. This lays a solid foundation for subsequent data analysis and machine learning model training. For example, in the energy consumption prediction model, accurate historical data can improve the prediction accuracy; in network anomaly detection, complete traffic data helps to identify potential security threats; and in personnel behavior analysis, coherent access control records can better depict the work patterns of employees and the building usage efficiency.

[0021] Step S103: Detect outliers in the complete dataset. An outlier detection algorithm based on statistical distribution is used. If a data point deviates from the mean by more than three standard deviations, it is marked as an outlier. The outlier is corrected using the median replacement method to obtain the cleaned data.

[0022] Read the original dataset using the Pandas library in Python and obtain the numerical attribute columns of the dataset. For each numerical attribute, calculate its mean using the mean() function of Pandas and calculate its standard deviation using the std() function. Traverse each data point in the dataset. For each numerical attribute, determine whether its value deviates from the mean of the corresponding attribute by more than 3 standard deviations. If so, mark the data point as an outlier. For the data points marked as outliers, obtain all the non-outlier values of the attribute where they are located. Calculate the median of these non-outlier values using the median() function of Pandas. Replace the value of the data point marked as an outlier with the median of the non-outlier values of the corresponding attribute to obtain the corrected data point. To further improve the accuracy of outlier detection and correction, unsupervised outlier detection algorithms such as IsolationForest or LocalOutlierFactor in the scikit-learn library can be used to re-judge the suspected outlier data points. Screen and correct the data points according to the outlier scores or outlier labels of the algorithms. Recombine the corrected data points and use the concat() function of Pandas to generate the cleaned dataset after outlier detection and correction. Save the cleaned dataset as a CSV file for subsequent data analysis and modeling tasks, such as machine learning modeling using scikit-learn or data visualization using Matplotlib. The improvement of data quality helps to improve the accuracy and reliability of subsequent tasks.

[0023] Exemplarily, reading the original dataset using the Pandas library is the first step in data processing, which lays the foundation for subsequent outlier detection and correction. Taking the energy consumption data of a smart campus as an example, a CSV file containing information such as the electricity consumption and water consumption of each building can be read. After obtaining the numerical attribute columns, calculating the mean and standard deviation are the keys to identifying outliers. For example, the average daily electricity consumption of an office building is 10,000 kWh, and the standard deviation is 2,000 kWh. These statistical indicators reflect the central tendency and dispersion degree of the data, providing a reference standard for judging outliers. Traversing the data points and determining whether they deviate from the mean by 3 standard deviations is a commonly used outlier detection method. In the above example, if the electricity consumption on a certain day exceeds 16,000 kWh or is less than 4,000 kWh, it will be marked as an outlier. This method is based on the characteristics of the normal distribution and can capture data points that deviate significantly from the normal range. However, it may also misjudge some reasonable but rare data, such as a significant increase in electricity consumption due to high temperatures in summer. For the data points marked as abnormal, replacing them with the median of non-outliers is a robust processing method. The median is not affected by extreme values and can better represent the central tendency of the data. For example, if the abnormally high electricity consumption is caused by a temporary large-scale event, using the median replacement can better reflect the daily electricity consumption situation. To further improve the detection accuracy, it is very necessary to use algorithms such as IsolationForest or LocalOutlierFactor for re-judgment. These algorithms can consider features in multiple dimensions, such as considering the relationship between electricity consumption, water consumption, and pedestrian flow at the same time. For example, the electricity consumption and pedestrian flow of a shopping mall may be relatively high on weekends, but they match each other and should not be simply judged as abnormal. IsolationForest identifies outliers by constructing decision trees, while LocalOutlierFactor detects outliers by calculating local density. These methods can capture more complex outlier patterns. Re-combine the corrected data and save it as a CSV file to prepare for subsequent data analysis and modeling tasks. For example, when performing energy consumption prediction, the cleaned data can improve the accuracy of the model. If there are outliers in the original data, such as abnormally high electricity consumption due to equipment failure, it may cause the prediction model to overestimate future energy consumption. By detecting and correcting outliers, these interference factors can be eliminated, enabling the model to better capture the real energy consumption pattern. In practical applications, this set of processes can help smart campus managers understand the energy usage situation more accurately and discover potential energy waste problems. For example, by comparing the energy consumption data of different buildings, it may be found that some buildings still have abnormally high electricity consumption during non-working hours, which may indicate that the equipment is not turned off in time or there is an energy leakage problem. By promptly discovering and handling these abnormal situations, the energy usage efficiency of the campus can be significantly improved, and the operating costs can be reduced. In addition, the cleaned data can also be used for abnormal event detection.For example, in a security system, by analyzing data such as the number of people flow and access control records, suspicious activities can be detected in a timely manner. If the number of people flow at a certain time point far exceeds the normal level and shows a significant difference from historical data, it may indicate potential security risks. In this way, data cleaning not only improves the data quality but also provides reliable support for the intelligent management and decision-making of the park.

[0024] Step S104: Input the cleaned data into the real-time processing system. The data processing tasks are split into multiple independent services using a microservices architecture. If the data volume exceeds the single-node processing capacity of 1 million records per second, the computing nodes are dynamically increased through an elastic scaling mechanism to ensure real-time processing performance.

[0025] Clean the original data according to the data cleaning component to remove noise and invalid data, improving the data quality. The high-quality cleaned data provides a reliable data foundation for subsequent real-time processing and reduces anomalies during the data processing. Buffer and distribute the cleaned data through the Kafka message queue to ensure high throughput and low latency in the real-time processing system. Design the real-time processing system based on the microservices architecture, splitting the data processing into multiple independent services, with each service responsible for a specific processing function, and realizing communication and collaboration between services through REST APIs. For real-time data aggregation and statistical analysis, use the Flink streaming computing framework, leveraging its window mechanism and state management capabilities to achieve real-time summarization and statistics of data. For real-time prediction and anomaly detection, classic machine learning algorithms such as decision trees and random forests can be selected to train historical data to build a prediction model for predicting real-time data and identifying anomalies. For monitoring the data processing performance, establish a monitoring system such as Prometheus to collect key metrics such as data throughput and latency. Set reasonable thresholds, trigger alarms when the metrics exceed the thresholds, and use Kubernetes for elastic scaling according to the data volume growth trend and processing requirements to adjust the resource configuration of the real-time processing cluster. The result data of the real-time processing is stored in the distributed time series database InfluxDB, which is suitable for storing time series data and supports high-concurrency writing and aggregation query analysis. Collect logs from each service node through the data collection agent Fluentd, record the metadata information of each link in the data processing, and establish data lineage. At the same time, set quality monitoring rules for key data fields to monitor the timeliness, integrity, accuracy, etc. of the data, and trigger alarm processing in a timely manner when data quality problems are found. After the above optimizations, the entire real-time data processing system can fully utilize the distributed computing power, ensuring real-time and accurate data while possessing high availability and elastic scaling capabilities, providing stable and efficient data support for business applications.

[0026] Exemplarily, the data cleaning component is the cornerstone of the real-time data processing system in the smart park. It improves data quality by removing noise and invalid data. For example, in energy consumption monitoring, sensors may generate abnormal readings due to faults. For instance, the instantaneous electricity consumption of an office building reaches ten times the normal level. The cleaning component will identify and eliminate such outliers to ensure the accuracy of subsequent analysis. The high-quality data after cleaning is buffered and distributed through the Kafka message queue. This method can effectively handle sudden increases in data traffic. For example, during large-scale events, the pedestrian flow in the park surges, resulting in a sharp increase in various types of data. Kafka can easily handle this high-concurrency situation and ensure the timely transmission and processing of data. The real-time processing system with a microservices architecture splits complex data processing tasks into multiple independent services. In the smart parking scenario, services such as parking space detection, license plate recognition, and billing services can be set up. This design makes the system easier to maintain and expand. When new functions need to be added, such as adding electric vehicle charging management, only new microservices need to be added without modifying the existing system. The Flink streaming computing framework plays an important role in real-time data aggregation and statistical analysis. Taking park energy management as an example, Flink can use a sliding time window to calculate the average electricity consumption of each building every 15 minutes, and at the same time record the historical peak electricity consumption through the state management function to achieve dynamic energy consumption monitoring. This real-time analysis ability enables park managers to detect abnormal energy consumption in a timely manner and quickly take energy-saving measures. In real-time prediction and anomaly detection, machine learning algorithms such as decision trees and random forests are widely used. Taking the park security system as an example, a model can be trained based on historical data to predict the normal pedestrian flow in different time periods and regions. When the real-time monitored pedestrian flow significantly deviates from the predicted value, an alarm is issued to remind security personnel to pay attention to possible security risks. Performance monitoring is the key to ensuring the stable operation of the real-time processing system. Monitoring systems such as Prometheus can track various performance indicators in real time. For example, in data center monitoring, an alarm can be triggered when the CPU usage exceeds 80% or the memory usage exceeds 90%. When the system load continues to increase, the processing cluster can be automatically scaled by Kubernetes. For example, processing nodes are automatically added during peak periods to ensure the response speed. As a distributed time series database, InfluxDB is very suitable for storing real-time data in the smart park. It can efficiently store and query a large amount of time series data, such as elevator operation status, air conditioner temperature settings, etc. This storage method facilitates subsequent data analysis in the time dimension. For example, analyzing the elevator usage frequency at different times to optimize the elevator scheduling strategy. Fluentd plays an important role in establishing data lineage. It can collect and integrate the logs of each service node, recording the whole process of data from collection, cleaning to processing. This data lineage helps with problem tracking and data tracing. For example, when an abnormal piece of data is found, the problem source can be quickly located through the lineage, whether it is a sensor failure or a problem in the data processing link.Through the above optimizations, the real-time processing system of the smart park can efficiently and accurately process massive amounts of data, providing timely and reliable decision-making support for park management. Whether it is energy management, security monitoring, or facility scheduling, this real-time processing system can provide powerful data support, driving the park towards a more intelligent and efficient direction.

[0027] Step S105: In the real-time processing system, extract features and normalize the energy consumption data, network usage data, and security incident data. Then, build a basic prediction model for each of these processed data, and train the model using linear regression, decision tree, and random forest algorithms to obtain multiple basic prediction results.

[0028] Obtain the energy consumption data, network usage data, and security incident data in the real-time processing system. For the obtained data, use the Pandas library in Python to clean the data, remove noise data such as missing values and outliers, and improve the data quality. The cleaned data will be used for subsequent feature extraction and model construction. According to the cleaned data, use the describe() function of Pandas to perform statistical analysis on the data, extract the key attributes that can reflect the data characteristics, and construct the feature vectors. Adopt the min-max normalization method and use sklearn.preprocessing.MinMaxScaler to normalize the feature vectors, map the features with different dimensions to the same scale of [0, 1], and eliminate the influence of dimensions. For the normalized energy consumption feature data, adopt the linear regression algorithm to construct the energy consumption data prediction model. Linear regression can depict the linear change trend of the energy consumption data. Use sklearn.linear_model.LinearRegression to fit the regression coefficients through the training data. When new energy consumption feature data is obtained in real time, substitute it into the regression equation to obtain the corresponding energy consumption prediction value. For the normalized network usage feature data, adopt the decision tree algorithm to construct the network usage data prediction model. The decision tree can depict the non-linear characteristics of the network usage data by recursively dividing the feature space. Use sklearn.tree.DecisionTreeRegressor to generate the decision tree through the training data. When new network usage feature data is obtained in real time, predict the network usage according to the judgment conditions of the decision tree. For the normalized security incident feature data, adopt the random forest algorithm to construct the security incident prediction model. The random forest can reduce the variance of a single decision tree and improve the stability of the prediction by integrating multiple decision trees. Use sklearn.ensemble.RandomForestClassifier to generate multiple decision trees through the training data. When new security incident feature data is obtained in real time, obtain the final security incident prediction result through the majority voting mechanism according to the judgment results of each decision tree. Weightedly fuse the prediction results of the linear regression model, decision tree model, and random forest model on the validation set. The weight coefficients are calculated according to the mean squared error and accuracy of each model on the validation set. The model with a small mean squared error and high accuracy will obtain a greater weight to get the final comprehensive prediction result. According to the comprehensive prediction result, judge whether there are energy consumption anomalies, network usage anomalies, and security incident risks. By analyzing the distribution of historical data, set the warning threshold. If the predicted value exceeds the preset threshold, use the logging module in Python to generate warning log information to prompt the management staff to handle it in time to avoid the problem from expanding.

[0029] Exemplarily, the core of the intelligent park real-time data processing system lies in the efficient processing and analysis of various types of data. Taking energy consumption data as an example, data cleaning is first performed through the Pandas library. Suppose there are outliers in the electricity consumption data of an office building, such as the electricity consumption suddenly reaching ten times the normal level during a certain period. The cleaning process will identify and remove such abnormal data to ensure the accuracy of subsequent analysis. The cleaned data undergoes statistical analysis to extract key features. For example, statistical indicators such as the average value and standard deviation of electricity consumption can be obtained through the describe() function, and these indicators can reflect the overall characteristics of the electricity consumption pattern. To eliminate the dimensional difference between different features, the maximum-minimum normalization method is adopted. For instance, the electricity consumption data originally in the range of 0 - 10000 kWh is mapped to the range of 0 - 1, making it comparable to the network usage data with a proportion in the range of 0 - 100%. In terms of energy consumption data prediction, a linear regression model can capture the overall trend of electricity consumption changing over time. For example, by analyzing historical data, the model may find that the electricity consumption on weekdays shows a linear growth trend, while it shows a downward trend on weekends. This model can provide basic electricity consumption predictions for park managers, helping to arrange power resources reasonably. The decision tree algorithm is used for the prediction of network usage data. Decision trees can handle non-linear relationships. For example, network traffic may suddenly increase during working hours and decrease sharply during lunch breaks. By constructing a decision tree, the network usage situation can be accurately predicted based on features such as time and the number of users, providing a basis for the IT department to allocate network resources. The random forest algorithm is used for security incident prediction. Suppose there are multiple security monitoring points in the park, and each point collects data such as the number of people and noise levels. The random forest can more accurately predict potential security risks by integrating the judgments of multiple decision trees. For example, when the number of people in a certain area increases abnormally and the noise level also rises, it may warn of a possible mass incident. Model fusion is a key step in improving prediction accuracy. By taking the weighted average of the prediction results of each model, the influence of different factors can be comprehensively considered. For example, in energy consumption prediction, if the mean squared error of the linear regression model is smaller, it will obtain a greater weight and thus play a greater role in the final prediction. By analyzing historical data, a reasonable warning threshold is set. For example, if the predicted electricity consumption exceeds 150% of the same period in history, a warning log will be generated to remind managers to pay attention to possible electricity consumption abnormalities and take energy-saving measures or troubleshoot faults in a timely manner. The value of this intelligent prediction and warning system lies in its foresight and comprehensiveness. It can not only predict the situation in a single field but also discover potential associated risks through data fusion. For example, when it is predicted that the network usage volume surges and at the same time the security incident risk increases, it may indicate a network security threat. This multi-dimensional analysis ability provides comprehensive decision-making support for park management, helping to improve the overall operation efficiency and security level of the park.

[0030] Step S106: Input multiple basic prediction results into an ensemble learning framework, and use the weighted average method to combine the basic models. If the prediction error of a certain model is lower than the preset threshold of 5%, increase its weight to obtain the ensemble prediction result.

[0031] Obtain the prediction results of multiple basic models for the test data, compare the prediction results with the true labels, and calculate the prediction error of each model. The prediction error can use common error metric indicators such as mean squared error and mean absolute error. Set a prediction error threshold, such as 5%. For models with a prediction error lower than the threshold, mark them as high-confidence models. The selection of the threshold can be adjusted according to the requirements of the specific task and the distribution of model performance. For high-confidence models, increase their weight coefficients in the ensemble learning framework. The adjustment of the weight coefficients can be done in an additive or multiplicative way, such as multiplying the weight coefficient by 2 or adding 2. The adjustment range of the weight coefficients can be set according to the confidence level of the model and the importance of the task. Use the weighted average method to combine the prediction results of each basic model. Let the prediction result of the i-th model be Pi and the weight coefficient be Wi, then the ensemble prediction result P = ∑(Wi×Pi) / ∑Wi. Through weighted average, the prediction results of high-confidence models will have a greater weight in the final ensemble prediction. Use the k-fold cross-validation method to evaluate the performance of the ensemble prediction result. Randomly divide the dataset into k subsets, each time select one subset as the test set, and the remaining subsets as the training set, conduct k times of training and testing, and take the average of the k evaluation metrics as the final performance evaluation result. Common evaluation metrics include accuracy, precision, recall, F1 value, etc. According to the results of cross-validation, adjust the weight coefficients of each basic model to optimize the ensemble learning framework. Methods such as grid search and random search can be used to optimize the weight coefficients and find the optimal weight combination that can maximize the evaluation metrics. Apply the optimized ensemble learning framework to the actual prediction task. For new input data, use the trained basic models to make predictions, and then perform weighted average according to the optimized weight coefficients to obtain the final ensemble prediction result. The ensemble prediction result combines the prediction capabilities of multiple models and can improve the accuracy and robustness of the prediction.

[0032] Exemplarily, the core of the ensemble learning framework lies in effectively integrating the prediction capabilities of multiple base models to improve the overall prediction accuracy. First, predictions are made on the test data and compared with the true labels to calculate the prediction error. Taking the energy consumption prediction in a smart park as an example, assume that the mean squared error of the linear regression model is 3%, that of the decision tree model is 6%, and that of the random forest model is 4%. Setting the prediction error threshold at 5%, the linear regression and random forest are marked as high-confidence models. Increasing the weight coefficients of the high-confidence models is to highlight their contributions in the ensemble prediction. For example, the weight coefficients of the linear regression and random forest are increased from 1 to 2 and 1.5 respectively. This adjustment reflects the differences in model performance and helps the final prediction result to rely more on the better-performing models. Weighted average is a simple and effective ensemble method. Assume that the linear regression predicts the electricity consumption during a certain period to be 1000 kWh with a weight of 2; the decision tree predicts 1200 kWh with a weight of 1; and the random forest predicts 1100 kWh with a weight of 1.5. Then the ensemble prediction result is (2×1000 + 1×1200 + 1.5×1100) / (2 + 1 + 1.5) ≈ 1077 kWh. This result combines the predictions of the three models and may be more robust than the prediction of a single model. k-fold cross-validation is a reliable method for evaluating model performance. Taking 5-fold cross-validation as an example, the dataset is divided into 5 parts, and each time 4 parts are used as the training set and 1 part as the test set. After 5 rounds of training and testing, 5 accuracy values are obtained, such as 92%, 94%, 93%, 91%, 95%, and the average value of 93% is taken as the final evaluation result. This method can reduce the contingency caused by a single division and provide a more objective performance evaluation. Adjusting the weight coefficients according to the cross-validation results is a key step in optimizing the ensemble framework. For example, through grid search, different weight combinations are tried: (2, 1, 1.5), (2.5, 1, 1.5), (2, 0.5, 2), etc., and finally the weight combination that maximizes the evaluation index is found. This process may find that further increasing the weight of the random forest to 2 and slightly reducing the weight of the decision tree to 0.8 can increase the F1 value of the ensemble model by 2 percentage points. When the optimized ensemble learning framework is applied to actual prediction, it can give full play to the advantages of each model. For example, when predicting network usage, the decision tree may perform well in capturing the usage patterns on weekdays and weekends, while the random forest is better at handling special situations such as holidays. Through reasonable weight allocation, the ensemble prediction result can maintain a high accuracy in different scenarios. The advantage of this ensemble method lies in its flexibility and robustness. It can not only improve the prediction accuracy but also reduce the risks that may be brought by a single model. For example, in the prediction of security incidents, a certain model may be particularly sensitive to specific types of anomalies, while another model performs better in other aspects. Through ensemble, stable prediction performance can be maintained in various situations, providing reliable decision support for the security management of the smart park.

[0033] Step S107: According to the integrated prediction results, generate an energy optimization strategy for energy consumption data, a traffic scheduling strategy for network usage data, and a risk warning strategy for security incident data, and form a comprehensive operation decision.

[0034] According to the integrated prediction results, the following steps are adopted to generate a comprehensive operation decision-making plan to achieve global optimization management of energy, network, and security: Data preprocessing: Perform preprocessing operations such as cleaning and normalization on energy consumption data, network usage data, and security incident data, and extract key features. Use dimensionality reduction algorithms such as principal component analysis (PCA) to compress the data dimension. Train machine learning models: Select machine learning algorithms suitable for time series prediction, such as long short-term memory network (LSTM), and train multiple models to predict the energy consumption demand, network traffic, and probability of security incidents in the future for a certain period. Use historical data to divide the training set and validation set, and perform cross-validation to select the optimal model. Generate an energy optimization strategy: According to the energy consumption data predicted by the LSTM model, combined with factors such as the energy-saving characteristics of equipment and electricity price ladders, use operational research optimization algorithms such as integer programming to formulate the power consumption priority and on / off time of equipment for each period, and form an energy-saving scheduling plan. Generate a traffic scheduling strategy: According to the network usage data predicted by the LSTM model, adopt heuristic algorithms such as ant colony optimization to solve multi-objective optimization problems, balance network load, reasonably allocate bandwidth resources, generate a traffic scheduling plan, and avoid link congestion. Generate a risk warning strategy: According to the security incident data predicted by the LSTM model, adopt the fuzzy comprehensive evaluation method to evaluate the possibility and impact degree of each risk event, and generate a risk level matrix. For high-risk events, formulate corresponding physical prevention, technical prevention, and human prevention measures. Integrate the three strategies: Construct an analytic hierarchy process (AHP) model, use the three strategies of energy optimization, traffic scheduling, and risk warning as the criterion layer, and each specific measure as the scheme layer. Determine the weights through pairwise comparison, and run the AHP model to obtain the strategy combination with the highest comprehensive score, forming a globally optimal operation decision-making plan. Strategy execution and dynamic adjustment: Convert the decision-making plan into control instructions and transmit them to the energy management and control system, network controller, etc. to achieve automated operations such as equipment on / off, traffic restriction, and security linkage. At the same time, establish a human-computer interaction interface to allow managers to adjust parameters and modify strategies according to the actual situation. Through the above steps, this plan uses machine learning algorithms to predict future trends, combines optimization and decision-making models to formulate global strategies, and is implemented in a way that combines automated execution and manual intervention, which can effectively improve energy use efficiency, ensure network performance, and respond to security risks, realizing intelligent operation management of facilities.

[0035] Exemplarily, data preprocessing is the foundation of the operation decision-making in the smart park. For example, the energy consumption data may contain outliers. For instance, the electricity consumption during a certain period suddenly soars to ten times the normal level, which may be caused by equipment failures. Identifying and processing these outliers through methods such as the median method can improve the accuracy of subsequent analysis. For network usage data, there may be records at different granularities, such as traffic data counted by seconds, minutes, or hours. Through data aggregation, all data is unified to the hourly level for subsequent modeling and analysis. Machine learning model training is the key to predicting future trends. The Long Short-Term Memory (LSTM) network is suitable for capturing long-term dependencies in time series data. In energy consumption prediction, LSTM can learn the differences in electricity consumption patterns between weekdays and weekends, as well as the special electricity consumption situations during holidays. For example, by analyzing historical data, the model may find that the electricity consumption in the park during the Spring Festival each year is 30% lower than normal, while it increases by 25% during the peak summer electricity consumption period compared to normal. These patterns are memorized by the LSTM network to improve the accuracy of future predictions. Energy optimization strategy generation is an important means to reduce operating costs. Integer programming algorithms can find the optimal solution considering various constraints. For example, a certain park has 100 air conditioners, each with a power of 2 kilowatts. During the peak electricity consumption period, the electricity price is 1.2 yuan per kilowatt-hour, while it is 0.5 yuan per kilowatt-hour during the off-peak period. Through integer programming, an air conditioner switching strategy can be formulated to maximize the use of off-peak electricity prices while ensuring comfortable indoor temperatures, such as precooling during the off-peak period and appropriately increasing the temperature setting during the peak period, thereby reducing the overall electricity cost. Flow scheduling strategy optimization helps improve network performance. The ant colony algorithm simulates the foraging behavior of ants and can effectively solve complex combinatorial optimization problems. In network traffic scheduling, different types of data packets can be regarded as different ant colonies, and bandwidth resources as food sources. Through iterative optimization, the best traffic allocation scheme can be found. For example, during the peak period of video conferencing, the algorithm may suggest allocating 80% of the bandwidth to video traffic, 15% to file transfer, and 5% to other applications to ensure the quality of the conference while taking into account the needs of other services. Risk warning strategy generation is the core of ensuring the safety of the park. The fuzzy comprehensive evaluation method can handle problems such as security risks that are difficult to precisely quantify. For example, for fire risk, factors such as temperature, humidity, and smoke concentration can be considered. Suppose the temperature in a certain area is 35°C, the relative humidity is 20%, and the smoke concentration is 10 ppm. According to the fuzzy rules set by expert experience, the fire risk level in this area may be "medium high", and preventive measures such as increasing the inspection frequency and turning on the automatic sprinkler system are recommended. The Analytic Hierarchy Process (AHP) is used to integrate multiple strategies and balance the needs of all aspects. When constructing the AHP model, the three aspects of energy, network, and security can be used as the first-level indicators, and their specific measures as the second-level indicators. The weights are determined by expert scoring, such as energy accounting for 40%, network accounting for 30%, and security accounting for 30%. Further refined, energy efficiency may account for 60% of the energy indicator, and cost accounts for 40%.In this way, a comprehensive decision-making plan considering various factors can be obtained. Finally, strategy execution and dynamic adjustment ensure the implementation effect of the plan. For example, according to the optimized energy strategy, an instruction may be issued to lower the air-conditioning temperature in the office area by 2°C for precooling from 22:00 to 6:00 the next day. At the same time, managers can adjust parameters at any time through the human-machine interaction interface. In case of special weather, the precooling strategy can be temporarily changed. This flexible management method not only ensures the efficiency improvement brought by automation but also retains the flexibility of human intervention, thus realizing the efficient operation of the smart park.

[0036] Step S108: Feed back the comprehensive operation decision to the park management system, monitor the decision execution situation through the fault tolerance mechanism of the microservice architecture. If a certain service node fails, automatically switch to the standby node, use heartbeat detection and data consistency verification to ensure stable operation, and record the changes of key indicators during the decision execution process.

[0037] According to the feedback of comprehensive operation decisions, obtain decision execution tasks from the message queue, decompose the tasks into multiple subtasks through the task orchestration module, and generate a directed acyclic graph (DAG) in combination with task dependencies. The consistent hashing algorithm is used to allocate subtasks to different microservice nodes, and virtual nodes are introduced into the mapping relationship between tasks and nodes to achieve the balance of task allocation. The running status of each microservice node is monitored in real time through the timed heartbeat detection mechanism. If a node failure or response timeout is found, it is removed from the consistent hashing ring, and the task allocation scheme is recalculated based on virtual nodes, and the tasks of the failed node are reassigned to other available nodes. For critical tasks and data, master-slave multi-copy fault-tolerant deployment is adopted, and the state synchronization and data consistency guarantee between replicas are achieved through the Raft protocol. Multiple replica nodes form a Raft cluster, and a Leader node is elected through voting. The Leader node receives and submits tasks, and the follower nodes asynchronously replicate the state of the Leader. When the Leader fails, a new Leader is automatically elected to ensure the high availability of critical tasks. The Raft protocol is based on the majority voting mechanism, and only the commit records confirmed by more than half of the nodes are considered consistent, avoiding the problems of split brain and data inconsistency. During the task execution process, a real-time monitoring system is built based on Prometheus and Grafana to collect resource metrics such as CPU, memory, and network of each node, as well as business metrics such as task completion rate and response time. For the collected time series data, the time series prediction tool Prophet open-sourced by Facebook is used for trend analysis and anomaly detection. The Prophet model decomposes the time series data into seasonal and trend components, constructs a generalized linear model, automatically fits the non-linear trend, and takes into account factors such as holiday effects to accurately predict the future indicator trends and identify abnormal fluctuations. When a key indicator is monitored to exceed the preset threshold, an alarm is triggered through the Alertmanager service, and operation and maintenance personnel are notified through multiple channels such as SMS, email, and enterprise WeChat for timely handling. At the same time, the preset emergency handling script is called to dynamically adjust parameters such as node resource quotas and task concurrency to achieve self-healing and defense against faults. After the task execution is completed, the key indicator data during the task execution process, such as completion rate, processing delay, and resource consumption, are transmitted to the data warehouse for offline analysis. Data cleaning and feature engineering are performed on the task execution data, and feature dimensionality reduction is performed through the Principal Component Analysis (PCA) algorithm to extract the behavior portraits of different decision tasks. The K-means clustering algorithm is used to automatically cluster decision tasks, and the clustering effect under different K values is evaluated through the Silhouette coefficient to obtain the optimal task clustering result.The clustering results are displayed through a visualization dashboard. Based on the clustering analysis results, operators can take targeted optimization measures for different types of decision-making tasks, adjust task generation and scheduling strategies, and continuously improve the efficiency and quality of decision execution. The clustering results are also fed back to the feature storage of the comprehensive operation platform, enriching the decision feature dimensions, optimizing the learning samples of the decision engine, and enhancing the decision-making intelligence level. By establishing a full-link closed-loop mechanism for task decomposition, node monitoring, fault recovery, and task analysis feedback, this solution can achieve high availability, high efficiency, and continuous optimization in the decision execution process, providing strong support and guarantee for comprehensive operation decision-making.

[0038] Exemplarily, the intelligent park operation decision execution system uses a message queue to obtain tasks and decomposes complex tasks into multiple subtasks through a task orchestration module. For example, the air-conditioning energy consumption optimization task may be decomposed into subtasks such as data collection, model training, and policy generation. A directed acyclic graph (DAG) is used to represent the dependencies between tasks to ensure the correctness of the execution order. The consistent hashing algorithm is used for task allocation. By mapping tasks and nodes to the hash ring, load balancing is achieved. Introducing virtual nodes can further improve the uniformity of allocation. Suppose there are 3 physical nodes, and 100 virtual nodes are introduced for each node. Then there are 300 points on the hash ring, greatly reducing the possibility of uneven data distribution. The node status is monitored through heartbeat detection. When a node failure is detected, it will be removed from the consistent hash ring and tasks will be reallocated. For example, if a node is responsible for processing video surveillance data and this node fails, its tasks will be immediately transferred to a neighboring healthy node to ensure service continuity. For critical tasks, the system uses the Raft protocol to achieve multi-copy deployment. Suppose 5 replica nodes are deployed. As long as 3 nodes are running normally, the system can keep working properly. The Raft protocol ensures data consistency by electing a Leader node. When the Leader node fails, the remaining nodes will quickly elect a new Leader to minimize service interruption time. The real-time monitoring system is built based on Prometheus and Grafana to collect various metric data. For example, it may be monitored that the CPU usage rate of a certain node suddenly soars to 95%, which is much higher than the normal level of 60%. At this time, the Prophet model will analyze historical data, predict future trends, and identify this as an abnormal fluctuation. When the metric exceeds the preset threshold, an alarm will be triggered. If it is detected that the network bandwidth usage rate continuously exceeds 80%, it will notify the operation and maintenance personnel through multiple channels and automatically execute a preset script, such as restricting the bandwidth usage of non-critical services, to ensure the normal operation of core services. After the task execution is completed, the data is analyzed offline. Through the PCA algorithm for feature dimensionality reduction, it may be found that the main influencing factors of the energy consumption optimization task are outdoor temperature and population density. The K-means clustering algorithm may classify decision-making tasks into three categories: high energy consumption, medium energy consumption, and low energy consumption. Operators can formulate differentiated energy-saving strategies based on this. This full-link closed-loop mechanism not only ensures the high availability and efficiency of decision execution but also continuously optimizes the decision-making quality through continuous data analysis and feedback. For example, it may be found that on Monday mornings, the electricity consumption in the park often shows a short-term peak. Based on this discovery, the decision-making engine can adjust the energy distribution strategy in advance to effectively smooth the electricity peak and reduce the overall operation cost.

[0039] Only some preferred embodiments of the present invention are listed above, but the present invention is not limited thereto, and many improvements and transformations can be made. As long as the improvements and transformations are made on the basis of the basic principle of the present invention, they shall be regarded as falling within the protection scope of the present invention.

Claims

1. A park management service method, characterized in that: include: Acquiring multi-source heterogeneous data, wherein the multi-source heterogeneous data includes energy usage data, network usage data, and security event data; Performing data cleaning, data missing value filling and data outlier processing on the multi-source heterogeneous data to obtain standardized data; Inputting the standardized data into a pre-built real-time processing system to obtain feature data; Inputting the feature data into a pre-built integrated learning framework, performing predictions through multiple basic models in the integrated learning framework, and obtaining integrated prediction results; Generate comprehensive operational decisions based on the integrated prediction results, including energy optimization strategies, traffic scheduling strategies, and risk warning strategies; The comprehensive operational decision is fed back to the park management system.

2. The method according to claim 1, characterized in that The data cleaning, data missing value filling and data outlier processing include: Convert the multi-source heterogeneous data into a unified format, remove redundant fields, format the timestamp, and obtain standardized data in a unified format; For the standardized data, an interpolation algorithm is used to fill in the missing values. When the data missing rate exceeds a preset threshold, the missing values ​​are supplemented by a weighted average method of neighboring data points; Anomaly detection is performed on the complete data set. When a data point deviates from the mean by more than three standard deviations, it is marked as an outlier and corrected by the median replacement method.

3. The method according to claim 1, characterized in that The real-time processing system adopts a microservice architecture to split the data processing task into multiple independent services. When the data volume exceeds the processing capacity of a single node, computing nodes are dynamically added through an elastic expansion mechanism, and heartbeat detection and data consistency verification are used.

4. The method according to claim 1, characterized in that The construction of the integrated learning framework includes: Based on the feature data, a linear regression model is trained to predict energy consumption data, a decision tree model is trained to predict network usage data, and a random forest model is trained to predict security event data, so as to obtain multiple basic prediction results; The multiple basic prediction results are input into the integrated learning framework, and the basic models are combined using the weighted average method. When the prediction error of a certain model is lower than a preset threshold, its weight coefficient is increased.

5. The method according to claim 1, characterized in that The energy optimization strategy is based on the energy consumption forecast results, combined with the energy-saving characteristics of the equipment and the electricity price ladder. Through the integer programming algorithm, the power consumption priority and on / off time of the equipment in each time period are formulated to generate an energy-saving scheduling plan.

6. The method according to claim 1, characterized in that The traffic scheduling strategy is based on the network usage prediction results, adopts the ant colony optimization algorithm, solves the multi-objective optimization problem, balances the network load, reasonably allocates bandwidth resources, and generates a traffic scheduling plan.

7. The method according to claim 1, characterized in that The risk warning strategy is based on the security event prediction results and adopts a fuzzy comprehensive evaluation method to evaluate the possibility and impact of each risk event, generate a risk level matrix, and formulate physical, technical and human defense measures for high-risk events.

8. The method according to claim 1, characterized in that The generation of the comprehensive operational decision adopts the hierarchical analysis method, which takes energy optimization, flow scheduling and risk warning strategies as the criterion layer and specific measures as the solution layer. The weights are determined by pairwise comparison, and the hierarchical analysis model is run to obtain the strategy combination with the highest comprehensive score.

9. A park management service device, based on the park management service method described in any one of claims 1 to 8, characterized in that: Data collection and preprocessing module, real-time processing system module, integrated learning framework module and comprehensive operation decision module; The data acquisition and preprocessing module is used to obtain multi-source heterogeneous data and perform data cleaning, data missing value filling and data outlier processing on the multi-source heterogeneous data to obtain standardized data; A real-time processing system module is used to input standardized data into a pre-built real-time processing system to obtain characteristic data; An integrated learning framework module is used to input feature data into a pre-built integrated learning framework, perform predictions through multiple basic models in the integrated learning framework, and obtain integrated prediction results; The comprehensive operation decision module is used to generate comprehensive operation decisions based on the integrated prediction results, including energy optimization strategy, traffic scheduling strategy and risk warning strategy.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the park management service method described in any one of claims 1 to 8 are implemented.