Data analysis method and system for smart city gateway system

By performing adaptive sampling, differentiated edge computing and context-aware compression in the smart city gateway system, and performing adaptive task scheduling and distributed load balancing, the problems of insufficient data transmission bandwidth and high processing delay in existing smart city systems are solved, and more efficient data processing and computing load optimization are achieved.

CN119484538BActive Publication Date: 2025-06-20GUANGDONG SHENCHUANG OPTOELECTRONICS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510069399.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-20
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The existing smart city systems have insufficient transmission bandwidth, high processing delays, excessive server load due to excessive data volume and centralized processing architecture, and lack of efficient hierarchical data processing and task dynamic scheduling mechanisms.

Method used

Through the smart city gateway system, the original data of IoT devices and sensors is adaptively sampled and initially classified, combined with differentiated edge computing and context-aware compression processing, adaptive task scheduling and distributed load balancing processing are performed, local task processing is realized and transmitted to the central server.

Benefits of technology

Reduces data transmission pressure, optimizes computing load, and improves the response speed and processing capabilities of smart city systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119484538B_ABST
    Figure CN119484538B_ABST
Patent Text Reader

Abstract

The present invention provides a data analysis method and system for a smart city gateway system. The method includes: adaptively sampling and preliminarily classifying the original data of Internet of Things devices and sensors through the smart city gateway system to generate a classified data set. Performing differential edge computing and context-aware compression processing on the classified data set to obtain processed local data and compressed data. Combining the gateway computing resource utilization rate, network bandwidth status, and central server load conditions, performing adaptive task scheduling and distributed load balancing to generate a local task list and a remote task list. Completing the local processing in the local task list through the gateway system to obtain local analysis results, and transmitting them together with the remote task list to the central server. By combining edge computing and distributed processing, the present invention realizes the reduction of data transmission pressure and the dynamic optimization of computing load, and improves the response speed and processing capacity of the smart city system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart cities, and particularly to a data analysis method and system for a smart city gateway system. Background Art

[0002] With the rapid development of smart cities, the application of Internet of Things (IoT) devices and sensors in urban management has become increasingly widespread. These devices provide key support for various fields such as traffic, environmental monitoring, and energy management by collecting various types of data during the operation of the city in real time. However, with the expansion of the scale of smart cities, the number of IoT devices has been increasing continuously, and the amount of raw data generated has grown explosively, bringing huge pressure to data transmission, storage, and processing.

[0003] Existing smart city systems usually rely on a central server to centrally process the data collected by IoT devices. Although this centralized architecture can achieve comprehensive analysis of global data, it is easily affected by bandwidth limitations and network latency during data transmission. Especially during peak periods, data loss or processing delays are likely to occur. In addition, due to the lack of effective utilization of gateway computing resources, existing methods cannot achieve efficient hierarchical processing of data and dynamic allocation of tasks, resulting in an overly high load on the central server and reducing the overall response speed and processing efficiency of the system. Summary of the Invention

[0004] The main objective of the present invention is to solve the technical problems in existing smart city systems, including insufficient transmission bandwidth, high processing latency, overloaded server, and lack of an efficient hierarchical data processing and dynamic task scheduling mechanism due to excessive data volume and centralized processing architecture;

[0005] In a first aspect of the present invention, a data analysis method for a smart city gateway system is provided. The data analysis method for the smart city gateway system includes:

[0006] Performing adaptive sampling and preliminary classification processing on the raw data from IoT devices and sensors in the smart city system through the smart city gateway system to obtain a classified data set;

[0007] Performing differential edge computing and context-aware compression processing on the classified data set through the smart city gateway system to obtain processed local data and compressed data;

[0008] Based on the processed local data and compressed data, and in combination with the gateway computing resource utilization rate, network bandwidth status, and central server load, performing adaptive task scheduling and distributed load balancing processing through the smart city gateway system to obtain a local task list and a remote task list;

[0009] The tasks in the local task list are locally processed through the smart city gateway system to obtain local analysis results, and the local analysis results and the task data in the remote task list are transmitted to the central server.

[0010] Optionally, in the first implementation manner of the first aspect of the present invention, the adaptive sampling and preliminary classification processing of the raw data from the Internet of Things devices and sensors in the smart city system through the smart city gateway system to obtain a classification data set includes:

[0011] The smart city gateway system monitors the data change rate of the Internet of Things devices and sensors in the smart city system in real time, and calculates a sampling frequency adjustment coefficient according to the data change rate;

[0012] Based on the sampling frequency adjustment coefficient, dynamic sampling is performed on the raw data, and the pre-trained decision tree model and rule set are used to classify and label the sampled raw data to obtain a classification result;

[0013] Generate metadata of the category identifier of the classification result, and attach the metadata to the raw data to form a classification data set.

[0014] Optionally, in the second implementation manner of the first aspect of the present invention, the dynamic sampling of the raw data based on the sampling frequency adjustment coefficient and the classification and labeling of the sampled raw data using the pre-trained decision tree model and rule set to obtain a classification result includes:

[0015] Interpolate the raw data according to the sampling frequency adjustment coefficient to obtain an interpolated data sequence;

[0016] Apply the sliding window technique to the interpolated data sequence to extract time series features and obtain a feature vector set;

[0017] Input the feature vector set into the pre-trained decision tree model to obtain a preliminary classification result, and adjust and label the preliminary classification result according to a predefined rule set to obtain a classification result.

[0018] Optionally, in the third implementation manner of the first aspect of the present invention, the classification data set includes emergency data, important data, and regular data;

[0019] The differential edge computing and context-aware compression processing of the metadata in the classification data set through the smart city gateway system to obtain the processed local data and compressed data includes:

[0020] The edge computing process is performed on the emergency data in the classified dataset through the smart city gateway system using a pre-configured expert system to obtain an emergency event analysis result, and the emergency event analysis result is used as local data;

[0021] Obtain the environmental parameters of the important data and the regular data, and construct corresponding context feature vectors according to the environmental parameters;

[0022] Using the context feature vectors, perform adaptive compression processing on the important data and the regular data to obtain the compressed important data and regular data, and use the compressed important data and regular data as compressed data.

[0023] Optionally, in the fourth implementation manner of the first aspect of the present invention, the using the context feature vectors to perform adaptive compression processing on the important data and the regular data to obtain the compressed important data and regular data includes:

[0024] Perform data type identification on the important data and the regular data according to the context feature vectors to obtain data type identifiers;

[0025] Based on the data type identifiers, select corresponding compression algorithms to perform preliminary compression on the important data and the regular data to obtain preliminary compression results;

[0026] Using a deep reinforcement learning model, with the context feature vectors and the preliminary compression results as inputs, dynamically adjust the preset compression parameters to obtain optimized compression parameters;

[0027] Use the optimized compression parameters to perform recompression processing on the preliminary compression results to obtain the compressed important data and regular data.

[0028] Optionally, in the fifth implementation manner of the first aspect of the present invention, the performing adaptive task scheduling and distributed load balancing processing by the smart city gateway system based on the processed local data and compressed data, combined with the gateway computing resource utilization rate, network bandwidth status, and central server load conditions, to obtain a local task list and a remote task list includes:

[0029] Obtain the data characteristics of the compressed data through the smart city gateway system, and calculate the resource requirement scores of the related tasks of each compressed data according to the data characteristics;

[0030] Combine the current computing resource utilization rate of the smart city gateway system, the available network bandwidth, the real-time obtained central server load conditions, and the resource requirement scores of the related tasks of each compressed data to construct a multi-dimensional decision matrix;

[0031] Using an improved particle swarm optimization algorithm, with the multi-dimensional decision matrix as the input, solve the multi-objective optimization problem of task allocation to obtain an initial task allocation plan;

[0032] Exchange load information and task characteristics with neighboring network nodes of the smart city gateway system through a distributed consensus algorithm, and perform task reallocation on the initial task allocation plan according to the exchange results to obtain an optimized allocation plan;

[0033] Divide the relevant tasks of the local data into the local task list, and divide the relevant tasks of the compressed data into the local task list and the remote task list according to the optimized allocation plan.

[0034] Optionally, in the sixth implementation manner of the first aspect of the present invention, the step of exchanging load information and task characteristics with neighboring network nodes of the smart city gateway system through a distributed consensus algorithm, and performing task reallocation on the initial task allocation plan according to the exchange results to obtain an optimized allocation plan includes:

[0035] Record the load information of the smart city gateway system and the task characteristics of the relevant tasks of the compressed data in a pre-established distributed ledger to generate current transaction data;

[0036] Broadcast the current transaction data to neighboring network nodes through a consensus algorithm, and receive the transaction data of neighboring network nodes to update the local distributed ledger;

[0037] Based on the updated distributed ledger, extract the latest network load status and task distribution information to construct a global resource view;

[0038] Apply a load balancing algorithm to utilize the global resource view and combine it with the initial task allocation plan to perform task reallocation to generate an optimized allocation plan.

[0039] The second aspect of the present invention provides a data analysis system for a smart city gateway system, and the data analysis system for the smart city gateway system includes:

[0040] A data classification module, configured to adaptively sample and preliminarily classify the original data from the Internet of Things devices and sensors in the smart city system through the smart city gateway system to obtain a classified data set;

[0041] A compression module, configured to perform differential edge computing and context-aware compression processing on the classified data set through the smart city gateway system to obtain processed local data and compressed data;

[0042] A task scheduling module, which is used to perform adaptive task scheduling and distributed load balancing processing through the smart city gateway system based on the processed local data and compressed data, in combination with the gateway computing resource utilization rate, network bandwidth status, and central server load conditions, to obtain a local task list and a remote task list;

[0043] A local processing module, which is used to locally process the tasks in the local task list through the smart city gateway system to obtain local analysis results, and transmit the local analysis results and the task data in the remote task list to the central server.

[0044] The above-mentioned data analysis method and system of the smart city gateway system adaptively sample and preliminarily classify the original data of Internet of Things devices and sensors through the smart city gateway system to generate a classified data set. Perform differential edge computing and context-aware compression processing on the classified data set to obtain processed local data and compressed data. Combine the gateway computing resource utilization rate, network bandwidth status, and central server load conditions to perform adaptive task scheduling and distributed load balancing to generate a local task list and a remote task list. Complete the local processing in the local task list through the gateway system to obtain local analysis results, and transmit them to the central server together with the remote task list. The present invention combines edge computing with distributed processing to reduce the data transmission pressure and dynamically optimize the computing load, improving the response speed and processing ability of the smart city system.

[0045] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.

[0046] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. Brief Description of the Drawings

[0047] Figure 1 It is a schematic diagram of the first embodiment of the data analysis method of the smart city gateway system in the embodiment of the present invention;

[0048] Figure 2 It is a schematic diagram of an embodiment of the data analysis system of the smart city gateway system in the embodiment of the present invention. Detailed Description of the Embodiment

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0050] As used in the embodiments of the present invention, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0051] To facilitate the understanding of this embodiment, first, a data analysis method for a smart city gateway system disclosed in the embodiments of the present invention will be introduced in detail. As Figure 1 shown, the method includes the following steps:

[0052] 101. Adaptive sampling and preliminary classification processing are performed on the raw data from the Internet of Things devices and sensors in the smart city system through the smart city gateway system to obtain a classified data set;

[0053] In one embodiment of the present invention, the adaptive sampling and preliminary classification processing of the raw data from the Internet of Things devices and sensors in the smart city system through the smart city gateway system to obtain a classified data set includes: real-time monitoring of the data change rate of the Internet of Things devices and sensors in the smart city system through the smart city gateway system, calculating a sampling frequency adjustment coefficient based on the data change rate; dynamically sampling the raw data based on the sampling frequency adjustment coefficient, and classifying and labeling the sampled raw data using a pre-trained decision tree model and rule set to obtain a classification result; generating metadata of the class identifier of the classification result and attaching the metadata to the raw data to form a classified data set.

[0054] Specifically, the implementation of a real-time data stream processing pipeline involves using high-performance stream processing frameworks such as Apache Flink or Apache Kafka Streams. This pipeline first creates dedicated data receivers for each type of device and sensor. For example, for traffic cameras, the RTSP protocol is used to receive video streams; for air quality sensors, periodic data packets are received via the MQTT protocol. Timestamp marking and source identification are performed immediately after data reception. The implementation of sliding windows uses time-driven window operations, and the window size is dynamically adjusted according to the data type. For example, a 5-minute window may be used for traffic data, while a 1-hour window can be used for air quality data. The calculation of statistical features includes using online algorithms to calculate the mean (cumulative sum divided by count), variance (Welford's algorithm), and kurtosis (moment-based calculation method). The evaluation of information entropy adopts the histogram method of data within the sliding window. The consideration of historical patterns is achieved by maintaining long-term (e.g., 24 hours) and short-term (e.g., 1 hour) data statistics and comparing the deviation of the current window from these statistics. Periodic detection uses the Fast Fourier Transform (FFT) to analyze data in the frequency domain. Emergency event detection is based on setting dynamic thresholds and is triggered when the data exceeds the average value calculated based on the past N cycles plus or minus 3 standard deviations. Finally, the calculation of the sampling frequency adjustment coefficient uses a piecewise function. For example, when the change rate is below the threshold T1, the coefficient is 0.5; it increases linearly between T1 and T2; when it is above T2, the coefficient is 2. The parameters of this function are set through historical data analysis and expert knowledge and are optimized regularly. The core of the adaptive sampling algorithm is a control loop that takes the sampling frequency adjustment coefficient as input and dynamically adjusts the sampling strategy for each data source. The algorithm first defines a baseline sampling frequency fbase, and then calculates the new sampling frequency fnew = fbase * k according to the adjustment coefficient k. To avoid drastic fluctuations in the sampling frequency, a smoothing factor α is introduced, and the actually applied sampling frequency f = α * fnew + (1 - α) * fold, where fold is the sampling frequency of the previous cycle. Considering physical limitations, a minimum sampling interval tmin (e.g., 100 ms) and a maximum sampling interval tmax (e.g., 1 hour) are set. In the actual sampling process, an accurate time control mechanism (such as the high-precision timer of Linux) is used to trigger sampling. The data quality control mechanism includes multiple steps: first, check the integrity of the data to ensure that all necessary fields are present; then verify the consistency of the data, such as ensuring that the timestamps are monotonically increasing; finally, perform validity checks, including range checks (e.g., temperature between -50°C and 50°C) and change rate checks (e.g., the temperature changes no more than 1°C per second). The handling of outliers adopts the Z-score method, and data points with an absolute Z-score greater than 3 are marked as potential outliers.The implementation of the decision tree model uses algorithms such as CART (Classification and Regression Trees). The depth of the tree is limited within 10 levels to prevent overfitting, and the Gini impurity is used as the splitting criterion. The training data of the decision tree comes from historical data, and the model is updated regularly (such as weekly) using the sliding window method to adapt to the changes in data distribution. The implementation of the rule set adopts a JSON-based rule description language, allowing experts to define complex condition combinations and corresponding processing logics. The rule engine uses the Rete algorithm to efficiently match data and rules, supporting real-time rule updates without stopping the system.

[0055] The implementation of the standardized coding system is based on a hierarchical coding structure. For example, the first two characters represent the data type (such as TF for traffic flow), and the next three digits represent the specific classification. The coding system is stored in a distributed key-value store (such as etcd), supporting real-time updates and queries. The metadata generation process uses a template engine (such as Jinja2) to dynamically select appropriate templates according to the data type. The metadata includes a UUID as the unique identifier, a timestamp in ISO8601 format, a classification identifier, a confidence score (a floating point number from 0 to 1), the model version number used, the processing node ID, etc. For the classification confidence, if it is the result of the decision tree, the purity of the leaf node is used as the confidence; if it is the result of rule matching, the confidence is calculated based on the number and priority of the matching rules. The metadata attachment process adopts different strategies for different types of data: for structured data (such as JSON or Avro format), a new "metadata" field is directly added to the original structure; for semi-structured data (such as XML), a new metadata element is added under the root element; for unstructured data (such as video streams), the metadata is serialized and added to the data packet as a custom header field. During the data transmission process, a compression algorithm (such as LZ4) is used to compress the metadata to reduce bandwidth occupancy. The final classified dataset is saved in a columnar storage format (such as Parquet) for subsequent rapid analysis and queries.

[0056] Furthermore, the original data is dynamically sampled based on the sampling frequency adjustment coefficient, and the pre-trained decision tree model and rule set are used to classify and label the sampled original data. The obtained classification results include: interpolating the original data according to the sampling frequency adjustment coefficient to obtain an interpolated data sequence; applying the sliding window technique to the interpolated data sequence to extract time series features and obtain a feature vector set; inputting the feature vector set into the pre-trained decision tree model to obtain a preliminary classification result, and adjusting and labeling the preliminary classification result according to the predefined rule set to obtain the classification result.

[0057] Specifically, the process of interpolating the original data according to the sampling frequency adjustment coefficient first requires determining the interpolation method and parameters. Considering the complexity and diversity of smart city data, it is necessary to adopt an adaptive interpolation strategy. For data with strong periodicity, such as traffic flow, cubic spline interpolation can be used to maintain the smoothness and continuity of the data. For data with strong mutability, such as emergency sensor data, piecewise linear interpolation is adopted to retain the spike characteristics of the data. The interpolation density is directly determined by the sampling frequency adjustment coefficient. Specifically, when implementing, the original sampling interval is divided by the adjustment coefficient to obtain a new sampling interval. For example, if the original sampling interval is 1 second and the adjustment coefficient is 1.5, the new sampling interval is 2 / 3 seconds. During the interpolation process, it is also necessary to consider the boundary conditions of the data and the handling of outliers. For the start and end parts of the time series, extrapolation is used for interpolation to avoid endpoint effects. For the detected outliers, a lower weight is given during interpolation to reduce their impact on the overall trend.

[0058] Specifically, applying the sliding window technique to the interpolated data sequence and extracting time series features is a key step in feature engineering. The size of the sliding window needs to be dynamically adjusted according to the characteristics of the data. Usually, a window size that can cover a complete cycle of the data is selected. For example, for daily traffic flow data, 24 hours can be selected as the window size. The sliding step of the window is determined according to the required feature accuracy. The smaller the step, the greater the computational effort, but the higher the time resolution of the features. Within each window, various statistical features and time-frequency domain features are extracted. Statistical features include mean, variance, skewness, kurtosis, quantiles, etc., which can describe the overall distribution of the data. Time domain features include autocorrelation coefficient, cross-correlation coefficient, etc., which are used to capture the time dependence of the data. Frequency domain features are obtained through the fast Fourier transform (FFT) or wavelet transform, including the amplitude and phase information of the main frequency components, which is very effective for identifying the periodic patterns of the data. In addition, some specific domain features can also be extracted, such as the congestion index of traffic data and the pollution index of environmental data. All these features are combined to form a high-dimensional feature vector set, and each vector corresponds to the feature description of a time window.

[0059] Specifically, the process of inputting the feature vector set into the pre-trained decision tree model to obtain the preliminary classification results involves model selection, training, and application. Considering the complexity of smart city data and the requirements of real-time processing, it is more appropriate to select random forest or gradient boosting decision trees (such as XGBoost) as the base model. These models not only have high classification accuracy, but also can handle high-dimensional features, and have fast training and prediction speeds. The training process of the model uses historical data and adopts the method of cross-validation to select the optimal hyperparameters, such as the depth of the tree, the minimum number of samples in the leaf node, etc. To cope with the dynamic changes in data distribution, the model adopts the online learning method and regularly updates the model parameters using new data samples. In actual application, each feature vector is input into the model, and the model outputs the probability distribution of each category. The preliminary classification result selects the category with the highest probability, and at the same time records this probability value as the confidence of the classification.

[0060] Specifically, the process of adjusting and marking the preliminary classification results according to the predefined rule set is the key step to integrate domain expert knowledge into the automatic classification system. The design of the rule set needs to consider multiple aspects, including regulatory requirements, historical experience, and handling of special situations. The rules are represented in a structured format based on JSON or XML, and each rule contains a condition part and an action part. The condition part can include the logical combination of multiple sub-conditions, such as "IF classification probability < 0.8 AND time BETWEEN 22:00 AND 06:00". The action part defines the operations to be taken when the conditions are met, such as adjusting the category, adding marks, or triggering an alarm. The rules are executed using a forward chaining inference engine, and each rule is evaluated one by one in the order of rule priority. During the adjustment process, not only the preliminary classification results are considered, but also the original data and context information are combined. For example, if the traffic flow is classified as "congested", but the surrounding roads are all unobstructed, the rule may adjust the classification to "abnormal event" and add the corresponding mark. In addition, the rule set also contains some meta-rules for handling classification conflicts or high uncertainty situations.

[0061] 102. Based on the standardized data structure, set the camera parameters and lighting environment parameters of the scene where the building block model is located respectively to obtain the scene rendering parameters;

[0062] In an embodiment of the present invention, the classified data set includes emergency data, important data, and regular data; the process of performing differential edge computing and context-aware compression processing on the metadata in the classified data set through the smart city gateway system to obtain processed local data and compressed data includes: performing edge computing processing on the emergency data in the classified data set through the smart city gateway system using a pre-configured expert system to obtain an emergency event analysis result, and using the emergency event analysis result as local data; obtaining the environmental parameters of the important data and the regular data, and constructing corresponding context feature vectors according to the environmental parameters; using the context feature vectors to perform adaptive compression processing on the important data and the regular data to obtain compressed important data and regular data, and using the compressed important data and regular data as compressed data.

[0063] Specifically, the process of the smart city gateway system performing edge computing processing on the emergency data in the classified data set using a pre-configured expert system first involves the construction and deployment of the expert system. This expert system is a rule-based inference engine that contains a large number of processing rules for various emergency situations. The construction of the rule base requires the participation of experts in multiple fields such as urban management and emergency response, and encodes their knowledge and experience into a structured rule set. Each rule includes a condition part and an action part. The condition part describes the scenario that triggers the rule, such as "IF traffic accident AND number of casualties > 5", and the action part defines the corresponding processing steps, such as "notify the nearest hospital and police station". The execution of the rules adopts a forward chaining inference mechanism, and the system will continuously evaluate the current situation and trigger the corresponding rules. To improve processing efficiency, the rules are sorted by priority, and the high-priority rules are evaluated first. The expert system also includes a fact base for storing all known information about the current situation, and this information will be continuously updated as new data arrives. During the processing, the expert system will generate a series of inference steps and decision suggestions, which constitute the emergency event analysis result. The analysis result not only includes the classification and severity assessment of the event, but also includes the recommended response measures and the predicted impact range. These results are stored in a structured format for easy quick reading and execution.

[0064] Specifically, the process of obtaining the environmental parameters of important data and regular data and constructing corresponding context feature vectors involves the integration of multiple data sources and feature extraction. To obtain the environmental parameters, it is first necessary to determine the relevant parameter sets, which usually include time (such as timestamps, day of the week, whether it is a holiday), location (such as GPS coordinates, administrative divisions), meteorological conditions (such as temperature, humidity, wind speed), status of surrounding facilities (such as traffic flow, population density), etc. These parameters come from different sensors and databases and need to be integrated through data fusion technology. The time parameter is directly obtained from the system clock and necessary conversions are performed, such as calculating whether it is a peak period. Location information may need to be obtained through a GPS module or network positioning service and then parsed through a Geographic Information System (GIS) to obtain a detailed location description. Meteorological conditions can be obtained through local weather stations or by calling weather service APIs. The status of surrounding facilities needs to query relevant real-time databases or other Internet of Things devices. All these raw parameters are standardized and normalized to ensure that data with different dimensions can be compared. Then, feature engineering techniques are used to extract higher-level features from these parameters, such as calculating the periodic features of time and the clustering features of location. Finally, all these features are combined into a unified feature vector, which comprehensively describes the environmental context when the data is generated.

[0065] Specifically, the process of adaptively compressing important data and regular data using the context feature vector is a process of implementing a dynamically adjusted compression strategy. First, based on the context feature vector, the system will select the compression algorithm that is most suitable for the current data and environment. For example, for data with strong periodicity, a compression method based on Fourier transform may be selected; for data with high spatial correlation, a compression method based on wavelet transform may be selected. The selection of the algorithm is implemented using a machine learning model, which is trained based on historical data and can predict the best compression algorithm according to the context features. After the algorithm is selected, the system will further adjust the compression parameters according to the context features. For example, when the network bandwidth is sufficient, a lower compression ratio can be selected to retain more details; when the bandwidth is limited, the compression ratio is increased. For important data, the system will adopt a more conservative compression strategy to ensure that key information is not lost. During the compression process, the system will also dynamically monitor the compression effect, including indicators such as the compression ratio and information retention degree, and adjust the compression strategy in real time according to these feedbacks. The compressed data will be appended with a metadata header that records the summary of the compression algorithm, parameter settings, and context features used for subsequent decompression and analysis. Finally, the compressed important data and regular data are packaged into a unified format for network transmission or storage. This adaptive compression method can not only significantly reduce the amount of data but also flexibly adjust according to the importance of the data and environmental conditions, ensuring efficient data processing and transmission on resource-constrained edge devices.

[0066] Furthermore, the process of adaptively compressing the important data and the regular data by using the context feature vector to obtain the compressed important data and regular data includes: identifying the data types of the important data and the regular data according to the context feature vector to obtain data type identifiers; selecting corresponding compression algorithms based on the data type identifiers to perform preliminary compression on the important data and the regular data to obtain preliminary compression results; using a deep reinforcement learning model, with the context feature vector and the preliminary compression results as inputs, to dynamically adjust the preset compression parameters to obtain optimized compression parameters; and using the optimized compression parameters to perform re-compression on the preliminary compression results to obtain the compressed important data and regular data.

[0067] Specifically, the process of identifying the data types of the important data and the regular data according to the context feature vector first involves constructing a multi-layer perceptron neural network model. This model takes the context feature vector as input, and the output layer corresponds to different data types. The ReLU activation function is used in the hidden layer of the model to capture the non-linear relationships between features. The model is trained using a supervised learning method with a large amount of labeled historical data. The cross-entropy loss function and the Adam optimizer are used during the training process to improve the generalization ability of the model. To handle the class imbalance problem, a class weight adjustment technique is adopted. After the model training is completed, by performing forward propagation on the new context feature vector, the probability distribution of each data type is obtained. The type with the highest probability is selected as the recognition result, and at the same time, this probability value is recorded as the confidence of the recognition. For cases where the confidence is lower than the preset threshold, the system will trigger an artificial review mechanism. The recognition result is output in the form of a data type identifier, which is a predefined code, such as "TF" representing traffic flow data, "AQ" representing air quality data, etc. This data type identifier not only contains the basic category of the data but also contains the specific attributes of the data, such as the periodicity, continuity, and other characteristics of the data.

[0068] Specifically, the process of selecting the corresponding compression algorithm based on the data type identifier and performing preliminary compression involves establishing an algorithm selection mapping table. This mapping table is constructed based on a large number of experiments and expert knowledge, mapping different data type identifiers to the most suitable compression algorithms. For example, for traffic flow data with strong periodicity, a compression algorithm based on Fourier transform may be selected; for image data with high spatial correlation, a compression algorithm based on wavelet transform may be selected; for time series data, differential coding or run-length coding may be selected. Algorithm selection also takes into account the importance level of the data. For important data, algorithms with high fidelity are preferred, while for regular data, the compression ratio is more emphasized. After the algorithm is selected, the system will perform compression using preset initial parameters. These initial parameters are obtained based on historical data statistics and can provide good compression effects in most cases. During the compression process, the system will record multiple metrics, including compression ratio, compression time, information entropy change, etc. These metrics together with the compressed data constitute the preliminary compression result.

[0069] Specifically, the process of dynamically adjusting compression parameters using a deep reinforcement learning model is a complex optimization problem. First, a deep Q-network (DQN) model is constructed. The state space of this model includes the context feature vector, the metrics of the preliminary compression result, and the current compression parameters. The action space is defined as the adjustment operations for each compression parameter, such as increasing or decreasing the quantization level, adjusting the transform coefficients, etc. The design of the reward function takes into account multiple objectives, including compression ratio, information retention, and computational resource consumption. These objectives are combined into a single reward value through weighted summation. Model training uses techniques such as experience replay and target network to improve the stability and efficiency of learning. In practical applications, the model receives the current state as input and outputs the optimal parameter adjustment actions. This process is iterative, and the compression effect is re-evaluated after each adjustment until the preset number of iterations is reached or the termination condition is met. To accelerate the convergence speed, the prioritized experience replay technique is used to preferentially learn those experiences that can bring significant improvements. The finally obtained optimized compression parameters not only adapt to the current data characteristics and environmental conditions but also achieve a good balance between the compression ratio and information retention.

[0070] Specifically, the process of re-compressing the preliminary compression result using optimized compression parameters is a refined adjustment of the preliminary compression. First, the system parses the optimized compression parameters, which may include quantization levels, transform coefficient thresholds, encoding method selections, etc. Then, the compression algorithm is reconfigured according to these parameters. For transform-based compression methods (such as wavelet transform or DCT), the truncation threshold of the coefficients may need to be adjusted; for prediction-based methods, the parameters of the prediction model may need to be adjusted; for entropy encoding methods, the encoding dictionary may need to be updated. During the re-compression process, the system monitors the compression effect in real time, including metrics such as compression ratio, distortion (such as mean squared error or structural similarity index SSIM), and compression speed. If it is found that certain parameter adjustments lead to a significant decline in the compression effect, the system will trigger a rollback mechanism to restore to a better-performing parameter setting. After re-compression is completed, the system generates a detailed compression report, including the final compression ratio, information retention degree evaluation, compression parameter configuration, etc. For important data, the system also saves a summary or key features of the original data for data recovery or verification when necessary.

[0071] 103. Adaptive sampling and preliminary classification processing are performed on the raw data from the Internet of Things devices and sensors in the smart city system through the smart city gateway system to obtain a classified data set;

[0072] In an embodiment of the present invention, the execution of adaptive task scheduling and distributed load balancing processing by the smart city gateway system based on the processed local data and compressed data, combined with the gateway computing resource utilization rate, network bandwidth status, and central server load conditions, to obtain a local task list and a remote task list includes: obtaining the data characteristics of the compressed data through the smart city gateway system, and calculating the resource requirement scores of the relevant tasks for each compressed data according to the data characteristics; combining the current computing resource utilization rate, available network bandwidth, and real-time obtained central server load conditions of the smart city gateway system, as well as the resource requirement scores of the relevant tasks for each compressed data, to construct a multi-dimensional decision matrix; using an improved particle swarm optimization algorithm, with the multi-dimensional decision matrix as the input, to solve the multi-objective optimization problem of task allocation to obtain an initial task allocation plan; exchanging load information and task characteristics with neighboring gateway nodes of the smart city gateway system through a distributed consensus algorithm, and reallocating tasks to the initial task allocation plan according to the exchange results to obtain an optimized allocation plan; dividing the relevant tasks of the local data into the local task list, and dividing the relevant tasks of the compressed data into the local task list and the remote task list according to the optimized allocation plan.

[0073] Specifically, the process of the smart city gateway system obtaining the data characteristics of compressed data and calculating the resource demand scores for related tasks involves multi-dimensional data analysis. First, the system parses the compressed data to extract key features, including data size, data type, compression ratio, data generation time, etc. For different types of data, the extracted features may vary. For example, for image data, information such as resolution and color depth will also be extracted; for time series data, features such as sampling rate and periodicity will be extracted. These features are obtained through predefined feature extraction functions, and the selection of the function is based on the data type. After obtaining the features, the system uses a machine learning model to predict the resource requirements for each related task. This model is a pre-trained regression model, such as a random forest or gradient boosting tree, with the input being the data features and the output being resource demand metrics such as CPU usage, memory occupancy, and estimated execution time. The training data for the model comes from historical task execution records and is updated regularly to adapt to new data patterns and task types. The predicted resource demand metrics are combined into a single resource demand score through weighted summation. The setting of the weights reflects the relative importance of different resources in the current system and can be dynamically adjusted according to the system configuration and operating conditions. Finally, each task related to the compressed data will obtain a resource demand score, which quantifies the complexity and resource consumption level of the task.

[0074] Specifically, the process of constructing a multi-dimensional decision matrix by combining the current state of the gateway system and the task resource demand scores is the core preparatory step for task scheduling. First, the system collects real-time operating state information, including CPU usage, memory occupancy rate, disk I / O speed, network bandwidth utilization, etc. This information is obtained through APIs provided by the operating system or specialized monitoring tools. At the same time, the system obtains the load situation of the central server through network probing or preset communication protocols, including the resource utilization rate of the server and the current task queue length. All the collected state information is standardized to unify data with different dimensions to the same scale. Then, the system combines this state information with the resource demand scores of each task calculated previously to construct a multi-dimensional decision matrix. Each row of this matrix represents a task to be assigned, and each column represents a decision factor, such as the resource demand score of the task, the available resources of the current gateway, network conditions, and the load of the central server. Each element in the matrix is normalized to ensure the comparability of data in different dimensions. In addition, the system also assigns a weight to each decision factor to reflect its importance in the decision-making process. This weight can be set by the system administrator or automatically adjusted through machine learning methods. The constructed multi-dimensional decision matrix provides a comprehensive decision-making basis for subsequent task assignment optimization.

[0075] Specifically, using the improved particle swarm optimization algorithm to solve the multi-objective optimization problem of task allocation is a complex computational process. First, the optimization objectives are defined, usually including minimizing the total execution time, balancing the load distribution, maximizing resource utilization, etc. Each objective has a corresponding objective function. Then, the particle swarm is initialized, and each particle represents a possible task allocation scheme. The position of the particle is encoded as a vector, and each element of the vector represents whether a task is assigned to be executed locally or remotely. At initialization, a group of particles is randomly generated to ensure coverage of the solution space. The improved particle swarm algorithm introduces an adaptive inertia weight and learning factors to improve the convergence speed and avoid being trapped in local optima. In each iteration, the fitness value of each particle is calculated, and this value is based on the weighted sum of multiple objective functions. Then, the individual best position and the global best position of each particle are updated. The velocity and position update formulas of the particle are modified to include a chaotic map to increase the randomness of the search. To handle the multi-objective problem, the concept of Pareto dominance is adopted, and a non-dominated solution set is maintained. The iterative process continues until the maximum number of iterations is reached or the convergence condition is satisfied. Finally, the algorithm outputs a set of non-dominated solutions on the Pareto front, and the system selects a balanced solution from them as the initial task allocation scheme.

[0076] Specifically, the process of exchanging load information and task characteristics with neighboring gateway nodes through a distributed consensus algorithm and performing task reallocation involves complex network communication and decision-making mechanisms. First, the system defines a standardized load information and task characteristic exchange protocol, including specifications such as data format, communication frequency, and security encryption. Each gateway node maintains a list of neighboring nodes and regularly broadcasts its own load status and the characteristics of pending tasks to these nodes. A distributed ledger based on blockchain technology is used to record and verify this information to ensure data consistency and immutability. When receiving information from neighboring nodes, the system uses the Byzantine fault tolerance algorithm for consensus verification to filter out possible false or incorrect information. The verified information is integrated into a global view, which contains the resource distribution and task situation of the entire network. Based on this global view, the system uses a distributed optimization algorithm to re-evaluate the initial task allocation scheme. This optimization process takes into account factors such as network topology, communication cost, and load balancing, with the goal of achieving optimal resource utilization across the entire network. The optimization algorithm uses a distributed version of the genetic algorithm or simulated annealing algorithm, and each node runs the algorithm independently, but coordinates the global optimization direction through periodic information exchange. Finally, the system obtains a task allocation scheme optimized at the network level.

[0077] Specifically, the process of dividing local data-related tasks into the local task list and dividing compressed data-related tasks into local and remote task lists according to the optimized allocation scheme is the final execution stage of task scheduling. First, the system identifies all tasks related to local data. These tasks usually have high priority and low latency requirements and are directly assigned to the local task list. For tasks related to compressed data, the system parses the optimized allocation scheme, which contains the recommended execution location (local or remote) for each task. The system checks each task one by one. For tasks assigned to be executed locally, they are added to the local task list; for tasks assigned to be executed remotely, they are added to the remote task list. During this process, the system also needs to consider the dependencies between tasks to ensure that mutually dependent tasks are reasonably allocated to avoid deadlocks or excessive communication overhead. For some complex tasks, the system may need to decompose the tasks, splitting large tasks into multiple subtasks and determining their execution locations separately. The construction of the task list also needs to consider priority sorting, usually using a multi-level feedback queue algorithm to manage the execution order of tasks. Finally, the system generates detailed task metadata, including task descriptions, resource requirements, estimated execution times, data dependencies, etc. These metadata are added to the corresponding task lists along with the tasks.

[0078] Further, the process of exchanging load information and task characteristics with neighboring network nodes of the smart city gateway system through a distributed consensus algorithm and reallocating tasks for the initial task allocation scheme according to the exchange results to obtain an optimized allocation scheme includes: recording the load information of the smart city gateway system and the task characteristics of the tasks related to the compressed data in a pre-established distributed ledger to generate current transaction data; broadcasting the current transaction data to neighboring network nodes through the consensus algorithm and receiving the transaction data of neighboring network nodes to update the local distributed ledger; based on the updated distributed ledger, extracting the latest network load status and task distribution information to construct a global resource view; applying a load balancing algorithm using the global resource view and combining the initial task allocation scheme to perform task reallocation to generate an optimized allocation scheme.

[0079] Specifically, the load information of the smart city gateway system and the task characteristics of the tasks related to compressed data are recorded in a pre-established distributed ledger. The process of generating the current transaction data involves multiple key steps. First, the system needs to collect and collate the load information of the current gateway, including key metrics such as CPU usage, memory occupancy, storage space, and network bandwidth utilization. This information is obtained through operating system-level monitoring tools or custom resource monitoring modules to ensure the real-time and accuracy of the data. At the same time, the system also needs to extract the characteristics of the tasks related to compressed data, such as task type, estimated execution time, resource requirements, etc. These characteristic data are derived from the previous task analysis and resource requirement assessment phases. The collected load information and task characteristics are then formatted into a predefined data structure, usually using lightweight data exchange formats such as JSON or Protocol Buffers. This standardized data structure facilitates the information exchange and processing between different gateway nodes. Next, the system encapsulates the formatted data into a transaction object, which contains meta-information such as data content, timestamp, and gateway node identifier. Finally, the system uses a private key to digitally sign the transaction data to ensure the integrity and non-repudiation of the data, and writes the signed transaction data into the pre-established distributed ledger. This distributed ledger is implemented based on blockchain technology, and each transaction is added to the chain as a new block, forming an immutable historical record.

[0080] Specifically, the process of broadcasting the current transaction data to neighboring gateway nodes through a consensus algorithm, receiving the transaction data of neighboring gateway nodes, and updating the local distributed ledger is a key step in achieving information synchronization between gateways. First, the system needs to maintain a list of neighboring gateway nodes, which is obtained through a network discovery protocol or pre-configuration. For each neighboring node, the system establishes a secure communication channel, usually using the TLS protocol to ensure encryption and authentication during the transmission process. Then, the system broadcasts the generated current transaction data to all neighboring nodes through these secure channels. The broadcast is carried out asynchronously to avoid system blocking caused by network latency. At the same time, the system also listens for broadcast messages from other nodes and receives their transaction data. The received transaction data first undergoes a verification process, including checking the validity of digital signatures, verifying the integrity of the transaction structure, and confirming the identity of the sending node, etc. The verified transaction data is temporarily stored in a pending queue. The system starts a consensus algorithm, such as Practical Byzantine Fault Tolerance (PBFT) or Raft, to determine the order and validity of these new transactions. After consensus is reached, the system adds these new transaction data to the local distributed ledger in the agreed order. If conflicts or invalid transactions are found during the consensus process, these transactions will be rejected and removed from the pending queue. The whole process ensures the consistency and real-time nature of information between gateway nodes, laying a foundation for the subsequent construction of the global resource view.

[0081] Specifically, based on the updated distributed ledger, the process of extracting the latest network load status and task distribution information to construct a global resource view involves data aggregation and analysis. First, the system needs to traverse the updated distributed ledger to extract the latest load information and task characteristics of all gateway nodes. This process needs to consider the timeliness of data, usually only processing the transaction data within a certain recent time window to ensure the real-time nature of the resource view. The extracted data undergoes normalization processing to unify the possibly different data formats reported by different gateway nodes. Next, the system performs aggregation analysis on the normalized data to calculate statistical indicators such as the resource utilization distribution and task density distribution of the entire network. This process may involve complex data mining algorithms, such as clustering analysis and anomaly detection, to identify hot spots or potential performance bottlenecks in the network. The analysis results are organized into a multi-dimensional data structure, reflecting the network topology, resource distribution, and task distribution. This data structure is usually represented in a graph form, where nodes represent gateways, edges represent connections between gateways, and various attribute information is attached to the nodes and edges. Finally, the system generates a visual global resource view based on this graph structure for easy human-computer interaction and decision support. This global resource view not only contains static resource distribution information but also dynamic resource utilization trends and task flow patterns, providing comprehensive data support for subsequent load balancing decisions.

[0082] Specifically, the process of using the load balancing algorithm to utilize the global resource view and combine it with the initial task allocation scheme for task reallocation to generate an optimized allocation scheme is the core of the overall system optimization. First, the system needs to select an appropriate load balancing algorithm. Considering the characteristics of the smart city gateway system, dynamic adaptive algorithms such as Weighted Round Robin or Least Connection are usually adopted. After selecting the algorithm, the system converts the global resource view into the input format required by the algorithm, including the current load, available resources, network connection status, etc. of each gateway node. Then, the system treats each task in the initial task allocation scheme as an object to be reallocated, and evaluates the feasibility and efficiency of its execution on different gateway nodes one by one. The evaluation process considers multiple factors, such as the matching degree between the resource requirements of the task and the available resources of the node, network transmission overhead, load balance between nodes, etc. Based on the evaluation results, the system calculates an optimal target node for each task. In this process, the system also needs to consider the dependency relationship between tasks and the data locality principle, and try to allocate related tasks to the same node or adjacent nodes as much as possible to reduce data transmission and improve execution efficiency. The result of the reallocation is continuously compared and adjusted with the current resource view to ensure that no node is overloaded. Finally, the system generates a new task allocation scheme, which contains detailed information such as the target execution node, estimated execution time, and resource requirements of each task.

[0083] 104. The tasks in the local task list are locally processed through the smart city gateway system to obtain local analysis results, and the local analysis results and the task data in the remote task list are transmitted to the central server.

[0084] In an embodiment of the present invention, the system uses a priority queue to manage local tasks and sorts them according to the urgency and resource requirements of the tasks. Subsequently, the system allocates the required computing resources for each task, such as CPU cores, memory, and storage space. During the task execution process, the system monitors the resource usage in real time and makes dynamic adjustments when necessary to optimize performance. For data processing tasks, the system applies pre-configured algorithms and models, such as time series analysis, anomaly detection, or pattern recognition, etc., to generate preliminary analysis results. These results are aggregated and feature-extracted through local data to form more refined local analysis results. At the same time, the system preprocesses and compresses the task data in the remote task list to reduce the transmission burden. Finally, the system packages the local analysis results and the preprocessed remote task data and transmits them to the central server using a secure communication protocol (such as TLS). This process ensures the efficient utilization of local resources and the secure transmission of data, providing high-quality data input that has been preliminarily processed and screened for the central server.

[0085] In this embodiment, the original data of Internet of Things devices and sensors is adaptively sampled and preliminarily classified through the smart city gateway system to generate a classified data set. Differential edge computing and context-aware compression processing are performed on the classified data set to obtain processed local data and compressed data. Combining the gateway computing resource utilization rate, network bandwidth status, and central server load, adaptive task scheduling and distributed load balancing are performed to generate a local task list and a remote task list. The local processing in the local task list is completed through the gateway system to obtain local analysis results, and the local analysis results and the remote task list are transmitted to the central server. By combining edge computing and distributed processing, the present invention realizes the reduction of data transmission pressure and the dynamic optimization of computing load, and improves the response speed and processing ability of the smart city system.

[0086] The data analysis method of the smart city gateway system in the embodiment of the present invention is described above. Next, the data analysis system of the smart city gateway system in the embodiment of the present invention will be described. Please refer to Figure 2 , an embodiment of the data analysis system of the smart city gateway system in the embodiment of the present invention includes:

[0087] A data classification module 201, configured to perform adaptive sampling and preliminary classification processing on the original data from Internet of Things devices and sensors in the smart city system through the smart city gateway system to obtain a classified data set;

[0088] A compression module 202, configured to perform differential edge computing and context-aware compression processing on the classified data set through the smart city gateway system to obtain processed local data and compressed data;

[0089] A task scheduling module 203, configured to perform adaptive task scheduling and distributed load balancing processing through the smart city gateway system based on the processed local data and compressed data, combined with the gateway computing resource utilization rate, network bandwidth status, and central server load, to obtain a local task list and a remote task list;

[0090] A local processing module 204, configured to perform local processing on the tasks in the local task list through the smart city gateway system to obtain local analysis results, and transmit the local analysis results and the task data in the remote task list to the central server.

[0091] In the embodiment of the present invention, the data analysis system of the smart city gateway system runs the data analysis method of the smart city gateway system. The data analysis system of the smart city gateway system adaptively samples and preliminarily classifies the original data of the Internet of Things devices and sensors through the smart city gateway system to generate a classified data set. Perform differential edge computing and context-aware compression processing on the classified data set to obtain processed local data and compressed data. Combine the gateway computing resource utilization rate, network bandwidth status, and central server load conditions to perform adaptive task scheduling and distributed load balancing to generate a local task list and a remote task list. Complete the local processing in the local task list through the gateway system to obtain a local analysis result, and transmit it and the remote task list to the central server. The present invention combines edge computing with distributed processing to reduce the data transmission pressure and dynamically optimize the computing load, improving the response speed and processing capacity of the smart city system.

[0092] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, devices, or units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0093] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0094] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A data analysis method for a smart city gateway system, characterized in that: The data analysis method of the smart city gateway system includes: The smart city gateway system adaptively samples and preliminarily classifies the raw data from the IoT devices and sensors in the smart city system to obtain a classified data set; the classified data set includes emergency data, important data and regular data; The emergency data in the classified data set is processed by edge computing using a preconfigured expert system through the smart city gateway system to obtain an emergency event analysis result, and the emergency event analysis result is used as local data; the environmental parameters of the important data and regular data are obtained, and a corresponding context feature vector is constructed according to the environmental parameters; the data type of the important data and regular data is identified according to the context feature vector to obtain a data type identifier; the corresponding compression algorithm is selected based on the data type identifier, and the important data and regular data are preliminarily compressed to obtain a preliminary compression result; the deep reinforcement learning model is used to dynamically adjust the preset compression parameters with the context feature vector and the preliminary compression result as input to obtain the optimized compression parameters; the preliminary compression result is recompressed using the optimized compression parameters to obtain compressed important data and regular data, and the compressed important data and regular data are used as compressed data; The data features of the compressed data are obtained through the smart city gateway system, and the resource requirement score of each compressed data-related task is calculated according to the data features; a multidimensional decision matrix is ​​constructed based on the current computing resource utilization rate, available network bandwidth and real-time central server load of the smart city gateway system, as well as the resource requirement score of each compressed data-related task; an improved particle swarm optimization algorithm is used to solve the multi-objective optimization problem of task allocation with the multidimensional decision matrix as input to obtain an initial task allocation plan; load information and task features are exchanged with neighboring gateway nodes of the smart city gateway system through a distributed consensus algorithm, and tasks are reallocated to the initial task allocation plan according to the exchange results to obtain an optimized allocation plan; the local data-related tasks are divided into a local task list, and the compressed data-related tasks are divided into a local task list and a remote task list according to the optimized allocation plan; The method of exchanging load information and task characteristics with the neighboring gateway nodes of the smart city gateway system through a distributed consensus algorithm, and redistributing tasks for the initial task allocation plan according to the exchange results to obtain an optimized allocation plan includes: recording the load information of the smart city gateway system and the task characteristics of the tasks related to the compressed data in a pre-established distributed ledger to generate current transaction data; broadcasting the current transaction data to the neighboring gateway nodes through a consensus algorithm, receiving the transaction data of the neighboring gateway nodes, and updating the local distributed ledger; extracting the latest network load status and task distribution information based on the updated distributed ledger to construct a global resource view; applying a load balancing algorithm to utilize the global resource view, combining the initial task allocation plan to redistribute tasks, and generate an optimized allocation plan; The tasks in the local task list are locally processed by the smart city gateway system to obtain local analysis results, and the local analysis results and the task data in the remote task list are transmitted to the central server.

2. The data analysis method of the smart city gateway system according to claim 1 is characterized in that: The smart city gateway system performs adaptive sampling and preliminary classification processing on the raw data from the IoT devices and sensors in the smart city system to obtain a classified data set including: The data change rate of IoT devices and sensors in the smart city system is monitored in real time through the smart city gateway system, and a sampling frequency adjustment coefficient is calculated according to the data change rate; Dynamically sampling the original data based on the sampling frequency adjustment coefficient, and classifying and marking the sampled original data using a pre-trained decision tree model and a rule set to obtain a classification result; Metadata of the category identification of the classification result is generated, and the metadata is attached to the original data to form a classification data set.

3. The data analysis method of the smart city gateway system according to claim 2 is characterized in that: The raw data is dynamically sampled based on the sampling frequency adjustment coefficient, and the sampled raw data is classified and labeled using a pre-trained decision tree model and a rule set, and the classification results obtained include: Performing interpolation processing on the original data according to the sampling frequency adjustment coefficient to obtain an interpolated data sequence; Applying a sliding window technique to the interpolated data sequence to extract time series features and obtain a feature vector set; The feature vector set is input into a pre-trained decision tree model to obtain a preliminary classification result, and the preliminary classification result is adjusted and marked according to a predefined rule set to obtain a classification result.

4. A data analysis system for a smart city gateway system, characterized in that: The data analysis system of the smart city gateway system includes: A data classification module is used to perform adaptive sampling and preliminary classification processing on the raw data from the IoT devices and sensors in the smart city system through the smart city gateway system to obtain a classified data set; the classified data set includes emergency data, important data and regular data; A compression module is used to perform edge computing processing on the emergency data in the classified data set through the smart city gateway system using a preconfigured expert system to obtain emergency event analysis results, and use the emergency event analysis results as local data; obtain environmental parameters of the important data and regular data, and construct corresponding context feature vectors based on the environmental parameters; identify the data types of the important data and regular data based on the context feature vectors to obtain data type identifiers; select a corresponding compression algorithm based on the data type identifier, perform preliminary compression on the important data and regular data, and obtain preliminary compression results; use a deep reinforcement learning model, take the context feature vector and the preliminary compression results as input, and dynamically adjust the preset compression parameters to obtain optimized compression parameters; use the optimized compression parameters to re-compress the preliminary compression results to obtain compressed important data and regular data, and use the compressed important data and regular data as compressed data; A task scheduling module is used to obtain the data features of the compressed data through the smart city gateway system, and calculate the resource requirement score of each compressed data-related task based on the data features; construct a multidimensional decision matrix based on the current computing resource utilization rate, available network bandwidth and real-time central server load of the smart city gateway system, as well as the resource requirement score of each compressed data-related task; use an improved particle swarm optimization algorithm, with the multidimensional decision matrix as input, to solve the multi-objective optimization problem of task allocation and obtain an initial task allocation plan; exchange load information and task features with neighboring gateway nodes of the smart city gateway system through a distributed consensus algorithm, and redistribute tasks for the initial task allocation plan based on the exchange results to obtain an optimized allocation plan; divide the local data-related tasks into a local task list, and divide the compressed data-related tasks into a local task list and a remote task list based on the optimized allocation plan; The method of exchanging load information and task characteristics with the neighboring gateway nodes of the smart city gateway system through a distributed consensus algorithm, and redistributing tasks for the initial task allocation plan according to the exchange results to obtain an optimized allocation plan includes: recording the load information of the smart city gateway system and the task characteristics of the tasks related to the compressed data in a pre-established distributed ledger to generate current transaction data; broadcasting the current transaction data to the neighboring gateway nodes through a consensus algorithm, receiving the transaction data of the neighboring gateway nodes, and updating the local distributed ledger; extracting the latest network load status and task distribution information based on the updated distributed ledger to construct a global resource view; applying a load balancing algorithm to utilize the global resource view, combining the initial task allocation plan to redistribute tasks, and generate an optimized allocation plan; The local processing module is used to locally process the tasks in the local task list through the smart city gateway system, obtain local analysis results, and transmit the local analysis results and the task data in the remote task list to the central server.

Citation Information

Patent Citations

  • Video analysis method and system based on edge cloud cooperative computing in smart rod scene

    CN114501037A