Prediction system based on multi-source data fusion indoor air quality index AQI
By fusing multi-source data and extracting features, an air quality feature matrix and relationship graph are constructed, core feature sequences are selected, and a neural network model is used to achieve accurate prediction of indoor air quality index. This solves the problems of single data and insufficient feature extraction in existing technologies, and improves prediction accuracy and response capability.
Patent Information
- Application Number
- CN202610060993.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-16
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2046-01-16
AI Technical Summary
Existing technologies for predicting indoor air quality index (AQI) suffer from problems such as single data sources, data heterogeneity, lack of effective integration methods, insufficient feature extraction, low accuracy and stability of prediction results, and difficulty in responding to sudden pollution sources.
A multi-source data fusion approach is adopted to obtain multi-dimensional data related to indoor air quality from multiple heterogeneous data sources. The data is integrated and preprocessed to build an indoor air quality data warehouse. An air quality feature matrix and relationship graph are generated through feature extraction and relationship construction modules. The core feature sequences of air quality are screened out and AQI prediction is performed using a neural network model.
It achieves comprehensive coverage of indoor air quality and accurate capture of dynamic trends, providing timely and reliable forecast results, supporting air quality regulation, and ensuring a healthy indoor environment.
Smart Images

Figure CN121543030A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of indoor air monitoring technology, specifically to a prediction system based on multi-source data fusion of the indoor air quality index (AQI). Background Technology
[0002] With the acceleration of urbanization, people are spending more time indoors, and indoor air quality directly affects human health. The Indoor Air Quality Index (AQI), as a crucial indicator for measuring indoor air quality, has become a key focus of the industry for accurate prediction. Currently, technologies related to AQI prediction have been initially developed, but many shortcomings still exist in practical applications.
[0003] At the data acquisition level, most existing technologies rely on a single data source for data collection, such as obtaining limited data on temperature, humidity, and particulate matter concentration through fixed indoor monitoring equipment. This single-source acquisition method is insufficient to comprehensively reflect the actual indoor air quality and is prone to bias in subsequent analysis results due to missing data dimensions. Furthermore, data heterogeneity exists between different data sources. Even when multi-source data is acquired, there is a lack of effective data integration methods to unify and transform data of different formats and types into usable analytical data, leaving a large amount of data idle and unable to fully realize its value.
[0004] In the feature processing stage, traditional methods often employ simple statistical analysis techniques, such as calculating basic statistics like averages and variances, when extracting features from collected air quality data. This fails to delve deeper into the underlying characteristics hidden within the data, resulting in extracted features that cannot accurately represent the changing patterns of indoor air quality. Furthermore, the lack of a systematic method for constructing relationships between features makes it difficult to clearly define the connections between different feature dimensions. This lack of a reliable basis for subsequent selection of core features makes it difficult to identify feature sequences that have a crucial impact on AQI prediction.
[0005] In terms of predictive model application, due to insufficient data integration and poor feature extraction and screening, existing predictive systems often fail to accurately capture the dynamic trends of indoor air quality when predicting AQI, resulting in low accuracy and stability of prediction results. For example, when sudden pollution sources appear indoors (such as harmful gases released by furniture or pollutants generated by human activities), existing systems struggle to respond quickly and adjust their predictive models, leading to significant deviations between predicted results and actual AQI values. This makes it difficult to provide effective reference information for indoor air quality control and to meet people's needs for a healthy indoor environment. Summary of the Invention
[0006] The purpose of this invention is to provide a prediction system for indoor air quality index (AQI) based on multi-source data fusion, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides a prediction system for indoor air quality index (AQI) based on multi-source data fusion, the system comprising:
[0008] The data acquisition module is used to acquire multidimensional data related to indoor air quality from multiple heterogeneous data sources, integrate the multidimensional data, and obtain an indoor air quality data warehouse.
[0009] The feature extraction module is used to extract features from the indoor air quality data warehouse to obtain an indoor air quality feature matrix.
[0010] The relationship building module is used to determine the dimensional relationship diagram of each dimension in the indoor air quality feature matrix, and to build a data relationship map related to indoor air quality through the dimensional relationship diagrams of each dimension.
[0011] The core screening module is used to extract air quality feature vectors of each dimension from the indoor air quality feature matrix, determine the air quality feature weights of each dimension based on the corresponding air quality feature vectors, and screen out the air quality core feature sequence based on all air quality feature weights and the data relationship graph.
[0012] The prediction module is used to predict the indoor air quality index (AQI) based on the core air quality feature sequence.
[0013] Preferably, the multidimensional data related to indoor air quality in the data acquisition module includes temperature data, humidity data, carbon dioxide data, particulate matter data, and volatile organic compound data.
[0014] Preferably, the data acquisition module integrates the multidimensional data to obtain an indoor air quality data warehouse, specifically including:
[0015] The multidimensional data is preprocessed to obtain preprocessed multidimensional data;
[0016] Semantic mapping is performed on the preprocessed multidimensional data to obtain indoor air quality-related data for each dimension.
[0017] An indoor air quality data warehouse is built using indoor air quality-related data from all dimensions.
[0018] Preferably, the feature extraction module performs feature extraction on the indoor air quality data warehouse to obtain an indoor air quality feature matrix, specifically including:
[0019] Feature extraction is performed on the indoor air quality-related data of each dimension in the indoor air quality data warehouse to obtain air quality feature vectors for each dimension.
[0020] An indoor air quality feature matrix is constructed based on air quality feature vectors of all dimensions.
[0021] Preferably, the determination of the dimensional relationship diagram for each dimension of the indoor air quality feature matrix in the relationship construction module specifically includes:
[0022] Select one dimension from all dimensions of the indoor air quality feature matrix and obtain the air quality feature vector of the selected dimension;
[0023] Determine the feature correlation between various indoor air quality-related features in the air quality feature vector of the selected dimension;
[0024] Based on the feature correlation between various indoor air quality-related features, a dimensional relationship diagram of the selected dimension is constructed, thereby obtaining the dimensional relationship diagram of each dimension in the indoor air quality feature matrix.
[0025] Preferably, the data relationship graph related to indoor air quality constructed in the relationship construction module through dimensional relationship graphs of various dimensions specifically includes:
[0026] Identify key indoor air quality-related features for each dimension, and then determine the feature correlation between these key indoor air quality-related features.
[0027] By connecting the dimensional relationship diagrams of each key indoor air quality-related feature based on the feature correlation between them, a data relationship map related to indoor air quality is obtained.
[0028] Preferably, the core screening module determines the air quality feature weights for each dimension based on the corresponding air quality feature vector, specifically including:
[0029] Obtain the scaling factor for each dimension;
[0030] For each dimension of the air quality feature vector, obtain the weights corresponding to each indoor air quality-related feature in the air quality feature vector;
[0031] The air quality feature weights of the corresponding dimensions of the air quality feature vector are determined by the weights and scale factors of each indoor air quality-related feature, thereby obtaining the air quality feature weights of each dimension.
[0032] Preferably, the core screening module, which filters out the air quality core feature sequence based on all air quality feature weights and the data relationship graph, specifically includes:
[0033] For each indoor air quality-related feature in the indoor air quality feature matrix, determine the air quality feature weight of the dimension in which the indoor air quality-related feature is located;
[0034] Extract the correlation of all features corresponding to the indoor air quality-related features from the data relationship graph;
[0035] The core air quality index of the indoor air quality related feature is determined by the air quality feature weight of the dimension in which the indoor air quality related feature is located and the corresponding correlation of all features, thereby obtaining the core air quality index of each indoor air quality related feature in the indoor air quality feature matrix.
[0036] All core air quality characteristics were selected based on the core air quality indicators related to various indoor air quality features;
[0037] Construct an air quality core feature sequence based on all the core air quality features.
[0038] Preferably, the prediction module predicts the indoor air quality index (AQI) based on the air quality core feature sequence by inputting the air quality core feature sequence into the prediction model to obtain the predicted value of the indoor air quality index (AQI).
[0039] Preferably, the prediction model is a neural network model.
[0040] Compared with the prior art, the beneficial effects of the present invention are:
[0041] By setting up a data acquisition module, multidimensional data related to indoor air quality can be obtained from multiple heterogeneous data sources and effectively integrated to form an indoor air quality data warehouse. Compared to traditional single-data source acquisition methods, this multi-source data acquisition and integration model can comprehensively cover various influencing factors of indoor air quality, such as particulate matter concentration, harmful gas content, temperature and humidity, and ventilation conditions. It breaks down data barriers between different data sources, transforms heterogeneous data into unified and usable data, fully activates the value of multi-source data, and provides a comprehensive and reliable data foundation for subsequent feature extraction, relationship construction, and AQI prediction, avoiding analytical biases caused by missing data dimensions or inconsistent data formats.
[0042] The feature extraction module enables deep feature extraction from air quality data in the data warehouse, generating an indoor air quality feature matrix. Unlike traditional simple statistical analysis-based feature extraction methods, this module can uncover hidden potential features in the data, such as the changing trends of air quality data over different time periods and the synergistic changes among different pollutants. The extracted feature matrix more accurately reflects the intrinsic attributes and changing patterns of indoor air quality, providing rich feature resources for subsequent feature analysis and core feature selection. This makes the subsequent analysis process more targeted and effective, avoiding the problem of poor analysis results caused by insufficient feature representation capabilities in traditional feature extraction methods.
[0043] The relationship building module determines the dimensional relationship diagram for each dimension in the feature matrix and constructs a data relationship graph based on these diagrams, clearly presenting the correlations between various air quality feature dimensions. This systematic relationship building approach can intuitively demonstrate the mutual influence and constraints between different features, such as the correlation between temperature changes and the rate of harmful gas volatilization, and the correlation between ventilation volume and particulate matter concentration. This helps staff deeply understand the role of each feature dimension in changes in indoor air quality, providing a clear relational basis for the selection of core features. It solves the problems of vague and difficult-to-organize feature relationships in traditional methods, making the subsequent core feature selection process more logical and scientific.
[0044] The core screening module extracts air quality feature vectors from various dimensions, determines the feature weights for each dimension based on these vectors, and then combines this with a data relationship graph to select the core air quality feature sequence. This screening process comprehensively considers the importance of each feature and the correlation between them, accurately identifying features that have a key impact on indoor AQI changes while eliminating irrelevant or secondary features. This effectively reduces the computational complexity of subsequent prediction models and ensures that the core feature sequence accurately reflects key changes in indoor air quality. Using the selected core feature sequence for AQI prediction reduces the interference of irrelevant features on the prediction results and improves the prediction model's ability to capture trends in indoor air quality changes.
[0045] The prediction module forecasts indoor air quality index (AQI) based on core feature sequences. Thanks to a solid foundation of prior data and excellent feature processing, this module can more accurately capture the dynamic changes in indoor air quality. Whether it's slow changes in air quality under normal conditions or rapid fluctuations caused by sudden pollution sources, it can respond quickly and make accurate predictions. The prediction results accurately reflect the actual trend of indoor AQI changes, providing timely and reliable reference information for the formulation of indoor air quality control measures. For example, when it is predicted that the AQI will exceed the safety threshold, ventilation and purification equipment can be activated in advance to control air quality, helping people improve their indoor air quality environment in a timely manner, ensuring the health and safety of indoor occupants, and better meeting people's needs for a healthy indoor living environment. Attached Figure Description
[0046] Figure 1 This is a time series diagram of the indoor air quality index (AQI) prediction system based on multi-source data fusion as described in this invention.
[0047] Figure 2 This is a flowchart of data integration in the data acquisition module;
[0048] Figure 3 A flowchart for constructing the dimensional relationship graph in the relationship construction module. Detailed Implementation
[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] Please see Figure 1 This invention provides a prediction system for indoor air quality index (AQI) based on multi-source data fusion. The system includes: a data acquisition module, a feature extraction module, a relationship construction module, a core screening module, and a prediction module.
[0051] The data acquisition module obtains multidimensional data related to indoor air quality from multiple heterogeneous data sources and integrates this data to form an indoor air quality data warehouse. The feature extraction module extracts features from the indoor air quality data warehouse to generate an indoor air quality feature matrix. The relationship construction module determines the dimensional relationship graph for each dimension in the indoor air quality feature matrix and constructs a data relationship graph related to indoor air quality based on these dimensional relationship graphs. The core selection module extracts air quality feature vectors for each dimension from the indoor air quality feature matrix, determines the air quality feature weights for each dimension based on the corresponding air quality feature vectors, and selects the core air quality feature sequence based on all air quality feature weights and the data relationship graph. The prediction module predicts the indoor air quality index (AQI) based on the core air quality feature sequence.
[0052] Example 1: See Figure 2 The data acquisition module is responsible for obtaining multi-dimensional monitoring information related to indoor air quality from various heterogeneous data sources. This information covers multiple key parameter dimensions, including temperature, humidity, carbon dioxide, particulate matter, and volatile organic compounds (VOCs). Temperature data typically comes from digital temperature sensors distributed throughout the indoor environment. These sensors collect ambient temperature readings at specific time intervals to form a time-series data stream. Humidity data is acquired through high-precision humidity sensors, which accurately reflect changes in the water vapor content in the air. Carbon dioxide concentration data is collected by non-dispersive infrared sensors, which quantitatively calculate gas concentration by detecting the degree to which carbon dioxide molecules absorb infrared light of a specific wavelength. Particulate matter monitoring data mainly comes from PM2.5 and PM10 sensors based on the laser scattering principle, which invert particulate matter mass concentration by measuring the scattering intensity of suspended particulate matter on a laser beam. VOCs data is collected by metal oxide semiconductor or photoionization detector sensors, which have broad-spectrum response characteristics to various volatile organic compounds. All of this data is uploaded to the central data processing unit in real time via wired or wireless transmission protocols. However, due to differences in sensor type, manufacturer specifications, and communication protocols, the raw data exhibits significant heterogeneity in terms of format, accuracy, and acquisition frequency.
[0053] The first step in integrating the aforementioned multidimensional monitoring data is to perform systematic preprocessing. This preprocessing aims to eliminate noise interference in the original data, fill in missing data, and correct outliers. The data cleaning stage uses digital filtering algorithms to smooth the original signals, with moving average filtering and Kalman filtering applied to eliminate random noise and transient impulse interference. For missing value handling, appropriate imputation strategies are selected based on the missing data mechanism. For random missing values with a low missing rate, time-series linear interpolation is used for imputation; for continuous missing segments, autoregressive models are used for predictive imputation. Outlier detection is achieved through statistical process control methods. First, the moving average and standard deviation of each parameter are calculated. Then, data points exceeding three times the standard deviation are marked as outliers and corrected or removed. After preprocessing, the multidimensional data achieves uniformity in timestamps, data precision, and sampling intervals, forming a clean and well-organized time-series dataset.
[0054] After data preprocessing, semantic mapping is required for the standardized multidimensional data. The core objective of semantic mapping is to establish the correspondence between raw sensor data and standardized indoor air quality parameters. This process is achieved through a predefined mapping rule base, which contains conversion relationships between various sensor output values and standard parameter values. For example, the raw voltage output value of a certain type of temperature sensor needs to be mapped to a Celsius temperature value using a linear conversion formula; the frequency signal output of a humidity sensor needs to be converted to a relative humidity percentage value using a lookup table; the analog voltage output of a carbon dioxide sensor needs to be mapped to ppm concentration units using a calibration curve provided by the manufacturer; the raw count data of a particulate matter sensor needs to be calculated into the μg / m³ concentration values of PM2.5 and PM10 using a mass concentration conversion algorithm; and the resistance change value of a volatile organic compound sensor needs to be converted to a TVOC concentration value using a polynomial fitting equation. The semantic mapping process also includes unit standardization, converting all parameter values to international standard units of measurement, and adding metadata tags to all data streams to identify parameter types, data sources, and quality levels.
[0055] Standardized indoor air quality data obtained through semantic mapping across various dimensions will be used to construct a structured indoor air quality data warehouse, built using a star schema design pattern. The fact table records the measured values of all parameters at each sampling time using timestamps as the primary key, while the dimension tables include auxiliary dimensions such as sensor device information, geographical location information, and environmental condition information. The physical storage of the data warehouse utilizes a columnar database to improve the compression efficiency and query performance of time-series data. Simultaneously, it establishes partitioned indexes based on time ranges and secondary indexes based on parameter types to support efficient multi-dimensional data retrieval. The data warehouse layer also includes a data quality monitoring module, which continuously monitors the integrity, consistency, and timeliness indicators of the incoming data, issuing early warning signals for abnormal data quality states. The resulting indoor air quality data warehouse not only provides storage and management functions for historical data but also offers unified data access services to upper-layer applications through standardized interfaces.
[0056] Example 2: The feature extraction module extracts features from the indoor air quality data warehouse. This process first processes the indoor air quality-related data for each dimension in the data warehouse. Feature extraction for the temperature data dimension includes not only the calculation of basic statistical features such as mean, variance, range, and percentiles, but also the construction of time-series features. The sliding window mean reflects short-term trend changes, the first-order difference sequence captures the instantaneous change amplitude, and the periodic components are extracted using Fourier transform to obtain daily and weekly periodic features. Autocorrelation features are used to quantify the memory effect and persistence characteristics of temperature data. A similar feature engineering strategy is used for the humidity data dimension. In addition to conventional statistics, the focus is on extracting the co-variance features of humidity and temperature, including the humidity-temperature ratio sequence, the humidity change lag correlation coefficient, and the dynamic calculation of the humidity saturation difference. These features can reflect the actual perceived effect of environmental humidity and the correlation information with thermal comfort.
[0057] Feature extraction for carbon dioxide data focuses more on the dynamic characteristics of concentration changes. In addition to basic statistical features, it constructs cumulative concentration features, calculates the percentage of hours with concentration exceeding the standard, designs a concentration change acceleration index to reflect the severity of concentration fluctuations, and extracts peak concentration features including peak frequency, duration, and rate of increase and decrease. Addressing the exponential growth characteristic often exhibited by carbon dioxide concentrations, it specifically designs exponential fitting residual features to quantify the deviation between actual values and the ideal exponential model, while constructing concentration recovery features to reflect the rate and efficiency of concentration decrease under ventilation conditions. These features together constitute a complete descriptive system for the dynamic behavior of carbon dioxide concentrations. Feature extraction for particulate matter data needs to distinguish the different characteristics of PM2.5 and PM10. Besides calculating the mass concentration statistical features of the two types of particulate matter separately, it focuses on constructing a series of ratio features: the PM2.5 / PM10 ratio reflects the source characteristics of particulate matter; particle size distribution features are reflected by the concentration ratio of different particle size ranges; and the cumulative effect of particulate matter is quantified by calculating the concentration-time integral value to quantify the exposure level. To address the frequent sudden peaks in particulate matter concentration, a peak feature set was designed, including peak identification, peak steepness calculation, and peak attenuation pattern classification. At the same time, correlation features between particulate matter concentration and meteorological conditions were constructed, such as the influence coefficient of humidity on particulate matter concentration and the correlation index between temperature inversion layer and concentration accumulation.
[0058] Feature extraction for volatile organic compound (VOC) data employs a multi-scale analysis method. In the time domain, it extracts the mean, variance, and peak values of TVOC concentration. In the frequency domain, it uses wavelet transform to extract concentration fluctuation patterns at different time scales. In terms of morphological features, it identifies the rising edge, plateau phase, and falling edge of the concentration curve. Considering the diverse nature of VOCs, a feature weighting mechanism is designed, assigning higher feature weights to highly toxic compounds, strengthening feature extraction for compounds with low odor thresholds, and constructing separate long-term accumulation features for persistent organic pollutants. Furthermore, considering the correlation between VOC concentration and indoor activity intensity, it constructs VOC release pattern features based on population density and activity type.
[0059] When constructing an indoor air quality feature matrix based on air quality feature vectors across all dimensions, it is necessary to address the issue of differences in feature scales. A hierarchical standardization method is employed to normalize various features: statistical features are standardized using Z-scores, ratio features using decimal scaling, and time-series features using maximum-minimum normalization. The row dimensions of the feature matrix correspond to the sampling points of the time series, while the column dimensions contain the concatenated vectors of all features. During matrix filling, a feature priority ranking mechanism is used to place more important features at the top of the matrix. The resulting indoor air quality feature matrix not only retains the information content of the original data but also enhances the data's representational capabilities through feature engineering, providing structured input data for subsequent relationship construction and core feature selection. The matrix storage employs a sparse matrix compression format to improve storage efficiency, while a feature index mapping table is established to enable rapid feature retrieval and updates.
[0060] Taking an open-plan office environment as an example for air quality monitoring, the feature extraction module needs to process indoor monitoring data from a continuous week. This environment deploys multiple types of sensors that collect data hourly. Temperature data recorded by the temperature sensors exhibits a typical diurnal cycle: the temperature gradually rises during the morning rush hour due to increased personnel and equipment operation, peaks at midday and then slightly declines, continuing to drop after the evening rush hour until reaching its lowest value the following morning. To address this pattern, the feature extraction process not only calculates the hourly temperature average but also pays particular attention to the rate of temperature increase (temperature change rate within the first two hours of the morning) and the rate of temperature decrease (temperature change rate within the first three hours of the evening). It also identifies the daily temperature fluctuation range (the difference between the highest and lowest values) and periods of stable temperature (the percentage of time when the rate of change is below a threshold). Humidity data shows the opposite trend to temperature changes: humidity is low in the morning due to ventilation system activation, gradually increases with human activity, respiration, and plant transpiration, and reaches a relatively stable state in the afternoon. Features extracted from the humidity data include absolute humidity values, relative humidity change curves, identification of humidity inflection points, and a sequence of humidity-temperature differences. Of particular note was the abnormal fluctuation in humidity data on Wednesday afternoon. This anomaly was identified during feature extraction, and the amplitude and duration of the fluctuations were calculated. Carbon dioxide concentration data clearly reflected patterns of human activity: concentrations accumulated continuously during working hours, peaked during meetings, and decreased during rest periods. Features extracted from this data included basic statistics, concentration accumulation rate, concentration decline rate, and duration of concentration exceeding limits. A team meeting on Thursday morning led to a significant increase in carbon dioxide concentration; the start time, peak concentration, and time required to return to normal were recorded during feature extraction. Particulate matter data included PM2.5 and PM10, with feature extraction revealing different patterns of change. PM2.5 concentrations showed short-term peaks during morning cleaning activities and lunchtime food delivery, while PM10 concentrations increased during periods of open ventilation. In addition to calculating the concentration statistics for each type of particulate matter, the extracted features included the PM2.5 / PM10 ratio curve, peak duration, and the correlation between particulate matter concentration and the status of open / closed doors and windows. Outdoor construction on Friday afternoon caused an anomaly in particulate matter concentration; this event was specifically recorded, and the concentration surge rate and natural sedimentation rate were extracted. The volatile organic compound (VOC) data exhibited a complex pattern of variation, with TVOC concentrations rising during morning cleaning agent use, showing slight fluctuations during printer and copier operation, and exhibiting a clear continuous release characteristic in areas with new furniture. Features extracted from this data included baseline concentration values, concentration change gradients, specific compound identification features, and correlation indicators between concentrations and indoor activities. The newly installed desk on Monday caused persistently high TVOC concentrations; this was recorded as a continuous release event, and release rate and decay constant features were extracted. After feature extraction across all dimensions, a feature matrix was constructed according to time alignment principles.The matrix's rows correspond to 168 time points, and the columns contain 56 feature indicators. Data standardization was employed during the feature matrix construction, transforming feature values of different dimensions to the same scale while preserving the relative relationships of the original data. Anomalous event time points are specifically marked in the matrix to provide reference information for subsequent analysis. The final feature matrix contains detailed feature information for each parameter while maintaining the integrity of the time series.
[0061] Example 3: See Figure 3 The process of determining the dimensional relationship diagram for each dimension in the indoor air quality feature matrix within the relationship construction module requires a systematic analysis of the intrinsic correlations between features. After selecting one dimension from all dimensions of the indoor air quality feature matrix, the air quality feature vector of that selected dimension forms the basis of the analysis. Taking the temperature dimension as an example, its feature vector contains more than ten specific feature terms, such as mean, variance, and moving average, each representing different aspects of the temperature data. When determining the feature correlation between these indoor air quality-related features, an information theory-based correlation measurement method is used, one of the core calculation methods being defined as follows:
[0062]
[0063] in: This represents the adjusted correlation between feature X and feature Y. and These are the values of features X and Y at the i-th sampling point, respectively. and is the mean of the corresponding feature, n is the total number of sampling points, r is the traditional Pearson correlation coefficient, and α is an adjustment coefficient used to control the strength of nonlinear correction. This calculation method retains the metric characteristics of linear correlation while smoothing extreme correlation values through the sigmoid function, avoiding weight bias in subsequent graph construction.
[0064] When constructing a dimensional relationship graph for selected dimensions based on the feature correlations between various indoor air quality-related features, a weighted undirected graph structure is used for representation: nodes in the graph represent each specific feature, edges represent the correlations between features, and the weight of the edge is the calculated absolute value of the feature correlation. A correlation threshold is set during graph construction, retaining only significant correlations (edges with an absolute correlation value exceeding 0.3). This filters out random correlations and highlights stable, strongly correlated feature pairs. For the temperature dimension, a high-strength positive correlation may be found between the moving average feature and the mean feature, while the variance feature and the extreme value feature show a moderate correlation. These relationships are accurately expressed in the graph as weighted edges. Each dimension independently constructs its own dimensional relationship graph following this process, forming a complete set including temperature, humidity, and carbon dioxide relationship graphs. Constructing an indoor air quality-related data relationship map through the dimensional relationship graphs of each dimension requires cross-dimensional correlation integration. A feature importance-based screening method is used to determine the key indoor air quality-related features for each dimension. The centrality indices of each feature within its respective dimension are calculated, including degree centrality (the number of edges directly connected to the feature), betweenness centrality (the frequency with which the feature appears on the shortest path), and feature variance contribution rate. The top 30% of features are selected as key features based on these indices. For the temperature dimension, moving averages and daily average fluctuations might be selected as key features; for the humidity dimension, relative humidity saturation and rate of change might be used. Furthermore, a cross-dimensional correlation analysis algorithm is employed to determine the correlation between key indoor air quality features. This algorithm not only calculates the linear correlation between features of different dimensions but also introduces time-delay cross-correlation analysis to capture correlations with time lags. For example, the daily average fluctuation feature of the temperature dimension might have a 6-hour lag correlation with the saturation feature of the humidity dimension; this time-delay correlation is accurately captured through sliding time window cross-correlation calculations. After calculating the correlation of all cross-dimensional feature pairs, a complete cross-feature correlation matrix is formed.
[0065] When connecting the dimensional relationship graphs of various dimensions based on the feature correlations between key indoor air quality-related features, graph fusion technology is employed. While maintaining the independent structure of each dimensional relationship graph, cross-graph edges are added between key feature nodes in different graphs, with the edge weight corresponding to the cross-feature correlation value. During the connection process, a cross-graph edge weight threshold is set, retaining only edges with significant cross-dimensional associations (weight values exceeding 0.25) to ensure the sparsity and interpretability of the final graph. For example, a moving average node in the temperature relationship graph may be connected to a rate of change node in the humidity relationship graph, with a weight value of 0.32; a concentration peak node in the carbon dioxide relationship graph may be connected to a release mode node in the volatile organic compound relationship graph, with a weight value of 0.28. These cross-graph edges truly achieve the organic connection of multi-dimensional features. This results in an indoor air quality-related data relationship graph, which is stored using a multi-layer graph structure: the bottom layer contains subgraphs of internal relationships within each dimension, and the upper layer contains cross-dimensional connection relationships. The graph is persistently stored using a graph database, supporting efficient feature association queries and path analysis. The resulting data relationship graph not only reveals the feature correlation structure within a single dimension, but more importantly, it discovers the complex interaction relationships between different air quality parameters, providing a comprehensive relationship network foundation for subsequent core feature selection. The graph's visualization employs a force-directed layout algorithm, clustering closely related features for display, thus intuitively showcasing the complex relationship patterns between multi-dimensional features.
[0066] Example 4: The process of selecting the core air quality feature sequence based on all air quality feature weights and data relationship graphs in the core screening module requires comprehensive consideration of the feature's own importance and its influence in the correlation network. For each indoor air quality-related feature in the indoor air quality feature matrix, a dimension weight allocation scheme is adopted when determining the air quality feature weight of the dimension in which the indoor air quality-related feature is located. Each dimension is assigned a different weight value based on its historical correlation strength with the Air Quality Index (AQI): temperature dimension weight is 0.18, humidity dimension weight is 0.15, carbon dioxide dimension weight is 0.22, particulate matter dimension weight is 0.25, and volatile organic compound (VOC) dimension weight is 0.20. These weight values are determined by analyzing the historical correlation between each dimension parameter and the AQI value, and are adjusted by experts to form the final weight allocation.
[0067] When extracting the correlation scores of all features corresponding to each indoor air quality-related feature in the data relationship graph, it is necessary to traverse all edges connected to the feature node in the graph. Taking the "moving average" feature of the temperature dimension as an example, this feature has a correlation score of 0.85 with the "daily average fluctuation" feature within the temperature dimension, a cross-dimensional correlation score of 0.42 with the "relative humidity saturation" feature of the humidity dimension, a correlation score of 0.31 with the "concentration change acceleration" feature of the carbon dioxide dimension, and a correlation score of 0.28 with the "PM2.5 / PM10 ratio" feature of the particulate matter dimension. These correlation scores are directly obtained from the edge weight attributes of the graph, forming a complete set of correlation scores for this feature. When determining the core air quality index of an indoor air quality-related feature based on the air quality feature weights of the dimension in which the feature is located and the corresponding correlation scores of all features, a weighted comprehensive evaluation method is used. The core index calculation formula is: the product of the dimension weight and the weighted average of all correlation scores, where the weight of the correlation score is its own numerical value (i.e., highly correlated correlation scores contribute more to the result). Continuing with the "moving average" feature as an example, its dimensional weight is 0.18. The weighted average of all correlations is (0.85×0.85+0.42×0.42+0.31×0.31+0.28×0.28) / (0.85+0.42+0.31+0.28)=0.57, and the final core indicator value is 0.18×0.57=0.103. Each feature's core indicator value is calculated using this method, forming a comprehensive feature importance assessment system. After obtaining the core air quality indicators for each indoor air quality-related feature in the indoor air quality feature matrix, these indicator values need to be standardized. All core indicator values are normalized to the [0,1] interval for easier comparison and threshold selection. The normalized core indicator values reflect the relative importance of each feature in the overall feature set; high-value features represent important positions in the multi-dimensional correlation network and originate from high-weight dimensions.
[0068] When selecting all core air quality features based on various indoor air quality-related characteristics, a dynamic threshold determination mechanism is employed. First, the mean and standard deviation of all core indicator values are calculated. The threshold is then set to the mean plus one standard deviation, automatically adapting to the feature distribution characteristics of different datasets. Features with core indicator values exceeding this threshold are marked as core features and included in subsequent sequence construction. This dynamic threshold method ensures the objectivity and adaptability of core feature selection, avoiding subjective bias from manually setting thresholds. When constructing the air quality core feature sequence based on all core air quality features, the values of these features need to be organized chronologically. A sliding window mechanism is used for sequence construction; the core feature values at each time point form a feature vector, arranged chronologically to form a multi-dimensional time series. Each element in the sequence contains the values of all core features at that time point, forming a regular data format for prediction. Refer to Table 1; to more clearly illustrate the core feature selection results, details of the core indicator calculations for some features are listed.
[0069] Table 1: Calculation of Core Feature Indicators
[0070] Feature Name Belonging Dimension Dimension weights Number of related features Average correlation Core Indicator Values Is it core? Temperature moving average temperature 0.18 4 0.57 0.103 yes <![CDATA[Acceleration of CO2 change]]> carbon dioxide 0.22 5 0.62 0.136 yes PM2.5 / PM10 ratio Particulate matter 0.25 3 0.48 0.120 yes VOC release mode VOC 0.20 2 0.35 0.070 no Humidity saturation humidity 0.15 4 0.41 0.062 no Peak frequency of particulate matter Particulate matter 0.25 6 0.58 0.145 yes
[0071] The final constructed air quality core feature sequence includes timestamp information and numerical values for all core features. The sequence length matches the time range of the original data, but the feature dimensionality is significantly reduced (from dozens of features to the number of core features). This sequence is stored in a standard time series format, with each time point containing a timestamp and an array of core feature values for direct use by the prediction module. The sequence data uses a columnar storage format to improve read efficiency, and a time-based index is established to support fast query access. The core feature sequence not only reduces data dimensionality but, more importantly, retains the most predictive information, providing optimized input data for subsequent air quality index predictions.
[0072] Example 5: The prediction module predicts the indoor Air Quality Index (AQI) based on the core air quality feature sequence. This requires inputting the sequence data into a trained prediction model. The core air quality feature sequence, as model input, contains historical data from multiple time steps. The input vector for each time step consists of selected core feature values, such as temperature moving average, CO2 acceleration, PM2.5 / PM10 ratio, and peak particulate matter frequency. These features are arranged chronologically to form a multi-dimensional time series input. The input sequence duration is set to 24 hours (one sampling point per hour), meaning the model needs to receive core feature data from 24 consecutive time steps to predict the AQI value at a future point in time. The prediction model employs a Long Short-Term Memory (LSTM) neural network structure, which comprises three main components: an input layer designed as a variable-length time series input interface, capable of receiving core feature sequences of different durations; an intermediate layer consisting of two stacked LSTM units, with the first LSTM containing 64 neurons for extracting short-term time patterns and the second LSTM containing 32 neurons for capturing long-term time dependencies; and an output layer that is fully connected, mapping the LSTM output to a single AQI prediction value. The network uses Dropout regularization to prevent overfitting, with a dropout rate of 0.2. A batch normalization layer is added after the LSTM layer to accelerate the training process and improve model stability.
[0073] The model is trained using a historical air quality dataset, with training data including indoor environmental monitoring data collected over the past six months. The dataset is divided into training, validation, and test sets in a 7:2:1 ratio. The training set is used for model parameter learning, the validation set for hyperparameter tuning and early stopping detection, and the test set for final performance evaluation. Mean squared error is used as the loss function during training. The Adam algorithm is selected as the optimizer with an initial learning rate of 0.001 and an exponential decay strategy. The batch size is set to 32 samples, and the maximum number of training epochs is set to 200. An early stopping mechanism is implemented, terminating training when the validation set loss no longer decreases after 10 consecutive epochs. The actual deployment of the prediction model requires a complete data pipeline. The real-time acquisition system continuously obtains the latest monitoring data from the sensor network. After data cleaning and feature extraction, the core feature values for the current time point are obtained. These feature values are added to a circular buffer to maintain the core feature sequence for the most recent 24 hours. The buffer data is updated every minute to ensure the timeliness of the input sequence. When forecasting is required, the system extracts the complete time series input from the buffer and feeds it into a pre-trained neural network model. The model inference process is executed on a dedicated hardware accelerator to ensure real-time performance. The model output is a predicted value for the Indoor Air Quality Index (AQI), with values ranging from 0 to 500 corresponding to different air quality levels. Post-processing of the forecast results includes numerical smoothing and range constraints. Moving average filtering is used to smooth continuous forecast results to eliminate random fluctuations while ensuring that the output value does not exceed the standard range of AQI. The system provides forecasting capabilities at three time scales: 1 hour, 3 hours, and 6 hours. Multi-step forecasting capabilities are achieved by adjusting the time length of the input sequence and the configuration of the output layer. The performance monitoring and maintenance mechanism of the forecasting system includes regular model retraining and prediction bias detection. Every two weeks, the model is incrementally trained using the latest collected data to adapt to changes in data distribution. A deviation alarm mechanism between predicted and actual measurements is established, triggering a model recalibration process when significant deviations occur consecutively. All forecast results and corresponding input data are recorded in a forecast log database. This data is used for subsequent model optimization and prediction accuracy analysis, forming a closed-loop model improvement system. The user interface layer provides various formats for displaying prediction results, including numerical displays, trend charts, and color-coded levels. It also supports historical prediction record queries and prediction accuracy statistics. The system provides an API interface allowing other applications to access prediction data, and supports the JSON data exchange standard for easy system integration and data sharing. The entire prediction module automates the process from data input to result output, providing reliable technical support for indoor air quality management and early warning.
[0074] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0075] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A prediction system for indoor air quality index (AQI) based on multi-source data fusion, characterized in that, The system comprises: a data acquisition module for acquiring multi-dimensional data related to indoor air quality from a plurality of heterogeneous data sources, performing data integration on the multi-dimensional data, and obtaining an indoor air quality data warehouse; a feature extraction module for performing feature extraction on the indoor air quality data warehouse, and obtaining an indoor air quality feature matrix; a relationship construction module for determining a dimension relationship graph of each dimension in the indoor air quality feature matrix, and constructing a data relationship graph atlas related to indoor air quality through the dimension relationship graphs of the respective dimensions; a core screening module for extracting an air quality feature vector of each dimension from the indoor air quality feature matrix, determining an air quality feature weight of each dimension according to the corresponding air quality feature vector, and screening an air quality core feature sequence based on all the air quality feature weights and the data relationship graph atlas; a prediction module for predicting an indoor air quality index (AQI) according to the air quality core feature sequence.
2. The prediction system for indoor air quality index (AQI) based on multi-source data fusion of claim 1, wherein, The multi-dimensional data related to indoor air quality in the data acquisition module includes temperature data, humidity data, carbon dioxide data, particulate matter data, and volatile organic compound data.
3. The prediction system for indoor air quality index (AQI) based on multi-source data fusion of claim 1, wherein, The data acquisition module performs data integration on the multi-dimensional data to obtain an indoor air quality data warehouse, specifically including: preprocessing the multi-dimensional data to obtain preprocessed multi-dimensional data; performing semantic mapping on the preprocessed multi-dimensional data to obtain indoor air quality related data of each dimension; constructing the indoor air quality data warehouse through the indoor air quality related data of all dimensions.
4. The prediction system for indoor air quality index (AQI) based on multi-source data fusion of claim 1, wherein, The feature extraction module performs feature extraction on the indoor air quality data warehouse to obtain an indoor air quality feature matrix, specifically including: performing feature extraction on the indoor air quality related data of each dimension in the indoor air quality data warehouse to obtain an air quality feature vector of each dimension; constructing the indoor air quality feature matrix according to the air quality feature vectors of all dimensions.
5. The prediction system for indoor air quality index (AQI) based on multi-source data fusion as claimed in claim 1, wherein, The relationship construction module determines the dimension relationship graph of each dimension in the indoor air quality feature matrix, specifically including: selecting a dimension from all dimensions of the indoor air quality feature matrix, and obtaining an air quality feature vector of the selected dimension; determining the feature correlation between each indoor air quality related feature in the air quality feature vector of the selected dimension; constructing the dimension relationship graph of the selected dimension according to the feature correlation between each indoor air quality related feature, and obtaining the dimension relationship graph of each dimension in the indoor air quality feature matrix.
6. The prediction system for indoor air quality index (AQI) based on multi-source data fusion of claim 1, wherein, The relationship construction module constructs a data relationship graph atlas related to indoor air quality through the dimension relationship graphs of the respective dimensions, specifically including: determining key indoor air quality related features of each dimension, and determining the feature correlation between each key indoor air quality related feature; connecting the dimension relationship graphs of each dimension according to the feature correlation between each key indoor air quality related feature, and obtaining the data relationship graph atlas related to indoor air quality.
7. The prediction system for indoor air quality index (AQI) based on multi-source data fusion of claim 1, wherein, The core screening module determines the air quality feature weight of each dimension according to the corresponding air quality feature vector, and specifically includes the following steps: Obtaining a scale factor of each dimension; For the air quality feature vector of each dimension, obtaining the weight corresponding to each indoor air quality related feature in the air quality feature vector; Determining the air quality feature weight of the corresponding dimension of the air quality feature vector through the weight corresponding to each indoor air quality related feature and the corresponding scale factor, and then obtaining the air quality feature weight of each dimension.
8. The prediction system for indoor air quality index (AQI) based on multi-source data fusion of claim 1, wherein, The core screening module screens the air quality core feature sequence based on all air quality feature weights and the data relationship graph, and specifically includes the following steps: For each indoor air quality related feature in the indoor air quality feature matrix, determining the air quality feature weight of the dimension where the indoor air quality related feature is located; Extracting all feature correlations corresponding to the indoor air quality related feature in the data relationship graph; Determining the air quality core index of the indoor air quality related feature through the air quality feature weight of the dimension where the indoor air quality related feature is located and the corresponding all feature correlations, and then obtaining the air quality core index of each indoor air quality related feature in the indoor air quality feature matrix; Screening all air quality core features according to the air quality core index of each indoor air quality related feature; Constructing an air quality core feature sequence according to all air quality core features.
9. The prediction system for indoor air quality index (AQI) based on multi-source data fusion of claim 1, wherein, The prediction module predicts the indoor air quality index AQI according to the air quality core feature sequence by inputting the air quality core feature sequence into a prediction model, and then obtaining the predicted value of the indoor air quality index AQI.
10. The prediction system for indoor air quality index (AQI) based on multi-source data fusion of claim 9, wherein, The prediction model is a neural network model.
Citation Information
Patent Citations
Tobacco plant number prediction method based on deep learning
CN119918023A
Multi-source data real-time fusion processing method and system of mobile intelligent device
CN120705826A
Method for cloud-based indoor air quality management and system therefor
KR102742952B1
Small-scale air quality index prediction method and system for city
WO2018214060A1