Multi-factor coupled subway ventilation air conditioning load and comfort degree prediction method and system
By employing a multi-factor coupled method for predicting the load and comfort of subway ventilation and air conditioning systems, and utilizing multi-factor correlation analysis and a dual-output neural network model, the problem of insufficient prediction accuracy and real-time performance in existing subway ventilation and air conditioning systems is solved. This method achieves efficient load and comfort prediction and supports intelligent scheduling and energy-saving operation.
Patent Information
- Application Number
- CN202511240685.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-09-02
AI Technical Summary
The existing operation strategies of subway ventilation and air conditioning systems rely on experience-based scheduling or expensive numerical simulation platforms, which are difficult to deploy efficiently across the entire network. Furthermore, the lack of systematic modeling of the relationship between multi-source data and control strategies makes it difficult to meet the actual operation and maintenance needs in terms of the accuracy and real-time performance of load and comfort predictions.
A multi-factor coupled method for predicting the load and comfort of subway ventilation and air conditioning is adopted. By acquiring multi-source historical data, preprocessing and feature derivation are performed, key features are extracted using multi-factor correlation analysis, and a dual-output neural network model is trained to achieve coordinated prediction of load and comfort.
It significantly improves the utilization efficiency and interpretability of multi-source heterogeneous data in air conditioning system modeling, enhances prediction accuracy and model interpretability, and provides direct and usable decision support for the intelligent scheduling and energy-saving operation of subway ventilation and air conditioning systems.
Smart Images

Figure CN120763495B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ventilation and air conditioning technology for rail transit, and in particular to a multi-factor coupled method and system for predicting the load and comfort of subway ventilation and air conditioning. Background Technology
[0002] Currently, the operation strategies of subway ventilation and air conditioning systems largely rely on experience-based scheduling rules or expensive numerical simulation platforms. Maintenance personnel set unit start-up and shutdown logic and chilled water supply temperatures based on historical experience. While simulation platforms can recreate physical processes, they require identifying numerous model configuration parameters and consuming significant computational resources, making efficient deployment across the entire network difficult. Meanwhile, although some studies have attempted to construct multivariate statistical models, these often focus on analyzing passenger flow, ambient temperature and humidity, or chilled and hot water systems individually. They lack systematic modeling of the coupling relationship between multi-source data and control strategies, resulting in insufficient accuracy and real-time performance in load and comfort predictions when dealing with sudden passenger flow peaks or strategy adjustments, failing to meet actual operation and maintenance needs. Summary of the Invention
[0003] To address the aforementioned technical problems, this invention provides a multi-factor coupled method and system for predicting subway ventilation and air conditioning load and comfort.
[0004] According to one aspect of the present invention, a multi-factor coupled method for predicting the load and comfort of subway ventilation and air conditioning is proposed, the method comprising:
[0005] Historical data from multiple sampling times of the ventilation and air conditioning system in subway stations are obtained. The historical data includes the number of people entering and exiting the station, temperature data, humidity data, water flow data, and equipment operation-related data.
[0006] The historical data is preprocessed;
[0007] Derived features are constructed based on preprocessed historical data, including total passenger flow, total passenger flow change rate, temperature deviation, and temperature deviation change rate.
[0008] Based on the multi-factor association analysis method, factor extraction and dimensionality reduction are performed on the multi-source feature matrix containing preprocessed historical data and derived features to screen out key features;
[0009] The key features are input into a prediction model based on a dual-output neural network for training;
[0010] Calculate the corresponding key features based on real-time data of the ventilation and air conditioning system of subway stations.
[0011] By inputting the key features corresponding to real-time data into the trained prediction model, the load and comfort prediction results of the ventilation and air conditioning system are obtained.
[0012] Furthermore, the preprocessing includes outlier detection and removal, and missing value filling; the total passenger flow in the derived features is the sum of the number of people entering and exiting the station at the same time; the temperature deviation includes the station hall temperature deviation and the platform temperature deviation, the station hall temperature deviation is the difference between the measured temperature of the station hall and the preset temperature threshold, and the platform temperature deviation is the difference between the measured temperature of the platform and the preset temperature threshold.
[0013] Furthermore, it also includes: after constructing the derived features, smoothing and standardizing the preprocessed historical data and derived features, specifically including: periodically encoding the sampling time corresponding to the historical data or derived features; distinguishing between weekdays and non-weekdays for the dates corresponding to the sampling time; and performing logarithmic transformation on the number of people entering and leaving the station in the preprocessed historical data and the total passenger flow in the derived features.
[0014] Furthermore, the multi-factor association analysis method for factor extraction and dimensionality reduction of the multi-source feature matrix containing preprocessed historical data and derived features includes:
[0015] For continuous and discrete variables in historical data or derived features, their correlation coefficients with load and comfort deviation are calculated respectively. The load is calculated based on temperature data and water flow data in the preprocessed historical data, and the comfort deviation is calculated based on temperature deviation in the derived features. If the absolute values of the two correlation coefficients of a certain continuous or discrete variable with load and comfort deviation are both lower than the first preset threshold, the continuous or discrete variable is removed to obtain a pre-screened feature set.
[0016] Factor analysis is performed on continuous variables in the pre-screened feature set, including: standardizing the continuous variables to obtain a standardized matrix; calculating the correlation coefficient matrix of the standardized matrix; performing feature decomposition, determining the number of factors, extracting the initial factor loading matrix, and orthogonally rotating the initial factor loading matrix; calculating the score coefficient matrix based on the correlation coefficient matrix and the orthogonally rotated factor loading matrix, and then combining it with the standardized matrix to calculate the factor score matrix.
[0017] Statistical mapping of discrete variables in the pre-screened feature set is performed based on the factor score matrix;
[0018] Based on the orthogonally rotated factor loading matrix, high-loading continuous variables are determined; based on the statistical mapping results, discrete variables significantly associated with the factors are determined; and the union of the high-loading continuous variables and the discrete variables significantly associated with the factors for each factor is taken as the final key feature.
[0019] Furthermore, the calculation of the correlation coefficients between continuous and discrete variables in historical data or derived features and load and comfort deviations includes: for continuous variables, calculating the Pearson correlation coefficient and Spearman rank correlation coefficient between the continuous variable and load and comfort deviations respectively; for discrete variables, calculating the point-bivariate correlation coefficient, Cramér's V coefficient, or Kendall's τ-b coefficient between the discrete variable and load and comfort deviations respectively.
[0020] The formula for calculating the load is:
[0021] ;
[0022] In the formula, This indicates the specific heat capacity of water; Indicates the density of water; This represents historical data on chilled water flow rate. , These represent the chilled water return temperature and supply temperature in historical data, respectively; t represents time.
[0023] The comfort deviation is the average of the temperature deviation in the concourse and the temperature deviation in the platform.
[0024] Furthermore, the formula for orthogonally rotating the initial factor loading matrix is as follows:
[0025] ;
[0026] In the formula, This represents the factor loading matrix after orthogonal rotation, and its elements are... , indicating the first The variable after rotation Loadings on each factor; Denotes the initial factor loading matrix, whose first factor loading is... Line number Column elements Indicates the first The variable in the first... Initial loadings on each factor; optimal rotation matrix , Let f(x) denote any orthogonal rotation matrix, and satisfy the following conditions: , express identity matrix This represents the total number of rows for a continuous variable. Indicates the number of factor columns to be retained.
[0027] Furthermore, the statistical mapping of discrete variables in the pre-screened feature set based on the factor score matrix includes:
[0028] Normality and homogeneity of variance tests are performed on the factor score vectors in the factor score matrix. If both tests pass, a one-way ANOVA is used to perform a significance test on the association between the factor score vectors to determine the p-value representing whether the differences in the means of each category group are significant, and the first effect size is calculated. Otherwise, the nonparametric Kruskal–Wallis H method is used to perform an association significance test on the factor score vectors to determine the p-value representing whether the differences in medians between the categories are significant, and the second effect size is calculated. ;
[0029] Corresponding to the one-way ANOVA method, if and only if the p-value representing whether the difference in means between the groups is significant is less than the second preset threshold and the first effect size is... When the value is greater than or equal to the third preset threshold, it is determined that the discrete variable and the factor score vector are significantly correlated in engineering.
[0030] Corresponding to the nonparametric Kruskal–Wallis H method, if and only if the p-value representing whether the difference in medians between the groups is significant is less than a second preset threshold and the second effect size... When the value is greater than or equal to the third preset threshold, it is determined that the discrete variable and the factor score vector are significantly correlated in engineering.
[0031] Wherein, the first effect size The calculation formula is:
[0032] ;
[0033] In the formula, Indicates the first The sample at the th Scores on each factor; Indicates the first The mean of each factor over the entire sample; Indicates the first The factor in the th Within a group: mean; n represents the total sample size; k represents the number of categories of the discrete variable; This represents the number of samples in the g-th group;
[0034] Second effect size The calculation formula is:
[0035] ;
[0036] In the formula, This represents the Kruskal-Wallis statistic.
[0037] Furthermore, the step of determining highly loaded continuous variables based on the orthogonally rotated factor loading matrix and determining discrete variables significantly associated with the factors based on the statistical mapping results includes: for each factor, determining the orthogonally rotated factor loading... The absolute values of the factors are sorted from largest to smallest, and the continuous variables corresponding to the 'a' factor loadings with the highest absolute values are extracted as key features; the first effect size... or second effect size Sort by size from largest to smallest and extract the first effect size. or second effect size The b largest discrete variables are taken as key features.
[0038] Furthermore, the loss function of the prediction model based on the dual-output neural network during the training process is:
[0039] ;
[0040] In the formula, and This represents the learnable task-related noise parameters; These represent the prediction errors for load and comfort deviations, respectively.
[0041] According to another aspect of the present invention, a multi-factor coupled subway ventilation and air conditioning load and comfort prediction system is proposed, the system comprising:
[0042] The data acquisition module is configured to acquire historical data from multiple sampling times of the ventilation and air conditioning system of the subway station. The historical data includes the number of people entering and exiting the station, temperature data, humidity data, water flow data, and equipment operation-related data.
[0043] A data preprocessing module, configured to preprocess the historical data;
[0044] The feature extraction module is configured to construct derived features based on preprocessed historical data. The derived features include total passenger flow, total passenger flow change rate, temperature deviation, and temperature deviation change rate. Based on the multi-factor association analysis method, factor extraction and dimensionality reduction are performed on the multi-source feature matrix containing preprocessed historical data and derived features to screen out key features.
[0045] The model training module is configured to input the key features into a prediction model based on a dual-output neural network for training.
[0046] The load and comfort prediction module is configured to calculate the corresponding key features based on real-time data of the ventilation and air conditioning system of the subway station; input the key features corresponding to the real-time data into the trained prediction model to obtain the load and comfort prediction results of the ventilation and air conditioning system.
[0047] The embodiments of the present invention have the following technical effects:
[0048] This invention provides a multi-factor coupled method and system for predicting the load and comfort of subway ventilation and air conditioning systems. It innovatively applies latent factor analysis to the dimensionality reduction of multi-source variables in ventilation and air conditioning systems, extracts key features from both physical and statistical perspectives, and then combines this with a neural network model featuring a shared hidden layer and a dual-output structure to achieve collaborative prediction of system load and comfort deviations. This invention effectively solves the problems of insufficient variable utilization and inadequate model generalization caused by single-task modeling in existing technologies, significantly improving the utilization efficiency and interpretability of multi-source heterogeneous data in air conditioning system modeling. This invention balances prediction accuracy, model interpretability, and engineering scalability, providing directly usable decision support for the intelligent scheduling and energy-saving operation of subway ventilation and air conditioning systems. Attached Figure Description
[0049] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0050] Figure 1 This is a flowchart of a multi-factor coupled method for predicting the load and comfort of subway ventilation and air conditioning, provided in an embodiment of the present invention.
[0051] Figure 2 This is another flowchart of a multi-factor coupled method for predicting the load and comfort of subway ventilation and air conditioning, provided in an embodiment of the present invention.
[0052] Figure 3 This is a schematic diagram of a multi-factor coupled subway ventilation and air conditioning load and comfort prediction system provided in an embodiment of the present invention. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0054] This invention proposes a multi-factor coupled method and system for jointly predicting the load and comfort of a subway station ventilation and air conditioning system. First, it systematically collects and preprocesses multi-dimensional monitoring information such as passenger flow data, fresh air temperature and humidity, chilled water / cooling water supply and return temperature difference, and unit operation strategies. At a unified time granularity, it addresses temporal irregularities and missing values through methods such as forward filling, interpolation, and moving average, and removes and completes outliers. Second, based on the original features, it constructs a series of derived indicators covering total passenger flow and passenger flow change rate, station hall / platform temperature deviation and its change rate, operation strategy indicators, and time labels. It eliminates dimensional and periodic differences through periodic coding and logarithmic transformation, providing high-density, interpretable feature input for subsequent modeling. The feature matrix is described, and latent common factors are extracted using factor analysis. Varimic rotation is then used to give the factors clear physical meaning. By combining correlation and statistical tests, the continuous or discrete variables that have the most significant response to load and comfort are included in the core factor space, achieving efficient dimensionality reduction and semantic mapping of multi-source variables. Then, based on the selected key feature set, a dual-output multi-task regression neural network model is constructed. The shared hidden layer is used to capture the common drivers of load and comfort, and the branch output layers predict the cooling load and station temperature deviation, respectively. An adaptive uncertainty weighted loss function is used to achieve dynamic weight balance between the two tasks.
[0055] This invention proposes a multi-factor coupled method for predicting the load and comfort of subway ventilation and air conditioning systems, such as... Figures 1-2 As shown, the method includes:
[0056] S1. Obtain historical data from multiple sampling times of the ventilation and air conditioning system of the subway station. The historical data includes the number of people entering and exiting the station, temperature data, humidity data, water flow data, and equipment operation-related data.
[0057] S2. Preprocess the historical data; the preprocessing includes outlier detection and removal, and missing value filling.
[0058] S3. Construct derived features based on preprocessed historical data, including total passenger flow, total passenger flow change rate, temperature deviation, and temperature deviation change rate;
[0059] S4. Based on the multi-factor association analysis method, factor extraction and dimensionality reduction are performed on the multi-source feature matrix containing preprocessed historical data and derived features to screen out key features;
[0060] S5. Input the key features into the prediction model based on the dual-output neural network for training;
[0061] S6. Calculate the key features based on real-time data of the subway station ventilation and air conditioning system;
[0062] S7. Input the key features corresponding to the real-time data into the trained prediction model to obtain the load and comfort prediction results of the ventilation and air conditioning system.
[0063] Furthermore, after constructing derived features in S3, the preprocessed historical data and derived features are smoothed and standardized. Specifically, this includes: periodically encoding the sampling time corresponding to the historical data or derived features; distinguishing between weekdays and non-weekdays for the dates corresponding to the sampling time; and performing logarithmic transformation on the number of people entering and leaving the station in the preprocessed historical data and the total passenger flow in the derived features.
[0064] The method begins with S1. In S1, historical data from multiple sampling times of the subway station's ventilation and air conditioning system are acquired. This historical data includes the number of people entering and exiting the station, temperature data, humidity data, water flow data, frequency data, and equipment operation-related data.
[0065] According to an embodiment of the present invention, raw data of the ventilation and air conditioning system of a subway station is collected. The raw data includes sampling time (date, specific timestamp), number of people entering the station, number of people exiting the station, temperature data (fresh air temperature, concourse temperature, platform temperature, chilled water supply and return water temperature, cooling water supply and return water temperature, supply air temperature, return air temperature), humidity data (fresh air humidity), water flow data (chilled water supply and return water flow, cooling water supply and return water flow), and equipment operation-related data (equipment such as chillers, water pumps, fans, cooling towers; related data include water pump operating frequency, fan operating frequency, equipment start / stop status, operating speed, operating strategy mode, etc.).
[0066] The original data may be obtained from different sources, and the time intervals during data collection from different sources may be inconsistent, resulting in different time intervals for the data obtained from each source. In irregularly sampled data, it is necessary to transform data with different time intervals into data with the same time interval (e.g., 15 minutes). Time series data is resampled to transform irregular time interval data into time series data with the same intervals. The following methods can be used: 1) Forward imputation: filling missing values with the values of the previous time point; 2) Linear interpolation: interpolating missing values based on the changing trend between time points to obtain a smooth time series; 3) Mean interpolation: filling missing values based on the average value of adjacent data points.
[0067] Then, S2 is executed, in which the historical data is preprocessed; the preprocessing includes outlier detection and removal, and missing value filling.
[0068] According to an embodiment of the present invention, after data resampling, outlier detection and removal are first performed (e.g., using the quartile rule, Laida criterion, etc.). For missing values resulting from outlier removal (or missing values present in the original data), a moving average method is used for imputation. The moving average method imputs missing values by calculating the average of the data within a rolling time window before and after the missing value. Its effectiveness depends on the choice of window size; smaller windows are sensitive to data changes but easily affected by noise; larger windows help smooth the data but may delay information capture. The window size is determined based on the rate of data change. A cross-validation method simulating missing values is used to evaluate the imputation effect under different window sizes. Specifically, a high-quality continuous variable time series segment is selected as a "reference complete dataset." Missing values are randomly introduced manually into the reference dataset, and the imputation effect (RMSE, CV) of the moving average method with different window sizes is evaluated to determine the optimal window size applicable to the variable. This optimal window size is then applied to imputation of the entire dataset.
[0069] Then, S3 is executed. In S3, derived features are constructed based on the preprocessed historical data. The derived features include total passenger flow, total passenger flow change rate, temperature deviation, and temperature deviation change rate.
[0070] According to an embodiment of the present invention, after data preprocessing is completed, the original features (obtained directly from monitoring) and derived features (calculated based on physical, behavioral or strategic logic) are systematically integrated in the same feature matrix to comprehensively characterize the impact of passenger flow disturbances, environmental changes, water system status and control strategies on load and comfort.
[0071] The original features can be further divided into: 1) Passenger flow, including the number of people entering and exiting the station; 2) Environmental parameters, including fresh air temperature and fresh air humidity; 3) Comfort measurement points, including station hall temperature and platform temperature; 4) Air conditioning water system, including chilled water supply temperature, return water temperature and flow rate, and cooling water supply temperature, return water temperature and flow rate; 5) Control strategy, including unit operating status, water supply temperature setpoint, fan frequency and mode indicator; 6) Time information, including hour, day of the week and holiday indicators.
[0072] Passenger entry and exit behavior is a significant source of heat load within subway stations. Therefore, in addition to the original number of passengers entering and exiting, it is necessary to construct derived features of total passenger flow and passenger flow change rate to capture the magnitude and trend of heat disturbance caused by passenger flow, thereby effectively enhancing the sensitivity and accuracy of load forecasting. Maintaining passenger comfort is the core objective of ventilation and air conditioning systems. Therefore, besides the original temperature measurement points, it is also necessary to calculate the temperature deviation in the station hall, the temperature deviation in the platform, and their change rate to measure the deviation and fluctuation rate between the actual temperature and the design target. Deviation indicators intuitively reflect the system's adjustment effect, while the change rate reveals the lag or overshoot behavior of the temperature response. Incorporating these features helps multi-task models not only predict load but also simultaneously predict comfort trends. Therefore, derived features include total passenger flow, total passenger flow change rate, temperature deviation, and temperature deviation change rate.
[0073] 1) Total Passenger Flow: The load within a subway station is primarily affected by the number of people. Passenger flow influences factors such as body heat dissipation, air pollution concentration, and fresh air demand. As the most direct "internal disturbance" factor, this variable is strongly positively correlated with air conditioning load and is one of the core driving variables. The formula for approximating the direct impact of the overall passenger volume on the heat load using total passenger flow is as follows:
[0074] Out(t)
[0075] In the formula, In(t) represents the number of people entering the station at time t; Out(t) represents the number of people leaving the station at time t.
[0076] 2) Passenger Flow Change Rate: The speed of dynamic change has a significant impact on the system's responsiveness and lag adjustment mechanism. The formula for reflecting the impact of a sudden increase or decrease in passenger flow on the system load using the passenger flow change rate is as follows:
[0077]
[0078] In the formula, This represents the total passenger flow at time t-1; It is a constant to prevent the denominator from being 0, and can be taken as 0.001.
[0079] 3) Temperature deviation includes station hall temperature deviation and platform temperature deviation; the station hall temperature deviation is:
[0080]
[0081] The platform temperature deviation is:
[0082]
[0083] In the formula, This represents the measured temperature in the station hall at time t. This represents the measured temperature at the platform at time t; This indicates the preset temperature threshold (e.g., 26 °C).
[0084] 4) The rate of change of deviation is:
[0085]
[0086] Furthermore, after constructing the derived features, the preprocessed historical data and derived features are smoothed and standardized, specifically including: periodically encoding the sampling time corresponding to the historical data or derived features; distinguishing between weekdays and non-weekdays for the dates corresponding to the sampling time; and performing logarithmic transformation on the number of people entering and leaving the station in the preprocessed historical data and the total passenger flow in the derived features.
[0087] According to embodiments of the present invention, some original features and derived features are smoothed and standardized to improve the model's ability to identify periodicity and magnitude differences.
[0088] 1) Periodic Encoding (Hourly): Sine and cosine functions are used to periodically encode time information, constructing feature variables that reflect time periodicity. These feature variables retain the periodic characteristics of time, making the model more easily able to identify the time dependence of load patterns. The encoding formula is as follows:
[0089]
[0090] In the formula, Hour_sin(t) represents the hour number (0-23) corresponding to the current time; Hour_sin(t) and Hour_cos(t) represent the sine and cosine encoded values of the hour.
[0091] 2) Weekday / Non-Workday Classification: Decomposing date information into a binary code for weekday / non-working day helps the model clearly distinguish the periodic differences in passenger flow and workload between weekdays and non-working days. Compared to representing it as a single value from 1 to 7, this binarization eliminates unnecessary misleading information based on weekday order. Binary features simplify the input dimensions and improve the model's ability to identify weekdays and non-working days. This approach ensures information integrity while also improving the efficiency and interpretability of regression predictions. The formula is as follows:
[0092]
[0093] In the formula, A Boolean variable indicating whether it is a working day.
[0094] 3) Logarithmic Transformation: A Logarithmic transformation is performed on the raw passenger flow data (number of people entering and exiting the station) and the total passenger flow in the derived features from the ventilation and air conditioning data. This prevents the significant difference in magnitude between passenger flow data and other data from affecting the analysis results, thereby improving the model's ability to learn the relationships between different features. Before performing the logarithmic transformation on the passenger flow data, zero values are smoothed (e.g., by adding 1), such as the number of people entering the station. .
[0095] The above steps preserve fundamental features such as passenger flow, environment, water system, and control strategies. They also construct derived indicators like passenger flow intensity, cooling capacity, temperature deviation, and their rate of change based on physical mechanisms and passenger behavior. Furthermore, periodic encoding and logarithmic transformation smooth out temporal and order-of-magnitude differences. This allows the model to intuitively capture the coupling relationships between multi-source data while avoiding interference from categories and extreme values in the learning process. These steps take the original fields as input and output a complete matrix containing original, derived, and transformed features, providing comprehensive and efficient input data for subsequent factor selection and regression prediction.
[0096] Then, S4 is executed. In S4, factor extraction and dimensionality reduction are performed on the multi-source feature matrix containing preprocessed historical data and derived features based on the multi-factor association analysis method to screen out key features.
[0097] According to an embodiment of the present invention, by systematically reducing the dimensionality and extracting factors from the multi-source feature matrix constructed in the above steps, redundant features are eliminated, multicollinearity is eliminated, and a small number of key features that can explain the load and comfort deviations to the greatest extent are selected, providing efficient input for subsequent regression models.
[0098] First, execute S41, pre-screening based on target correlation, including: calculating the correlation coefficients between continuous and discrete variables in historical data or derived features and load and comfort deviation, respectively. The load is calculated based on temperature and water flow data in the pre-processed historical data, and the comfort deviation is calculated based on temperature deviation in the derived features. If the absolute values of the two correlation coefficients of a certain continuous or discrete variable with load and comfort deviation are both lower than a first preset threshold, then the continuous or discrete variable is removed to obtain a pre-screened feature set.
[0099] Specifically, this step aims to eliminate noisy features that are not significantly related to load and comfort deviations. The original variable set contains the original observed features and the derived features constructed; subsequent correlation analysis, standardization, and factor mapping are all carried out within this extended feature space.
[0100] For each continuous variable (including the number of people entering and exiting the station after feature transformation, the time information after periodic transformation, fresh air temperature, fresh air humidity, station hall temperature, platform temperature, chilled water supply and return water temperature and flow rate, cooling water supply and return water temperature and flow rate, water pump operating frequency, fan frequency, supply air temperature, return air temperature, total passenger flow, and passenger flow change rate), calculate its Pearson correlation coefficient and Spearman's rank correlation coefficient with load and comfort deviation respectively;
[0101] For discrete variables, for binary variables (including equipment start / stop status and binary encoded date information), calculate their point-binary correlation coefficients with load and comfort deviations respectively; for multi-category nominal variables (including strategy patterns in equipment operation-related data), calculate their Cramér's V coefficients with load and comfort deviations respectively; for multi-category ordinal variables (including equipment operation levels in equipment operation-related data, such as 'fan level: low / medium / high'), calculate their Kendall's τ-b coefficients with load and comfort deviations respectively.
[0102] Among them, Pearson correlation coefficient is used to measure the degree of linear correlation between two continuous variables; Spearman rank correlation coefficient is more suitable for continuous variables that do not meet the normal distribution or have a non-linear relationship; point-bivariate correlation coefficient is used for correlation analysis between binary variables and continuous variables; Cramér's V coefficient is used to test the correlation between multi-category nominal variables and target variables; Kendall's τ-b coefficient is used to assess the correlation between multi-category ordinal variables and target variables.
[0103] It should be noted that the comfort deviation to be predicted in this invention refers to the average of the temperature deviation in the concourse and the temperature deviation in the platform. Since the load value in a subway ventilation and air conditioning system is difficult to determine directly, this invention uses the instantaneous cooling capacity on the chilled water side to characterize the load value as follows:
[0104]
[0105] In the formula, t represents time; This indicates the specific heat capacity of water; Indicates the density of water; This indicates the chilled water flow rate, measured in m³ / s. , These represent the chilled water return / supply water temperatures, respectively.
[0106] If the absolute value of the correlation coefficient between a certain feature and the deviation from comfort in the correlation test is lower than a first preset threshold, such as 0.3, then the feature is considered to have no significant association with any objective and is removed from the candidate set, ultimately forming the pre-screening set C. This initial screening step, through strict correlation threshold control, ensures that subsequent analysis focuses on variables closely related to system performance, thereby effectively reducing the input size and noise interference.
[0107] Furthermore, all continuous variables in the pre-screened set C are first mean-centered and then divided by the standard deviation to transform them into a standard normal form with a mean of 0 and a variance of 1. This eliminates differences in the dimensions, amplitude, and scale of the continuous variables, and outputs a standardized matrix. Discrete variables retain their original encoding and are not standardized.
[0108] Then, S42 is executed to perform factor analysis on the continuous variables in the pre-screened feature set, including: standardizing the continuous variables to obtain a standardized matrix; calculating the correlation coefficient matrix of the standardized matrix; performing feature decomposition, determining the number of factors, extracting the initial factor loading matrix, and orthogonally rotating the initial factor loading matrix; calculating the score coefficient matrix based on the correlation coefficient matrix and the orthogonally rotated factor loading matrix, and then combining it with the standardized matrix to calculate the factor score matrix; S43 is to perform statistical mapping on the discrete variables in the pre-screened feature set based on the factor score matrix. The statistical mapping refers to the process of mapping discrete variables to the latent factor space using statistical testing methods, including: firstly, performing normality tests and homogeneity of variance tests on the factor score vectors respectively; if both tests pass, then using one-way ANOVA to perform association significance tests to determine the p-value characterizing whether the differences in the means of each category group are significant, and calculating the first effect size. Otherwise, the nonparametric Kruskal–Wallis H method is used to perform an association significance test to determine the p-value representing whether the differences in medians between the groups are significant, and the second effect size is calculated. ; Corresponding to the use of one-way ANOVA, if and only if the p-value representing whether the difference in the means of each category group is significant is less than the second preset threshold and the first effect size When the value is greater than or equal to the third preset threshold, the discrete variable is determined to be significantly associated with the factor score vector in engineering terms; corresponding to the use of the nonparametric Kruskal–Wallis H method, if and only if the p-value representing whether the difference in medians among the class groups is significant is less than the second preset threshold and the second effect size is... When the value is greater than or equal to the third preset threshold, it is determined that the discrete variable and the factor score vector are significantly correlated in engineering; S44, based on the factor loading matrix after orthogonal rotation, determine the high-load continuous variable; based on the statistical mapping results, determine the discrete variable that is significantly correlated with the factor; take the union of the high-load continuous variable and the discrete variable that is significantly correlated with the factor for each factor as the final key feature.
[0109] S42~S44 are used to identify potential common factors driving the two objectives in the data and map all variables (including discrete variables) to the factor space to construct the system driving structure. The specific steps of S42 are as follows.
[0110] For the input data, i.e., the standardized matrix of the pre-selected set C Where n represents the sample size (e.g., 30 days × 96 15-minute points = 2880), and p represents the number of continuous variables after pre-screening. First, calculate the correlation coefficient matrix (the correlation coefficients between standardized continuous variables) as follows:
[0111]
[0112] In the formula, , representing variables and The Pearson correlation coefficient.
[0113] Then, feature decomposition is performed, as follows:
[0114]
[0115] In the formula, , representing eigenvalues, ; This represents the eigenvector matrix.
[0116] Then, the number of factors is determined as follows: based on - Kaiser criterion preserves eigenvalues Factors;
[0117] Select the smallest The value makes the cumulative variance proportion If the Kaiser criterion conflicts with the principle of cumulative variance ratio, the number of factors m is gradually increased until the cumulative variance ratio is ≥70%.
[0118] Then, the initial factor loading matrix L is extracted as follows:
[0119]
[0120] In the formula, This represents a submatrix (dimension p×m) composed of the first m columns of eigenvectors. This represents a diagonal submatrix consisting of the first m eigenvalues. The factor loading matrix is used to quantify the linear relationship between "each latent factor" and "each observed variable".
[0121] Then, the initial factor loading matrix is orthogonally rotated to obtain the rotated factor loading matrix. :
[0122]
[0123] In the formula, , To represent an arbitrary orthogonal rotation matrix, it must satisfy the following condition: , Denotes the optimal rotation matrix, which is the matrix that satisfies the orthogonality condition. In the matrix, the matrix that maximizes the Varimax objective function; the initial factor loading matrix. The Middle Line number Column elements Indicates the first The variable in the first... Initial loadings on each factor; The elements are , indicating the first The variable after rotation Loadings on the factor; Indicates the total number of consecutive variables (number of rows); Indicates the number of factors retained (number of columns); express The identity matrix is used to guarantee The orthogonality. Two terms in the above optimal rotation matrix: It is the first The second moment of the square of the column load, It is the square of its mean.
[0124] Varimax rotation adjusts the factor loading matrix while maintaining the independence (orthogonality) of the factors, resulting in a sparser structure for each column (i.e., each factor): only a few variables have high loadings, while others are close to zero. This provides higher explanatory power, making it easier to explain "what a certain factor represents." For example, before rotation, if factor 1 and factor 2 have roughly equal loadings on all variables, making it unclear which variable is more important, after Varimax rotation, factor 1's loadings on fresh air temperature and humidity become very high, while others are very low ⇒ this can be interpreted as environmental factors; factor 2's loadings on fan frequency and water supply setpoint temperature increase ⇒ this can be interpreted as operational strategy factors.
[0125] In factor analysis, each latent factor has an influence on several original variables; this influence is the factor loading of that variable on that factor.
[0126] Then, the score coefficient matrix is calculated based on the correlation coefficient matrix and the rotated factor loading matrix as follows:
[0127]
[0128] Then the factor score matrix for:
[0129]
[0130] In the formula, This represents the score of the i-th sample on the j-th factor, so each sample has a set of lengths... The score.
[0131] The specific steps of S43 are as follows: Determine whether the discrete variable is associated with a certain factor. Discrete variable These are the labels used to group samples (e.g., "day of the week," "equipment status"). If a discrete variable truly influences a factor, then the factor score distributions under different categories (groups) will differ significantly. To determine whether a discrete variable is associated with a latent factor, the following steps are performed for each discrete variable and each latent factor:
[0132] Assume the discrete variable has k categories, and the factor score vector is... Factor score vector for the entire sample Perform a normality test: if The Shapiro-Wilk test yielded the following results. ,like The Kolmogorov-Smirnov test yielded the following results. ;judge Whether the overall distribution is approximately normal serves as a prerequisite for subsequent selection of parametric tests (ANOVA);
[0133] Factor score vector for the entire sample Classified by discrete variables After grouping, a homogeneity of variance test was performed:
[0134] The Levene test yielded the following results. Determining whether the variances of each category group are equal is one of the prerequisites of ANOVA.
[0135] if and This describes the factor score vector of the entire sample. If the distribution is approximately normal and the variances are homogeneous, then use one-way ANOVA to test the significance of the association; otherwise, use the nonparametric Kruskal–Wallis H test to test the significance of the association.
[0136] After performing the selected association significance test, the association test results are obtained. Value: ANOVA gives This represents whether the "differences between group means" are significant; Kruskal–Wallis provides... This indicates whether the "difference in median values between groups" is significant. Only if the selected... (here represent or When considering discrete variables and factor score vectors, it is assumed that... They are statistically significantly correlated.
[0137] However, even Even with a large sample size, small differences can become significant, so the effect size is calculated simultaneously. For the ANOVA test, the first effect size η² is calculated as follows:
[0138]
[0139] In the formula, Indicates the first The sample at the th Scores on each factor; Indicates the first The mean of each factor over the entire sample; Indicates the first The mean of each factor in the g-th group; n represents the total sample size; k represents the number of categories of the discrete variable; Let g represent the number of samples in the g-th group. If and only if... and If the value is greater than or equal to the second preset threshold (for example, 0.01), then the discrete variable is considered to be related to the factor score vector. It has a significant connection in engineering.
[0140] For the Kruskal-Wallis H test, the first effect size ε² is calculated as follows:
[0141]
[0142] in, This is the Kruskal-Wallis statistic. It is valid if and only if... and When the value is ≥ the third preset threshold (the third preset threshold is, for example, 0.01), the discrete variable is considered to be related to the factor score vector. It has a significant connection in engineering.
[0143] Continuous variables explain the factor structure through rotated high factor loadings. High factor loadings are continuous variables that have significant factor loadings on a particular factor; typically, a factor loading whose absolute value is greater than a certain threshold (such as 0.5 or 0.6) is considered a high factor loading. Discrete variables participate in factor stratification analysis through statistical mapping, enabling the integrated analysis of mixed variables within a unified latent factor framework.
[0144] The specific steps of S44 are as follows: For each factor, select the most representative and physically interpretable continuous or discrete variables to construct a key feature set for prediction and pattern mining. For each factor, orthogonally rotate the factor loadings... Sort the absolute values of factors from largest to smallest, and extract the continuous variables corresponding to the 'a' factor loadings with the highest absolute values as key features (where 'a' is, for example, 2); for the first effect size... or second effect size Sort by size from largest to smallest and extract the first effect size. or second effect size The largest b discrete variables are selected as key features; the union of the selected variables corresponding to all factors is used to construct the final key feature set. .
[0145] As an example, the final key feature set Includes: 1) Continuous variables including fresh air temperature, fresh air humidity, chilled water supply and return water temperature and flow rate, and station hall / platform temperature; 2) Discrete variables including equipment operating status (Boolean type) and equipment control mode (categorical variable), etc.
[0146] This step systematically filters a subset of variables that are significantly related to load and comfort goals from the original feature set, serving as effective input for subsequent modeling and analysis. Irrelevant features are initially eliminated through correlation analysis; at the latent factor level, factor analysis reveals the implicit driving dimensions of system operation, and combined with load assessment and statistical tests, continuous features that significantly contribute to the factors and correlated discrete variables are incorporated into the final key feature set, thereby optimizing the input dimensions and enhancing semantic interpretability.
[0147] Then, S5 is executed, in which the key features are input into a prediction model based on a dual-output neural network for training.
[0148] According to an embodiment of the present invention, a multi-task regression model is constructed to simultaneously predict the cooling load and comfort deviation of the subway ventilation and air conditioning system. By sharing input features and jointly optimizing the two output objectives, the collaborative perception of the system state and the dual-objective prediction capability are realized, thereby improving the global characterization capability of subway operation conditions and the model generalization performance.
[0149] Before constructing the regression model, all discrete variables are preprocessed by encoding to transform them into numerically quantifiable input terms, thereby unifying the input space. Continuous variables have already been standardized in the above steps, so the input data can be directly concatenated to form the training sample matrix. Where n is the number of samples and d is the dimension of the core variable; the output is set to a dual-task structure, where task 1 is to predict and calculate the cooling load. Task 2 is to predict comfort deviations. The dual-output neural network structure is used as the basic framework. The dual-output neural network differs from the traditional single-task network in both structure and training objectives: it sets a shared hidden layer after the input layer to uniformly couple and encode multi-source heterogeneous features, and then branches at the end to output two predictions: cooling load and comfort deviation.
[0150] The dual-output neural network structure is as follows. Input data, after passing through multiple shared hidden layers, is mapped to an intermediate representation vector ℎ, used to capture common features affecting both types of targets. This part has the following functions: fusing multi-source variable information: the shared hidden layers couple and encode input information from different physical systems (such as water systems and passenger flow systems); mining the correlation between multiple targets: cross-variables that affect both load and comfort (such as passenger flow and platform temperature and humidity) can establish a unified representation in the shared space; improving modeling generalization ability: through structural regularization, redundant parameter learning between tasks is reduced, improving model stability. Compared to ordinary networks that only minimize a single loss, this invention introduces an adaptive weighted loss function based on task uncertainty in the branch output layer. Branch 1 predicts the cooling load Q, and the output... Branch 2 predicts the comfort deviation D, and outputs... During training, a joint loss function is constructed to perform end-to-end optimization of the two objectives. This loss function has the following form:
[0151]
[0152] in, and For learnable task-related noise parameters, the weights are dynamically adjusted through an adaptive weighting strategy based on task uncertainty. This strategy automatically learns the optimal weight allocation according to the inherent uncertainty characteristics of each task, without the need for manual specification of hyperparameters, which can effectively improve the overall modeling accuracy and stability. These are the prediction errors for cooling load and comfort, respectively (the loss function type can be selected from typical indices such as MSE and MAE).
[0153] This adaptive loss function has the following characteristics: each training step dynamically adjusts according to changes in task uncertainty. and The value of the task weight is updated, and the parameters of the shared hidden layer and branch layer are updated accordingly. The task weight is negatively correlated with the task noise level; the higher the noise (the higher the uncertainty), the lower the task weight.
[0154] The model optimizes noise parameters By adjusting network parameters, an adaptive balance can be achieved for the collaboration ratio among multiple tasks.
[0155] The aforementioned multi-task modeling structure, based on a shared hidden layer and branched output layer, is used to simultaneously predict cooling load Q and comfort deviation D. It should be noted that, in addition to the aforementioned neural network structure, it is also compatible with other multi-output modeling strategies, including but not limited to: multi-output regression trees; multi-objective gradient boosting models; and other machine learning models with shared input and task-specific output capabilities. The aforementioned alternative structures can replace or extend neural network models without changing the multi-task learning objectives and input / output variable configurations, making them suitable for different computing power conditions or scenario requirements, and possessing good engineering adaptability and scalability.
[0156] Then, steps S6 and S7 are executed. In step S6, the corresponding key features are calculated based on the real-time data of the subway station's ventilation and air conditioning system. In step S7, the key features corresponding to the real-time data are input into the trained prediction model to obtain the load and comfort prediction results of the ventilation and air conditioning system.
[0157] According to an embodiment of the present invention, in practical applications, the original data corresponding to the input key feature set of the above-trained prediction model is collected, and the corresponding derived features are calculated to construct the key features corresponding to the real-time data; the key features are input into the above-trained prediction model based on a dual-output neural network to obtain the load and comfort prediction results of the ventilation and air conditioning system.
[0158] In summary, this invention uses factor analysis to map different variables into interpretable physical latent factors, and proposes a two-stage screening mechanism of statistical testing and factor loading to achieve unified integrated modeling of cross-type variables (continuous / discrete). Based on this, a multi-task neural network structure is constructed, which automatically learns the shared and differential information flow between load and comfort prediction tasks using shared hidden layers. By dynamically adjusting the weights of the multi-task loss function through learnable noise parameters, end-to-end adaptive optimization is achieved without the need for manual specification of task priorities, thereby improving prediction accuracy and stability.
[0159] Another embodiment of the present invention proposes a multi-factor coupled subway ventilation and air conditioning load and comfort prediction system, such as... Figure 3 As shown, the system includes:
[0160] The data acquisition module 310 is configured to acquire historical data of the ventilation and air conditioning system of the subway station at multiple sampling times. The historical data includes the number of people entering and leaving the station, temperature or humidity data, water flow data, and equipment operation-related data.
[0161] Data preprocessing module 320 is configured to preprocess the historical data;
[0162] The feature extraction module 330 is configured to construct derived features based on preprocessed historical data. The derived features include total passenger flow, total passenger flow change rate, temperature deviation, and temperature deviation change rate. Based on the multi-factor association analysis method, factor extraction and dimensionality reduction are performed on the multi-source feature matrix containing preprocessed historical data and derived features to screen out key features.
[0163] The model training module 340 is configured to input the key features into a prediction model based on a dual-output neural network for training.
[0164] The load and comfort prediction module 350 is configured to calculate the corresponding key features based on real-time data of the ventilation and air conditioning system of the subway station; input the key features corresponding to the real-time data into the trained prediction model to obtain the load and comfort prediction results of the ventilation and air conditioning system.
[0165] It should be noted that the function of the multi-factor coupled subway ventilation and air conditioning load and comfort prediction system described in this embodiment can be explained by the aforementioned multi-factor coupled subway ventilation and air conditioning load and comfort prediction method. Therefore, for the parts not described in detail in the system embodiment, please refer to the above method embodiment.
[0166] It should also be noted that the terminology used in this invention is for describing specific embodiments only and is not intended to limit the scope of this application. As shown in this specification, unless the context clearly indicates otherwise, words such as "a," "an," "an," and / or "the" do not specifically refer to the singular and may include the plural. The terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element.
[0167] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting metro ventilation and air conditioning load and comfort with multi-factor coupling, characterized in that, The method comprises the following steps: obtaining historical data of a subway station ventilation and air conditioning system at multiple sampling time points, wherein the historical data comprises the number of passengers entering and leaving the station, temperature data, humidity data, water flow data, and equipment operation related data; preprocessing the historical data; constructing derived features based on the preprocessed historical data, wherein the derived features comprise total passenger flow, total passenger flow change rate, temperature deviation, and temperature deviation change rate; extracting and reducing dimensions of a multi-source feature matrix comprising the preprocessed historical data and the derived features based on a multi-factor correlation analysis method to screen out key features; including: for continuous variables and discrete variables in the historical data or the derived features, calculating their correlation coefficients with load and comfort deviation, wherein the load is calculated according to the temperature data and the water flow data in the preprocessed historical data, and the comfort deviation is calculated according to the temperature deviation in the derived features; if the absolute values of the two correlation coefficients of a continuous variable or a discrete variable with the load and the comfort deviation are both lower than a first preset threshold, the continuous variable or the discrete variable is removed, and a pre-screened feature set is obtained; performing factor analysis on the continuous variables in the pre-screened feature set, including: standardizing the continuous variables to obtain a standardized matrix; calculating a correlation coefficient matrix of the standardized matrix; and performing feature decomposition, determining the number of factors, extracting an initial factor loading matrix, and performing orthogonal rotation on the initial factor loading matrix; calculating a score coefficient matrix according to the correlation coefficient matrix and the factor loading matrix after orthogonal rotation, and then calculating a factor score matrix in combination with the standardized matrix; statistically mapping the discrete variables in the pre-screened feature set based on the factor score matrix; including: performing normality test and homogeneity of variance test on the factor score vectors in the factor score matrix respectively, if both tests pass, performing correlation significance test on the factor score vectors using single factor variance analysis method to determine the p value representing whether the difference of mean values of each category group is significant, and calculating a first effect size ; otherwise, performing correlation significance test on the factor score vectors using non-parametric Kruskal-Wallis H method to determine the p value representing whether the difference of median values of each category group is significant, and calculating a second effect size ; corresponding to the one-way analysis of variance method, the p-value representing whether the difference between the means of the groups of categories is significant is less than a second preset threshold value and the first effect size is greater than or equal to a third preset threshold value greater than or equal to a third preset threshold value, it is determined that the discrete variable is significantly related to the factor score vector in engineering. corresponding to the non-parametric Kruskal-Wallis H method, the p-value characterizing whether the difference between the medians of the groups of categories is significant is less than a second preset threshold and the second effect size when greater than or equal to a third preset threshold, it is determined that the discrete variable is significantly associated with the factor score vector in engineering determining high-load continuous variables based on the factor loading matrix after orthogonal rotation; determining discrete variables that are significantly associated with factors based on the statistical mapping result; and taking the union of the high-load continuous variables and the discrete variables that are significantly associated with the factors of each factor as the final key features; inputting the key features into a prediction model based on a double-output neural network for training; calculating corresponding key features based on real-time data of the subway station ventilation and air conditioning system; inputting the key features corresponding to the real-time data into the trained prediction model to obtain load and comfort prediction results of the ventilation and air conditioning system.
2. The multi-factor coupled subway ventilation and air conditioning load and comfort prediction method according to claim 1, characterized in that, The preprocessing comprises detection and removal of outliers and filling of missing values; the total passenger flow in the derived features is the sum of the number of passengers entering and leaving the station at the same time; the temperature deviation comprises station hall temperature deviation and platform temperature deviation, wherein the station hall temperature deviation is the difference between the measured station hall temperature and a preset temperature threshold, and the platform temperature deviation is the difference between the measured platform temperature and a preset temperature threshold. 3.The method of claim 2, wherein, Further comprising: after constructing the derived features, smoothing and standardizing the preprocessed historical data and the derived features, specifically including: periodically encoding the sampling time points corresponding to the historical data or the derived features; distinguishing and encoding the dates corresponding to the sampling time points as weekdays or non-weekdays; and performing logarithmic transformation on the number of passengers entering and leaving the station in the preprocessed historical data and the total passenger flow in the derived features.
4. The multi-factor coupled subway ventilation and air conditioning load and comfort prediction method according to claim 1, characterized in that, The correlation coefficients of the continuous variables and the discrete variables in the historical data or the derived features with the load and the comfort deviation are calculated respectively, including: for the continuous variables, Pearson correlation coefficients and Spearman rank correlation coefficients of the continuous variables with the load and the comfort deviation are calculated respectively; for the discrete variables, point biserial correlation coefficients or Cramér's V coefficients or Kendall's τ-b coefficients of the discrete variables with the load and the comfort deviation are calculated respectively; The calculation formula of the load is: ; wherein, Cp represents the specific heat capacity of water; ρ represents the density of water; Qf represents the chilled water flow rate data in the historical data; Tf and Ts represent the chilled water return temperature and the supply water temperature in the historical data, respectively; and t represents time. The comfort deviation is the average of the station hall temperature deviation and the platform temperature deviation.
5. The multi-factor coupled subway ventilation and air conditioning load and comfort prediction method according to claim 1, characterized in that, The formula for orthogonal rotation of the initial factor loading matrix is: ; wherein denotes the orthogonal rotated factor loading matrix, whose element denotes the load of the th variable on the th factor after rotation; denotes the initial factor loading matrix, whose element denotes the initial load of the th variable on the th factor; the optimal rotation matrix , denotes an arbitrary orthogonal rotation matrix, and satisfies , denotes an arbitrary orthogonal rotation matrix, and satisfies , denotes the identity matrix, denotes the total number of rows of continuous variables; denotes the number of reserved factor columns.
6. The multi-factor coupled subway ventilation and air conditioning load and comfort prediction method according to claim 1, wherein, the first effect quantity The formula for calculating is: ; In the formula, Indicates the first The sample at the th Scores on each factor; Indicates the first The mean of each factor over the entire sample; Indicates the first The factor in the th Within a group: mean; n represents the total sample size; k represents the number of categories of the discrete variable; This represents the number of samples in the g-th group; The second effect quantity The calculation formula is: ; In the formula, indicates the Kruskal-Wallis statistic.
7. The multi-factor coupled subway ventilation and air conditioning load and comfort prediction method according to claim 6, characterized in that, The high-load continuous variables are determined based on the factor loading matrix after the orthogonal rotation; The discrete variables significantly associated with the factors are determined based on the statistical mapping result, including: For each factor, the orthogonally rotated factor loadings The absolute values of the factors are sorted from largest to smallest, and the continuous variables corresponding to the 'a' factor loadings with the highest absolute values are extracted as key features; the first effect size... or second effect size Sort by size from largest to smallest and extract the first effect size. or second effect size The b largest discrete variables are taken as key features.
8. The multi-factor coupled subway ventilation and air conditioning load and comfort prediction method according to claim 1, characterized in that, The loss function of the prediction model based on the double-output neural network in the training process is: ; wherein and denote the learnable task-dependent noise parameters; denote the prediction errors of the load and comfort bias, respectively.
9. A multi-factor coupled subway ventilation and air conditioning load and comfort prediction system, characterized in that, including: The data acquisition module is configured to acquire historical data of a subway station ventilation and air conditioning system at multiple sampling times, including the number of people entering and leaving the station, temperature data, humidity data, water flow data, and equipment operation related data; The data preprocessing module is configured to preprocess the historical data; The feature extraction module is configured to construct derived features based on the preprocessed historical data, including total passenger flow, total passenger flow change rate, temperature deviation, and temperature deviation change rate; The multi-factor correlation analysis method is used to extract and reduce the dimension of the multi-source feature matrix containing the preprocessed historical data and the derived features to screen out key features; including: The correlation coefficients of the continuous variables and the discrete variables in the historical data or the derived features with the load and the comfort deviation are calculated respectively, including: for the continuous variables, Pearson correlation coefficients and Spearman rank correlation coefficients of the continuous variables with the load and the comfort deviation are calculated respectively; for the discrete variables, point biserial correlation coefficients or Cramér's V coefficients or Kendall's τ-b coefficients of the discrete variables with the load and the comfort deviation are calculated respectively; if the absolute values of the two correlation coefficients of a continuous variable or a discrete variable with the load and the comfort deviation are both lower than a first preset threshold, the continuous variable or the discrete variable is removed, and a pre-screened feature set is obtained; The factor analysis of the continuous variables in the pre-screened feature set includes: standardizing the continuous variables to obtain a standardized matrix; calculating the correlation coefficient matrix of the standardized matrix; performing feature decomposition, determining the number of factors, extracting the initial factor loading matrix, and performing orthogonal rotation on the initial factor loading matrix; calculating the score coefficient matrix according to the correlation coefficient matrix and the factor loading matrix after the orthogonal rotation, and then calculating the factor score matrix combined with the standardized matrix; The statistical mapping of the discrete variables in the pre-screened feature set is based on the factor score matrix; including: performing normality test and homogeneity of variance test on the factor score vectors in the factor score matrix respectively, if both tests pass, performing correlation significance test on the factor score vectors using single factor variance analysis method to determine the p value representing whether the difference of mean values of each category group is significant, and calculating a first effect size ; otherwise, performing correlation significance test on the factor score vectors using non-parametric Kruskal-Wallis H method to determine the p value representing whether the difference of median values of each category group is significant, and calculating a second effect size ; corresponding to the one-way analysis of variance method, the p-value representing whether the difference between the means of the groups of categories is significant is less than a second preset threshold value and the first effect size is greater than or equal to a third preset threshold value when the p-value representing whether the difference between the means of the groups of categories is significant is less than a second preset threshold value and the first effect size is greater than or equal to a third preset threshold value, it is determined that the discrete variable is significantly related to the factor score vector in engineering corresponding to the non-parametric Kruskal-Wallis H method, the p-value characterizing whether the difference between the medians of the groups of categories is significant is less than a second preset threshold and the second effect size greater than or equal to a third preset threshold, it is determined that the discrete variable is significantly associated with the factor score vector in engineering The high-load continuous variables are determined based on the factor loading matrix after the orthogonal rotation; the discrete variables significantly associated with the factors are determined based on the statistical mapping result; the union of the high-load continuous variables and the discrete variables significantly associated with the factors of each factor is taken as the final key features; The model training module is configured to input the key features into a prediction model based on a double-output neural network for training. The load and comfort degree prediction module is configured to calculate corresponding key features based on real-time data of the subway station ventilation and air conditioning system; and input the real-time data corresponding key features into the trained prediction model to obtain the load and comfort degree prediction result of the ventilation and air conditioning system.
Citation Information
Patent Citations
Urban rail transit passenger flow prediction method based on multi-factor algorithm
CN120087527A
Energy-saving and carbon-reducing multi-objective optimization method and device for cold source system of high-speed rail station
CN120403042A