Energy consumption prediction method for newly-built subway station
By constructing an initial energy consumption analysis data set, environmental visualization model and a multi-scale geographic weighted regression total passenger flow prediction model, combined with the SVR model and the Hippo algorithm, the accuracy of energy consumption prediction of new subway stations is solved, and accurate prediction of energy consumption and energy conservation and emission reduction support for new stations is achieved.
Patent Information
- Application Number
- CN202510555127.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional methods cannot accurately predict the energy consumption of new subway stations, mainly because the new station lacks historical energy consumption and passenger flow data, and the degradation of equipment performance leads to significant differences in the energy consumption characteristics from existing stations.
By constructing the initial energy consumption analysis data set, using Spearman correlation coefficient to analyze influencing factors, obtain environmental point of interest data, build an environmental visualization model, reclassify target points of interest, establish a multi-scale geographic weighted regression in and outbound total passenger flow prediction model, and train an SVR model, and combine the Hippo algorithm to optimize the kernel coefficient and penalty coefficient to achieve energy consumption prediction.
Accurately predicting the energy consumption of newly built stations, providing data support for energy conservation and emission reduction, and improving the accuracy and reliability of energy consumption prediction.
Smart Images

Figure CN120494165A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of urban rail transit energy consumption prediction, and in particular relates to an energy consumption prediction method for a newly built subway station. Background Art
[0002] As a vital component of urban public transportation, subways' operational energy consumption is directly linked to a city's energy consumption and carbon emissions. With the acceleration of urbanization and the continuous expansion of subway networks, the number of newly built subway stations is increasing annually. At the same time, due to increasingly stringent energy conservation and emission reduction requirements, energy consumption forecasting for newly built stations has become a key focus.
[0003] Traditional methods for predicting energy consumption at subway stations rely primarily on historical energy consumption and passenger flow data, using statistical analysis or machine learning models. However, the lack of historical energy consumption and passenger flow data for newly built subway stations makes these traditional methods inapplicable. Furthermore, equipment in existing stations may degrade over time, significantly differing in energy consumption characteristics from those of newly built stations. Therefore, a method for accurately predicting energy consumption at newly built subway stations is urgently needed. Summary of the Invention
[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method for predicting energy consumption of newly built subway stations, which solves the problem of difficulty in accurately predicting energy consumption of newly built subway stations.
[0005] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0006] The present invention provides a method for predicting energy consumption of a newly built subway station, comprising the following steps:
[0007] S1. Based on the Spearman correlation coefficient and the energy consumption prediction time scale, the initial station energy consumption analysis dataset is constructed based on historical energy consumption data and energy consumption impact data selected from the building scale data, equipment system data, operation data, and environmental data of existing stations;
[0008] S2. Obtain the target environment point of interest data of the station through the map API open platform and build an environmental visualization model of the station;
[0009] S3. Reclassify and partially eliminate target points of interest in the station's environmental visualization model, and build a target total passenger flow prediction model based on the total passenger flow in and out of the station;
[0010] S4. Construct a station energy consumption prediction dataset based on the initial station energy consumption analysis dataset and the total passenger flow predicted by the target in-and-out passenger flow prediction model, and train an SVR model based on the station energy consumption prediction dataset and the Hippo algorithm to obtain a trained station energy consumption prediction model;
[0011] S5. Obtain the energy consumption impact data of the newly built station and the total passenger flow in and out of the station predicted by the target total passenger flow prediction model, input the energy consumption impact data and the total passenger flow in and out of the station into the trained station energy consumption prediction model, and obtain the energy consumption prediction result of the newly built station.
[0012] The beneficial effects of the present invention are as follows: the present invention provides a method for predicting energy consumption of a newly built subway station, which constructs an initial station energy consumption analysis data set by quantifying the building scale data, equipment system data, operation data and environmental data of the station, and combining the historical energy consumption data of the station, thereby providing a basis for accurately predicting the total passenger flow in and out of the station and the energy consumption of the station; the present invention obtains the target interest point data of the station, thereby realizing the construction of an environmental visualization model of the station within a fixed distance buffer range of the station, thereby providing a basis for accurately analyzing the total passenger flow in and out of the station, and visually analyzing the geographical distribution of the station environment; the present invention reclassifies and partially eliminates the target interest points in the environmental visualization model of the station, and combines the station's entry and exit data to obtain the station's target interest point data. The total passenger flow out of the station is used to construct a target in-and-out passenger flow prediction model. The model can accurately predict the average total passenger flow in and out of the station based on the target point of interest data of the station, providing a basis for accurately predicting the energy consumption of the station; the present invention constructs a station energy consumption prediction dataset based on the initial station energy consumption analysis dataset and the predicted total passenger flow in and out of the station, and trains the SVR model based on the station energy consumption prediction dataset and the Hippo algorithm. The trained station energy consumption prediction model can accurately predict the energy consumption of newly built subway stations based on the energy consumption impact data of the newly built stations and the predicted total passenger flow in and out of the station; the present invention can provide energy consumption data support for further accurate control of energy conservation and emission reduction of subway stations and cities by accurately predicting the energy consumption of newly built stations.
[0013] Furthermore, the S1 includes the following steps:
[0014] S11. Using hourly units as a time scale, obtain building scale data, equipment system data, operation data, environmental data, and historical energy consumption data of existing stations as sample data for energy consumption analysis of the stations;
[0015] S12. Use the Spearman correlation coefficient to analyze the impact relationship between building scale data, equipment system data, operation data, environmental data and historical energy consumption data, and obtain the building scale energy consumption impact coefficient, equipment system energy consumption impact coefficient, operation energy consumption impact coefficient and environmental energy consumption impact coefficient;
[0016] The calculation expression of the Spearman correlation coefficient is as follows:
[0017]
[0018] Where ρ represents the Spearman correlation coefficient, Σ· represents the sum, D represents the difference in data order between the influencing factor data and the historical energy consumption data, and N represents the number of samples of energy consumption analysis sample data of existing stations. The Spearman correlation coefficient ranges from -1 to 1. The influencing factor data include building scale data, equipment system data, operation data, and environmental data.
[0019] S13. Select a coefficient between -0.4 and -1 or 0.4 and 1 among the building scale energy consumption impact coefficient, equipment system energy consumption impact coefficient, operation energy consumption impact coefficient, and environmental energy consumption impact coefficient as the target influencing factor coefficient;
[0020] S14. Selecting energy consumption analysis sample data of the station corresponding to the target impact coefficient as energy consumption impact data;
[0021] S15. According to the energy consumption prediction time scale, select the energy consumption impact data and historical energy consumption data under the corresponding time, and construct the initial energy consumption analysis data set of the station.
[0022] Furthermore, the S2 includes the following steps:
[0023] S21. According to the URL link format and based on the location of the station, write the interface key, point of interest type, center point coordinates and search radius as the target range to be crawled;
[0024] S22, substituting the target data to be crawled into the Python crawler program, crawling to obtain a number of target points of interest data, wherein the target point of interest data includes the unique ID, name, category, location information and latitude and longitude coordinates of the point of interest;
[0025] S23. If the same point of interest exists in different categories, all duplicate target point of interest data are found using the unique ID of the point of interest as an identifier, and the target point of interest data in the most matching category is retained. The remaining duplicate target point of interest data are deleted to obtain the cleaned target point of interest data to form a target point of interest dataset.
[0026] S24. Construct an environmental visualization model of the station based on the latitude and longitude coordinates of the station and the target point of interest data set.
[0027] Furthermore, the S24 includes the following steps:
[0028] S241. Add layers corresponding to the station and the target POI data in the target POI dataset in the QGIS geographic information system.
[0029] S242, using the same projection module to set the projection of the layer corresponding to the station and the layer corresponding to each target point of interest data, and setting a fixed distance buffer using the layer corresponding to the station;
[0030] S243, using a spatial query module to filter and obtain target point of interest data within a fixed distance buffer;
[0031] S244 , associating the target point of interest data within the fixed distance buffer with its corresponding layer to obtain a number of target points of interest, and generating an environmental visualization model of the station based on each target point of interest.
[0032] Furthermore, the S3 includes the following steps:
[0033] S31, reclassifying the target points of interest in the environment visualization model into commercial service category, residential category, office category, landscape category, and public service category;
[0034] S32, calculating the variance inflation coefficients between target points of interest in the commercial service, residential, office, scenic, and public service categories and the total passenger flow in and out of the station, and eliminating some target points of interest so that the variance inflation coefficients are less than 10;
[0035] The calculation expression of the variance inflation factor is as follows:
[0036]
[0037] Among them, VIF represents the variance inflation factor, R i Represents the negative correlation coefficient of the regression analysis between the i-th type of target interest point and the other types of target interest points, where VIF is greater than 1, i = 1, 2, 3, 4, 5, where i = 1, the first type of target interest point refers to the commercial service type target interest point, i = 2, the second type of target interest point refers to the residential type target interest point, i = 3, the third type of target interest point refers to the office type target interest point, i = 4, the fourth type of target interest point refers to the landscape type target interest point, and i = 5, the fifth type of target interest point refers to the public service type target interest point;
[0038] S33. Obtain the total passenger flow in and out of the station through the subway automatic ticket vending and checking system as the dependent variable, and obtain the average value of the total passenger flow in and out of the station as the descriptive statistical result of the dependent variable;
[0039] S34, taking the POI density, the distance from the bus station to the city center, the number of bus stops, and the road network density of the target points of interest of commercial service, residential, office, scenic, and public service categories within the corresponding range of the environmental visualization model as independent variables;
[0040] S35. Based on the independent variables and dependent variables, a multi-scale geographically weighted regression model for predicting total passenger flow in and out of the station is constructed;
[0041] The calculation expression of the multi-scale geographically weighted regression total passenger flow prediction model is as follows:
[0042]
[0043] Among them, y j represents the predicted average total passenger flow in and out of station j, represents the constant coefficient of multi-scale geographically weighted regression, u j represents the longitude coordinate of station j, v j represents the dimensional coordinates of station j, represents the spatial weight of the kth independent variable at station j, x jk represents the independent variable corresponding to station j, ε j represents the random error corresponding to station j when estimating the passenger flow of entry and exit of the station using multi-scale geographically weighted regression, m represents the total number of independent variables, k represents the kth independent variable, and D k represents the optimal bandwidth used when determining the spatial weight of the k-th independent variable, where k = 1, 2, ..., m;
[0044] S36. The bi-square function is used as the spatial weight function of the multi-scale geographically weighted regression in- and out-station total passenger flow prediction model, and the modified Chi-Shen criterion is used as the spatial weight allocation evaluation model. The sum of squares of the convergence factor is used as the model parameter convergence evaluation indicator. The optimal adaptive bandwidth is selected to obtain the target in- and out-station total passenger flow prediction model.
[0045] Furthermore, the S36 includes the following steps:
[0046] S361. Use the bi-square function as the spatial weight allocation function for the multi-scale geographically weighted regression total passenger flow prediction model;
[0047] S362, selecting the modified Chi Xin pool criterion as the spatial weight allocation evaluation model, and selecting the sum of squares of the convergence factor as the model parameter convergence evaluation indicator;
[0048] The calculation expression of the spatial weight evaluation model is as follows:
[0049]
[0050] Among them, AIC c represents the modified Chixinchi criterion value, AIC represents the Chixinchi criterion value, n represents the total number of stations on the same subway track, represents the maximum likelihood estimate of the random error variance, trace(S) represents the trace of the smoothing matrix S, and RSS represents the residual sum of squares;
[0051] The calculation expression of the model parameter convergence evaluation index is as follows:
[0052]
[0053] Among them, SOC f represents the sum of squares of convergence factors, represents the predicted average total passenger flow in and out of station j during the current iteration. represents the predicted average total passenger flow in and out of station j in the previous iteration;
[0054] S363, iteratively assigning the spatial weights of the respective variables in the multi-scale geographically weighted regression inbound and outbound passenger flow prediction model based on the spatial weight allocation function, and stopping the iteration when the change in the model parameter convergence evaluation index is less than a preset convergence range, and calculating the modified red-signal pool criterion value through the spatial weight evaluation model;
[0055] S364, repeatedly executing S363 a first preset number of times to obtain a number of modified red-signal pool criterion values and the spatial weights of the respective variables in the multi-scale geographically weighted regression in- and out-station total passenger flow prediction model corresponding to the modified red-signal pool criterion values;
[0056] S365, obtaining the optimal adaptive bandwidth based on the spatial weights of the respective variables in the multi-scale geographically weighted regression total passenger flow prediction model corresponding to the minimum modified red signal pool criterion value;
[0057] S366. Determine the spatial weights of the respective variables in the multi-scale geographically weighted regression inbound and outbound passenger flow prediction model based on the optimal adaptive bandwidth to obtain the target inbound and outbound passenger flow prediction model.
[0058] Furthermore, the S4 includes the following steps:
[0059] S41. Performing one-hot encoding or normalization processing on the building scale data, equipment system data, operation data, and environmental data in the initial station energy consumption analysis data set to obtain pre-processed energy consumption impact data;
[0060] S42. Combining the pre-processed energy consumption impact data at the same time scale with the total passenger flow in and out of the station predicted by the target total passenger flow prediction model based on the energy consumption impact data, the data is used as energy consumption prediction feature input data, and the historical energy consumption data corresponding to the energy consumption impact data is used as the true value of the energy consumption prediction;
[0061] S43. Constructing a station energy consumption prediction dataset based on each energy consumption prediction feature input data and the corresponding energy consumption prediction true value;
[0062] S44. Select a kernel function of the SVR model and initialize the kernel coefficient and penalty coefficient of the SVR model;
[0063] S45. Let the objective function be to minimize the mean square error between the predicted energy consumption value of the station and the true value of the predicted energy consumption;
[0064] S46. According to the station energy consumption prediction data set and the objective function, the SVR model is trained, and the kernel coefficient and penalty coefficient of the SVR model are optimized using the Hippo optimization algorithm to obtain a trained station energy consumption prediction model.
[0065] Furthermore, the S46 includes the following steps:
[0066] S461. Based on the station energy consumption prediction dataset and the objective function, the SVR model is trained and the hippo population is initialized. The kernel coefficient and penalty coefficient of each set of SVR models are used as a set of candidate solutions, corresponding to a hippo in the hippo population.
[0067] S462, calculating the mean square error of the station energy consumption prediction value corresponding to the SVR model based on each set of candidate solutions as the Hippo fitness value;
[0068] S463. According to the search, defense, and escape behaviors of the hippopotamus, the kernel coefficient and penalty coefficient of the SVR model are adjusted, the corresponding hippopotamus fitness value is calculated, and the candidate solution corresponding to the minimum hippopotamus fitness value is taken as the global optimal solution;
[0069] S464. Repeat S463 to S463 several times until the objective function converges to a preset threshold, and use the SVR model corresponding to the global optimal solution as the trained station energy consumption prediction model.
[0070] Other advantages of the present invention will be analyzed in more detail in subsequent embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0072] Figure 1 This is a flowchart of the steps of a method for predicting energy consumption of a newly built subway station in an embodiment of the present invention. DETAILED DESCRIPTION
[0073] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0074] like Figure 1 As shown, in one embodiment of the present invention, the present invention provides an energy consumption prediction method for a newly built subway station, comprising the following steps:
[0075] S1. Based on the Spearman correlation coefficient and the energy consumption prediction time scale, the initial station energy consumption analysis dataset is constructed based on historical energy consumption data and energy consumption impact data selected from the building scale data, equipment system data, operation data, and environmental data of existing stations;
[0076] The S1 comprises the following steps:
[0077] S11. Using hourly units as a time scale, obtain building scale data, equipment system data, operation data, environmental data, and historical energy consumption data of existing stations as sample data for energy consumption analysis of the stations;
[0078] In this plan, building scale data includes the station building area, number of entrances and exits, and total number of underground floors; equipment system data includes air-conditioning system equipment model, air-conditioning system equipment distribution area, lighting system equipment model, lighting system equipment distribution area, elevator system equipment model, and number of elevator system equipment; operation data includes station operating hours, total passenger flow in and out of the station, and station transfer types; environmental data includes commercial building passenger flow, residential area population, office building passenger flow, scenic spot passenger flow, and healthcare passenger flow; historical energy consumption data is the historical electricity consumption of the station.
[0079] S12. Use the Spearman correlation coefficient to analyze the impact relationship between building scale data, equipment system data, operation data, environmental data and historical energy consumption data, and obtain the building scale energy consumption impact coefficient, equipment system energy consumption impact coefficient, operation energy consumption impact coefficient and environmental energy consumption impact coefficient;
[0080] The calculation expression of the Spearman correlation coefficient is as follows:
[0081]
[0082] Where ρ represents the Spearman correlation coefficient, Σ· represents the sum, D represents the difference in data order between the influencing factor data and the historical energy consumption data, and N represents the number of samples of energy consumption analysis sample data of existing stations. The Spearman correlation coefficient ranges from -1 to 1. The influencing factor data include building scale data, equipment system data, operation data, and environmental data.
[0083] In this scheme, when the absolute value of the Spearman correlation coefficient is in [0.0, 0.2], it indicates that the correlation strength between the corresponding data and the station energy consumption is no correlation; when the absolute value of the Spearman correlation coefficient is in (0.2, 0.4], it indicates that the correlation strength between the corresponding data and the station energy consumption is weak correlation; when the absolute value of the Spearman correlation coefficient is in (0.4, 0.6], it indicates that the correlation strength between the corresponding data and the station energy consumption is medium correlation; when the absolute value of the Spearman correlation coefficient is in (0.6, 0.8], it indicates that the correlation strength between the corresponding data and the station energy consumption is strong correlation; when the absolute value of the Spearman correlation coefficient is in (0.8, 1.0], it indicates that the correlation strength between the corresponding data and the station energy consumption is strong correlation. In this scheme, other data with at least medium correlation strength between the energy consumption analysis sample data of existing stations and the historical energy consumption data are selected to construct a station energy consumption analysis dataset in combination with the historical energy consumption data.
[0084] S13. Select a coefficient between -0.4 and -1 or 0.4 and 1 among the building scale energy consumption impact coefficient, equipment system energy consumption impact coefficient, operation energy consumption impact coefficient, and environmental energy consumption impact coefficient as the target influencing factor coefficient;
[0085] S14. Selecting energy consumption analysis sample data of the station corresponding to the target impact coefficient as energy consumption impact data;
[0086] S15. According to the energy consumption prediction time scale, select the energy consumption impact data and historical energy consumption data under the corresponding time, and construct the initial energy consumption analysis data set of the station.
[0087] In this scheme, the energy consumption prediction time scale is usually set to daily or monthly. When the energy consumption prediction time scale is daily, the corresponding station energy consumption analysis dataset is a day-level dataset. When the energy consumption prediction time scale is monthly, the corresponding station energy consumption analysis dataset is a monthly dataset.
[0088] S2. Obtain the target environment point of interest data of the station through the map API open platform and build an environmental visualization model of the station;
[0089] The S2 comprises the following steps:
[0090] S21. According to the URL link format and based on the location of the station, write the interface key, point of interest type, center point coordinates and search radius as the target range to be crawled;
[0091] In this scenario, the types of POIs are commercial buildings, residential buildings, office buildings, scenic spots, and healthcare service points;
[0092] S22, substituting the target data to be crawled into the Python crawler program, crawling to obtain a number of target points of interest data, wherein the target point of interest data includes the unique ID, name, category, location information and latitude and longitude coordinates of the point of interest;
[0093] In this solution, some of the crawled target POI data may belong to multiple categories at the same time, resulting in some target POI data appearing repeatedly in different categories. Therefore, the crawled target POI data needs to be cleaned;
[0094] S23. If the same point of interest exists in different categories, all duplicate target point of interest data are found using the unique ID of the point of interest as an identifier, and the target point of interest data in the most matching category is retained. The remaining duplicate target point of interest data are deleted to obtain the cleaned target point of interest data to form a target point of interest dataset.
[0095] In this embodiment, in order to facilitate the visual analysis of the target interest point data, the target interest point data are reclassified into commercial service category, residential category, office category, landscape category and public service category according to the name and category of the target interest point data, and the latitude and longitude coordinates of each target interest data are converted into the WGS1984 coordinate system.
[0096] S24. Construct an environmental visualization model of the station based on the latitude and longitude coordinates of the station and the target point of interest data set.
[0097] The S24 includes the following steps:
[0098] S241. Add layers corresponding to the station and the target POI data in the target POI dataset in the QGIS geographic information system.
[0099] In this embodiment, the format of the target point of interest data is Shapefile format, GeoJSON format or CSV format; if it is in Shapefile format or GeoJSON format, the layer is directly added by selecting Add vector layer; if it is in CSV format, select Add delimited text layer and set the latitude and longitude fields according to the latitude and longitude coordinates.
[0100] S242, using the same projection module to set the projection of the layer corresponding to the station and the layer corresponding to each target point of interest data, and setting a fixed distance buffer using the layer corresponding to the station;
[0101] In this example, a fixed distance buffer is set to 800 meters, and the same projection module is used to facilitate the QGIS geographic information system to successfully display the spatial relationship between the station and the points of interest around the station.
[0102] S243, using a spatial query module to filter and obtain target point of interest data within a fixed distance buffer;
[0103] S244 , associating the target point of interest data within the fixed distance buffer with its corresponding layer to obtain a number of target points of interest, and generating an environmental visualization model of the station based on each target point of interest.
[0104] S3. Reclassify and partially eliminate target points of interest in the station's environmental visualization model, and build a target total passenger flow prediction model based on the total passenger flow in and out of the station;
[0105] The S3 comprises the following steps:
[0106] S31, reclassifying the target points of interest in the environment visualization model into commercial service category, residential category, office category, landscape category, and public service category;
[0107] S32, calculating the variance inflation coefficients between target points of interest in the commercial service, residential, office, scenic, and public service categories and the total passenger flow in and out of the station, and eliminating some target points of interest so that the variance inflation coefficients are less than 10;
[0108] The calculation expression of the variance inflation factor is as follows:
[0109]
[0110] Among them, VIF represents the variance inflation factor, R i Represents the negative correlation coefficient of the regression analysis between the i-th type of target interest point and the other types of target interest points, where VIF is greater than 1, i = 1, 2, 3, 4, 5, where i = 1, the first type of target interest point refers to the commercial service type target interest point, i = 2, the second type of target interest point refers to the residential type target interest point, i = 3, the third type of target interest point refers to the office type target interest point, i = 4, the fourth type of target interest point refers to the landscape type target interest point, and i = 5, the fifth type of target interest point refers to the public service type target interest point;
[0111] In this scheme, the value of the variance inflation coefficient is always greater than 1. The closer the variance inflation coefficient is to 1, the lighter the multicollinearity between various target interest points: when the variance inflation coefficient is less than 10, there is no multicollinearity; when the variance inflation coefficient is between greater than or equal to 10 and less than 100, there is strong multicollinearity; when the variance inflation coefficient is greater than or equal to 100, there is serious multicollinearity, and the target interest points need to be eliminated.
[0112] S33. Obtain the total passenger flow in and out of the station through the subway automatic ticket vending and checking system as the dependent variable, and obtain the average value of the total passenger flow in and out of the station as the descriptive statistical result of the dependent variable;
[0113] In this scheme, the average value of the total in- and out-of-station passenger flow is the average value of the passenger flow under the time scale of the scheme's preset energy consumption forecast.
[0114] S34, taking the POI density, the distance from the bus station to the city center, the number of bus stops, and the road network density of the target points of interest of commercial service, residential, office, scenic, and public service categories within the corresponding range of the environmental visualization model as independent variables;
[0115] In this embodiment, the corresponding range of the environment visualization model refers to a fixed distance buffer range with the station as the center point.
[0116] S35. Based on the independent variables and dependent variables, a multi-scale geographically weighted regression model for predicting total passenger flow in and out of the station is constructed;
[0117] The calculation expression of the multi-scale geographically weighted regression total passenger flow prediction model is as follows:
[0118]
[0119] Among them, y j represents the predicted average total passenger flow in and out of station j, represents the constant coefficient of multi-scale geographically weighted regression, u j represents the longitude coordinate of station j, v j represents the dimensional coordinates of station j, represents the spatial weight of the kth independent variable at station j, x jk represents the independent variable corresponding to station j, ε j represents the random error corresponding to station j when estimating the passenger flow of entry and exit of the station using multi-scale geographically weighted regression, m represents the total number of independent variables, k represents the kth independent variable, and D k represents the optimal bandwidth used when determining the spatial weight of the k-th independent variable, where k = 1, 2, ..., m;
[0120] S36. The bi-square function is used as the spatial weight function of the multi-scale geographically weighted regression in- and out-station total passenger flow prediction model, and the modified Chi-Shen criterion is used as the spatial weight allocation evaluation model. The sum of squares of the convergence factor is used as the model parameter convergence evaluation indicator. The optimal adaptive bandwidth is selected to obtain the target in- and out-station total passenger flow prediction model.
[0121] In this scheme, the bi-square function is a spatial weight function that can assign weights based on the distance between spatial units. When the distance is less than the bandwidth, the weight decreases as the distance increases; when the distance is greater than the bandwidth, the weight is truncated to zero. When using the bi-square function, selecting an adaptive bandwidth means that for each spatial unit, a suitable bandwidth value will be calculated based on the density or distribution of the data points around it. The advantage of doing this is that it can better capture the local characteristics of spatial data and improve the accuracy of spatial analysis. Combined with the modified red signal pool criterion, the optimal bandwidth type can be selected.
[0122] The modified Chi Xinchi criterion is a model selection criterion that considers the balance between model goodness of fit and complexity. By comparing the AICc values of the multi-scale geographically weighted regression inbound and outbound passenger flow prediction model under different bandwidth types, the optimal adaptive bandwidth can be selected. Adaptive bandwidth can more accurately reflect the local characteristics of spatial data, thereby improving the accuracy and reliability of spatial analysis. By allowing each spatial unit to have a different bandwidth value, adaptive bandwidth enhances the flexibility of the inbound and outbound passenger flow prediction model, making it more adaptable to spatial data of different types and distributions.
[0123] The S36 includes the following steps:
[0124] S361. Use the bi-square function as the spatial weight allocation function for the multi-scale geographically weighted regression total passenger flow prediction model;
[0125] S362, selecting the modified Chi Xin pool criterion as the spatial weight allocation evaluation model, and selecting the sum of squares of the convergence factor as the model parameter convergence evaluation indicator;
[0126] The calculation expression of the spatial weight evaluation model is as follows:
[0127]
[0128] Among them, AIC c represents the modified Chixinchi criterion value, AIC represents the Chixinchi criterion value, n represents the total number of stations on the same subway track, represents the maximum likelihood estimate of the random error variance, trace(S) represents the trace of the smoothing matrix S, and RSS represents the residual sum of squares;
[0129] In this scheme, the predicted average total inbound and outbound passenger flow of station j in the multi-scale geographically weighted regression inbound and outbound passenger flow estimation model is equal to the product of the smoothing matrix S and the actual daily average total inbound and outbound passenger flow of station j; the trace of the matrix is defined as the sum of the elements on the main diagonal; the smaller the bandwidth, the stronger the localization of the model, the more detailed the data fit, the more complex the smoothing matrix S, and the larger the trace of the smoothing matrix; the larger the bandwidth, the smoother the model, and the smaller the trace of the smoothing matrix.
[0130] The calculation expression of the model parameter convergence evaluation index is as follows:
[0131]
[0132] Among them, SOC f represents the sum of squares of convergence factors, represents the predicted average total passenger flow in and out of station j during the current iteration. represents the predicted average total passenger flow in and out of station j in the previous iteration;
[0133] S363, iteratively assigning the spatial weights of the respective variables in the multi-scale geographically weighted regression inbound and outbound passenger flow prediction model based on the spatial weight allocation function, and stopping the iteration when the change in the model parameter convergence evaluation index is less than a preset convergence range, and calculating the modified red-signal pool criterion value through the spatial weight evaluation model;
[0134] S364, repeatedly executing S363 a first preset number of times to obtain a number of modified red-signal pool criterion values and the spatial weights of the respective variables in the multi-scale geographically weighted regression in- and out-station total passenger flow prediction model corresponding to the modified red-signal pool criterion values;
[0135] S365, obtaining the optimal adaptive bandwidth based on the spatial weights of the respective variables in the multi-scale geographically weighted regression total passenger flow prediction model corresponding to the minimum modified red signal pool criterion value;
[0136] S366. Determine the spatial weights of the respective variables in the multi-scale geographically weighted regression inbound and outbound passenger flow prediction model based on the optimal adaptive bandwidth to obtain the target inbound and outbound passenger flow prediction model.
[0137] S4. Construct a station energy consumption prediction dataset based on the initial station energy consumption analysis dataset and the total passenger flow predicted by the target in-and-out passenger flow prediction model, and train an SVR model based on the station energy consumption prediction dataset and the Hippo algorithm to obtain a trained station energy consumption prediction model;
[0138] The S4 comprises the following steps:
[0139] S41. Performing one-hot encoding or normalization processing on the building scale data, equipment system data, operation data, and environmental data in the initial station energy consumption analysis data set to obtain pre-processed energy consumption impact data;
[0140] In this scheme, categorical variables such as station type and system equipment model are processed with one-hot encoding, and numerical variables such as building area, equipment distribution area, and passenger flow are processed with Min-Max normalization;
[0141] S42. Combining the pre-processed energy consumption impact data at the same time scale with the total passenger flow in and out of the station predicted by the target total passenger flow prediction model based on the energy consumption impact data, the data is used as energy consumption prediction feature input data, and the historical energy consumption data corresponding to the energy consumption impact data is used as the true value of the energy consumption prediction;
[0142] S43. Constructing a station energy consumption prediction dataset based on each energy consumption prediction feature input data and the corresponding energy consumption prediction true value;
[0143] S44. Select a kernel function of the SVR model and initialize the kernel coefficient and penalty coefficient of the SVR model;
[0144] S45. Let the objective function be to minimize the mean square error between the predicted energy consumption value of the station and the true value of the predicted energy consumption;
[0145] S46. According to the station energy consumption prediction data set and the objective function, the SVR model is trained, and the kernel coefficient and penalty coefficient of the SVR model are optimized using the Hippo optimization algorithm to obtain a trained station energy consumption prediction model.
[0146] The S46 includes the following steps:
[0147] S461. Based on the station energy consumption prediction dataset and the objective function, the SVR model is trained and the hippo population is initialized. The kernel coefficient and penalty coefficient of each set of SVR models are used as a set of candidate solutions, corresponding to a hippo in the hippo population.
[0148] S462, calculating the mean square error of the station energy consumption prediction value corresponding to the SVR model based on each set of candidate solutions as the Hippo fitness value;
[0149] S463. According to the search, defense, and escape behaviors of the hippopotamus, the kernel coefficient and penalty coefficient of the SVR model are adjusted, the corresponding hippopotamus fitness value is calculated, and the candidate solution corresponding to the minimum hippopotamus fitness value is taken as the global optimal solution;
[0150] S464. Repeat S463 to S463 several times until the objective function converges to a preset threshold, and use the SVR model corresponding to the global optimal solution as the trained station energy consumption prediction model.
[0151] S5. Obtain the energy consumption impact data of the newly built station and the total passenger flow in and out of the station predicted by the target total passenger flow prediction model, input the energy consumption impact data and the total passenger flow in and out of the station into the trained station energy consumption prediction model, and obtain the energy consumption prediction result of the newly built station.
[0152] In this embodiment, the energy consumption impact data of the newly built station is obtained according to the methods of S11 to S14. By processing the energy consumption impact data within the fixed distance buffer range of the newly built station according to the methods of S2 to S32 and inputting it into the target inbound and outbound passenger flow prediction model, the total inbound and outbound passenger flow of the newly built station can be predicted within the time scale of the preset energy consumption prediction plan.
[0153] In a practical example, the station energy consumption prediction data set is divided into an 80% training set and a 20% test set. The SVR model is trained using the training set through the S44 to S46 methods to obtain a trained station energy consumption prediction model. The trained station energy consumption prediction model is used to predict the energy consumption of the feature data in the test set to obtain the energy consumption prediction results of the test data. When evaluating the performance of the trained station energy consumption prediction model, the performance evaluation indicators include root mean square error (RMSE), average decision percentage error (MAPE), mean absolute error (MAE), and determination coefficient (R). 2 , to obtain the performance of the trained station energy consumption prediction model.
[0154] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.
Claims
1. A method for predicting energy consumption of a new subway station, characterized in that: The steps include: S1. Based on the Spearman correlation coefficient and the energy consumption prediction time scale, the initial station energy consumption analysis dataset is constructed based on historical energy consumption data and energy consumption impact data selected from the building scale data, equipment system data, operation data, and environmental data of existing stations; S2. Obtain the target environment point of interest data of the station through the map API open platform and build an environmental visualization model of the station; S3. Reclassify and partially eliminate target points of interest in the station's environmental visualization model, and build a target total passenger flow prediction model based on the total passenger flow in and out of the station; S4. Construct a station energy consumption prediction dataset based on the initial station energy consumption analysis dataset and the total passenger flow predicted by the target in-and-out passenger flow prediction model, and train an SVR model based on the station energy consumption prediction dataset and the Hippo algorithm to obtain a trained station energy consumption prediction model; S5. Obtain the energy consumption impact data of the newly built station and the total passenger flow in and out of the station predicted by the target total passenger flow prediction model, input the energy consumption impact data and the total passenger flow in and out of the station into the trained station energy consumption prediction model, and obtain the energy consumption prediction result of the newly built station.
2. The energy consumption prediction method for a new subway station according to claim 1 is characterized in that: The S1 comprises the following steps: S11. Using hourly units as a time scale, obtain building scale data, equipment system data, operation data, environmental data, and historical energy consumption data of existing stations as sample data for energy consumption analysis of the stations; S12. Use the Spearman correlation coefficient to analyze the impact relationship between building scale data, equipment system data, operation data, environmental data and historical energy consumption data, and obtain the building scale energy consumption impact coefficient, equipment system energy consumption impact coefficient, operation energy consumption impact coefficient and environmental energy consumption impact coefficient; The calculation expression of the Spearman correlation coefficient is as follows: Where ρ represents the Spearman correlation coefficient, Σ· represents the sum, D represents the difference in data order between the influencing factor data and the historical energy consumption data, and N represents the number of samples of energy consumption analysis sample data of existing stations. The Spearman correlation coefficient ranges from -1 to 1. The influencing factor data include building scale data, equipment system data, operation data, and environmental data. S13. Select a coefficient between -0.4 and -1 or 0.4 and 1 among the building scale energy consumption impact coefficient, equipment system energy consumption impact coefficient, operation energy consumption impact coefficient, and environmental energy consumption impact coefficient as the target influencing factor coefficient; S14. Selecting energy consumption analysis sample data of the station corresponding to the target impact coefficient as energy consumption impact data; S15. According to the energy consumption prediction time scale, select the energy consumption impact data and historical energy consumption data under the corresponding time, and construct the initial energy consumption analysis data set of the station.
3. The energy consumption prediction method for a new subway station according to claim 1 is characterized in that: The S2 comprises the following steps: S21. According to the URL link format and based on the location of the station, write the interface key, point of interest type, center point coordinates and search radius as the target range to be crawled; S22, substituting the target data to be crawled into the Python crawler program, crawling to obtain a number of target points of interest data, wherein the target point of interest data includes the unique ID, name, category, location information and latitude and longitude coordinates of the point of interest; S23. If the same point of interest exists in different categories, all duplicate target point of interest data are found using the unique ID of the point of interest as an identifier, and the target point of interest data in the most matching category is retained. The remaining duplicate target point of interest data are deleted to obtain the cleaned target point of interest data to form a target point of interest dataset. S24. Construct an environmental visualization model of the station based on the latitude and longitude coordinates of the station and the target point of interest data set.
4. The energy consumption prediction method for a new subway station according to claim 3 is characterized in that: The S24 includes the following steps: S241. Add layers corresponding to the station and the target POI data in the target POI dataset in the QGIS geographic information system. S242, using the same projection module to set the projection of the layer corresponding to the station and the layer corresponding to each target point of interest data, and setting a fixed distance buffer using the layer corresponding to the station; S243, using a spatial query module to filter and obtain target point of interest data within a fixed distance buffer; S244 , associating the target point of interest data within the fixed distance buffer with its corresponding layer to obtain a number of target points of interest, and generating an environmental visualization model of the station based on each target point of interest.
5. The energy consumption prediction method for a new subway station according to claim 4 is characterized in that: The S3 comprises the following steps: S31, reclassifying the target points of interest in the environment visualization model into commercial service category, residential category, office category, landscape category, and public service category; S32, calculating the variance inflation coefficients between target points of interest in the commercial service, residential, office, scenic, and public service categories and the total passenger flow in and out of the station, and eliminating some target points of interest so that the variance inflation coefficients are less than 10; The calculation expression of the variance inflation factor is as follows: Among them, VIF represents the variance inflation factor, R i Represents the negative correlation coefficient of the regression analysis between the i-th type of target interest point and the other types of target interest points, where VIF is greater than 1, i = 1, 2, 3, 4, 5, where i = 1, the first type of target interest point refers to the commercial service type target interest point, i = 2, the second type of target interest point refers to the residential type target interest point, i = 3, the third type of target interest point refers to the office type target interest point, i = 4, the fourth type of target interest point refers to the landscape type target interest point, and i = 5, the fifth type of target interest point refers to the public service type target interest point; S33. Obtain the total passenger flow in and out of the station through the subway automatic ticket vending and checking system as the dependent variable, and obtain the average value of the total passenger flow in and out of the station as the descriptive statistical result of the dependent variable; S34, taking the POI density, the distance from the bus station to the city center, the number of bus stops, and the road network density of the target points of interest of commercial service, residential, office, scenic, and public service categories within the corresponding range of the environmental visualization model as independent variables; S35. Based on the independent variables and dependent variables, a multi-scale geographically weighted regression model for predicting total passenger flow in and out of the station is constructed; The calculation expression of the multi-scale geographically weighted regression total passenger flow prediction model is as follows: Among them, y j represents the predicted average total passenger flow in and out of station j, represents the constant coefficient of multi-scale geographically weighted regression, u j represents the longitude coordinate of station j, v j represents the dimensional coordinates of station j, represents the spatial weight of the kth independent variable at station j, x jk represents the independent variable corresponding to station j, ε j represents the random error corresponding to station j when estimating the passenger flow of entry and exit of the station using multi-scale geographically weighted regression, m represents the total number of independent variables, k represents the kth independent variable, and D k represents the optimal bandwidth used when determining the spatial weight of the k-th independent variable, where k = 1, 2, ..., m; S36. The bi-square function is used as the spatial weight function of the multi-scale geographically weighted regression in- and out-station total passenger flow prediction model, and the modified Chi-Shen criterion is used as the spatial weight allocation evaluation model. The sum of squares of the convergence factor is used as the model parameter convergence evaluation indicator. The optimal adaptive bandwidth is selected to obtain the target in- and out-station total passenger flow prediction model.
6. The energy consumption prediction method for a new subway station according to claim 5, characterized in that: The S36 includes the following steps: S361. Use the bi-square function as the spatial weight allocation function for the multi-scale geographically weighted regression total passenger flow prediction model; S362, selecting the modified Chi Xin pool criterion as the spatial weight allocation evaluation model, and selecting the sum of squares of the convergence factor as the model parameter convergence evaluation indicator; The calculation expression of the spatial weight evaluation model is as follows: Among them, AIC c represents the modified Chixinchi criterion value, AIC represents the Chixinchi criterion value, n represents the total number of stations on the same subway track, represents the maximum likelihood estimate of the random error variance, trace(S) represents the trace of the smoothing matrix S, and RSS represents the residual sum of squares; The calculation expression of the model parameter convergence evaluation index is as follows: Among them, SOC f represents the sum of squares of convergence factors, represents the predicted average total passenger flow in and out of station j during the current iteration. represents the predicted average total passenger flow in and out of station j in the previous iteration; S363, iteratively assigning the spatial weights of the respective variables in the multi-scale geographically weighted regression inbound and outbound passenger flow prediction model based on the spatial weight allocation function, and stopping the iteration when the change in the model parameter convergence evaluation index is less than a preset convergence range, and calculating the modified red-signal pool criterion value through the spatial weight evaluation model; S364, repeatedly executing S363 a first preset number of times to obtain a number of modified red-signal pool criterion values and the spatial weights of the respective variables in the multi-scale geographically weighted regression in- and out-station total passenger flow prediction model corresponding to the modified red-signal pool criterion values; S365, obtaining the optimal adaptive bandwidth based on the spatial weights of the respective variables in the multi-scale geographically weighted regression total passenger flow prediction model corresponding to the minimum modified red signal pool criterion value; S366. Determine the spatial weights of the respective variables in the multi-scale geographically weighted regression inbound and outbound passenger flow prediction model based on the optimal adaptive bandwidth to obtain the target inbound and outbound passenger flow prediction model.
7. The energy consumption prediction method for a new subway station according to claim 6, characterized in that: The S4 comprises the following steps: S41. Performing one-hot encoding or normalization processing on the building scale data, equipment system data, operation data, and environmental data in the initial station energy consumption analysis data set to obtain pre-processed energy consumption impact data; S42. Combining the pre-processed energy consumption impact data at the same time scale with the total passenger flow in and out of the station predicted by the target total passenger flow prediction model based on the energy consumption impact data, the data is used as energy consumption prediction feature input data, and the historical energy consumption data corresponding to the energy consumption impact data is used as the true value of the energy consumption prediction; S43. Constructing a station energy consumption prediction dataset based on each energy consumption prediction feature input data and the corresponding energy consumption prediction true value; S44. Select a kernel function of the SVR model and initialize the kernel coefficient and penalty coefficient of the SVR model; S45. Let the objective function be to minimize the mean square error between the predicted energy consumption value of the station and the true value of the predicted energy consumption; S46. According to the station energy consumption prediction data set and the objective function, the SVR model is trained, and the kernel coefficient and penalty coefficient of the SVR model are optimized using the Hippo optimization algorithm to obtain a trained station energy consumption prediction model.
8. The energy consumption prediction method for a new subway station according to claim 7, characterized in that: The S46 includes the following steps: S461. Based on the station energy consumption prediction dataset and the objective function, the SVR model is trained and the hippo population is initialized. The kernel coefficient and penalty coefficient of each set of SVR models are used as a set of candidate solutions, corresponding to a hippo in the hippo population. S462, calculating the mean square error of the station energy consumption prediction value corresponding to the SVR model based on each set of candidate solutions as the Hippo fitness value; S463. According to the search, defense, and escape behaviors of the hippopotamus, the kernel coefficient and penalty coefficient of the SVR model are adjusted, the corresponding hippopotamus fitness value is calculated, and the candidate solution corresponding to the minimum hippopotamus fitness value is taken as the global optimal solution; S464. Repeat S463 to S463 several times until the objective function converges to a preset threshold, and use the SVR model corresponding to the global optimal solution as the trained station energy consumption prediction model.