New line passenger flow prediction method based on semi-supervised Conv-Transform deep learning model

By adopting the semi-supervised learning Conv-Transformer model in the new line passenger flow prediction, the passenger flow is decomposed into trend flow and scale flow, and using data augmentation methods, the problems of insufficient data and insufficient temporal correlation capture in the existing technology are solved, achieving higher prediction accuracy.

CN120197741APending Publication Date: 2025-06-24BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510134447.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing new line passenger flow prediction model has insufficient data and space-time correlation capture, resulting in low prediction accuracy.

Method used

The Conv-Transformer deep learning model based on semi-supervised learning is adopted to improve the prediction effect by obtaining POI data and decomposing it into trend flow and scale flow, and combining data enhancement methods.

Benefits of technology

It improves the prediction accuracy of new line incoming flow, solves the problem of insufficient data, and achieves better spatial and temporal correlation capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197741A_ABST
    Figure CN120197741A_ABST
Patent Text Reader

Abstract

The invention discloses a new line passenger flow prediction method based on a semi-supervised Conv-Transform deep learning model. The method comprises the following steps: acquiring POI data for a target line; the POI data are input into a trained convolution Transform model, a trend flow and a scale flow are obtained, the scale flow reflects scale features of the passenger flow volume of the station, the trend flow reflects trend features of the passenger flow volume of the station, and training of the convolution Transform model adopts a semi-supervised learning mode; and combining the trend flow and the scale flow to obtain a passenger flow prediction result. According to the method, the accuracy of predicting the new line passenger flow is improved, and the demand for training data is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of passenger flow prediction, and more specifically, to a new line passenger flow prediction method based on a semi-supervised Conv-Transformer deep learning model. Background Art

[0002] Due to the rapid development of urban rail transit and the expansion of the original subway lines, it has become increasingly important to accurately predict the passenger flow of newly opened subway lines and subway extension lines. Due to challenges such as the lack of historical data and the long prediction time span, existing studies have problems with insufficient accuracy when predicting the inbound flow of new lines.

[0003] In the prior art, the models for predicting the inbound flow of new lines can be roughly divided into three categories: traditional statistics-based models, machine learning-based models, and deep learning-based models. Traditional statistics-based models include autoregressive integrated moving average model (ARIMA), fuzzy C-means clustering (FCM), geographically weighted regression (GWR), etc. For example, the ARIMA model is used to analyze the historical passenger flow data of subway lines with similarities, understand their seasonal, trend, and periodic characteristics, and predict the passenger flow after the opening of new lines. FCM is an extension of the C-means clustering method and can handle data ambiguity more flexibly. For example, the GWR model is used to analyze the station site characteristics around new lines, understand the geographical attributes, surrounding environment, and passenger flow of different station sites, evaluate the factors affecting the passenger flow change after the opening of new lines, and thus predict the inbound flow of new lines. However, using traditional statistics-based models often results in large information losses, thus being unable to fully capture the spatio-temporal characteristics of data, which undoubtedly limits the prediction accuracy.

[0004] In recent years, passenger flow prediction models based on machine learning and deep learning have also been widely used in urban rail transit systems. In the field of machine learning, some researchers have proposed a two-step prediction model to predict the passenger flow demand of stations on subway extension lines, determine the decisive features from the potential influencing factors of passenger flow demand, and predict the passenger flow of new lines. Although machine learning-based models are superior to traditional models, they are still difficult to capture the spatio-temporal correlation of complex data well.

[0005] Deep learning-based models can capture complex temporal and spatial correlations from data, improving the prediction accuracy across the network. With the rise of attention mechanisms and Transformer architectures, attention-based neural network frameworks have been widely used because they can extract global features of passenger flow. For example, some researchers have proposed an end-to-end deep learning framework based on attention mechanisms that can perform multi-step predictions for all stations in a large-scale subway system simultaneously. This end-to-end model consists of an encoder network and a decoder network and can effectively model sequences of different lengths. However, deep learning-based models are often data-driven models that require a large amount of data for prediction, and the lack of historical data in new line prediction greatly affects the prediction performance of such models.

[0006] After analysis, the current new line prediction models mainly have the following problems: The traditional method using manual surveys consumes a large amount of human and material resources, and the traditional method uses linear methods for clustering prediction, which cannot effectively extract non-linear features, thus reducing the prediction accuracy; Although the machine learning-based models have improved the prediction accuracy of new line prediction to a certain extent, they usually cannot well extract and process high-dimensional problems and complex spatio-temporal correlations during the prediction process; Although the deep learning-based models can extract the spatio-temporal correlations of passenger flow data to a certain extent, their training often requires a large amount of data, and the existing data volume for new line prediction is too small to meet the needs of deep learning training. Summary of the Invention

[0007] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a new line passenger flow prediction method based on a semi-supervised Conv-Transformer deep learning model. The method includes:

[0008] For the target line, obtain POI data;

[0009] Input the POI data into the trained convolutional Transformer model to obtain a trend flow and a scale flow, where the scale quantity reflects the scale characteristics of the passenger flow at the station, and the trend flow reflects the trend characteristics of the passenger flow at the station. The convolutional Transformer model is trained using semi-supervised learning;

[0010] Combine the trend flow and the scale flow to obtain the passenger flow prediction result.

[0011] Compared with the prior art, the advantages of the present invention are as follows: an effective deep learning framework is introduced, which is based on semi-supervised learning and deep learning neural networks and is used to predict the inbound flow of new-line stations. The prediction of the inbound flow is decomposed into trend features and scale features, and the data set is expanded by means of data augmentation. Finally, a prediction model based on the attention mechanism and the deep attention module is used for prediction, improving the prediction effect of the new-line inbound flow. Verified by experiments, the present invention has better experimental performance and higher prediction accuracy.

[0012] Other features and advantages of the present invention will become clear from the following detailed description of exemplary embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings incorporated in and constituting a part of this specification illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.

[0014] Figure 1 is a flowchart of a new-line passenger flow prediction method based on a semi-supervised Conv-Transformer deep learning model according to an embodiment of the present invention;

[0015] Figure 2 is a schematic diagram of a semi-supervised deep learning prediction framework according to an embodiment of the present invention;

[0016] Figure 3 is a schematic diagram of the semi-supervised learning iteration process according to an embodiment of the present invention;

[0017] Figure 4 is a schematic diagram of the overall Conv-Transformer model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0018] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present invention.

[0019] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way serves as a limitation on the present invention, its application or use.

[0020] Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods and devices should be regarded as part of the specification.

[0021] In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0022] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof is not required in subsequent figures.

[0023] The present invention proposes a deep learning framework based on semi-supervised learning and attention mechanism, or called Semi-Conv-Transformer. This framework decomposes the inbound flow prediction task into trend flow prediction and scale quantity prediction, and completes the new line prediction task. This framework extracts the non-linear relationship between the spatial features of POIs and the inbound flow of the new line, and applies a data augmentation method based on semi-supervised learning to effectively solve the problems caused by insufficient data, thereby improving the prediction accuracy of the new line prediction, which can help the subway operator reasonably allocate resources, such as arranging appropriate train schedules, adjusting train timetables, and reasonably allocating personnel, so as to improve the operation efficiency and reduce the operation cost.

[0024] Generally speaking, the present invention divides the inbound flow prediction of the new line into three parts: using semi-supervised learning for data augmentation to expand the inbound flow dataset; using the convolutional Transformer (Conv-Transformer) model to predict the trend features and scale features after the inbound flow is decomposed respectively; combining the trend features and scale features of the inbound flow to obtain the required inbound flow prediction data.

[0025] See Figure 1 As shown, the provided new line passenger flow prediction method based on the semi-supervised Conv-Transformer deep learning model includes: step S110, for the target line, obtain POI data; step S120, input the POI data into the trained convolutional Transformer model to obtain the trend flow and the scale flow, where the scale quantity reflects the scale feature of the station passenger flow, and the trend flow reflects the trend feature of the station passenger flow, and the convolutional Transformer model is trained in a semi-supervised learning manner; step S130, combine the trend flow and the scale flow to obtain the passenger flow prediction result.

[0026] In the following, first, the scientific problems to be solved are defined in detail. Then, the overall model structure proposed is shown. Next, the semi-supervised learning and deep learning network used are described in detail. Finally, the AFC data of a new line opening in a certain urban rail transit network in China is used as the dataset for experiments, and compared with several other new line prediction models to verify the effectiveness and superiority of the present invention.

[0027] I. Problem Definition

[0028] The present invention provides relevant definitions for predicting the incoming flow of new lines.

[0029] Specifically, is expressed as the passenger flow of all N stations within all D time intervals in T days. The element represents the passenger flow of the i th station during the time interval j on the k th day.

[0030] The scale of a station can be used to describe the scale characteristics of the incoming flow of the station. It is obtained by adding up the incoming flow of the station in a day, and the definition of the scale is as follows:

[0031]

[0032] Among them, represents the scale of all N stations within all D days. The element represents the scale of the i th station on the j th day. The scale represents the passenger flow scale of the station in a day, which reflects the passenger flow scale characteristics of the station.

[0033] The trend flow of a station can be used to describe the trend characteristics of the station. It is obtained by dividing the incoming flow of the station in a day by the scale of the station on that day, and is expressed as:

[0034]

[0035] represents the trend flow of all N stations within all D days for all T time intervals. The element represents the trend flow of the i th station during the time interval j on the k th day. The trend flow represents the passenger flow trend of the station in a day, reflecting the passenger flow trend characteristics of the station.

[0036] In a geographic information system, a POI (Point of Interest) represents a data point of interest, and each POI data contains its name, code, location, and category information. For example, the number of POIs of different categories within a radius of 1000 meters around each station is counted. Therefore, can be expressed as all stations N in all categoriesC (The number of POIs in category C, representing how many categories of POIs there are) the number of POIs. Element Indicates the i th station's number of POIs in the j category.

[0037] In subsequent applications, POIs are the input of the model, and the scale quantity and trend flow are the outputs of the model. In the prediction of new line passenger flow, POIs can be used to represent the scale of different types of buildings and land use types near the station.

[0038] II. Model Structure

[0039] The prediction framework designed by the present invention is as shown in Figure 2 . After the given POI data, this framework adopts the method of semi-supervised learning for data augmentation, and then uses the augmented dataset to train the convolutional Transformer model. After obtaining the parameters of the convolutional Transformer model, the trend flow and scale quantity required for predicting the POIs in the test set are imported. The extended POI data is used as the training input of the convolutional Transformer model, while the extended passenger flow data is used as the training output.

[0040] The Semi-Conv-Transformer model needs to be used for scale quantity prediction and trend flow prediction respectively, and the scale quantity prediction and trend flow prediction are independent of each other. Each station site is regarded as a sample. In the following symbols, i indicates the i th station site. When using the Semi-Conv-Transformer architecture for scale quantity prediction, the input of the framework is a one-dimensional vector of POIs ( represents the scale POI of the i-th site, and C represents the total number of POI categories), and the output is a one-dimensional scale quantity vector (D represents the number of prediction days). When using the Semi-Conv-Transformer architecture for trend flow prediction, the input of the framework is a one-dimensional vector of trend POIs ( represents the trend POI of the i-th site), and the output is a one-dimensional vector of trend flow ( represents the one-dimensional trend flow vector of the i-th site. Subsequently, this vector will be integrated into a two-dimensional form as an intermediate vector) and then this one-dimensional vector is reshaped into a two-dimensional vector ( represents the trend flow vector of the i-th site, which is the final vector).

[0041] III. Semi-Supervised Learning

[0042] The semi-supervised learning process for expanding the dataset is an iterative process, as Figure 3 shown. In each iteration, samples with pseudo-labels are first generated, and then these samples with pseudo-labels are used to expand the original training set after the previous iteration. The original training set and the expanded training set are compared under a fixed validation set, and the training set will obtain corresponding evaluation metrics, which can reflect the performance of the training set. If the expanded training set performs better than the original training set in the fixed validation set, the expanded training set will be used as the training set for the next iteration. If the expanded training set performs worse than the original training set in the fixed validation set, the training set will remain unchanged. When the number of iterations reaches 1000 or the sample size of the dataset reaches the required scale, the iterative process ends.

[0043] The semi-supervised learning process is used to expand the dataset to make up for the problem of insufficient data training that occurs during the training of trend flow and scale data. Combining Figure 2 as shown, X n is POI data, and X w is the aggregated POI data of X n . X G is the set of pseudo-features generated from X n . X Gw is the aggregated pseudo-feature of X G . Y G1 and Y G2 are the pseudo-labels obtained after inputting X G and X Gw into the XGBoost model and the MLR model.

[0044] IV. Conv-Transformer Framework

[0045] In one embodiment, a Conv-Transformer model framework for passenger flow prediction is proposed. Refer to Figure 4 shown, where Figure 4 (a) is a schematic diagram of the multi-head attention mechanism, Figure 4 (b) is a schematic diagram of the Transformer layer, Figure 4 (c) is the Conv-Transformer framework. As Figure 4 (c) shown, the entire Conv-Transformer model includes N Transformer layers, a one-dimensional convolutional layer, and three fully connected layers (i.e., linear layers). The Transformer layers are used to extract the global features of the POI, the convolutional layer is used to extract the local features of the POI, and the fully connected layers are used to learn the complex relationship between the POI data and the passenger flow data after feature extraction. The final output uses residual skip connections to solve the problems of gradient explosion and gradient disappearance. The model uses the expanded dataset X e (Xe Take the extended POI dataset (denoted as) as input, and the corresponding extended dataset Y e (Y e represents the extended passenger flow dataset) as output.

[0046] The Transformer is used to process sequential data and extract non - linear features. This framework uses a re - defined Transformer layer. As shown in Figure 4 (b), the Transformer layer consists of three parts, namely three parallel one - dimensional convolutional layers, one multi - head self - attention layer, and one feed - forward layer. The input of the Transformer layer forms a residual connection with the output of the multi - head self - attention and the output of the feed - forward layer to eliminate the impact of gradient explosion and gradient vanishing. The one - dimensional convolutional layer uses three parallel convolutional blocks to convert the single - channel input data into three - channel data, initially realizing the extraction of local features of POIs. The multi - head self - attention layer outputs the three - channel data as single - channel data and extracts the global features of POIs. The multi - head attention mechanism is shown in Figure 4 (a). For each attention head h = {1, …, H}, the query matrix is defined as the key matrix and the value matrix where In particular, d m is a preset hyperparameter of the multi - head self - attention layer, Each attention head is defined as:

[0047]

[0048] where i represents the i - th Transformer layer. By concatenating each in each attention head the output after the attention module can be obtained Then, the multi - head self - attention layer uses a linear layer to align the hidden dimension with the backbone model and outputs to the feed - forward layer, where C is the length of the aligned one - dimensional vector. The feed - forward layer is a fully - connected layer and is also the last layer of a single Transformer layer. This layer keeps the dimension unchanged. The Transformer layer extracts the global features of POIs and outputs the extracted data to the convolutional layer, as shown in Figure 4 (c).

[0049] V. Evaluation

[0050] To further verify the effectiveness of the present invention, the used model was evaluated, including several parts: dataset, model configuration, evaluation metrics, baseline model, and experimental results.

[0051] 1. Dataset Description

[0052] The prediction of the new line inbound flow requires separate prediction of the trend flow and the scale. The prediction of the trend flow and the scale is independent of each other. The inbound flow dataset for the scale can be different from the inbound flow dataset used for the trend flow. For example, the AFC card data and the corresponding POI data of a certain urban rail transit are used as the original dataset.

[0053] The dataset for scale prediction is the AFC card data and POI data of the new line stations from December 2, 2019 to December 8, 2019, with a total of 20 stations. The dataset for trend flow prediction is the AFC card data and POI data of all stations from November 16, 2020 to November 22, 2020, with a total of 61 stations. The test dataset is the AFC card data and POI data from June 7, 2021 to June 13, 2021, with a total of 18 stations.

[0054] All AFC data covers seven days a week, from Monday to Sunday. The AFC card data will be converted into inbound flow, scale, and trend flow data. The time range of the inbound flow is from 6 am to 11 pm, and the time granularity is 60 minutes. Therefore, there are 17 time intervals per day. The dimension of the inbound flow for each station is (7, 17), representing the inbound flow for each time interval in 7 days. The dimension of the trend flow for each station is also (7, 17), representing the trend flow for each time interval in 7 days. The scale is a one-dimensional vector of length 7, representing the scale for each station in 7 days.

[0055] 2. Model Configuration

[0056] In the experiment, the model can be implemented using PyTorch. In the scale prediction task, the XGBoost-MLR model is adopted for semi-supervised learning extension. The main model for scale prediction takes POI as the input. The main model consists of a Transformer layer with three heads, where d m is set to 8. After passing through the Transformer layer, the output is input into a one-dimensional convolutional layer with five channels filled with 0s and a convolutional kernel size of 6. Then, through two fully connected layers, the output dimensions are 256 and 512 neurons respectively. The final output layer consists of 7 neurons. Residual skip connections connect the initial POI data with the final result to eliminate gradient explosion and gradient disappearance.

[0057] For the trend flow, the XGBoost-KNN model is adopted for semi-supervised learning extension. The trend flow model consists of a Transformer layer with three heads, where d mSet to 8. The trend POI is processed through the Transformer layer to obtain intermediate data. The intermediate data is input into a one-dimensional convolutional layer with five channels filled with 0s, and the convolutional kernel size is 6. Then, through three fully connected layers, the output dimensions are 256, 512, and 119 neurons respectively. The initial trend POI data is connected using residual skip connections and added to the output of the final fully connected layer to obtain the trend flow data. At this time, the trend flow has not been separated by day and needs to be reorganized by day to obtain the required 7×17 trend flow matrix.

[0058] 3. Evaluation Metrics

[0059] The root mean square error (RMSE), weighted mean absolute percentage error (WMAPE), and mean absolute error (MAE) are applied to evaluate the performance of the model. The definitions are as follows:

[0060]

[0061] where y i is the predicted value, is the true value, and m is the total length of the input sequence. In addition, the mean square error (MSE) can be used as the loss function for each model:

[0062]

[0063] where y i and represent the predicted value and the true value respectively, and m is the total length of the input sequence.

[0064] 4. Comparison with Benchmark Models

[0065] The benchmark models obtained from the experiments are divided into benchmark models for scale quantity prediction and trend flow prediction. The experimental results are shown in Table 1, where E-Transformer represents the model of the present invention.

[0066] Table 1: Comparison Results of Benchmark Models

[0067]

[0068] As can be seen from Table 1, compared with the existing models, the prediction effect of the Conv-Transformer model based on the augmented dataset is the best. This phenomenon indicates that the Conv-Transformer model that can extract the global characteristics of passenger flow can better meet the needs of new line prediction and thus obtain better prediction results.

[0069] (1) Benchmark Models for Scale Quantity

[0070] The scale quantity benchmark model includes a classifier model, a CNN neural network, a Res neural network, and a Transformer neural network.

[0071] The classifier model uses the k-means algorithm to cluster the inflow traffic of the training set, and then uses the POI and inflow traffic in the training set to train the SVM classification model. According to the POI of the test set, the category of the test set sites can be obtained.

[0072] The CNN neural network predicts the scale quantity, taking the scale POI as the input and the scale quantity as the output. The network includes a one-dimensional convolutional layer with 5 channels and 6 kernel sizes, and two fully connected layers with 256 and 512 neurons in the output dimension. The output dimension of the last layer is 7.

[0073] The Res neural network uses residual connections to obtain the scale quantity, which is added to the CNN neural network. The skip residual connection uses a fully connected layer with 7 neurons to align the dimensions.

[0074] The Transformer neural network takes the scale POI as the input and the scale quantity as the output. A Transformer layer is added between the input and the convolutional layer. The Transformer layer includes a Transformer layer with three heads, and d m is set to 8.

[0075] (2) Benchmark model of trend flow

[0076] The trend flow benchmark model includes a classifier model, a CNN neural network, a Res neural network, and a Transformer neural network.

[0077] The classifier model in the trend flow uses a method similar to that of the scale quantity to obtain the category of the test set sites.

[0078] The CNN neural network, Res neural network, and Transformer neural network benchmark models of the trend flow model are roughly the same as the corresponding benchmark models in the scale quantity model. The only difference is that the output dimension of the last layer is 119.

[0079] (3) Analysis of benchmark experiment results

[0080] From the results of the benchmark experiments, the classifier model based on machine learning has the worst experimental effect, which is because the machine learning method cannot extract the spatio-temporal correlation of passenger flow data well. At the same time, it can be seen that the model trained on the dataset expanded by semi-supervised learning has a better effect than the model trained on the original dataset, which indicates that data expansion has a good promotion effect on the prediction of the model. Finally, the prediction effect of the Conv-Transformer model based on the expanded dataset is the best, which indicates that the Conv-Transformer model that can extract the global features of passenger flow can better meet the needs of new line prediction.

[0081] 5. Experimental Results

[0082] (1) Ablation Experiments

[0083] To test whether there are redundant parts in the trend flow model and the scale quantity model, ablation experiments were conducted on these two models respectively. For the scale quantity ablation experiment, some modules in the scale quantity model were removed and their effects were tested, as shown in Table 2. Letters A to D correspond to Semi, CNN, Res, and Trans respectively.

[0084] Table 2: Results of Ablation Experiments

[0085]

[0086] For the trend flow ablation experiment, the same components as in the scale quantity ablation experiment were removed, and the meanings of the letters are the same as those in the scale quantity ablation experiment. The specific experimental details are as follows:

[0087] Semi-Conv-Transformer (A): Remove the semi-supervised module.

[0088] Semi-Conv-Transformer (B): Remove the convolutional layer.

[0089] Semi-Conv-Transformer (C): Remove the residual connection.

[0090] Semi-Conv-Transformer (D): Remove the Transformer layer.

[0091] According to the experimental results, it can be concluded that no matter which module is removed, the trend flow model and the scale quantity model will show worse performance. This phenomenon indicates that the corresponding module has a positive impact on improving the accuracy of new line prediction.

[0092] (2) Independent Prediction Experiments

[0093] Table 3 shows the advantages of separately predicting the trend flow and scale volume, where the performance of traditional methods is inferior to that of the in-bound flow prediction methods based on semi-supervised learning and deep learning. The effect of directly using the in-bound flow data as input and output for training is inferior to that of separately predicting the trend flow and scale volume.

[0094] Table 3: Experimental Results of Separate Prediction

[0095] RMSE MAE WMAPE Traditional model 323 163.47 1.62 Inflow model 130.56 84.07 0.83 Our model 98.06 64.13 0.63

[0096] As can be seen from Table 3, compared with the present invention, the prediction methods corresponding to the traditional models in the prior art and the method of directly predicting using the in-bound flow have poor prediction effects.

[0097] In summary, compared with the prior art, the present invention has the following advantages:

[0098] 1) In the new line prediction task, the present invention decomposes the new line passenger flow prediction into trend features and scale features, predicts them separately, and finally integrates them to achieve more accurate inflow prediction.

[0099] 2) In the new line prediction task, the present invention performs data augmentation on the data set through co-training, and solves the problem of insufficient data volume in the prediction of newly extended subway lines.

[0100] 3) The present invention uses the Conv-Transformer model based on convolutional neural network and attention mechanism to extract the non-linear features of passenger flow data, further improving the accuracy of in-bound flow prediction, which has important guiding significance for engineering practice.

[0101] 4) The present invention is particularly applicable to the prediction of in-bound flow of new lines, while the publicly disclosed technologies focus on the passenger flow prediction of the subway network and other traffic modes under normal conditions, or the short-term passenger flow prediction of urban rail transit in case of emergencies.

[0102] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0103] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0104] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0105] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.

[0106] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer - readable program instructions.

[0107] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions comprises a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0108] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0109] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. As is well known to those skilled in the art, implementations by hardware, by software, and by a combination of software and hardware are equivalent.

[0110] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A new line passenger flow prediction method based on a semi-supervised Conv-Transformer deep learning model, comprising the following steps: For the target route, obtain POI data; Inputting the POI data into a trained convolutional Transformer model to obtain a trend flow and a scale flow, wherein the scale reflects the scale characteristics of the passenger flow of the station, and the trend flow reflects the trend characteristics of the passenger flow of the station, and the convolutional Transformer model is trained using a semi-supervised learning method; The trend flow and the scale flow are combined to obtain a passenger flow prediction result.

2. The method according to claim 1, characterized in that The convolutional Transformer model includes a Transformer layer, a one-dimensional convolutional layer and a fully connected layer. The Transformer layer is used to extract the global features of POI data, the one-dimensional convolutional layer is used to extract the local features of POI data, the fully connected layer is used to learn the relationship between POI data and passenger flow data, and a residual jump connection is provided between the output and input of the convolutional Transformer model.

3. The method according to claim 2, characterized in that The Transformer layer includes a one-dimensional convolutional layer, a multi-head self-attention layer and a feedforward layer, and the input of the Transformer layer forms a residual connection with the output of the multi-head self-attention layer and the output of the feedforward layer. The one-dimensional convolutional layer is set to three parallel one-dimensional convolutional layers, which are used to convert the input single-channel data into three-channel data and extract local features of POI data. The multi-head self-attention layer is used to output the three-channel data into single-channel data and extract global features of POI data.

4. The method according to claim 3, characterized in that: In the multi-head self-attention layer, for each attention head h = {1, ..., H}, define the query matrix Key Matrix Sum Matrix Each attention head is defined as: in, d m is the preset hyperparameter of the multi-head self-attention layer, Represents the i-th Transformer layer, by concatenating each get The multi-head self-attention layer uses a linear layer to Align with the backbone model and output To the feed-forward layer, C is the length of the aligned one-dimensional vector.

5. The method according to claim 1, characterized in that Training the convolutional Transformer model includes: In each iteration, samples with pseudo labels are first generated, and then the samples with pseudo labels are used to expand the original training set after the previous iteration to obtain an expanded training set; Under the validation set, the performance of the original training set and the extended training set are compared based on the set evaluation index, and then the original training set or the extended training set is selected as the training set for the next iteration according to the comparison result.

6. The method according to claim 4, characterized in that The feed-forward layer is a fully connected layer.

7. The method according to claim 1, characterized in that The scale quantity is expressed as: in, represents the size of all N stations in all D days, the elements Represents the scale of the i-th station on the j-th day.

8. The method according to claim 1, characterized in that The trend flow is expressed as: in, represents the trend flow of all T time intervals for all N stations in all D days, with the elements represents the trend flow of the i-th station at the time interval k on the j-th day.

9. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer device comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.