Subway carriage seat availability prediction method and device and computer equipment
Through high-dimensional spatiotemporal feature mapping and random forest model, passenger behavior is analyzed in segments using camera data in subway cars, which solves the accuracy and transparency of subway car seat availability prediction, and improves subway carriage efficiency and passenger satisfaction.
Patent Information
- Application Number
- CN202510332978.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-22
AI Technical Summary
The passenger capacity distribution between subway train cars is uneven, and the existing methods cannot accurately predict the individual behavior of passengers in the car, resulting in poor passenger travel experience and low loading efficiency. The existing models are low in transparency and interpretability and insufficient trust.
Using a high-dimensional spatiotemporal feature mapping model and a random forest model, image data is obtained through multiple cameras, passenger behavior is analyzed in segments, dynamic prediction models are constructed, and subway seat availability is predicted.
Accurate modeling and dynamic prediction of the usability of subway seats is realized, the prediction accuracy and passenger experience are improved, and the transparency and interpretability of the model are enhanced.
Smart Images

Figure CN120354263A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent transportation, and particularly relates to a method, device and computer equipment for predicting the availability of seats in subway carriages. Background Art
[0002] The passenger capacity distribution between subway train carriages is uneven. On the one hand, it will affect the travel experience and satisfaction of passengers. On the other hand, it will harm the operation efficiency of the subway system. At present, the information service provided by subway operation for passengers only reflects the rough congestion level, lacking the specific representation of the availability of seats in the carriage. Existing methods generally obtain data from outside the train or the overall passengers, which cannot fully reflect the congestion experience of passengers in the car and ignore the consideration of the individual behaviors of passengers in the carriage. Existing models have low transparency and interpretability. For example, deep learning methods use "black box" algorithms, and the models are complex and difficult to understand. At the same time, there is also a lack of quantitative analysis of the importance of features and their interactions, which affects the trust of the model in practical applications. Summary of the Invention
[0003] The purpose of the present invention is to use a high-dimensional spatio-temporal feature mapping model to comprehensively analyze the dynamic behavior patterns of passengers to predict the availability of subway seats, which has the advantage of high prediction accuracy.
[0004] The first object of the present invention is to provide a method for predicting the availability of seats in subway carriages. The prediction method includes the following steps: Step S10: Receive image data sent by multiple cameras in each subway carriage during the operation between two adjacent stations; Step S20: Divide the operation period between two adjacent stations into three time series stages consisting of a departure stage, a mid-way operation stage, and a stage approaching the station, and based on the three time series stages and the image data of each subway carriage, obtain a plurality of acquisition data sets corresponding to each subway carriage during the operation between two adjacent stations. Each acquisition data set includes the number of passengers, the posture of each passenger in the departure stage, and a plurality of data subsets, where each data subset is used to represent N action features of a passenger in three time series stages respectively, and N is a positive integer; Step S30: Input all data subsets of each acquisition data set into a prediction model to obtain the behavior prediction results of each passenger in each subway carriage. The prediction model is a trained random forest model, and the behavior prediction results include getting off at the station and continuing to ride; Step S40: Obtain the number of seats in each subway carriage, and based on the number of seats in each subway carriage, the number of passengers in each subway carriage, the posture of each passenger in the departure stage, and the behavior prediction results of each passenger, predict the available number of seats in each subway carriage.
[0005] In a specific embodiment, the method for training the random forest model includes: Step 1, obtain a model training data set, where the model training data set includes a plurality of sub-data sets corresponding one by one to multiple subway carriages. Each sub-data set includes a plurality of data packets with the same number as the number of passengers in each subway carriage. Each data packet includes input data for representing M action features of a passenger in the three time series stages respectively and output data for representing the behavior result of a passenger after arriving at the station. There is a one-to-one correspondence between the input data and the output data in each data packet, where M is a positive integer and M is greater than N; Step 2, taking the subway carriage as a unit, split the model training data set into a training set and a test set; Step 3, first train the random forest model using the training set, and then verify the trained random forest model using the test set to obtain a trained random forest model.
[0006] In a specific embodiment, Step 3 includes: Step (a), based on the training set, train a random forest classifier using the Bootstrap resampling method to construct an initial random forest model including multiple decision trees; Step (b), perform the first training on the initial random forest model using the training set to obtain the best hyperparameter combination, and use the random forest model with the best hyperparameter combination as the optimized random forest model. Among them, the first training tries all possible hyperparameter combinations in the initial random forest model through grid search and uses K-fold cross-validation or OOB scoring to evaluate the performance of each hyperparameter combination to determine the best hyperparameter combination; Step (c), perform the second training on the optimized random forest model using the training set to improve the model accuracy to obtain a completely trained random forest model; Step (d), evaluate the completely trained random forest model using the test set. When the evaluation index meets the preset conditions, use the random forest model obtained in Step (c) as the trained random forest model. When the evaluation index does not meet the preset conditions, return to Step (b) to further optimize the hyperparameter combination. The evaluation index is selected from at least one of accuracy, recall rate, F1 score, and AUC.
[0007] In a specific embodiment, the obtaining steps of each data subset in Step S20 include: Step (1), based on the image data corresponding to the current subway carriage, obtain the action features of the current passenger in each time series stage in the current subway carriage; Step (2), compare the action features of the current passenger in each time series stage with the pre-determined N action features respectively to obtain the data subset corresponding to the current passenger. The data subset is a binary sequence and the total number of columns is 3N.
[0008] In a specific implementation manner, in step (2), the method for determining the N action features includes: based on the trained random forest model, using SHAP to evaluate to obtain the contribution value of each action feature among the M action features to the passenger getting-off result; based on the magnitude of the contribution value of each action feature to the passenger getting-off result, determining the N action features among the M action features, and the contribution value of each action feature among the N action features is greater than or equal to the first preset contribution value.
[0009] In a specific implementation manner, in step (2), the method for determining the N action features includes: based on the trained random forest model, using SHAP to evaluate to obtain the contribution value of each action feature among the M action features to the passenger getting-off result; based on the magnitude of the contribution value of each action feature to the passenger getting-off result, obtaining multiple target action features with positive contribution values among the M action features; combining the multiple target action features in pairs to obtain multiple action feature combinations; based on the trained random forest model, using SHAP to evaluate to obtain the contribution value of each action feature combination to the passenger getting-off result; based on the magnitude of the contribution value of each action feature combination to the passenger getting-off result, determining the N action features among the multiple target action features, and the N action features are composed of the action features in the action feature combinations whose contribution values are greater than the second preset contribution value.
[0010] In a specific implementation manner, the steps for obtaining the available number of seats in each subway car include: Step (1), based on the subway train system or manual statistics, obtaining the total seat capacity of the current subway car; Step (2), based on the number of passengers in the current car and the postures of each passenger at the departure stage, obtaining the number of standing passengers and the number of sitting passengers in the current car; Step (3), based on the postures of each passenger at the departure stage and the behavior prediction results of each passenger, obtaining the number of passengers expected to get off among the sitting passengers and the number of passengers expected to get off among the standing passengers in the current car; Step (4), calculating the available number of seats NSP in the current car according to the preset formula, and the preset formula is:
[0011]
[0012] where S is the total seat capacity of the current subway car, P is the number of passengers in the current subway car, P seated is the number of sitting passengers in the current car, P standing is the number of standing passengers in the current car, E is the predicted value of the number of passengers getting off in the current car, E seated is the number of passengers expected to get off among the sitting passengers, E standing is the number of passengers expected to get off among the standing passengers.
[0013] The second object of the present invention is to provide a subway car seat availability prediction device, and the prediction device includes the following steps: a receiving module for receiving image data sent by a plurality of cameras in each subway car during the operation between two adjacent stations; an obtaining module for dividing the operation period between two adjacent stations into three time sequence stages consisting of a departure stage, a mid-operation stage, and a stage approaching the station, and obtaining a plurality of acquisition data sets corresponding to each subway car during the operation between two adjacent stations based on the image data of each subway car, each of the acquisition data sets including the number of passengers, the postures of each passenger in the departure stage, and a plurality of data subsets, wherein each data subset is used to represent N action characteristics of a passenger in three time sequence stages respectively, and N is a positive integer; a first prediction module for inputting all the data subsets of each acquisition data set into a prediction model to obtain a behavior prediction result of each passenger in each subway car, the prediction model being a trained random forest model, and the behavior prediction result including getting off at the station and continuing to ride; a second prediction module for obtaining the number of seats in each subway car, and predicting the available number of seats in each subway car based on the number of seats in each subway car, the number of passengers in each subway car, the postures of each passenger in the departure stage, and the behavior prediction result of each passenger.
[0014] The third object of the present invention is to provide a computer device, and the computer device includes: a processor and a memory, wherein a computer program is stored in the memory, and when the processor executes the computer program, the computer device implements the method described above.
[0015] The fourth object of the present invention is to provide a computer-readable storage medium, and when the instructions in the computer-readable storage medium are executed by a processor, the processor executes the method described above.
[0016] The beneficial effects of the present invention at least include:
[0017] The present invention provides a method for predicting the availability of seats in a subway car, and the prediction method includes the following steps: Step S10, receiving image data sent by multiple cameras in each subway car during the operation between two adjacent stations; Step S20, dividing the operation period between the two adjacent stations into three time series stages consisting of a departure stage, a mid-operation stage, and a stage approaching the station, and based on the three time series stages and the image data of each subway car, obtaining a plurality of acquisition data sets corresponding to each subway car during the operation between the two adjacent stations, each of the acquisition data sets including the number of passengers, the postures of each passenger in the departure stage, and a plurality of data subsets, wherein each data subset is used to represent N action features of a passenger in each of the three time series stages, and N is a positive integer; Step S30, inputting all the data subsets of each acquisition data set into a prediction model to obtain the behavior prediction results of each passenger in each subway car, the prediction model being a trained random forest model, and the behavior prediction results including getting off at the station and continuing to ride; Step S40, obtaining the number of seats in each subway car, and predicting the available number of seats in each subway car based on the number of seats in each subway car, the number of passengers in each subway car, the postures of each passenger in the departure stage, and the behavior prediction results of each passenger; The prediction method provided by the present invention constructs a high-dimensional spatio-temporal feature mapping model, segments the passenger behavior into a departure stage, a mid-operation stage, and a stage approaching the station, extracts action features, and combines with a random forest model driven by information entropy for prediction, realizing accurate modeling and dynamic prediction from the actions of passengers to the state of seat resources, and is applicable to the seat resource management in an intelligent rail transit system.
[0018] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The following will refer to the drawings for a further detailed description of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic flow chart of the steps of the method for predicting the availability of seats in a subway car provided by an embodiment of the present invention;
[0020] Figure 2 It is the ROC curve of the training set and the test set corresponding to the optimized random forest model;
[0021] Figure 3 It is the learning curve of the training set and the validation set corresponding to the optimized random forest model;
[0022] Figure 4 It is the validation curve of the training set and the validation set corresponding to the optimized random forest model;
[0023] Figure 5 It is the performance of the training set of the prediction model provided by the present invention;
[0024] Figure 6 The performance of the test set of the prediction model provided by the present invention;
[0025] Figure 7 The module diagram of the subway car seat availability prediction device provided by an embodiment of the present invention. Detailed implementation manners
[0026] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0027] Please refer to Figure 1 , according to the first aspect of the present invention, the present invention provides a method for predicting the availability of seats in a subway car, and the prediction method includes the following steps:
[0028] Step S10: Receive image data sent by multiple cameras in each subway car during the operation between two adjacent stations.
[0029] In the present invention, image data is obtained through multiple cameras installed in the subway car. The number of cameras can be considered to be large enough and the installation positions are diverse to ensure that the postures and movement characteristics of each passenger in the subway car can be collected.
[0030] Step S20: Divide the operation period between two adjacent stations into three time sequence stages consisting of a departure stage, a mid-operation stage, and an approaching station stage, and based on the three time sequence stages and the image data of each subway car, obtain a plurality of acquisition data sets corresponding to each subway car during the operation between two adjacent stations. Each acquisition data set includes the number of passengers, the posture of each passenger in the departure stage, and a plurality of data subsets, where each data subset is used to represent N action characteristics of a passenger in three time sequence stages respectively, and N is a positive integer.
[0031] In the present invention, the start time during the operation between two adjacent stations refers to the moment when the subway departs from the departure station, and the end time refers to the moment when the subway stops running at the destination station. The departure station and the destination station are the stations of two adjacent stations. For easy understanding, by way of example, the two adjacent stations are Station A and Station B respectively, and the operation period between two adjacent stations refers to the moment when the subway departs from Station A to the moment when the subway stops at Station B.
[0032] Preferably, the durations of the subway departure stage, the mid-operation stage, and the approaching station stage are all the same. It can be understood that if the operation time between two adjacent stations is 9 minutes, then the durations of the subway departure stage, the mid-operation stage, and the approaching station stage are all 3 minutes.
[0033] In an embodiment of the present invention, the durations of the subway departure stage, the mid - journey operation stage, and the approaching - station stage are the same. In other embodiments, the duration of the subway departure stage can also be set to 1 minute, 2 minutes, 3 minutes, etc., and the duration of the approaching - station stage can be set to 1 minute, 2 minutes, 3 minutes, etc.
[0034] In the present invention, the posture of each passenger in the subway departure stage is the posture of each passenger at the last moment of the subway departure stage. For ease of understanding, by way of example, assume that the subway departs from Station A at 9:00 and arrives at Station B at 9:09. Then, from 9:00 to 9:03 is the subway departure stage, and the posture of each passenger in the subway departure stage corresponds to the posture of each passenger at the last second of 9:03.
[0035] In the present invention, the postures of the passengers include standing and sitting postures.
[0036] In an alternative embodiment, the steps for obtaining each data subset in step S20 include:
[0037] Step (1): Based on the image data corresponding to the current subway carriage, obtain the action features of the current passengers in each time - series stage in the current subway carriage.
[0038] Step (2): Compare the action features of the current passengers in each time - series stage with N pre - determined action features respectively to obtain the data subset corresponding to the current passengers. The data subset is a binary sequence and the total number of columns is 3N.
[0039] In the present invention, the binary sequence is composed of a row of binary - digit data.
[0040] For ease of understanding, by way of example, assume that N is 5, and the 5 action features are A1, A2, A3, A4, A5. The data in this row includes 15 columns, and the specific arrangement is: A1, A2, A3, A4, A5, A1, A2, A3, A4, A5, A1, A2, A3, A4, A5. Among them, the 1st to 5th columns are used to represent the action features of passengers in the subway departure stage, the 6th to 10th columns are used to represent the action features of passengers in the subway operation stage, and the 11th to 15th columns are used to represent the action features of passengers in the approaching - station stage. If in the subway departure stage, the action features of the passengers include A1 and A3, in the subway operation stage, the action feature of the passengers is A5, and in the approaching - station stage, the action features of the passengers include A2 and A3, then the data in this row is specifically: 1 0 1 0 0 0 0 0 0 1 0 1 1 0 0, where 1 indicates that the passenger has the corresponding action feature and 0 indicates that the passenger does not have the corresponding action feature.
[0041] In an alternative embodiment, the N action features include spatial movement features, device interaction features, and interpersonal interaction features.
[0042] In an alternative embodiment, the N action features include spatial movement features and device interaction features.
[0043] In an alternative embodiment, the spatial movement features include moving forward towards the interior of the carriage, turning around, moving towards the carriage exit, sitting down, and getting up.
[0044] In an alternative embodiment, the device interaction features include placing luggage, picking up luggage, putting the phone back, using the phone, reading, watching the information display screen, making a call, and sleeping.
[0045] In an alternative embodiment, the interpersonal interaction features include talking to a companion and leaning against a companion.
[0046] In the present invention, the above-mentioned spatial movement features only list some of the spatial movement features, the above-mentioned device interaction features only list some of the device interaction features, and the above-mentioned interpersonal interaction features only list some of the interpersonal interaction features, including but not limited to the features listed above.
[0047] In an alternative embodiment, the N action features are determined based on the contribution values of the spatial movement features, device interaction features, and interpersonal interaction features to the result of the passenger getting off the train.
[0048] In the present invention, based on the trained random forest model, SHAP is used to evaluate and obtain the contribution value of each action feature to the result of the passenger getting off the train or the contribution value of the combination of multiple action features to the result of the passenger getting off the train.
[0049] More preferably, the N action features include moving towards the carriage exit, turning around, getting up, picking up luggage, and putting the phone back.
[0050] More preferably, the N action features include moving towards the carriage exit, turning around, getting up, picking up luggage, using the phone, putting the phone back, and watching the information display screen.
[0051] Step S30: Input all data subsets of each of the collected data sets into the prediction model to obtain the behavior prediction results of each passenger in each subway carriage. The prediction model is a trained random forest model, and the behavior prediction results include getting off the train and continuing to ride.
[0052] In the present invention, if the number of passengers in each subway carriage is P, then the data input into the random forest model is a matrix composed of P rows × 3N columns of data, and each row of data is composed of binary bit data, that is, a binary sequence.
[0053] In the present invention, the output of the random forest model is whether passenger i gets off at the station or continues to ride. By predicting whether each passenger gets off at the station, the number of passengers getting off in each subway car can be obtained, and thus the seat availability can be predicted based on the number of passengers getting off.
[0054] In an alternative embodiment, the training method of the random forest model includes the following steps:
[0055] Step 1: Obtain a model training data set. The model training data set includes a plurality of sub-data sets corresponding one-to-one to multiple subway cars. Each sub-data set includes a plurality of data packets with the same number as the number of passengers in each subway car. Each data packet includes input data representing M action features of a passenger in the three time series stages respectively and output data representing the behavior result of a passenger after arriving at the station. The input data and the output data in each data packet have a one-to-one correspondence, where M is a positive integer and M is greater than N.
[0056] In the present invention, the M action features include spatial movement features, device interaction features, and interpersonal interaction features. The spatial movement features include moving forward towards the interior of the car, turning around, moving towards the car exit, taking a seat, and getting up. The device interaction features include placing luggage, picking up luggage, putting the phone back, using the phone, reading, watching the information display screen, making a call, sleeping. The interpersonal interaction features include talking to a companion and leaning against a companion.
[0057] In the present invention, the M action features include 14 action features: moving forward towards the interior of the car, turning around, moving towards the car exit, taking a seat, getting up, placing luggage, picking up luggage, putting the phone back, using the phone, reading, watching the information display screen, making a call, sleeping, and talking to a companion.
[0058] In the present invention, the input data can be understood as the X value, and the output data can be understood as the Y value. Each data packet is composed of the X value and the Y value, and the X value and the Y value have a corresponding relationship.
[0059] In the present invention, the X value is a binary sequence composed of 0 and 1, and the length of the sequence is 3M; the Y value is 0 or 1.
[0060] In the present invention, each sub-data set can be represented by a matrix. Then, the number of rows of the sub-data set is the same as the number of passengers, and the number of columns is 3M + 1.
[0061] For easy understanding, by way of example, assume that in the departure stage, passenger No. 01 has only the action feature M2 among the M action features. Then, among the M values corresponding to the departure stage in the binary sequence corresponding to passenger No. 01, only the element corresponding to the action feature M2 takes the value 1, and the elements corresponding to the other M - 1 action features all take the value 0.
[0062] It is understandable that the X value is the same as the data subset exemplified above, except that the number of action features may be different. Therefore, the acquisition of the X value can also refer to the example description of the data subset above.
[0063] In the present invention, when the Y value is 0, it indicates that the passenger continues to ride, and when the Y value is 1, it indicates that the passenger gets off at the station.
[0064] The data of three data packets are shown in tabular form. Please refer to Table 1.
[0065] Table 1 shows the three data packets in tabular form
[0066]
[0067] Step 2: Split the model training data set into a training set and a test set in units of subway carriages.
[0068] In the present invention, when dividing the training data set and the test data set, the meaning of taking the subway carriage as a unit is that all the data packets corresponding to the same subway carriage either belong to the training data set or belong to the test data set, that is, all the data packets corresponding to the same subway carriage will not be split into the training data set and the test data set.
[0069] For easy understanding, an example is given. Suppose the number of subway carriages is 10, and the number of passengers in each subway carriage is the same. Based on the training data set and the test data set being divided in a ratio of 8:2, then randomly select 8 sub-data sets corresponding to 8 subway carriages as the training data set, and the other 2 sub-data sets corresponding to 2 subway carriages as the test data set.
[0070] Suppose the number of subway carriages is 10, and the number of passengers in each subway carriage is different. Based on the training data set and the test data set being divided in a ratio of 8:2, then select A sub-data sets corresponding to A subway carriages as the training data set, and the other B sub-data sets corresponding to B subway carriages as the test data set, A + B = 10. Then the ratio of the number of data packets included in the A sub-data sets to the number of data packets included in the B sub-data sets is approximately equal to 8:2.
[0071] It is understandable that because it is in units of subway carriages and the number of passengers in each subway carriage is different, it is relatively difficult to ensure that the number of data packets in the training data set and the number of data packets in the test data set are exactly 8:2. It can only be ensured that when dividing, the ratio of the two is made as close to 8:2 as possible.
[0072] Step 3: Use the training set to train the random forest model, and use the test set to test the trained random forest model to obtain the trained random forest model.
[0073] In an alternative embodiment, step three includes:
[0074] Step (a), based on the training set, training a random forest classifier using the Bootstrap resampling method to construct an initial random forest model including multiple decision trees.
[0075] In the present invention, the number of decision trees is 100 to 200.
[0076] Step (b), using the training set to perform the first training on the initial random forest model, obtaining the best hyperparameter combination, and using the random forest model with the best hyperparameter combination as the optimized random forest model. Among them, the first training attempts all possible hyperparameter combinations in the initial random forest model through grid search, and uses K-fold cross-validation or OOB scoring to evaluate the performance of each hyperparameter combination to determine the best hyperparameter combination.
[0077] In the present invention, the metrics obtained by K-fold cross-validation include accuracy, F1, and AUC, etc.
[0078] In the present invention, the best hyperparameter combination has optimal K-fold cross-validation metrics and a controllable overfitting degree.
[0079] In the present invention, 10-fold cross-validation is used to evaluate the performance of each hyperparameter combination.
[0080] In this embodiment, step (b) uses K-fold cross-validation to evaluate the performance of each hyperparameter combination. In other embodiments, step (b) can also use OOB scoring to evaluate the performance of each hyperparameter combination. The OOB scoring is specifically: during the training process, automatically estimate the data not drawn corresponding to each decision tree, compare the predicted output and the true label, and calculate metrics such as accuracy, recall, and F1.
[0081] In the present invention, the hyperparameter combination of the optimized random forest model is shown in Table 2 in detail.
[0082] Table 2 Hyperparameter combination of the optimized random forest model
[0083]
[0084] In the present invention, the prediction performance of the optimized random forest model is shown in Figures 2 to 4 , where Figure 2 is the ROC curve of the training set and the test set, Figure 3 is the learning curve corresponding to the training set and the validation set, Figure 4 is the validation curve corresponding to the training set and the validation set. From Figures 2 to 4It can be seen that the optimized random forest model has good prediction performance.
[0085] It should be noted that in the present invention, the validation set is part of the data in the training set.
[0086] Step (c): Use the training set to perform a second training on the optimized random forest model to improve the model accuracy, and obtain a trained random forest model.
[0087] In step (b), since the K-fold cross-validation method is adopted and only part of the data in the training set is used in the first training, after obtaining the best parameter combination, using all the data in the training set to perform a second training on the optimized random forest model can further improve the accuracy of the model.
[0088] Step (d): Use the test set to evaluate the trained random forest model. When the evaluation index meets the preset conditions, use the random forest model obtained in step (c) as the trained random forest model. When the evaluation index does not meet the preset conditions, return to step (b) to further optimize the hyperparameter combination, where the evaluation index is selected from at least one of accuracy, recall rate, F1 score, and AUC.
[0089] In the present invention, that the evaluation index meets the preset conditions means that the accuracy in the evaluation index is greater than or equal to the preset accuracy, and / or the recall rate is greater than or equal to the preset recall rate, and / or the F1 score is greater than or equal to the preset F1 score, and / or the AUC is greater than or equal to the preset AUC value.
[0090] It should be noted that the accuracy, recall rate, F1 score (the harmonic mean of precision and recall rate), and AUC (the area under the ROC curve) are all calculated using the formulas disclosed in the prior art.
[0091] Step S40: Obtain the number of seats in each subway car, and predict the available number of seats in each subway car based on the number of seats in each subway car, the number of passengers in each subway car, the posture of each passenger at the departure stage, and the prediction result of each passenger's behavior.
[0092] It can be understood that the availability of seats is related to the available number of seats. When a seat is unavailable, the available number of seats does not increase. When a seat is available, the available number of seats is incremented by 1.
[0093] In an optional implementation manner, the steps for obtaining the available number of seats in each subway car include:
[0094] Step (1): Based on the subway train type or manual statistics, obtain the number of seats in the current subway car.
[0095] It is understandable that the current subway car is any one of the multiple subway cars of the subway, that is, it can be one of the first subway car, the second subway car... the last subway car.
[0096] In the present invention, the subway train type is mainly type B. According to on-site observation and the actual seat facility distribution of a single type B car, the maximum total seat capacity set for a single car is 40 people.
[0097] Step (2): Based on the number of passengers in the current car and the posture of each passenger at the departure stage, obtain the number of standing passengers and the number of sitting passengers in the current car.
[0098] In the present invention, the postures of passengers include standing and sitting postures.
[0099] Step (3): Based on the posture of each passenger at the departure stage in the current car and the behavior prediction result of each passenger, obtain the number of passengers expected to get off among the sitting passengers and the number of passengers expected to get off among the standing passengers in the current car.
[0100] Step (4): Calculate the available seat number NSP in the current car according to a preset formula.
[0101] The preset formula is:
[0102]
[0103] Among them, S is the total seat capacity of the current car, P is the number of passengers in the current car, P seated is the number of sitting passengers in the current car, P standing is the number of standing passengers in the current car, E is the predicted value of the number of passengers getting off in the current car, E seated is the number of passengers expected to get off among the sitting passengers, E standing is the number of passengers expected to get off among the standing passengers.
[0104] In actual situations, passengers who do not need to get off in the car can be given priority to obtain available seats, that is, the original standing passengers have priority over the passengers boarding at the next stop to obtain seats. At the same time, there may be situations where even when there are vacant seats, some passengers still choose to stand. This may be because these passengers are only a short distance from their destinations, so it is considered that these passengers will not choose to sit down even when there are more vacant seats available; except for these passengers, it is assumed that all standing passengers who do not need to get off have an equal chance to sit down when there are vacant seats.
[0105] In the present invention, when predicting how many seats will be available at the next stop, the relationship between the number of passengers P in the current carriage and the maximum seat capacity S needs to be considered separately. When P ≤ S, there are extra empty seats in addition to ensuring that all passengers have seats, that is, after some seated passengers get off, E seated the number of available seats will directly increase on the basis of the original empty seats. When P > S, the standing passengers who do not need to get off can sit down first. When the remaining empty seats are more than the number of standing passengers still in the carriage, the extra empty seats will provide more seat options for the passengers on the platform; when the number of passengers getting off at the next stop is less than the number of standing passengers in the current carriage, there will be no available seats for the passengers at downstream stops, and the index result will be displayed as 0.
[0106] In the present invention, the MAE index (Mean Absolute Error), MSE index (Mean Squared Error), and MAPE index (Mean Absolute Percentage Error) of the prediction model are shown in Table 3.
[0107] Table 3 Error analysis indicators for available seats and passengers getting off
[0108]
[0109]
[0110] As can be seen from Table 3, the average absolute difference between the predicted and actual available seats is 0.83 seats in the training set and 1.30 seats in the test set, which proves the reliability of the model. For the passenger getting off count, the MSE decreases from 5.94 in training to 3.40 in testing, indicating that the model captures key patterns more effectively on unseen data. The SMAPE for available seats slightly increases from 4.93% (training) to 6.93% (testing), and the passenger getting off count slightly increases from 6.28% to 8.14%, demonstrating a robust generalization performance.
[0111] Please refer to Figure 5 and Figure 6 , where Figure 5 is the performance of the training set of the prediction model provided by the present invention, Figure 6 is the performance of the test set of the prediction model provided by the present invention. From Figure 5 and Figure 6 it can be seen that the actual available seat data and the predicted available seat numbers are equal or have a small difference in most carriages, and only a small number of carriages have a large difference. These differences may be caused by the complexity of passenger behavior and the model's inability to account for factors such as different levels of transport congestion, individual passenger preferences, and journey duration.
[0112] In the present invention, the N action features in step S20 can be determined based on a trained random forest model. When training the random forest model, the collected data should preferably cover all the action features of each passenger. When predicting the behavior result of a passenger, only the action features with a relatively high correlation with the behavior prediction result of the passenger need to be obtained. In this way, the amount of image data processing can be reduced, and at the same time, the prediction accuracy can be improved.
[0113] In an alternative embodiment, the method for determining the N action features in step S20 includes:
[0114] Step 1: Based on the trained random forest model, use SHAP to evaluate and obtain the contribution value of each of the M action features to the result of the passenger getting off the vehicle.
[0115] Step 2: Based on the magnitude of the contribution value of each action feature to the result of the passenger getting off the vehicle, determine the N action features among the M action features, where the contribution value of each action feature in the N action features is greater than or equal to a first preset contribution value.
[0116] The first preset contribution value can be 0, 0.05, 0.08, 0.1, 0.2, etc.
[0117] In the present invention, the M action features include walking forward facing the interior of the carriage, turning around, moving towards the carriage exit, taking a seat, getting up, placing luggage, picking up luggage, putting the phone back, using the phone, reading, watching the information display screen, making a call, sleeping, and talking to a companion, a total of 14 action features.
[0118] When the first preset contribution value is 0.05, the N action features determined by the above determination method include moving towards the carriage exit, turning around, getting up, picking up luggage, and putting the phone back.
[0119] In the present invention, the Shapley value is used to quantify the contribution of each action feature to the result of the passenger getting off the vehicle.
[0120] In an alternative embodiment, the method for determining the N action features in step S20 includes:
[0121] Step 1: Based on the trained random forest model, use SHAP to evaluate and obtain the contribution value of each of the M action features to the result of the passenger getting off the vehicle.
[0122] Step 2: Based on the magnitude of the contribution value of each action feature to the result of the passenger getting off the vehicle, obtain multiple target action features with positive contribution values among the M action features.
[0123] Step 3: Combine the multiple target action features in pairs to obtain multiple action feature combinations.
[0124] Step 4: Based on the trained random forest model, SHAP is used to evaluate and obtain the contribution value of each action feature combination to the passenger getting-off result.
[0125] Step 5: Based on the magnitude of the contribution value of each action feature combination to the passenger getting-off result, the N action features are determined among multiple target action features, and the N action features are composed of action features in which the contribution value of the action feature combination is greater than the second preset contribution value.
[0126] The second preset contribution value can be 0.1, 0.15, 0.2, 0.25, 0.3, 0.5, etc.
[0127] In the present invention, the M action features include walking forward facing the interior of the carriage, turning around, moving towards the carriage exit, sitting down, getting up, placing luggage, picking up luggage, putting the mobile phone back, using the mobile phone, reading, watching the information display screen, making a call, sleeping, and talking with a companion, a total of 14 action features.
[0128] When the second preset contribution value is 0.1, the N action features determined by the above determination method include moving towards the carriage exit, turning around, getting up, picking up luggage, using the mobile phone, putting the mobile phone back, and watching the information display screen.
[0129] In the present invention, the interaction effect function is used to quantify the synergistic effect of features to obtain the contribution value of the action feature combination.
[0130] Please refer to Figure 7, according to the second aspect of the present invention, the present invention also provides a subway car seat availability prediction device 100. The prediction device 100 includes the following steps: a receiving module 101 for receiving image data sent by a plurality of cameras in each subway car during the operation between two adjacent stations; an obtaining module 102 for dividing the operation period between two adjacent stations into three time series phases consisting of a departure phase, a mid-operation phase, and an approaching station phase, and obtaining a plurality of acquisition data sets corresponding to each subway car during the operation between two adjacent stations based on the image data of each subway car. Each acquisition data set includes the number of passengers, the posture of each passenger in the departure phase, and a plurality of data subsets, where each data subset is used to represent N action features of a passenger in each of the three time series phases, and N is a positive integer; a first prediction module 103 for inputting all the data subsets of each acquisition data set into a prediction model to obtain the behavior prediction results of each passenger in each subway car. The prediction model is a trained random forest model, and the behavior prediction results include getting off at the station and continuing to ride; a second prediction module 104 for obtaining the number of seats in each subway car and predicting the available number of seats in each subway car based on the number of seats in each subway car, the number of passengers in each subway car, the posture of each passenger in the departure phase, and the behavior prediction results of each passenger.
[0131] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described device and modules can refer to the content in the foregoing method embodiments and will not be repeated here.
[0132] According to the third aspect of the present invention, the present invention also provides a computer device Figure 7 The described prediction device can be arranged in this computer device. The computer device includes a processor and a memory. A computer program is stored in the memory. When the processor executes the computer program, the computer device implements the method described above in the claims. The computer device further includes a multimedia component, a memory, a processor, a communication interface, and a bus. Among them, the memory, the processor, and the communication interface are communicatively connected to each other through the bus. And this computer device can include multiple processors to facilitate the implementation of the functions of the above different modules by different processors.
[0133] Among them, the processor is used to control the overall operation of the subway car seat availability prediction device to complete all or part of the steps in the above subway car seat availability prediction method.
[0134] The memory is used to store various types of data to support the operation of the subway car seat availability prediction device. The stored data includes instructions for any application or method operating on the subway car seat availability prediction device, as well as application-related data such as contact data, sent and received messages, pictures, audio, video, and so on. The memory includes any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0135] The multimedia component includes a screen and an audio component. The screen is preferably a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component includes a microphone for receiving external audio signals. The received audio signals are further stored in the memory or sent via the communication component. The audio component also includes at least one speaker for outputting audio signals. The I / O interface provides an interface between the processor and other interface modules, and the above other interface modules include a keyboard, a mouse, and buttons. The buttons are virtual buttons or physical buttons. The communication component is used for wired or wireless communication between the subway car seat availability prediction device and other devices. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G, or 4G, or a combination of one or more of them. Accordingly, the communication component includes: a Wi-Fi module, a Bluetooth module, and an NFC module.
[0136] The communication interface uses a transceiver module such as, but not limited to, a transceiver to implement communication between the computer device and other devices or communication networks. For example, the communication interface can be any one or any combination of the following devices: a network interface (such as an Ethernet interface), a wireless network card, and other devices with network access functions.
[0137] The bus may include a path for transmitting information between various components of the computer device (such as the memory, the processor, and the communication interface).
[0138] A communication path is established between each of the above computer devices through a communication network. Each computer device is used to implement part of the functions of the prediction method provided in the embodiments of the present application. Any computer device can be a computer device in a cloud data center (e.g., a server), or a computer device in an edge data center, etc.
[0139] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product for providing the data synchronization cloud service includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer device, the processes or functions of the prediction method provided in the embodiments of the present application are implemented in whole or in part.
[0140] The computer device can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium stores computer program instructions for providing the data synchronization cloud service.
[0141] Optionally, the subway car seat availability prediction device can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above subway car seat availability prediction method.
[0142] According to the fourth aspect of the present invention, this embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the subway car seat availability prediction method are implemented.
[0143] The readable storage medium is a readable storage medium capable of storing program code, specifically a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0144] An embodiment of the present application also provides a computer program product containing instructions. When the computer program product runs on a computer, it causes the computer to execute the prediction method provided by the embodiment of the present application.
[0145] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0146] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field of the present invention, without departing from the concept of the present invention, several simple deductions and substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for predicting the availability of subway car seats, characterized in that, The prediction method includes the following steps: Step S10: Receive image data sent by multiple cameras in each subway car during the operation between two adjacent stations; Step S20: Divide the operation period between the two adjacent stations into three time series stages consisting of a departure stage, an intermediate operation stage, and a stage approaching the station, and based on the three time series stages and the image data of each subway car, obtain multiple acquisition data sets corresponding one by one to each subway car during the operation between the two adjacent stations. Each acquisition data set includes the number of passengers, the postures of each passenger in the departure stage, and multiple data subsets, where each data subset is used to represent N action features of a passenger in the three time series stages respectively, and N is a positive integer; Step S30: Input all the data subsets of each acquisition data set into a prediction model to obtain the behavior prediction results of each passenger in each subway car. The prediction model is a trained random forest model, and the behavior prediction results include getting off at the station and continuing to ride; Step S40: Obtain the number of seats in each subway car, and based on the number of seats in each subway car, the number of passengers in each subway car, the postures of each passenger in the departure stage, and the behavior prediction results of each passenger, predict the available number of seats in each subway car.
2. The subway car seat availability prediction method according to claim 1, wherein The training method of the random forest model includes: Step 1: Obtain a model training data set. The model training data set includes multiple sub-data sets corresponding one by one to multiple subway cars. Each sub-data set includes multiple data packets with the same number as the number of passengers in each subway car. Each data packet includes input data for representing M action features of a passenger in the three time series stages respectively and output data for representing the behavior result of a passenger after arriving at the station. There is a one-to-one correspondence between the input data and the output data in each data packet, where M is a positive integer and M is greater than N; Step 2: Split the model training data set into a training set and a test set in units of subway cars; Step 3: Use the training set to train the random forest model, and use the test set to evaluate the trained random forest model to obtain a trained random forest model.
3. The subway car seat availability prediction method according to claim 2, wherein The following is included in Step 3: Step (a): Based on the training set, train a random forest classifier using the Bootstrap resampling method to construct an initial random forest model including multiple decision trees; Step (b): Use the training set to perform the first training on the initial random forest model to obtain the best hyperparameter combination, and use the random forest model with the best hyperparameter combination as the optimized random forest model. Among them, the first training tries all possible hyperparameter combinations in the initial random forest model through grid search, and uses K-fold cross-validation or OOB scoring to evaluate the performance of each hyperparameter combination to determine the best hyperparameter combination; Step (c): Use the training set to perform the second training on the optimized random forest model to improve the model accuracy and obtain a completed trained random forest model; Step (d), evaluate the trained random forest model using the test set. When the evaluation metrics meet the preset conditions, use the random forest model obtained in step (c) as the trained random forest model. When the evaluation metrics do not meet the preset conditions, return to step (b) to further optimize the hyperparameter combination. The evaluation metrics are selected from at least one of accuracy, recall, F1-score, and AUC.
4. The subway car seat availability prediction method according to claim 2 or 3, characterized in that The obtaining steps of each data subset in step S20 include: Step (1), based on the image data corresponding to the current subway car, obtain the action features of the current passenger in each time series stage in the current subway car; Step (2), compare the action features of the current passenger in each time series stage with N predetermined action features respectively, and obtain the data subset corresponding to the current passenger. The data subset is a binary sequence and the total number of columns is 3N.
5. The subway car seat availability prediction method according to claim 4, characterized in that, In step (2), the determination method of the N action features includes: Based on the trained random forest model, use SHAP to evaluate and obtain the contribution value of each of the M action features to the passenger getting off result; Based on the magnitude of the contribution value of each action feature to the passenger getting off result, determine the N action features among the M action features. The contribution value of each of the N action features is greater than or equal to the first preset contribution value.
6. The subway car seat availability prediction method according to claim 4, wherein In step (2), the determination method of the N action features includes: Based on the trained random forest model, use SHAP to evaluate and obtain the contribution value of each of the M action features to the passenger getting off result; Based on the magnitude of the contribution value of each action feature to the passenger getting off result, obtain multiple target action features with positive contribution values among the M action features; Combine the multiple target action features in pairs to obtain multiple action feature combinations; Based on the trained random forest model, use SHAP to evaluate and obtain the contribution value of each action feature combination to the passenger getting off result; Based on the magnitude of the contribution value of each action feature combination to the passenger getting off result, determine the N action features among the multiple target action features. The N action features are composed of the action features in the action feature combinations with contribution values greater than the second preset contribution value.
7. The subway car seat availability prediction method according to claim 1, characterized in that The obtaining steps of the available seat number in each subway car include: Step (1), based on the subway train type or manual statistics, obtain the total seat capacity of the current subway car; Step (2), based on the number of passengers in the current car and the posture of each passenger at the departure stage, obtain the number of standing passengers and the number of sitting passengers in the current car; Step (3), based on the posture of each passenger at the departure stage and the behavior prediction result of each passenger, obtain the number of passengers expected to get off among the sitting passengers and the number of passengers expected to get off among the standing passengers in the current car; Step (4), calculate the available seat number NSP in the current car according to the preset formula. The preset formula is: Among them, S is the total seat capacity of the current subway car, P is the number of passengers in the current subway car, P seated is the number of passengers sitting in the current car, P standing is the number of passengers standing in the current car, E is the predicted value of the number of passengers getting off at the current car, E seated is the number of passengers expected to get off among those sitting, E standing is the number of passengers expected to get off among those standing.
8. A subway car seat availability prediction device, characterized in that The prediction device includes the following steps: A receiving module, configured to receive image data sent by multiple cameras in each subway car during the operation between two adjacent stations; An acquisition module, configured to divide the operation period between two adjacent stations into three time sequence phases including a departure phase, an intermediate operation phase, and an approaching station phase, and acquire a plurality of acquisition data sets corresponding to each subway car during the operation period between two adjacent stations based on the image data of each subway car. Each of the acquisition data sets includes the number of passengers, the postures of each passenger in the departure phase, and a plurality of data subsets, where each data subset is used to represent N action features of a passenger in three time sequence phases respectively, and N is a positive integer; A first prediction module, configured to input all data subsets of each acquisition data set into a prediction model to obtain a behavior prediction result of each passenger in each subway car. The prediction model is a trained random forest model, and the behavior prediction result includes getting off at the station and continuing to ride; A second prediction module, configured to obtain the number of seats in each subway car, and predict the available number of seats in each subway car based on the number of seats in each subway car, the number of passengers in each subway car, the postures of each passenger in the departure phase, and the behavior prediction result of each passenger.
9. A computer device, characterized in that, The computer device includes: a processor and a memory. A computer program is stored in the memory. When the processor executes the computer program, the computer device implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor, the processor executes the method according to any one of claims 1 to 7.