Airport security check queuing number prediction method based on KAN and K-means clustering method

Through the KAN regression model and K-means clustering method, combined with multi-source data characteristics, a personalized sub-model library was constructed, which solved the problem of single features and poor abnormal adaptability in the prediction of the number of people queuing at airport security inspection, and achieved high precision and high robustness prediction effects.

CN120597237APending Publication Date: 2025-09-05BEIJING TECH & BUSINESS UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510977655.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The existing airport security inspection queuing prediction method lacks the ability to integrate multi-source data, is difficult to capture the interaction between complex features, and is insufficient to adapt to unexpected situations, resulting in low prediction accuracy and weak model generalization ability.

Method used

The KAN regression model and K-means clustering method were used to construct a multi-feature data set, and the KAN regression model was initially predicted and identified abnormalities. The data was divided into normal and abnormal data sets using K-means clustering, and XGBoost, RandomForest, LSTM, and CNN models were trained respectively to build a personalized sub-model library for accurate prediction.

Benefits of technology

It enhances the comprehensive perception of influencing factors, improves the adaptability to nonlinear and non-stationary scenarios, significantly improves the prediction accuracy and generalization capabilities of the model, and is better than the existing mainstream methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597237A_ABST
    Figure CN120597237A_ABST
Patent Text Reader

Abstract

The invention discloses an airport security check queuing people number prediction method based on a KAN and K-means clustering method, and relates to the field of airport security check queuing people number prediction, and the method comprises the steps: constructing multi-feature data, and carrying out the preprocessing; predicting each piece of preprocessed sample data through a KAN regression model, dividing normal and abnormal data sets, and training a KAN classification model to obtain a classifier; clustering the normal and abnormal data sets, training and testing each clustered cluster by using a plurality of models, and retaining the model with the best prediction effect; and classifying the prediction data by using a classifier, further determining a clustering center with the highest similarity of the prediction data, and performing prediction by using a corresponding prediction model to obtain a final security check queuing number prediction result. The method not only effectively improves the prediction precision of the number of queuing people in security check, but also provides important technical support for intelligent operation and maintenance, personnel scheduling optimization and dynamic resource configuration of an airport, and has remarkable social benefits and economic values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of airport security check queue number prediction, and more particularly to an airport security check queue number prediction method based on KAN and K-means clustering methods. Background Art

[0002] As a complex, capital-intensive operational system, airport management and operations rely heavily on accurate data. In particular, inaccurate long-term or short-term forecasts can lead to significant losses in security checkpoint management. Long-term forecast errors can mislead strategic decision-making and inhibit future growth; short-term forecast discrepancies can lead to disruptions in on-site resource scheduling, such as improper allocation of manpower and equipment. This can lead to passenger delays, reduced operational efficiency, increased operating costs, and even chain reactions such as flight delays and passenger dissatisfaction. Therefore, accurately predicting security checkpoint queues has become a crucial issue in intelligent airport management.

[0003] However, current airport security personnel scheduling and channel configuration often rely on managers' experience and judgment, lacking a scientific and effective forecasting basis. This experience-based management approach is passive and lagging in the face of volatile passenger flows and unexpected situations. Therefore, developing a set of technologies that can accurately predict security queues in real time would have significant practical significance and application value for improving airport operational efficiency, optimizing resource allocation, and strengthening security. Despite the emergence of numerous research results in traffic flow and crowd forecasting in recent years, and the experimental application of some algorithms to airport passenger flow forecasting, several deficiencies remain in the field of accurately predicting security queues. First, most existing methods rely on single time series models, such as ARIMA and traditional LSTM, which utilize only historical queue data for modeling. These methods lack the ability to integrate multiple data sources that influence security passenger flow, resulting in limited model input information and difficulty capturing the interactions between complex features. Second, existing research generally ignores "abnormal" data, such as sudden queue anomalies caused by holiday surges, sudden weather changes, and concentrated flight delays. Such data is often simply treated as noise and eliminated, resulting in the model's inability to respond to real-world emergencies. This modeling method not only reduces the prediction accuracy, but also limits the generalization and expansion capabilities of the model.

[0004] Therefore, how to develop a prediction method that can integrate multi-dimensional data features, identify and distinguish abnormal situations, and have high robustness and high precision to meet the security queue prediction needs in the complex dynamic environment of airports is an urgent problem that technical personnel in this field need to solve. Summary of the Invention

[0005] In view of this, the present invention provides a method for predicting the number of people in airport security queues based on KAN and K-means clustering methods to solve the problems existing in the current prediction of the number of people in security queues, such as single features, poor adaptability to anomalies and insufficient model generalization ability.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] The present invention discloses a method for predicting the number of people queuing for airport security inspection based on KAN and K-means clustering methods, comprising:

[0008] Step 1: Collect airport flight data, security check channel data, terminal traffic data, security check queue data, and weather data as the initial data set, and construct multi-feature data;

[0009] Step 2: Preprocessing each feature data in the multi-feature data to obtain a preprocessed data set;

[0010] Step 3: Predict each sample data in the preprocessed data set using the KAN regression model, classify the data with a prediction error of less than or equal to 10% as normal data, and classify the data with a prediction error of more than 10% as abnormal data, to obtain a normal data set and an abnormal data set;

[0011] Step 4: Using the sample data in the normal data set as positive samples and the sample data in the abnormal data set as negative samples, training the KAN classification model to obtain an anomaly detection classifier;

[0012] Step 5: Clustering the normal data set and the abnormal data set respectively by using the K-means clustering algorithm to obtain multiple clusters and determine the cluster center of each cluster;

[0013] Step 6: For each cluster divided after clustering, use XGBoost, RandomForest, LSTM, and CNN models for training and testing respectively, and retain the model with the best prediction effect as the optimal prediction model for the cluster;

[0014] Step 7: Obtain the predicted data and classify it using the anomaly detection classifier. Based on the classification results, compare the predicted data with the cluster centers of the normal data set or the abnormal data set, determine the cluster center with the highest similarity, and use the prediction model trained by the corresponding cluster cluster to perform prediction to obtain the final prediction result of the number of people in the security check queue.

[0015] Furthermore, the step 1 specifically includes:

[0016] For each data item in the initial data set at any time, extract data slices at different times in the past 24 hours, and perform statistics on each data slice as numerical statistical features; use the number of security checkpoints per minute in the past 24 hours as a time series to extract time series features; and use the number of security checkpoint queues in the next 30 minutes as the prediction target;

[0017] The numerical statistical features and the time series features of each data at any moment are used as sample features, and the prediction target is used as a sample label to obtain a sample data; the initial data set is sampled and features are extracted at fixed time intervals to obtain a plurality of sample data as the multi-feature data.

[0018] Furthermore, the preprocessing includes missing value processing and data standardization processing;

[0019] The missing value processing specifically includes: for samples with half or more features missing, directly delete them; for samples with less than half of the features missing, use the feature mean imputation method to fill in the missing values ​​of each feature in the multi-feature data, and fill in the missing values ​​with the mean value of the corresponding feature data within the adjacent 2 hours;

[0020] The standardization process specifically includes: using a maximum and minimum value standardization method to scale the value of each feature in the multi-feature data to the interval [0, 1].

[0021] Furthermore, the prediction process of the KAN regression model specifically includes:

[0022] Convert the sample features into feature vectors and input them into the KAN regression model;

[0023] The KAN regression model transforms all features in the feature vector through a univariate nonlinear transformation function to obtain a new feature vector;

[0024] Then the KAN regression model performs a linear weighted combination on the new feature vectors to obtain several intermediate combination results;

[0025] Then, the corresponding output function is applied to each intermediate combination result, and the final prediction result is obtained after summing up.

[0026] Furthermore, the KAN classification model completes the binary classification task by learning the difference between normal samples and abnormal samples in the feature space. The anomaly detection classifier f classify (x i ) is expressed as:

[0027]

[0028] Among them, x iRepresents the characteristics of the i-th sample. The output value is 1, which means that the sample is a normal sample. The output value is 0, which means that the sample is an abnormal sample. normal represents the normal sample set, x anomalous Represents an abnormal sample set.

[0029] Furthermore, in step 5, when clustering the normal data set and the abnormal data set, the goal of the K-means clustering algorithm is to minimize the intra-class squared error. The formula is:

[0030]

[0031] Among them, k is the number of clusters, C l represents the lth cluster, μ l is the center of the lth cluster, and x is the sample point vector;

[0032] The specific formula for determining the cluster center of each cluster is:

[0033]

[0034] Among them, |C l | represents the number of samples in the lth cluster, It represents the sum of the vectors of all samples in the cluster.

[0035] It should be noted that although the objective function includes the cluster center μ l , while the cluster center itself depends on the current sample cluster assignment, but this does not constitute a logical contradiction. K-means uses an iterative optimization method to solve this problem: in each iteration, the algorithm first fixes the cluster center, calculates the distance from each sample to each center, and updates the sample cluster label accordingly; then, based on the new sample clustering results, recalculates the position of each cluster center. This process continues until the objective function Convergence. It can be seen that the objective function and the cluster center calculation formula together constitute the core optimization process of the K-means algorithm. The two are solved alternately through an iterative process and are logically self-consistent.

[0036] Finally, after clustering is completed, the normal data set and the abnormal data set can be represented as multiple subclusters and corresponding cluster centers respectively:

[0037]

[0038] Among them, x normal and x anomalous are normal data sets and abnormal data sets respectively, k and m are the number of clusters of normal samples and abnormal samples respectively, and μ represents the cluster center of a cluster in each class.

[0039] Furthermore, the step 6 specifically includes:

[0040] For the normal sample set and abnormal sample set that have completed K-means clustering in step 5, each cluster subset is treated as an independent data subset and trained and evaluated using prediction models. The prediction models used include: XGBoost, Random Forest, LSTM, and CNN;

[0041] For each subcluster C l , let its corresponding data be x i , the actual number of people queuing for security check is y i , train several prediction models respectively, and calculate the prediction value of each model Among them, M j Represents several trained prediction models, x i is a subcluster C l The i-th sample in ;

[0042] The root mean square error, mean absolute error, mean absolute percentage error, and coefficient of determination were used to evaluate the prediction performance of each model. The calculation formula is as follows:

[0043]

[0044] Among them, n l Represents cluster C l The number of samples in For model M j The predicted value of the kth sample in the lth cluster subset, is the true value of the sample, the mean of the true label

[0045] Let the model M under the lth cluster be j The four indicators are: define whether each model performs best in the four indicators. If a model achieves the best performance in a certain indicator, it will be scored 1 point on that indicator. The minimization indicators are RMSE, MAE, and MAPE, and the maximization indicator is R 2 , calculate the comprehensive score of each model respectively, the formula is:

[0046]

[0047] in, is the comprehensive score of the jth model in the lth cluster, I(·) is the indicator function, which takes the value of 1 if the condition is true, otherwise it is 0; They represent the root mean square error, mean absolute error, mean absolute percentage, and determination coefficient of the j-th model in the l-th cluster, respectively. p represents the p-th comparison model participating in the evaluation.

[0048] When multiple models have the same score, the model with the smallest root mean square error is selected as the optimal model; otherwise, the model with the largest comprehensive score is selected as the optimal model.

[0049] When multiple models have the same score, the model with the smallest root mean square error is selected as the optimal model; otherwise, the model with the largest comprehensive score is selected as the optimal model.

[0050] Furthermore, the step 7 specifically includes:

[0051] Get the predicted data, for the input sample to be predicted x new ,First, the anomaly detection classifier is used to perform a binary classification operation to ,determine the category to which it belongs;

[0052] If the prediction is a positive sample, it is classified as normal data and the similarity is compared with the cluster center corresponding to the normal data set; if the prediction is a negative class, it is classified as abnormal data and the similarity is compared with the cluster center corresponding to the abnormal data set; the similarity is cosine similarity, and the formula is:

[0053]

[0054] in, Represents the sample to be predicted and the lth cluster center The cosine similarity of

[0055] Select and input data x new The cluster center μ with the largest cosine similarity * , in, Represents the set of all cluster centers corresponding to the positive or negative classification category; use μ * The optimal prediction model trained by the corresponding cluster Make a final prediction on the input data to get the predicted value of the number of people queuing for security check

[0056] It can be seen from the above technical solution that, compared with the prior art, the present invention discloses a method for predicting the number of people in the airport security check queue based on KAN and K-means clustering method, which has the following beneficial effects:

[0057] (1) A feature set that integrates multi-source information, including flight information, queue openness, traffic status, etc., is constructed to enhance the model's comprehensive perception of influencing factors and overcome the shortcomings of traditional models with single input features and weak expressiveness.

[0058] (2) The innovative introduction of the KAN model for preliminary prediction and anomaly identification avoids simply treating emergencies as noise elimination, improves the model's adaptability to nonlinear and non-stationary scenarios, and has stronger scenario generalization capabilities.

[0059] (3) K-means clustering is used for data hierarchical management, normal data and abnormal data are clustered separately, and a personalized sub-model library is constructed to achieve "on-demand prediction" and "model adaptation", which significantly improves the prediction accuracy.

[0060] Verified by the application of field data, the proposed method outperforms existing mainstream methods across multiple evaluation metrics, demonstrating strong practicality and potential for widespread adoption. Overall, this method not only effectively improves the accuracy of security check queue predictions but also provides important technical support for intelligent airport operations, optimized personnel scheduling, and dynamic resource allocation, resulting in significant social and economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0062] Figure 1 This is a schematic diagram of the overall process provided by the present invention. DETAILED DESCRIPTION

[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0064] The embodiment of the present invention discloses a method for predicting the number of people queuing for airport security check based on KAN and K-means clustering method, such as Figure 1 Shown, including:

[0065] Step 1: Collect airport flight data, security check channel data, terminal traffic data, security check queue data, and weather data as the initial data set, and construct multi-feature data;

[0066] Step 2: Preprocess each feature data in the multi-feature data to obtain a preprocessed data set;

[0067] Step 3: Use the KAN regression model to predict each sample data in the preprocessed data set, classify the data with a prediction error less than or equal to 10% as normal data, and classify the data with a prediction error greater than 10% as abnormal data, to obtain normal data sets and abnormal data sets;

[0068] Step 4: Take the sample data in the normal data set as positive samples and the sample data in the abnormal data set as negative samples, train the KAN classification model, and obtain the anomaly detection classifier;

[0069] Step 5: Use the K-means clustering algorithm to cluster the normal data set and the abnormal data set respectively, divide them into multiple clusters and determine the cluster center of each cluster;

[0070] Step 6: For each cluster divided after clustering, use XGBoost, RandomForest, LSTM, and CNN models for training and testing respectively, and retain the model with the best prediction effect as the optimal prediction model for the cluster;

[0071] Step 7: Obtain the predicted data and classify it using the anomaly detection classifier. Based on the classification results, compare the predicted data with the cluster centers of the normal data set or the abnormal data set, determine the cluster center with the highest similarity, and use the prediction model trained with the corresponding cluster cluster to make predictions to obtain the final prediction results of the number of people in the security check queue.

[0072] In a specific embodiment, step 1 specifically includes:

[0073] For each data point in the initial data set at any moment, extract data slices at different times in the past 24 hours and perform statistics on each data slice as numerical statistical features; use the number of security checkpoints per minute in the past 24 hours as a time series to extract time series features; and use the number of security checkpoint queues in the next 30 minutes as the prediction target;

[0074] The numerical statistical characteristics and time series characteristics of each data at any moment are used as sample features, and the predicted target is used as the sample label to obtain a sample data; the initial data set is sampled and features are extracted at fixed time intervals to obtain several sample data as multi-feature data.

[0075] Specifically, step 1 involves data collection. Comprehensive airport security-related information, including flight data, security checkpoint channel data, terminal traffic data, security checkpoint queue data, and weather data, is obtained from the airport department as an initial dataset. Data from different times within the past 24 hours are extracted from this initial dataset and sliced ​​and counted to form numerical statistical features. The number of people going through security checks every minute in the past 24 hours is used as a time series, and time series features are extracted. The number of people waiting in security checkpoints within the next 30 minutes is used as a prediction target. Specifically, the features at the current moment are constructed using the historical data from the past 24 hours, and the target value is the number of people waiting in security checkpoints from the current moment to the next 30 minutes. (For example, if the current moment is 8:00, the numerical statistical features are slices of the initial dataset from 8:00 the previous day to the current moment, and the time series features are the sequence of people going through security checkpoints from 8:00 the previous day to the current moment. Relevant temporal features are extracted from this time series.) This data is then rolled over at one-minute intervals (e.g., 8:00, 8:01, etc., each serving as a sample). The numerical statistical features, time series features, and prediction targets are integrated to construct the multi-feature data of the present invention.

[0076] The multi-feature data packet has two parts of features, one of which is numerical statistical features extracted from flight data, security check channel data, traffic data in front of the terminal, security check queue data and weather data. Such features include: flight departure time, number of flight pre-purchased tickets, departing aircraft flight number, departing aircraft terminal, departing aircraft flight type (domestic or international flight), number of open security check channels, number of people passing through the security check channel, time for passengers to pass through the security check channel, number of people entering the terminal, number of people leaving the terminal, time for security check passengers to pass through the gate, security check gate number, gender of security check passengers, flight number of security check passengers, temperature, wind speed, visibility, cloud cover, air pressure, and humidity data. Another part of the features comes from the time series of the number of people queuing for security checks with a granularity of minutes. The number of people queuing for security checks in the past 24 hours is used to construct a time series of the number of people queuing for security checks with an interval of one minute and a length of 1440. Based on this time series, time series features are extracted, including: autocorrelation characteristics, seasonal autocorrelation characteristics, trend and cycle characteristics, volatility characteristics, stationarity test characteristics, nonlinearity and uncertainty characteristics, structural characteristics, sequence morphology characteristics, change rate characteristics, fitting model parameter characteristics, cycle characteristics, and related statistics of trend terms, seasonal terms and residual terms extracted by STL decomposition.

[0077] The above features are constructed based on the following principles: trend and cycle modeling uses Holt-Winters model parameters and STL decomposition to accurately characterize the trend changes and cyclical patterns of the number of queues within a day (inter-day); seasonal characteristics combine seasonal autocorrelation, seasonal partial autocorrelation, cycle number, and seasonal intensity to improve the ability to identify high-frequency intraday repetitive behaviors (for example, the morning and evening peaks of airport queues within a day can be regarded as a kind of seasonality); volatility and stability are judged through ARCH or GARCH models, unit root tests, and Hurst exponents to determine whether the series has heteroskedasticity and long memory characteristics; structural complexity and change behavior descriptions such as crossing points, flat sections, peak degree, entropy, etc. are used to characterize the local structure, abnormal fluctuations, and change complexity of the time series of the number of people in security check queues; lagged correlation descriptions such as multi-order autocorrelation and partial autocorrelation characteristics are used to describe time dependence, assisting the model in grasping the influence relationship between current values ​​and historical values.

[0078] In a specific embodiment, the preprocessing includes missing value processing and data normalization processing;

[0079] Missing value processing specifically includes: directly deleting samples with half or more features missing; using feature mean imputation for samples with less than half of the features missing, filtering the missing values ​​of each feature in the multi-feature data and filling them with the mean of the corresponding feature data within the previous 2 hours. Specifically, suppose the missing value of the j-th feature in the original data set at the i-th sample, and the value after missing value processing is the mean of the feature in the previous 2 hours. Among them, x k,j It represents the kth valid sample of the jth feature in the past 2 hours, where n is the number of samples.

[0080] Normalization processing specifically includes: using the maximum and minimum value normalization method to scale the value of each feature in the multi-feature data to the range [0, 1]. Specifically, suppose the jth feature of the i-th sample in the original data set is replaced by the normalized value calculated by the maximum and minimum value method after outlier processing. Among them, min(x j ) and max(x j ) are the minimum and maximum values ​​of the j-th feature in the original dataset, respectively.

[0081] In a specific embodiment, the prediction process of the KAN regression model specifically includes:

[0082] The sample features are converted into feature vectors and input into the KAN regression model. The feature vector is recorded as x i , where xin Represents the nth feature of the i-th sample of the input vector, which is input into the KAN model as a training sample or prediction sample of the model.

[0083] The KAN regression model transforms all the features in the feature vector through a univariate nonlinear transformation function to obtain a new feature vector. i , all its features will go through a univariate nonlinear transformation function φ k (x ik ) to obtain the new eigenvector z i , the formula is:

[0084] z ik =φ k (x ik ),k=1,2,…,n;

[0085]

[0086] Among them, φ k (·) represents the learnable nonlinear mapping function used for the kth input feature in the KAN model, z ik Representative feature x ik The result after nonlinear transformation. k It is usually composed of the third-order B-spline function or other adaptive functions in KAN, and its parameters can be automatically learned through backpropagation.

[0087] Then the KAN regression model performs a linear weighted combination of the new feature vectors to obtain several intermediate combination results, the formula is:

[0088]

[0089] Among them, a jk Represents the linear weight from the kth feature map result to the jth combination unit, u ij is the intermediate output of the i-th sample in the j-th combination unit.

[0090] During the training phase, the model constructs the loss function of the KAN model by minimizing the error between the predicted value and the true value. The formula is:

[0091]

[0092] Among them, y i represents the actual number of people queuing for security check in the i-th sample, represents the predicted value of the KAN model, and N is the total number of samples.

[0093] Then apply the corresponding output function to each intermediate combination result and sum them up to get the final prediction result The formula is:

[0094]

[0095] Among them, ψ j (·) represents the output transformation function of the jth intermediate unit, which can be a linear or nonlinear function, and m is the total number of output transformation functions of the intermediate units.

[0096] After sorting out the above formula, we get the final prediction formula of the KAN regression model:

[0097]

[0098] The specific implementation of the abnormal classification process is: by calculating the relative error between the true value of the sample and the predicted value of the KAN model, the samples with a relative error greater than 10% are classified as abnormal samples, and the other samples are regarded as normal samples. Finally, the divided data set x split It can be expressed as:

[0099]

[0100] where x normal is a normal sample set, x anomalous is an abnormal sample set.

[0101] In a specific embodiment, the KAN classification model completes the binary classification task by learning the difference between normal samples and abnormal samples in the feature space. The anomaly detection classifier f classify (x i ) is expressed as:

[0102]

[0103] Among them, x i Represents the characteristics of the i-th sample. The output value is 1, which means that the sample is a normal sample. The output value is 0, which means that the sample is an abnormal sample. normal represents the normal sample set, x anomalous Represents the abnormal sample set. By training the KAN model on an existing labeled dataset (normal samples and abnormal samples), it is able to distinguish different categories.

[0104] In a specific embodiment, in step 5, the goal of the K-means clustering algorithm is to minimize the intra-class squared error in the process of clustering the normal data set and the abnormal data set. The formula is:

[0105]

[0106] Among them, k is the number of clusters, C l represents the lth cluster, μl is the center of the lth cluster, and x is the sample point vector;

[0107] The specific formula for determining the cluster center of each cluster is:

[0108]

[0109] Among them, |C l | represents the number of samples in the lth cluster, It represents the sum of the vectors of all samples in the cluster.

[0110] It should be noted that although the objective function includes the cluster center μl, and the cluster center itself depends on the current sample cluster assignment, this does not constitute a logical contradiction. K-means uses an iterative optimization method to solve this problem: in each iteration, the algorithm first fixes the cluster center, calculates the distance from each sample to each center, and updates the sample cluster label accordingly; then, based on the new sample clustering results, recalculates the position of each cluster center. This process continues until the objective function is Convergence. It can be seen that the objective function and the cluster center calculation formula together constitute the core optimization process of the K-means algorithm. The two are solved alternately through an iterative process and are logically self-consistent.

[0111] Finally, after clustering is completed, the normal data set and the abnormal data set can be represented as multiple subclusters and corresponding cluster centers respectively:

[0112]

[0113] Among them, x normal and x anomalous are normal data sets and abnormal data sets respectively, k and m are the number of clusters of normal samples and abnormal samples respectively, and μ represents the cluster center of a cluster in each class.

[0114] In a specific embodiment, for the normal sample set and abnormal sample set that have completed K-means clustering in step 5, each cluster subset is used as an independent data subset and trained and evaluated using a prediction model. The prediction models used include: XGBoost, Random Forest, LSTM, and CNN.

[0115] For each subcluster C l , let its corresponding data be x i , the actual number of people queuing for security check is y i , train several prediction models respectively, and calculate the prediction value of each model

[0116] Among them, Mj represents several trained prediction models (M1=XGBoost, M2=Random Forest, M3=LSTM, M4=CNN), x i is a subcluster C l The i-th sample in ;

[0117] In order to evaluate the prediction performance of each model, the following four evaluation indicators are used: root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE) and coefficient of determination (R 2 ), which is calculated as follows:

[0118]

[0119]

[0120] Among them, n l Represents cluster C l The number of samples in For model M j The predicted value of the kth sample in the lth cluster subset, is the true value of the sample, the mean of the true label

[0121] Keep the model with the best prediction effect, specifically: record the model M under the lth cluster j The four indicators are: define whether the performance of each model is optimal in the four indicators. If a model achieves the best performance in a certain indicator, it will be scored 1 point on that indicator. The minimization indicators are RMSE, MAE, and MAPE, and the maximization indicator is R 2 The root mean square error, mean absolute error, mean absolute percentage error, and coefficient of determination were used to evaluate each model. Based on the evaluation results, the comprehensive score of each model was calculated using the following formula:

[0122]

[0123] in, is the comprehensive score of the jth model in the lth cluster, I(·) is the indicator function, which takes the value of 1 if the condition is true, otherwise it is 0; They represent the root mean square error, mean absolute error, mean absolute percentage, and determination coefficient (R 2 ), p represents the pth comparison model involved in the evaluation;

[0124] When multiple models have the same score, the model with the smallest root mean square error is selected as the optimal model; otherwise, the model with the largest comprehensive score is selected as the optimal model.

[0125] In a specific embodiment, step 7 specifically includes:

[0126] Get the predicted data, for the input sample to be predicted x new ,First, the anomaly detection classifier is used to perform a binary ,classification operation to determine the category to which it belongs;

[0127] If the prediction is a positive sample, it is classified as normal data and the similarity is compared with the cluster center corresponding to the normal data set; if the prediction is a negative class, it is classified as abnormal data and the similarity is compared with the cluster center corresponding to the abnormal data set; the similarity is cosine similarity, and the formula is:

[0128]

[0129] in, Represents the sample to be predicted and the lth cluster center The cosine similarity of

[0130] Select and input data x new The cluster center μ with the largest cosine similarity * , in, Represents the set of all cluster centers corresponding to the positive or negative classification category; use μ * The optimal prediction model trained by the corresponding cluster Make a final prediction on the input data to get the predicted value of the number of people queuing for security check

[0131] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0132] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for predicting the number of people queuing for airport security checkpoints based on KAN and K-means clustering method, characterized by: include: Step 1: Collect airport flight data, security check channel data, terminal traffic data, security check queue data, and weather data as the initial data set, and construct multi-feature data; Step 2: Preprocessing each feature data in the multi-feature data to obtain a preprocessed data set; Step 3: Predict each sample data in the preprocessed data set using the KAN regression model, classify the data with a prediction error of less than or equal to 10% as normal data, and classify the data with a prediction error of more than 10% as abnormal data, to obtain a normal data set and an abnormal data set; Step 4: Using the sample data in the normal data set as positive samples and the sample data in the abnormal data set as negative samples, training the KAN classification model to obtain an anomaly detection classifier; Step 5: Clustering the normal data set and the abnormal data set respectively by using the K-means clustering algorithm to obtain multiple clusters and determine the cluster center of each cluster; Step 6: For each cluster divided after clustering, use XGBoost, RandomForest, LSTM, and CNN models for training and testing respectively, and retain the model with the best prediction effect as the optimal prediction model for the cluster; Step 7: Obtain the predicted data and classify it using the anomaly detection classifier. Based on the classification results, compare the predicted data with the cluster centers of the normal data set or the abnormal data set, determine the cluster center with the highest similarity, and use the prediction model trained by the corresponding cluster cluster to perform prediction to obtain the final prediction result of the number of people in the security check queue.

2. The method for predicting the number of people waiting in queues for airport security inspection based on KAN and K-means clustering method according to claim 1, characterized in that: The step 1 specifically includes: For each data item in the initial data set at any time, extract data slices at different times in the past 24 hours, and perform statistics on each data slice as numerical statistical features; use the number of security checkpoints per minute in the past 24 hours as a time series to extract time series features; and use the number of security checkpoint queues in the next 30 minutes as the prediction target; The numerical statistical features and the time series features of each data at any moment are used as sample features, and the prediction target is used as a sample label to obtain a sample data; the initial data set is sampled and features are extracted at fixed time intervals to obtain a plurality of sample data as the multi-feature data.

3. The method for predicting the number of people waiting in queues for airport security inspection based on KAN and K-means clustering method according to claim 1, characterized in that: The preprocessing includes missing value processing and data standardization processing; The missing value processing specifically includes: for samples with half or more features missing, directly delete them; for samples with less than half of the features missing, use the feature mean imputation method to fill in the missing values ​​of each feature in the multi-feature data, and fill in the missing values ​​with the mean value of the corresponding feature data within the adjacent 2 hours; The standardization process specifically includes: using a maximum and minimum value standardization method to scale the value of each feature in the multi-feature data to the interval [0, 1].

4. The method for predicting the number of people queuing for airport security inspection based on KAN and K-means clustering method according to claim 1, characterized in that: The prediction process of the KAN regression model specifically includes: Convert the sample features into feature vectors and input them into the KAN regression model; The KAN regression model transforms all features in the feature vector through a univariate nonlinear transformation function to obtain a new feature vector; Then the KAN regression model performs a linear weighted combination on the new feature vectors to obtain several intermediate combination results; Then, the corresponding output function is applied to each intermediate combination result, and the final prediction result is obtained after summing up.

5. The method for predicting the number of people queuing for airport security inspection based on KAN and K-means clustering method according to claim 1, characterized in that: The KAN classification model completes the binary classification task by learning the difference between normal samples and abnormal samples in the feature space. The anomaly detection classifier f classify (x i ) is expressed as: Among them, x i Represents the characteristics of the i-th sample. The output value is 1, which means that the sample is a normal sample. The output value is 0, which means that the sample is an abnormal sample. normal represents the normal sample set, x anomalous Represents an abnormal sample set.

6. The method for predicting the number of people queuing for airport security inspection based on KAN and K-means clustering method according to claim 1, characterized in that: In step 5, the goal of the K-means clustering algorithm is to minimize the intra-class square error in the process of clustering the normal data set and the abnormal data set. The formula is: Among them, k is the number of clusters, C l represents the lth cluster, μ l is the center of the lth cluster, and x is the sample point vector; The specific formula for determining the cluster center of each cluster is: Among them, |C l | represents the number of samples in the lth cluster, represents the sum of the vectors of all samples in the cluster; Finally, after clustering is completed, the normal data set and the abnormal data set can be represented as multiple subclusters and corresponding cluster centers respectively: Among them, x normal and x anomalous are normal data sets and abnormal data sets respectively, k and m are the number of clusters of normal samples and abnormal samples respectively, and μ represents the cluster center of a cluster in each class.

7. The method for predicting the number of people queuing for airport security inspection based on KAN and K-means clustering method according to claim 1, characterized in that: The step 6 specifically includes: For the normal sample set and abnormal sample set that have completed K-means clustering in step 5, each cluster subset is treated as an independent data subset and trained and evaluated using prediction models. The prediction models used include: XGBoost, Random Forest, LSTM, and CNN; For each subcluster C l , let its corresponding data be x i , the actual number of people queuing for security check is y i , train several prediction models respectively, and calculate the prediction value of each model Among them, M j Represents several trained prediction models, x i is a subcluster C l The i-th sample in ; The root mean square error, mean absolute error, mean absolute percentage error, and coefficient of determination were used to evaluate the prediction performance of each model. The calculation formula is as follows: Among them, n l Represents cluster C l The number of samples in For model M j The predicted value of the kth sample in the lth cluster subset, is the true value of the sample, the mean of the true label Let the model M under the lth cluster be j The four indicators are: define whether each model performs best in the four indicators. If a model achieves the best performance in a certain indicator, it will be scored 1 point on that indicator. The minimization indicators are RMSE, MAE, and MAPE, and the maximization indicator is R 2 , calculate the comprehensive score of each model respectively, the formula is: in, is the comprehensive score of the jth model in the lth cluster, I(·) is the indicator function, which takes the value of 1 if the condition is true, otherwise it is 0; They represent the root mean square error, mean absolute error, mean absolute percentage, and determination coefficient of the j-th model in the l-th cluster, respectively. p represents the p-th comparison model participating in the evaluation. When multiple models have the same score, the model with the smallest root mean square error is selected as the optimal model; otherwise, the model with the largest comprehensive score is selected as the optimal model.

8. The method for predicting the number of people queuing for airport security inspection based on KAN and K-means clustering method according to claim 1, characterized in that: The step 7 specifically includes: Get the predicted data, for the input sample to be predicted x new ,First, the anomaly detection classifier is used to perform a binary classification operation to ,determine the category to which it belongs; If the prediction is a positive sample, it is classified as normal data and the similarity is compared with the cluster center corresponding to the normal data set; if the prediction is a negative class, it is classified as abnormal data and the similarity is compared with the cluster center corresponding to the abnormal data set; the similarity is cosine similarity, and the formula is: in, Represents the sample to be predicted and the lth cluster center The cosine similarity of Select and input data x new The cluster center μ with the largest cosine similarity * , in, Represents the set of all cluster centers corresponding to the positive or negative classification category; use μ * The optimal prediction model trained by the corresponding cluster Make a final prediction on the input data to get the predicted value of the number of people queuing for security check