Air travel data acquisition system based on Internet technology
By adopting federated learning in the air travel data acquisition system and using local devices for data processing and model training, the problems of insufficient data analysis and privacy leakage are solved, and the accuracy of personalized recommendations and real-time decision-making is improved, ensuring data security and cross-platform collaborative analysis.
Patent Information
- Application Number
- CN202510359430.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing aviation travel data acquisition system is insufficient in the depth and intelligence of data analysis and mining, and cannot extract valuable information in a timely and accurate manner. The cross-platform data sharing and collaborative analysis are not ideal, and there is a risk of user privacy leakage.
Using federated learning, by performing data processing and model training on local devices, using data preparation and allocation modules, local model training modules, aggregation modules, optimization modules, and model evaluation and deployment modules, we establish a global neural network model to ensure data security and privacy protection while improving personalized recommendation and demand prediction capabilities.
It improves the depth and breadth of data analysis capabilities and AI technology applications, improves the accuracy of personalized recommendations and real-time decision-making, ensures data security and comply with compliance requirements.
Smart Images

Figure CN120296247A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of air travel data collection, and more specifically, to an air travel data collection system based on Internet technology. Background Art
[0002] With the rapid development of Internet technology, major platforms and airlines in the air travel industry have gradually relied on advanced data collection and processing technologies to optimize operation management, improve user experience, and drive market decisions. The current air travel data collection system can collect a large amount of data from multiple channels, including flight information, booking platforms, social media, and user feedback. Through the storage and preliminary analysis of these data, the air travel platform can provide services such as a certain degree of personalized recommendation, flight schedule optimization, and demand prediction; however, the existing data collection system relies on a centralized data processing architecture. Although a large amount of data can be obtained, its data analysis ability and intelligent level still need to be improved, especially in the fields of complex demand prediction, personalized recommendation, and real-time decision-making.
[0003] Although the existing systems have achieved certain results in data collection and storage, there are still obvious deficiencies in the depth and intelligence of data analysis and mining; traditional analysis methods cannot efficiently process massive and diverse air travel data, resulting in the inability to timely and accurately extract valuable information or trends. Especially in the implementation of key functions such as personalized recommendation, real-time decision-making, and demand prediction, it often relies on traditional statistical analysis methods, and the effect is limited; on the other hand, due to the data analysis of the existing system often relying on a centralized processing architecture, the effect of cross-platform data sharing and collaborative analysis is not ideal, and there is a risk of user privacy leakage when processing sensitive data; therefore, the existing technology urgently needs to be improved to enhance its data analysis ability and further improve the intelligent analysis and cross-platform cooperation ability on the premise of ensuring user privacy. Summary of the Invention
[0004] In view of the technical problems existing in the prior art, the present invention provides an air travel data collection system based on Internet technology, by setting up a data preparation and distribution module, a local model training module, an aggregation module, an optimization module, and a model evaluation and deployment module. Using methods such as calculating gradients, weighted averaging, and evaluating losses to continuously optimize the global model, thereby enhancing the cross-platform personalized recommendation and demand prediction functions. In this process, all participants can obtain an efficient and secure model through joint training without exposing the original data, so as to solve the problems proposed in the above background art.
[0005] The technical solution for the present invention to solve the above technical problems is as follows: It includes a data preparation and distribution module, a local model training module, an aggregation module, an optimization module, and a model evaluation and deployment module;
[0006] Data Preparation and Allocation Module: Each participating party collects local data based on the user's historical behavior, establishes a local dataset, and divides their respective local datasets into training sets, validation sets, and test sets without sharing the original data, thereby providing a necessary update basis for the global neural network model of the central server;
[0007] Local Model Training Module: Each participating party first initializes the sub-neural network model locally, then uses the training set of its local dataset to train the sub-neural network model, and calculates the gradient using the gradient descent algorithm to update the parameters of the sub-neural network model;
[0008] Aggregation Module: Send the gradients calculated by each participating party to the central server, and use the weighted average formula to update the global neural network model;
[0009] Optimization Module: The central server sends the updated global neural network model to each participating party. Each participating party uses the new global neural network model for local training to update the sub-neural network model. By evaluating the validation set loss value of the global neural network model, it is determined whether to continue the iteration;
[0010] Model Evaluation and Deployment Module: Use the test set to evaluate the accuracy of the global neural network model. If the accuracy of the global neural network model meets the requirements, it is deployed to each participating party's platform for real-time recommendation or demand prediction.
[0011] In a preferred embodiment, in the data preparation and allocation module, each participating party specifically refers to airlines, airports, and travel platforms.
[0012] In a preferred embodiment, preprocess the data of each participating party, including missing value filling, feature normalization, and data cleaning.
[0013] In a preferred embodiment, the local dataset is represented as:
[0014] D k =(x1,y1),(x2,y2),...,(x n ,y n );
[0015] where k represents the kth participating party, D k represents the dataset of the kth participating party, and (x i ,y i ) represents a sample, which includes the input feature x i and the corresponding label y i .
[0016] In a preferred embodiment, in the local model training module, the calculation formula of the gradient descent algorithm is:
[0017]
[0018] Where, represents the gradient update of the sub-neural network model parameters of the k-th participant, θ represents the parameters of the sub-neural network model, represents calculating the average gradient of the k-th participant based on the training set in its local dataset, n k represents the number of training set samples of the k-th participant, x i represents the input feature of the i-th training set sample, y i represents the target output of the i-th training set sample, represents the gradient of the i-th training set sample with respect to the sub-neural network model parameters.
[0019] In a preferred embodiment, in the aggregation module, the weighted average formula is:
[0020]
[0021] Where, θ k represents the local sub-neural network model parameters of the k-th participant, n k represents the number of training set samples of the k-th participant, N represents the total number of training samples of all participants, K represents the total number of participants, θ global represents the parameters of the global neural network model.
[0022] In a preferred embodiment, in the optimization module, the calculation formula for the validation set loss value of the global neural network model is:
[0023]
[0024] Where, L global represents the loss function value of the global neural network model, N represents the total number of training samples of all participants, n k represents the number of training set samples of the k-th participant, K represents the total number of participants, D k represents the local dataset of the k-th participant, which contains a set of samples (x i , y i ), where x i is the input feature, y i is the target label, (x i , y i ) represents a single sample in the dataset of the k-th participant, x i is the feature of this sample, y i is the label of this sample, L(θglobal , (x i , y i )) represents the loss function.
[0025] In a preferred embodiment, in the optimization module, the specific steps for determining whether to continue iteration are as follows:
[0026] S1. If the loss value is lower than the preset threshold, terminate the training;
[0027] S2. If the loss value is higher than the preset threshold, repeat the execution of local training, update and upload of the sub-neural network model, aggregation, and distribution, and continuously optimize the global model until convergence.
[0028] In a preferred embodiment, the optimization module encrypts the updates of the sub-neural network models of each participant using the differential privacy mechanism.
[0029] In a preferred embodiment, in the model evaluation and deployment module, the calculation formula for evaluating the accuracy of the global neural network model is:
[0030]
[0031] Among them, Accuracy represents the accuracy of the global neural network model on the test set, N test represents the total number of samples in the test set, D test represents the data set of the test set, which contains a set of samples (x i , y i ), where x i is the input feature, y i is the true label, (x i , y i ) represents a sample in the test set, x i is the input feature, y i is the true label, represents the predicted output of the global neural network model for the sample x i , that is, the label predicted by the global neural network model, represents the indicator function.
[0032] In a preferred embodiment, it specifically includes the following steps:
[0033] The beneficial effects of the present invention are as follows: By adopting the method of federated learning and performing data processing and model training on local devices, the present invention solves the problems of data privacy and security, and can improve the accuracy of personalized recommendation, real-time decision-making, and demand prediction; in the aviation and travel industry, federated learning allows different participants to use local data for collaborative training and jointly optimize the model, thereby improving the data analysis ability and the depth and breadth of the application of AI technology, ensuring data security and compliance requirements, and promoting the progress of intelligent decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flowchart of the method of the present invention;
[0035] Figure 2 It is a block diagram of the system structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0037] In the description of the present application, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality" means two or more, unless otherwise specifically defined.
[0038] In the description of the present application, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present application is not necessarily construed as being more preferred or having more advantages than other embodiments. In order for any person skilled in the art to implement and use the present invention, the following description is given. In the following description, details are set forth for purposes of explanation. It should be understood that those skilled in the art can recognize that the present invention can be implemented without using these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope that conforms to the principles and features disclosed in the present application.
[0039] Embodiment 1
[0040] This embodiment provides as Figure 1-2A travel data collection system based on Internet technology is shown as follows, specifically including: a data preparation and distribution module, a local model training module, an aggregation module, an optimization module, and a model evaluation and deployment module;
[0041] Data preparation and distribution module: Each participant collects local data according to the user's historical behavior, establishes a local data set, and divides their respective local data sets into training sets, validation sets, and test sets without sharing the original data, so as to provide a necessary update basis for the global neural network model of the central server;
[0042] Local model training module: Each participant first initializes a sub-neural network model locally, then uses the training set of their respective local data sets to train the sub-neural network model, and uses the gradient descent algorithm to calculate the gradient and update the parameters of the sub-neural network model;
[0043] Aggregation module: Send the gradients calculated by each participant to the central server, and use the weighted average formula to update the global neural network model;
[0044] Optimization module: The central server sends the updated global neural network model to each participant. Each participant uses the new global neural network model for local training, updates the sub-neural network model, and determines whether to continue iterating by evaluating the validation set loss value of the global neural network model;
[0045] Model evaluation and deployment module: Use the test set to evaluate the accuracy of the global neural network model to ensure the generalization ability of the global neural network model. If the accuracy of the global neural network model meets the requirements, deploy it to each participant platform for real-time recommendation or demand prediction.
[0046] In this embodiment, specifically, it should be noted that for the data preparation and distribution module, each participant specifically refers to airlines, airports, and travel platforms; preprocess the data of each participant, including missing value filling, feature normalization, and data cleaning;
[0047] The local data set is represented as:
[0048] D k =(x1,y1),(x2,y2),...,(x n ,y n );
[0049] where k represents the kth participant, D k represents the data set of the kth participant, and (x i ,y i ) represents a sample, which includes the input feature x i and the corresponding label y i , where the feature xi The description or representation of the input data, labeled y i Refers to the feature x i The corresponding target output, usually used in supervised learning tasks to indicate the true class or value of each sample.
[0050] In this embodiment, specifically, it should be noted that for the local model training module, the calculation formula of the gradient descent algorithm is:
[0051]
[0052] Where, Represents the gradient update of the sub - neural network model parameters of the k - th participant. The gradient is the derivative of the sub - neural network model parameters with respect to the loss function, used to guide how to adjust the sub - neural network model parameters. θ represents the parameters of the sub - neural network model, usually the weights and biases in the neural network, and the specific selection can be determined according to actual needs. Represents calculating the average gradient of the k - th participant based on the training set in its local dataset, n k Represents the number of training set samples of the k - th participant, x i Represents the input feature of the i - th training set sample, y i Represents the target output of the i - th training set sample, Represents the gradient of the i - th training set sample with respect to the sub - neural network model parameters, specifically the partial derivative of the loss function with respect to the model parameters;
[0053] In addition, when the local model training module selects a machine learning model, in addition to the neural network model of this application, a linear regression model can also be selected. Among them, the advantages of the neural network model are: 1. Strong adaptability: The neural network can handle complex non - linear relationships and is suitable for processing highly complex and high - dimensional data (such as multiple factors in user behavior data); 2. Strong expressive ability: It can automatically learn complex patterns from data, such as images, speech, text, or large - scale user behavior data; 3. Deep learning applications: If the task involves fields such as recommendation systems, image recognition, natural language processing, etc., the neural network is usually more effective than the linear regression model;
[0054] The advantages of the linear regression model are as follows: 1. Simple and efficient: The linear regression has low computational cost, is easy to implement and understand, and is suitable for simple tasks and small datasets; 2. Fast training speed: Due to the simplicity of the model, it requires less computational resources and is suitable for rapid iteration; 3. Strong interpretability: The linear regression model is easier to interpret and understand, which helps to understand the relationship between features and the target. The disadvantages of the linear regression model are as follows: 1. Limited applicability: The linear regression assumes a linear relationship between features and the target. Therefore, for complex non-linear problems, the linear regression model performs poorly; 2. Unable to capture complex patterns: For complex non-linear patterns such as user behavior data, the performance of linear regression is inferior to that of neural networks.
[0055] Therefore, based on the above comparison, the reason for preferentially adopting the neural network model in this application is that when it comes to user behavior analysis, the neural network model is often more suitable because it can handle complex non-linear relationships and can learn more refined patterns from a large amount of data.
[0056] In this embodiment, specifically, the aggregation module needs to be explained. The weighted average formula is:
[0057]
[0058] where θ k represents the local sub-neural network model parameters of the k-th participating party, n k represents the number of training set samples of the k-th participating party, N represents the total number of training samples of all participating parties, K represents the total number of participating parties, that is, how many local models participate in the training, and θ global represents the parameters of the global neural network model, that is, the parameters after weighted average based on the local model parameters of all participating parties;
[0059] This formula represents the weighted average calculated according to the local model parameters θ k of each participating party according to its number of training samples n k . The contribution of each participating party to the global neural network model is weighted, and the weight is the proportion of the data volume of this participating party in the total data volume
[0060] In this embodiment, specifically, the optimization module needs to be explained. The formula for calculating the validation set loss value of the global neural network model is:
[0061]
[0062] where L global represents the loss function value of the global neural network model, which represents the average loss of the global neural network model on all data and is usually used to measure the performance of the model. N represents the total number of training samples of all participating parties, and n krepresents the number of training set samples of the k-th participant, K represents the total number of participants, that is, how many local models participate in the training, D k represents the local dataset of the k-th participant, which contains a set of samples (x i , y i ), where x i is the input feature, y i is the target label, (x i , y i ) represents a single sample in the dataset of the k-th participant, x i is the feature of this sample, y i is the label of this sample, L(θ global , (x i , y i )) represents the loss function, indicating the loss when using the global neural network model parameter θ global to predict the sample (x i , y i ). In this application, the loss function is selected from any one of the mean squared error (MSE) and cross entropy;
[0063] The specific steps to determine whether to continue iteration are as follows:
[0064] S1. If the loss value is lower than the preset threshold, terminate the training;
[0065] S2. If the loss value is higher than the preset threshold, repeat the execution of local training, sub-neural network model update and upload, aggregation and distribution, and continuously optimize the global model until convergence;
[0066] During the entire federated learning process, ensure that the privacy of user data is not leaked; to prevent data leakage, the optimization module uses the differential privacy mechanism to encrypt the sub-neural network model updates of each participant, ensuring that the final model update can only be obtained after aggregation.
[0067] In this embodiment, specifically, it is necessary to explain the model evaluation and deployment module. The calculation formula for evaluating the accuracy of the global neural network model is:
[0068]
[0069] Among them, Accuracy represents the accuracy of the global neural network model on the test set, indicating the proportion of correctly predicted samples in the total samples, N test represents the total number of samples in the test set, that is, the size of the test dataset, D test represents the test dataset, which contains a set of samples (x i , y i ), where x i is the input feature, yi is the true label, (x i , y i ) represents a sample in the test set, where x i is the input feature and y i is the true label. denotes the predicted output of the global neural network model for the sample x i , that is, the label predicted by the global neural network model. denotes the indicator function. If the prediction of the global neural network model equals the true label y i , the value of the indicator function is 1 (indicating a correct prediction); otherwise, the value is 0 (indicating an incorrect prediction). This is a binary function used to mark whether the prediction is correct.
[0070] It should be noted that in the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0071] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0072] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0073] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes.
[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 or more boxes.
[0075] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0076] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. An air travel data collection system based on Internet technology, characterized in that, Specifically, it includes: A data preparation and distribution module, a local model training module, an aggregation module, an optimization module, and a model evaluation and deployment module; Data preparation and distribution module: Each participating party collects local data according to the user's historical behavior, establishes a local data set, and divides their respective local data sets into a training set, a validation set, and a test set without sharing the original data, so as to provide a necessary update basis for the global neural network model of the central server; Local model training module: Each participating party first initializes the sub-neural network model locally, then uses the training set of their respective local data sets to train the sub-neural network model, and calculates the gradient using the gradient descent algorithm to update the parameters of the sub-neural network model; Aggregation module: Send the gradients calculated by each participating party to the central server, and use the weighted average formula to update the global neural network model; Optimization module: The central server sends the updated global neural network model to each participating party. Each participating party uses the new global neural network model for local training, updates the sub-neural network model, and judges whether to continue iterating by evaluating the validation set loss value of the global neural network model; Model evaluation and deployment module: Use the test set to evaluate the accuracy of the global neural network model. If the accuracy of the global neural network model meets the requirements, it will be deployed to each participating party platform for real-time recommendation or demand prediction; 2. The air travel data acquisition system based on Internet technology according to claim 1, wherein: In the data preparation and distribution module, each participating party specifically refers to airlines, airports, and travel platforms.
3. The air travel data acquisition system based on Internet technology according to claim 2, wherein: Perform preprocessing operations on the data of each participating party, including missing value filling, feature normalization, and data cleaning.
4. A travel data collection system based on Internet technology according to claim 3, characterized in that: The local data set is expressed as: D k =(x1,y1),(x2,y2),...,(x n ,y n ); Among them, k represents the kth participant, and D k represents the dataset of the kth participant, and (x i , y i ) represents a sample, which contains the input feature x i and the corresponding label y i .
5. The air travel data collection system based on Internet technology according to claim 4, characterized in that: In the local model training module, the calculation formula of the gradient descent algorithm is: Among them, ▽θ k represents the gradient update of the sub-neural network model parameters of the k-th participant, θ represents the parameters of the sub-neural network model, represents calculating the average gradient of the k-th participant based on the training set in its local dataset, n k represents the number of training set samples of the k-th participant, x i represents the input feature of the i-th training set sample, y i represents the target output of the i-th training set sample, ▽θL(θ,(x i ,y i ) represents the gradient of the i-th training set sample with respect to the sub-neural network model parameters.
6. The air travel data collection system based on Internet technology according to claim 5, characterized in that: In the aggregation module, the weighted average formula is: Among them, θ k represents the local sub-neural network model parameters of the k-th participant, n k represents the number of training set samples of the k-th participant, N represents the total number of training samples of all participants, K represents the total number of participants, θ global represents the parameters of the global neural network model.
7. The air travel data collection system based on Internet technology according to claim 6, characterized in that: In the optimization module, the calculation formula of the validation set loss value of the global neural network model is: Among them, L global represents the loss function value of the global neural network model, N represents the total number of training samples of all participants, and n k represents the number of training set samples of the k-th participant, K represents the total number of participants, and D k represents the local dataset of the k-th participant, which contains a set of samples (x i , y i ), where x i is the input feature, y i is the target label, (x i , y i ) represents a single sample in the dataset of the k-th participant, x i is the feature of this sample, y i is the label of this sample, and L(θ global , (x i , y i )) represents the loss function.
8. An air travel data collection system based on Internet technology according to claim 7, characterized in that: In the optimization module, the specific steps for judging whether to continue iterating are: S1. If the loss value is lower than the preset threshold, terminate the training; S2. If the loss value is higher than the preset threshold, repeat the execution of local training, sub-neural network model update and upload, aggregation, and distribution, and the global model is continuously optimized until convergence.
9. An air travel data collection system based on Internet technology according to claim 8, characterized in that: The optimization module uses the differential privacy mechanism to encrypt the update of the sub-neural network model of each participating party.
10. An air travel data collection system based on Internet technology according to claim 9, characterized in that: In the model evaluation and deployment module, the calculation formula for evaluating the accuracy of the global neural network model is: Among them, Accuracy represents the accuracy of the global neural network model on the test set, N test represents the total number of samples in the test set, D test represents the data set of the test set, which contains a set of samples (x i , y i ), where x i is the input feature, y i is the true label, (x i , y i ) represents a sample in the test set, x i is the input feature, y i is the true label, represents the predicted output of the global neural network model for the sample x i , that is, the label predicted by the global neural network model, represents the indicator function.