Video monitoring and ADS-B fusion enhanced scene trajectory prediction method and device
By integrating video surveillance and ADS-B data, a situation information representation model for multi-source heterogeneous data is constructed, and LSTM is used for trajectory prediction, the problem of traditional systems relying on a single information source is solved, the prediction accuracy and system robustness are improved, and the support capabilities of airport management and flight safety decisions are enhanced.
Patent Information
- Application Number
- CN202411882289.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional airport scene situation awareness systems rely on a single information source, and have problems such as blind spots in monitoring, insufficient information, insufficient data processing flexibility, insufficient prediction accuracy, and single point failure risk.
Through the fusion of video surveillance and ADS-B data, using timestamp synchronization, coordinate conversion, data association technology and comparison learning algorithms, a situation information representation model for multi-source heterogeneous data is constructed, and a long and short-term memory network (LSTM) is used for trajectory prediction.
It improves the accuracy of trajectory prediction and the robustness of the system, enhances the support capabilities of airport scene management and flight safety decision-making, and promotes information sharing and decision-making optimization.
Smart Images

Figure CN120047482A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aviation data processing, and in particular to a method for predicting scene trajectories enhanced by fusion of video surveillance and ADS-B, and also to a device for predicting scene trajectories enhanced by fusion of video surveillance and ADS-B. Background Art
[0002] In recent years, with the rapid development of the aviation industry, my country's total air transport volume has ranked second in the world, with 39 airports with a passenger volume of more than 10 million, and 10 of the top 50 airports in the world in terms of throughput. The rapid growth of transport volume has posed a huge challenge to the airport's security capabilities. As airports become more busy, on-site security incidents are also on the rise, mainly reflected in the increasing number of security incidents and increasingly prominent flight delays.
[0003] Real-time situational awareness of the airport flight zone is extremely critical to maintaining the operational safety of the airport. It is also a prerequisite for achieving scene trajectory prediction. It involves real-time monitoring and analysis of various dynamic and static elements of the airport scene. However, traditional airport scene situational awareness systems often rely on a single information source, such as radar systems, ground surveillance equipment, or video surveillance. Although this single-modal detection method can provide the necessary monitoring information to a certain extent, it also has some significant disadvantages: (1) Monitoring blind spots: Since single-modal detection systems rely on a single type of sensor or data source, they may not be able to fully cover all key areas within the airport flight zone, resulting in monitoring blind spots. For example, severe weather conditions may affect the detection effect of radar, while video surveillance may be limited by field of view and lighting conditions. (2) A single data source may not provide enough information to accurately identify and judge complex scenes and events. For example, in areas with dense aircraft, vehicles, and personnel, single visual monitoring may find it difficult to distinguish different targets and behaviors. (3) Single-modal detection systems may lack flexibility and adaptability in data processing and analysis. Since it can only process one type of data, it may not be able to adjust and optimize the algorithm in a timely manner when faced with changing environments and complex scenarios, thus affecting the accuracy and real-time nature of situational awareness. (4) Insufficient prediction accuracy: When dealing with the interaction of dynamic and static elements, a single-modal system may not be able to accurately capture all factors that affect the trajectory, such as the mutual influence between aircraft and the movement of ground service vehicles, which are crucial for accurately predicting the trajectory. (5) When a single information source fails or fails, the entire system may face greater risks due to the lack of other data sources to supplement or back up. Summary of the invention
[0004] To overcome the deficiencies of the prior art, the technical problem to be solved by the present invention is to provide a method for fusing video surveillance and ADS-B to enhance surface trajectory prediction, which can improve the accuracy and reliability of airport surface situation awareness, and then accurately predict the operation trajectories of aircraft and other transportation vehicles, providing strong technical support for airport surface management and flight safety.
[0005] The technical solution of the present invention is as follows: This method for fusing video surveillance and ADS-B to enhance surface trajectory prediction includes the following steps:
[0006] (1) Synchronize video surveillance data and ADS-B data using timestamps, and achieve spatio-temporal alignment of different modality data through coordinate transformation;
[0007] (2) Determine the corresponding relationships between different data sources through data association technology, and construct positive and negative sample pairs;
[0008] (3) Use the contrastive learning algorithm to pre-train multi-source heterogeneous data to extract and fuse key features, and construct an efficient situation information representation model;
[0009] (4) Use the pre-trained model as the basis for trajectory feature extraction, and input it into the long short-term memory network LSTM for trajectory prediction of the airport surface.
[0010] The beneficial effects of the present invention are as follows:
[0011] (1) Improve the prediction accuracy: By making full use of the fusion of video surveillance and ADS-B data, this method can effectively improve the prediction accuracy. The comprehensive utilization of multi-modal data and the application of the contrastive learning algorithm enable the system to obtain rich information from different perspectives and more accurately capture the movement trajectories of aircraft and other transportation vehicles during the prediction process, thus effectively improving the prediction accuracy and reliability.
[0012] (2) Enhance the robustness of the system: By utilizing the complementarity of multi-source heterogeneous data, the system can better cope with environmental changes and sensor errors, thus ensuring the stability and reliability of the system in various complex situations. Even when some data sources are interfered with or fail, the system can accurately maintain the perception and prediction capabilities of the airport surface, ensuring the safety and continuity of airport operations.
[0013] (3) Promote information sharing and decision optimization: By providing accurate and reliable situation awareness capabilities, this method provides strong technical support for airport surface management and flight safety decision-making. Each participating party can share more accurate and consistent situation information, thereby improving the timeliness and effectiveness of decision-making, promoting the optimal allocation of airport resources, and enhancing the overall operation efficiency.
[0014] There is also provided a video surveillance and ADS-B fusion enhanced scene trajectory prediction device, which includes:
[0015] A synchronization and alignment module, configured to synchronize video surveillance data and ADS-B data using timestamps, and achieve spatio-temporal alignment of different modality data through coordinate transformation;
[0016] A sample pair construction module, configured to determine the corresponding relationship between different data sources through data association technology and construct positive and negative sample pairs;
[0017] A characterization model construction module, configured to pre-train multi-source heterogeneous data using a contrastive learning algorithm to extract and fuse key features, and construct a situation information characterization model;
[0019] A trajectory prediction module, configured to use the pre-trained model as the basis for trajectory feature extraction and input it into a long short-term memory network (LSTM) for trajectory prediction of the airport scene. Brief Description of the Drawings
[0020] Figure 1 Shown is a flowchart of the video surveillance and ADS-B fusion enhanced scene trajectory prediction method according to the present invention.
[0021] Figure 2 Shown is a flowchart of the spatio-temporal alignment method of step (1) of the video surveillance and ADS-B fusion enhanced scene trajectory prediction method according to the present invention.
[0022] Figure 3 Shown is a schematic diagram of multi-modal fusion based on contrastive learning. Detailed Embodiments
[0023] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0024] In order to make the description of the present disclosure more detailed and complete, the following provides an illustrative description of the embodiments and specific examples of the present invention; however, this is not the only form of implementing or applying the specific embodiments of the present invention. The embodiments cover the features of multiple specific embodiments and the method steps and their sequences for constructing and operating these specific embodiments. However, other specific embodiments can also be used to achieve the same or equivalent functions and step sequences.
[0025] As Figure 1 shown, this video surveillance and ADS-B fusion enhanced scene trajectory prediction method includes the following steps:
[0026] (1) Synchronize video surveillance data and ADS-B data using timestamps, and achieve spatio-temporal alignment of different modality data through coordinate transformation;
[0027] (2) Determine the corresponding relationships between different data sources through data association technology, and construct positive and negative sample pairs;
[0028] (3) Use contrastive learning algorithms to pre-train multi-source heterogeneous data to extract and fuse key features, and construct an efficient situation information representation model;
[0029] (4) Use the pre-trained model as the basis for trajectory feature extraction, and input it into the long short-term memory network LSTM for trajectory prediction of the airport surface.
[0030] The beneficial effects of the present invention are as follows:
[0031] (1) Improve the prediction accuracy: By making full use of the fusion of video surveillance and ADS-B data, this method can effectively improve the prediction accuracy. The comprehensive utilization of multi-modal data and the application of contrastive learning algorithms enable the system to obtain rich information from different perspectives and more accurately capture the movement trajectories of aircraft and other vehicles during the prediction process, thus effectively improving the prediction accuracy and reliability.
[0032] (2) Enhance the robustness of the system: By leveraging the complementarity of multi-source heterogeneous data, the system can better cope with environmental changes and sensor errors, thus ensuring the stability and reliability of the system under various complex conditions. Even when some data sources are disturbed or fail, the system can accurately maintain the perception and prediction capabilities of the airport surface, ensuring the safety and continuity of airport operations.
[0033] (3) Facilitate information sharing and decision optimization: By providing accurate and reliable situation awareness capabilities, this method provides strong technical support for airport surface management and flight safety decision-making. Each participating party can share more accurate and consistent situation information, thereby improving the timeliness and effectiveness of decision-making, promoting the optimal allocation of airport resources, and enhancing the overall operation efficiency.
[0034] Preferably, as Figure 2 shown, the step (1) includes the following sub-steps:
[0035] (1.1) Convert different format times to the same time axis, collect data with timestamps from multiple data sources, and convert the data time to a unified time format;
[0036] (1.2) Perform time synchronization by calculating the time difference between data source timestamps and comparing it with a threshold to achieve data comparison and analysis;
[0037] (1.3) Perform coordinate transformation and mapping, map to the pixel coordinate axis of the image using coordinate transformation, and perform target matching.
[0038] Preferably, the step (1.2) includes the following sub-steps:
[0039] (1.2.1) Convert the data source timestamps of the two modalities to the same format, and set the alignment threshold to Δt;
[0040] (1.2.2) Calculate the difference δt between the two modality timestamps;
[0041] (1.2.3) Compare the time difference δt with the threshold Δt;
[0042] (1.2.4) When the time difference is less than or equal to the threshold, align the data of different modalities in time; if the time difference is greater than the threshold, the two modalities cannot be aligned.
[0043] Preferably, in the step (2), the data that has completed time alignment and coordinate transformation is matched and associated, and the mutually corresponding multi-modal data forms positive sample pairs, and the rest form negative sample pairs, preparing for subsequent data fusion.
[0044] Preferably, the step (3) includes the following sub-steps:
[0045] (3.1) Construct a single-modal feature extraction tool, use a specific model to extract features from the single-modal information, use a model based on a convolutional neural network to extract features from the image data, construct an encoding network model to extract features from the ADS-B position data, and process the extracted single-modal feature maps into vectorized representations;
[0046] (3.2) By matching the data pairs from the same semantic space, make the distance between the data pairs within the same semantic space as small as possible, and the distance between the data pairs in different semantic spaces as large as possible.
[0047] Preferably, as Figure 3 shown, the step (3.2) includes the following sub-steps:
[0048] (3.2.1) Use a projection head to project the extracted single-modal feature vectors to obtain feature map data of the same dimension for comparison in the embedding space;
[0049] (3.2.2) Perform similarity calculation, calculate the similarity matrix logits between different modality features, and the similarity calculation usually uses cosine similarity. The similarity matrix logits is:
[0050]
[0051] (3.2.3) Calculate the contrastive loss. Use the contrastive loss function to compare the similarity between positive sample pairs and negative sample pairs, convert the similarity matrix into a probability distribution form, and use targets to represent its probability distribution:
[0052]
[0053] Among them, the formula of the softmax function is as follows:
[0054]
[0055] Among them, z i is the original unnormalized score, e is the base of the natural logarithm, and tem represents the temperature,
[0056] is a hyperparameter, and the smoothness of the distribution output by the model is controlled by adjusting the temperature;
[0057] (3.2.4) Use the cross-entropy loss function to calculate the loss values of different modalities respectively, and average the loss values of each
[0058] modality to obtain the final loss function;
[0059] (3.2.5) Backpropagation optimization: Update the model parameters according to the loss function to optimize the matching degree of similarity.
[0060] Preferably, step (4) includes the following sub-steps:
[0061] (4.1) Construct a feature set of a time series, arrange the trajectory features extracted in several consecutive frames or within a period of time in chronological order to form a feature sequence;
[0062]
[0063] (4.2) Load the ADSB feature extraction model pre-trained by the contrastive learning algorithm in step (3);
[0064]
[0065] (4.3) Extract the trajectory features, input the position information of ADSB into the loaded feature
[0066] extraction model to extract the key features for trajectory prediction;
[0067] (4.4) Design and configure an LSTM network, which can process time series data and predict future trajectory points. The LSTM network is designed to capture long-term dependencies to adapt to the time series characteristics of trajectory prediction;
[0068] (4.5) Train the LSTM network using the constructed feature sequences and known trajectory data. During the training process, the LSTM network learns how to predict future trajectory points based on the input feature sequences.
[0069] future trajectory points;
[0070] (4.6) After the LSTM network is trained, input the latest trajectory features into the network to predict the next trajectory points. The recursive nature of the LSTM allows it to consider historical information and generate
[0071] continuous trajectory predictions;
[0072] (4.7) Analyze the trajectory predicted by the LSTM network to evaluate its accuracy and reliability, and evaluate the prediction performance by comparing it with the actual trajectory data.
[0073] Those of ordinary skill in the art can understand that all or part of the steps in implementing the above-described embodiment methods can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes the steps of the above-described embodiment methods, and the storage medium can be: ROM / RAM, magnetic disk, optical disk, memory card, etc. Therefore, corresponding to the method of the present invention, the present invention also simultaneously includes a video surveillance and ADS-B fusion enhanced scene trajectory prediction device, which is usually represented in the form of functional modules corresponding to the steps of the method. The device includes:
[0074] A synchronization and alignment module configured to synchronize video surveillance data and ADS-B data using timestamps and achieve spatio-temporal alignment of different modality data through coordinate transformation;
[0075] A sample pair construction module configured to determine the corresponding relationships between different data sources through data association techniques and construct positive and negative sample pairs;
[0076] A feature model construction module configured to pre-train multi-source heterogeneous data using a contrastive learning algorithm to extract and fuse key features and construct a situation information feature model;
[0078] A trajectory prediction module configured to use the pre-trained model as the basis for trajectory feature extraction and input it into a long short-term memory network (LSTM) for trajectory prediction of the airport scene.
[0079] Preferably, the synchronization and alignment module performs the following sub-steps:
[0080] (1.1) Convert different format times to the same time axis, collect data with timestamps from multiple data sources, and convert the data time to a unified time format;
[0081] (1.2) Perform time synchronization, calculate the time difference between the timestamps of the data sources and compare it with a threshold value to achieve data comparison and analysis;
[0082] (1.3) Perform coordinate transformation and mapping, map to the pixel coordinate axes of the image using coordinate transformation for target matching.
[0083] Preferably, the characterization model construction module performs the following sub-steps:
[0084] (3.1) Construct a single-modal feature extraction tool, use a specific model to extract features from the single-modal information, use a model based on a convolutional neural network to extract features from the image data, construct an encoding network model for the position data of ADS-B to extract features, and process the extracted single-modal feature maps into vectorized representations;
[0085] (3.2) By matching the data pairs from the same semantic space, make the distance between the data pairs within the same semantic space as small as possible, while the distance between the data pairs in different semantic spaces as large as possible.
[0086] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for enhancing scene trajectory prediction by integrating video surveillance and ADS-B, characterized in that: It includes the following steps: (1) Use timestamps to synchronize video surveillance data and ADSB data, and achieve spatial and temporal alignment of different modal data through coordinate transformation; (2) Determine the correspondence between different data sources through data association technology and construct positive and negative sample pairs; (3) Use contrastive learning algorithms to pre-train multi-source heterogeneous data to extract and fuse key features and build an efficient situation information representation model; (4) The pre-trained model is used as the basis for trajectory feature extraction and input into the long short-term memory network (LSTM) for trajectory prediction of airport scenes.
2. The method for enhancing scene trajectory prediction by fusion of video surveillance and ADS-B according to claim 1 is characterized in that: The step (1) comprises the following sub-steps: (1.1) Convert different time formats to the same timeline, collect data with timestamps from multiple data sources, and convert the data time to a unified time format; (1.2) Perform time synchronization by calculating the time difference between the timestamps of the data sources and comparing it with the threshold to achieve data comparison and analysis; (1.3) Perform coordinate conversion and mapping, and use coordinate conversion to map to the pixel coordinate axis of the image for target matching.
3. The method for enhancing scene trajectory prediction by fusion of video surveillance and ADS-B according to claim 2 is characterized in that: The step (1.2) comprises the following sub-steps: (1.2.1) Convert the data source timestamps of the two modalities to the same format and set the alignment threshold to Δt; (1.2.2) Calculate the difference δt between the two modal timestamps; (1.2.3) Compare the time difference δt with the threshold Δt; (1.2.4) When the time difference is less than or equal to the threshold, the data of different modes are time-aligned; if the time difference is greater than the threshold, the two modes cannot be aligned.
4. The method for enhancing scene trajectory prediction by fusion of video surveillance and ADS-B according to claim 3 is characterized in that: In the step (2), the data after time alignment and coordinate conversion are matched and associated, and the corresponding multimodal data form positive sample pairs, and the rest form negative sample pairs, in preparation for subsequent data fusion.
5. The method for enhancing scene trajectory prediction by fusion of video surveillance and ADS-B according to claim 4 is characterized in that: The step (3) comprises the following sub-steps: (3.1) Construct a single-modal feature extraction tool, use a specific model to extract features from single-modal information, use a convolutional neural network-based model to extract features from image data, build a coding network model to extract features from ADS-B location data, and process the extracted single-modal feature map into a vectorized representation; (3.2) By matching data pairs from the same semantic space, the distance between data pairs in the same semantic space is made as small as possible, while the distance between data pairs in different semantic spaces is made as large as possible.
6. The method for enhancing scene trajectory prediction by fusion of video surveillance and ADS-B according to claim 5 is characterized in that: The step (3.2) comprises the following sub-steps: (3.2.1) Use the projection head to project the extracted unimodal feature vector to obtain feature map data of the same dimension for comparison in the embedding space; (3.2.2) Perform similarity calculation and calculate the similarity matrix logits between different modal features. The similarity calculation usually uses cosine similarity. The similarity matrix logits is: (3.2.3) Perform contrast loss calculation, use contrast loss function to compare the similarity between positive sample pairs and negative sample pairs, convert the similarity matrix into probability distribution form, and use targets to represent its probability distribution: The formula of the softmax function is as follows: where z i is the original unnormalized score, e is the base of the natural logarithm, tem represents the temperature, which is a hyperparameter that controls the smoothness of the distribution of the model output by adjusting the temperature; (3.2.4) Use the cross entropy loss function to calculate the loss values of different modes respectively, and average the loss values of each mode to obtain the final loss function; (3.2.5) Back-propagation optimization: According to the loss function, the model parameters are updated to optimize the matching degree of similarity.
7. The method for enhancing scene trajectory prediction by fusion of video surveillance and ADS-B according to claim 6 is characterized in that: The step (4) comprises the following sub-steps: (4.1) Construct a time series feature set, and arrange the trajectory features extracted from several consecutive frames or a period of time in chronological order to form a feature sequence; (4.2) loading the ADSB feature extraction model pre-trained by the contrastive learning algorithm in step (3); (4.3) Extract trajectory features, input the location information of ADSB into the loaded feature extraction model, and extract key features for trajectory prediction; (4.4) Design and configure an LSTM network that can process time series data and predict future trajectory points. The LSTM network is designed to capture long-term dependencies to adapt to the time series characteristics of trajectory prediction; (4.5) Use the constructed feature sequence and known trajectory data to train the LSTM network. During the training process, the LSTM network learns how to predict future trajectory points based on the input feature sequence; (4.6) After the LSTM network training is completed, the latest trajectory features are input into the network to predict the next trajectory points. The recursive nature of LSTM allows it to take into account historical information and generate continuous trajectory predictions; (4.7) The trajectory predicted by the LSTM network is analyzed to evaluate its accuracy and reliability, and the prediction performance is evaluated by comparing it with the actual trajectory data.
8. Video surveillance and ADS-B fusion enhanced scene trajectory prediction device, characterized by: It includes: A synchronization and alignment module, which is configured to synchronize video surveillance data and ADSB data using timestamps and achieve time-space alignment of different modal data through coordinate transformation; A sample pair construction module, which is configured to determine the correspondence between different data sources through data association technology and construct positive and negative sample pairs; A representation model building module configured to use a contrastive learning algorithm to pre-train multi-source heterogeneous data to extract and fuse key features and build a situation information representation model; The trajectory prediction module is configured to use a pre-trained model as the basis for trajectory feature extraction and input it into the long short-term memory network LSTM for trajectory prediction of the airport scene.
9. The video surveillance and ADS-B fusion enhanced scene trajectory prediction device according to claim 8, characterized in that: The synchronization and alignment module performs the following sub-steps: (1.1) Convert different time formats to the same timeline, collect data with timestamps from multiple data sources, and convert the data time to a unified time format; (1.2) Perform time synchronization by calculating the time difference between the timestamps of the data sources and comparing it with the threshold to achieve data comparison and analysis; (1.3) Perform coordinate conversion and mapping, and use coordinate conversion to map to the pixel coordinate axis of the image for target matching.
10. The video surveillance and ADS-B fusion enhanced scene trajectory prediction device according to claim 9, characterized in that: The characterization model building module performs the following sub-steps: (3.1) Construct a single-modal feature extraction tool, use a specific model to extract features from single-modal information, use a convolutional neural network-based model to extract features from image data, build a coding network model to extract features from ADS-B location data, and process the extracted single-modal feature map into a vectorized representation; (3.2) By matching data pairs from the same semantic space, the distance between data pairs in the same semantic space is made as small as possible, while the distance between data pairs in different semantic spaces is made as large as possible.
Citation Information
Cited By
Airport operation situation dynamic prediction method and device, equipment, storage medium and program product
CN120688694A
Biological diversity evaluation method and device based on space-time intelligent service calculation engine
CN120873197A
Low-altitude aircraft state monitoring method and system based on 5G base station iron tower
CN120913453A