A virtual listing method for aircraft based on ADS-B information and time-space matching

By adopting a detection network based on ADS-B information and time-space matching in aircraft virtual listing technology, the problem of delay and accuracy of positioning information is solved, and high-precision and stable virtual listing effect in complex scenarios is achieved.

CN114881946BActive Publication Date: 2025-05-02YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210442194.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-25
Publication Date
2025-05-02
Estimated Expiration
2042-04-25

AI Technical Summary

Technical Problem

The existing aircraft virtual listing technology has problems such as delay and reduced accuracy in airport scene monitoring, which leads to shaking or suddenly disappearing of listing information. It depends on the accuracy of the target detection algorithm, and cannot effectively deal with problems such as occlusion and posture changes in complex scenarios.

Method used

A detection network based on ADS-B information and space-time matching is adopted. By obtaining panoramic surveillance video, ground real-time and ADS-B information for preprocessing, multi-scale features are extracted using ResNet and Mix-FPN, and a priori region generation network and space-time matching network are combined to perform target detection and tracking and matching to ensure that the aircraft's virtual listing information can be displayed accurately and stably in complex scenarios.

Benefits of technology

It improves the detection accuracy and efficiency of aircraft in panoramic videos, enhances the detection effect of the detection network under long-term monitoring, improves the detection performance under problems such as occlusion and posture changes, and improves the accuracy and real-timeness of virtual listings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114881946B_ABST
    Figure CN114881946B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for virtual placarding of aircraft based on ADS-B information and time-space matching. The method optimizes the detection result through the time-space matching module, improves the detection effect of the algorithm under the conditions of occlusion, intermittent entry and exit, motion posture change, etc. in the airport panoramic scene, and improves the placarding identification accuracy of the aircraft in the airport panoramic scene; combines with a tracker to continuously optimize the detection and tracking of the initial detection target, and improves the reliability of the algorithm in aircraft detection identification when the positioning information has problems such as delay or disappearance, and improves the phenomenon of inaccurate placarding identification caused by frequent ID exchange when multiple targets coincide; efficiently updates and judges the aircraft state, and uses the state decay function to represent the effective state of the aircraft. Combined with a multi-target tracker and a matching strategy, the placarding identification problem when the aircraft disappears or the state changes in certain cases is solved, and finally, while ensuring real-time performance, the continuous and effective placarding identification of multiple aircraft under panoramic scene monitoring can be stabilized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of aviation management, and in particular to an aircraft virtual listing method based on ADS-B information and time-space matching. Background Art

[0002] With the rapid development of domestic civil aviation, more and more large-scale airports have been established all over the country. At the same time, in recent years, ICAO has launched a new scene monitoring system "Advanced Scene Motion Guidance Control System", which requires the detection and identification of moving targets in the airport scene. Under the new system requirements, panoramic video monitoring technology with multi-angle monitoring image stitching technology and aircraft virtual listing technology play an increasingly important role in airport scene monitoring. Through panoramic video monitoring of airport scenes, comprehensive monitoring of the complete area of ​​current large-scale airports can be achieved. Combined with virtual listing technology, the relevant information of aircraft can be continuously monitored and identified on the monitoring video, which effectively improves the efficiency of airport monitoring and command work, and has extremely high practical application value in the monitoring system of airport scenes.

[0003] The existing aircraft virtual placards can be divided into two types based on the projection of airport and aircraft positioning information and the placarding based on video target detection and positioning information from the perspective of technical implementation. The method based on positioning information projection mainly relies on the matching of prior positioning information such as ADS-B (Automatic Dependent Surveillance-Broadcast) automatic dependent surveillance broadcast technology. ADS-B is a technology vigorously launched by the Civil Aviation Administration in recent years, which can obtain detailed information such as the flight number, longitude and latitude, speed, and heading of the aircraft. This implementation method is highly dependent on the accuracy and reliability of the positioning system. By converting and projecting the acquired real-world coordinates into the corresponding two-dimensional coordinates of the panoramic surveillance video, the relevant information of the aircraft is displayed on the positioned video screen, thereby completing the virtual placarding of the aircraft. However, since ADS-B is sent at intervals and may be affected by relevant factors such as the environment, resulting in a decrease in the delay and accuracy of the positioning information, these problems may cause the mapped coordinate information to be unreliable, and the actual display will cause problems such as shaking and sudden disappearance of the placard information.

[0004] The implementation method based on video target detection and positioning information is to use the target detection algorithm to match the detected aircraft with the positioning information, and finally display the aircraft virtual plaque information on the video screen according to the positioned aircraft. This method overcomes the limitations of the unstable positioning system and can display the plaque information in real time. However, when the positioning information is delayed or missing, the detected aircraft will disappear due to the lack of actual positioning information and will not be able to continuously display the plaque information. At the same time, this method depends on the accuracy of the panoramic video target detection algorithm. In particular, many current target detection algorithms are not specifically designed for airport panoramic scenes, and cannot effectively deal with weather changes, aircraft posture changes, motion staggered occlusion, multi-scale and other issues at the airport scene. Especially in panoramic videos, these problems are more prominent, seriously affecting the implementation and display of virtual plaque information.

[0005] In summary, due to the complexity of airport panoramic scene monitoring, the current aircraft virtual signage technology still has some of the above-mentioned defects, and it is necessary to explore new strategies and practical algorithms to solve the many challenges of aircraft virtual signage technology in airport scenes. Summary of the invention

[0006] In view of the above-mentioned deficiencies in the prior art, the present invention provides an aircraft virtual listing method based on ADS-B information and time-space matching.

[0007] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0008] A method for virtual listing of aircraft based on ADS-B information and time-space matching, comprising the following steps:

[0009] S1. Obtain and preprocess the panoramic surveillance video, the ground truth of the corresponding frame, and the ADS-B information, perform redundant segmentation on the high-resolution panoramic image, and convert the ground truth of the corresponding frame and the ADS-B coordinates accordingly, and name and save the preprocessed data;

[0010] S2, taking the preprocessed data in S1 as input, using ResNet to extract features, combining the proposed hybrid feature pyramid Mix-FPN to obtain multi-scale features, and using the prior region generation network Priori RPN to generate predicted candidate boxes and prior candidate boxes, and using the region of interest pool ROI Pooling to obtain features for category determination and bounding box regression;

[0011] S3, sending the features obtained in S2 to the proposed spatiotemporal matching network, matching them with the historical target features in turn, matching the results according to the matching strategy, taking the prediction result with the highest average matching degree as the output of the detection network, and maintaining and updating the target template;

[0012] S4, using the tracker to perform similarity matching between the target tracking detection result of the current frame and the prediction result of the detection network, and updating and maintaining the target state of the tracker according to the judgment detection result;

[0013] S5. Perform continuous aircraft multi-target position tracking and matching based on the detection frame result of step S4 to determine whether the positioning information is missing. If so, reacquire the positioning information. If not, virtually label and display the aircraft based on the aircraft status and the detection frame result.

[0014] The beneficial effects of the above scheme are:

[0015] 1) A detection network based on ADS-B and time-space matching was invented, which improved the detection accuracy and efficiency of aircraft in panoramic videos, greatly improved the detection effect of the detection network under long-term monitoring, improved detection problems such as occlusion, posture changes, and frequent entry and exit, and effectively improved the accuracy and real-time performance of virtual listings.

[0016] 2) By using status monitors and tracker matching strategies, aircraft on the scene are continuously tracked, which effectively solves the problem of incorrect display of virtual signage information when multiple aircraft overlap or are misplaced. At the same time, it overcomes the problem of the signage algorithm being highly dependent on positioning information, which can effectively improve the stability of aircraft detection identification and greatly improve the practicality and reliability of the signage algorithm.

[0017] 3) The reliability judgment and correction of the aircraft detection results are completed according to the state matching strategy of the positioning information, and the state monitor is used to protect the detection results, which effectively improves the accuracy of the aircraft plate identification detection results and the adaptability in different scenarios, and improves the robustness of the algorithm.

[0018] Furthermore, the S1 specifically includes the following steps:

[0019] S11, converting the coordinate information of ADS-B into the coordinates of the real world, and performing coordinate mapping through the homography matrix according to the arrangement position of the monitoring camera to obtain the corresponding coordinates of each aircraft in the high-resolution panoramic video image;

[0020] S12, cutting the input high-resolution panoramic image and the corresponding bottom live image through a sliding window, and cutting an area with an overlapping ratio of 1 / 4 as the predicted input image data;

[0021] S13, performing translation transformation on the corresponding ADS-B coordinate information using the cropped upper left corner coordinates to obtain the ADS-B coordinate information corresponding to the pre-processed image.

[0022] The beneficial effect of the above further solution is that, without adding additional equipment conditions, the existing equipment information can be fully utilized to realize the positioning of the aircraft, while achieving efficient panoramic image preprocessing.

[0023] Furthermore, the S2 specifically includes:

[0024] S21, using the preprocessed data obtained in S1 as input, extracting features from the input image through ResNet, and processing the features at different levels through Mix-FPN to obtain multi-scale features for prediction;

[0025] S22, Priori RPN generates candidate boxes through the original candidate generation network RPN, and generates a priori target area probability map and corresponding label information through the ADS-B positioning information to generate additional candidate boxes;

[0026] S23, the multi-scale features of the candidate box area are used as the region of interest as the prediction feature, and the target category judgment and detection box are obtained through Soft-NMS processing. During training, the complete intersection-over-union loss function CIOU Loss and the binary cross-entropy loss function binary cross-entropy Loss are used as the loss functions for bounding box regression and category prediction;

[0027] The beneficial effect of the above further scheme is to obtain the shallow and deep features of the input through the ResNet network, communicate and fuse the feature information of different levels through Mix-FPN, and propose a priori RPN network to improve the original RPN network, and generate a priori target frame through ADS-B information to assist in improving the detection effect. At the same time, the proposed spatiotemporal matching module is used to perform similar matching on the predicted target in the corresponding embedding space. Through the target template update and maintenance strategy, the detection effect of the detection network for difficult targets is further improved, and the detection accuracy under occlusion, posture change, frequent entry and exit, etc. is improved.

[0028] Considering the positional relationship between the aircraft prediction boxes, traditional regression prediction can only compare the intersection and union of IOUs, and cannot distinguish between overlapping or mutually contained situations, which will reduce the accuracy of the network prediction for overlapping or close-range targets. In order to further improve the detection performance in cases of mutual occlusion, motion overlap, etc.

[0029] Furthermore, the multi-scale features used for prediction in S21 are expressed as:

[0030]

[0031] Among them, Conv represents the convolution layer, Down represents downsampling, and Up represents upsampling. represents the feature concatenation of the channel dimension, Represents the multi-scale features used for prediction, F1, F2, F3 are the output features of the FPN network, expressed as:

[0032]

[0033] Conv is the identifier convolution layer, Down ×2 is a downsampling operation, and {X1,X2,X3} are the output features of different levels in ResNet.

[0034] The beneficial effect of the above further scheme is that it completes the exchange of feature information at different levels, realizes the fusion of shallow features and deep features, enhances the abstract semantic features of shallow features, and also supplements the missing spatial information in deep features, propagating and enhancing the feature information of different receptive fields.

[0035] Furthermore, in S22, a probability map of the target area and corresponding marking information are generated according to the ADS-B information, an anchor point is added to the probability center, and anchor shapes of 1:1, 2:1, and 3:1 are used as prior candidate boxes of the RPN network, respectively. The original candidate box and the prior candidate box are mapped to the multi-scale feature map through coordinates for subsequent category prediction and frame regression.

[0036] The beneficial effect of the above further scheme is that the accuracy of the candidate frame predicted by the RPN network is increased by using the prior candidate frame as auxiliary information, and the prior anchor shape of the aircraft is used as the prior candidate frame size, which further improves the accuracy of the predicted candidate frame and effectively improves the network's detection effect in complex scenes.

[0037] Furthermore, the S3 specifically includes:

[0038] S31, the original RPN and prior candidate box prediction results in the prior region generation network are combined in the multi-scale feature The features of the mapped area are sent to the spatiotemporal matching module in sequence, and the one with the highest average similarity is selected as the output result of the detection network;

[0039] S32: Update the target template according to the final detection result of the detection network, and maintain the embedding vector in the template according to the update strategy.

[0040] The beneficial effect of the above further scheme is that it improves the detection effect of the detection network for targets in different environments, achieves more accurate detection and identification of targets in complex environments, and especially has a better enhancement effect on the detection of targets that appear continuously, further improving the long-term detection effect of the algorithm.

[0041] Furthermore, the prediction result of the target detection network in S31 is matched with the features of the prior candidate box and the predicted candidate box through the spatiotemporal matching network, and the spatiotemporal matching module is used to match the target template, and the one with the highest average similarity is selected as the final detection result. The calculation expression is:

[0042] Max(SPM(f1),SPM(f2))

[0043] Where f1 and f2 represent the mapping features of the prior candidate box and the predicted candidate box in the feature The mapping features on the target template are Max, which means the target corresponding to the target template with the highest average similarity is selected as the output result. SPM means the spatiotemporal matching module. The calculation expression is:

[0044] DisMax(AvgPool(Conv(f)))

[0045] Where Conv represents the convolution layer, AvgPool represents the global average pooling, DisMax represents the cosine distance matching module, and the calculation expression is:

[0046] Max(Dis(f1,P),Dis(f2,P))

[0047] Among them, Max represents the candidate box with the maximum value as the output result of the detection network, Dis represents the cosine distance matching module, and the calculation expression is:

[0048]

[0049] Where f represents the embedding vector of the input feature, i represents the target template, and x i ,y i , z i Represents the embedding vector of the target in the i-th target template.

[0050] The beneficial effect of the above further scheme is that it improves the detection effect of targets in complex environments, especially in cases of occlusion, similar background color, dark environment, etc., which has a significant enhancement effect on the detection results, further improving the algorithm's continuous detection capability in actual scenes.

[0051] Furthermore, in S32, the prediction result obtained in S32 is used to modify and replace the embedding vector in the target template according to the update strategy.

[0052] The feature embedding vector corresponding to the prediction result is added to the target template set, and the number of embedding vectors in the target template corresponding to the prediction result is determined. When the number of embedding vectors in the target template set is greater than the specified threshold, the embedding vector with the highest matching degree is removed from the target template.

[0053] The beneficial effect of the above further scheme is to improve the network's detection accuracy of aircraft in the current scene, especially for problems such as motion posture changes, occlusion and lighting changes. It has a better enhancement effect on the detection of continuously appearing targets, and improves the detection effect of the algorithm in practical applications.

[0054] Furthermore, in S4, the detection network result and the tracker result are matched with the target template in turn through the spatiotemporal matching module, and the one with the highest average similarity is selected as the final detection result. The calculation expression is:

[0055] Max(SPM(D),SPM(T))

[0056] Where D represents the features corresponding to the detection results of the target detection network, and T represents the features corresponding to the results of the tracker.

[0057] Establish a target tracking management mechanism to implement operations such as target initialization, update, and termination of tracking. According to the final output results, combined with historical tracking data, use IOU and similarity calculation to make the final association of the trajectory, calculate the similarity s between the tracking result and the final detection result, and use IOU as a supplementary evaluation indicator to calculate the IOU value r of the two target boxes. Based on the calculation of the correlation between the two, the calculation expression is:

[0058]

[0059] Where D i , T i Respectively represent the current frame detection result and tracking target, Represents the IOU between the two, and Dis represents the cosine distance.

[0060] According to the association calculation formula, each trajectory T in the trajectory set T is i And each detection result D in the final detection result set j Perform correlation calculation to obtain the correlation matrix M. Use the Hungarian algorithm to optimize the correlation matrix M for the final correlation. When there is a test result D without correlation k When a target is associated with a new target, it is initialized as a new tracking target, and the status and information of the associated target are updated, further improving the detection effect of multi-target tracking.

[0061] The beneficial effect of the above further scheme is that it improves the tracking and detection effect, and comprehensively utilizes the tracker and detection network to optimize the identification of target detection, and further enhances the detection ability of fixed targets, especially for targets that frequently enter and exit and disappear from occlusion.

[0062] Furthermore, in S5

[0063] When the positioning information is missing, the sign information is displayed according to the positioning detection result, and the missing time of the positioning information is maintained. The credibility of the prediction result is proportionally reduced according to the missing time, and the positioning information can be obtained by flushing;

[0064] When positioning information exists, the current positioning information and the detection result are matched, and the historical positioning information and historical detection information of the aircraft are obtained from the status monitor for correlation calculation. If there is no matching object, the preset vacancy length is used for display; if it exists, it is weighted with the matching result of the current information to obtain the credibility of the current detection result, complete the update and correction of the detection result, and use the status monitor to maintain and update the current detection result.

[0065] The beneficial effect of the above further scheme is to use historical matching information to enhance the correlation between long-term series predictions, effectively avoid the problem of repeated ID allocation caused by objects repeatedly entering and exiting the screen or detection errors, and improve the accuracy of virtual listings. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a schematic diagram of the process of the present invention.

[0067] Figure 2 Schematic diagram of panoramic video frames and corresponding Ground Truth used in the present invention.

[0068] Figure 3 Schematic diagram of the target detection network structure of the present invention.

[0069] Figure 4 Schematic diagram of the spatiotemporal matching module of the present invention.

[0070] Figure 5 Schematic diagram of the Mix-FPN module of the present invention.

[0071] Figure 6 It is a schematic diagram of the detection and tracking joint module of the present invention.

[0072] Figure 7 It is a schematic diagram of the display result of the aircraft plaque representation of the present invention. DETAILED DESCRIPTION

[0073] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.

[0074] This article will use the following abbreviations, which are explained as follows:

[0075]

[0076] A method for virtual listing of aircraft based on ADS-B information and time-space matching, as follows Figure 1 As shown, the following steps are included:

[0077] S1. Obtain and preprocess the panoramic surveillance video, the ground truth of the corresponding frame, and the ADS-B information, perform redundant segmentation on the high-resolution panoramic image, and convert the ground truth of the corresponding frame and the ADS-B coordinates accordingly, and name and save the preprocessed data;

[0078] In this embodiment, given the panoramic surveillance video, the Ground Truth of the corresponding frame, and the ADS-B information, the ADS-B coordinate information is first converted into the coordinates of the real world, and the coordinates are mapped through the homography matrix according to the arrangement position of the camera to complete the coordinate transformation of the two-dimensional plane. T , X are the corresponding points in the real world and the video image row, that is, The coordinate transformation formula is:

[0079] X T =HX

[0080] Where H is the homography matrix obtained by calculating the corresponding points between the real world coordinates and the image coordinates using the calibration plate:

[0081]

[0082] in They are the coordinates of the corresponding points in the real world and the two-dimensional video screen, and an overdetermined equation of the form Ak=0 is constructed through multiple corresponding points, and finally the homography matrix corresponding to the video screen is obtained. Subsequently, the coordinate conversion can be completed through the homography matrix corresponding to the monitoring screen.

[0083] Preprocess the acquired high-resolution panoramic video images and the corresponding ground truth, and name and save the preprocessed images;

[0084] The input high-resolution panoramic image and the corresponding Ground Truth are preprocessed, cut into specified sizes and named and saved according to the set naming convention to facilitate network training and testing. The panoramic image and the corresponding Ground Truth are as follows: Figure 2The image is cut by sliding the window, and the overlap ratio of the cropping size is set to 1 / 4 of the area to avoid the object being cut into different areas in advance and ensure the integrity of the object in each area. The cutting method is as follows Figure 2 shown.

[0085] The traditional NMS non-maximum suppression algorithm compares the IOU between the current detection box and the highest scoring detection box with a threshold, and simply returns the confidence of the detection box above the threshold to zero, which will cause targets with large overlapping areas to be missed. The NMS formula is:

[0086]

[0087] where b i is the i-th detection box, M is the detection box with the highest IOU with the GroundTruth box, and a i is the confidence of the detection box after IOU matching, N t Pre-set threshold hyperparameters.

[0088] Considering the phenomenon that aircrafts of different positions and sizes often block each other in panoramic scenes, the present invention uses Soft-NMS to perform non-maximum suppression on the predicted detection frame, effectively improving the detection accuracy of overlapping aircraft and thus the stability of virtual signage during subsequent real position matching. The Soft-NMS formula is:

[0089]

[0090] in Represents the Gaussian function, making it a continuous formula. The Gaussian function is used to decay the scores of adjacent detection frames. The larger the overlapping area, the faster the decay, so as to ensure that the final result can effectively remove duplicate detection frames and improve the detection accuracy of interlaced coverage targets. Finally, all cut areas are mapped and restored to the corresponding original images to complete the detection of panoramic video images.

[0091] To address the problem of repeated detection of objects in overlapping areas, the Soft-NMS non-maximum suppression method is used on the prediction results of splicing recovery to globally deduplicate the bounding box predictions of each area to avoid false detection errors caused by repeated detection.

[0092] S2, taking the preprocessed data in S1 as input, using ResNet to extract features, combining the proposed hybrid feature pyramid (Mix-FPN) to obtain multi-scale features, and generating predicted candidate boxes and prior candidate boxes through the prior region generation network (Priori RPN), and using the region of interest pool (ROI Pooling) to obtain features for category determination and bounding box regression;

[0093] In this embodiment, ResNet is used to obtain features at different levels, and Mix-FPN is used to obtain multi-scale features. The overall structure is as follows: Figure 3 In order to increase the information exchange between features of different scales and fully integrate features at different levels, Mix-FPN is proposed. By sending the features obtained by the backbone network into the multi-scale features obtained by the FPN module, each scale feature is fused with other features respectively, enriching the semantic information of different scales, as shown in Figure 4 shown.

[0094] Features at different levels are scaled through upsampling and downsampling, and then the features are connected using channel splicing operations. The data is reduced in dimension through 1×1 convolution, introducing more nonlinearity and improving the generalization of the network for feature extraction, thereby completing feature fusion at different scales.

[0095] Assume that the multi-scale features obtained by the features of different levels in the encoding network through the FPN network are F1, F2, and F3, then the calculation expression of the output feature is:

[0096]

[0097] in, Represents multi-scale features for prediction, Conv represents convolutional layer, Down represents downsampling, and Up represents upsampling. Represents the feature concatenation of the channel dimension.

[0098] Mix-FPN is used to exchange feature information at different levels, integrate shallow features with deep features, enhance the abstract semantic features of shallow features, supplement the missing spatial information in deep features, propagate and enhance deep position information, and use the output features of Mix-FPN as prediction features to perform regression of prediction boxes and categories.

[0099] Considering the positional relationship between the aircraft prediction frames, traditional regression prediction can only perform IOU intersection and union comparison, and cannot distinguish between overlapping or mutually contained situations, which will reduce the accuracy of the network for overlapping or close-range target prediction. In order to further improve the detection performance in cases of mutual occlusion, motion overlap, etc., the present invention uses CIOULoss as the loss function calculation of the border, while considering the positional relationship between the overlapping area and the center point of the border. When there is border inclusion, the Loss can be calculated by measuring the border distance. The loss function is:

[0100]

[0101] Where Distance_2 represents the Euclidean distance between the predicted center point and the true center point, Distance_C represents the diagonal distance of the minimum circumscribed rectangle, and α represents the parameter for measuring the consistency of the window width:

[0102]

[0103] where w gt ,h gt Indicates the width and height of the GroundTruth border, w p ,h p Represents the width and height of the predicted bounding box.

[0104] Since the present invention is only designed for two categories, background and target, subsequent target matching uses target tracker and historical matching information for judgment, so only the binary cross-entropy loss function binarycross-entropyLoss is needed for the predicted box category:

[0105]

[0106] where y i is the real category, is the predicted box category.

[0107] S3, send the features obtained in S2 to the proposed spatiotemporal matching network, match them with the historical target features in turn, match the results according to the matching strategy, take the prediction result with the highest average matching degree as the output of the detection network, and maintain and update the target template. Specifically,

[0108] S31, the original RPN and prior candidate box prediction results in the prior region generation network are combined in the multi-scale feature The features of the mapped area are sent to the spatiotemporal matching module in sequence, and the one with the highest average similarity is selected as the output result of the detection network;

[0109] In this embodiment, in order to improve the problems of aircraft posture change and appearance similarity, a spatiotemporal matching network is proposed. At the same time, in order to improve the abstract expression ability of features and more accurately express the features of the matching prediction target, the minimum size feature map F3 in Mix-FPN is used to map the prediction target box, and the features of the area mapped by the prediction box on the feature map are used as the feature Target of the prediction target. The features are converted into fixed-dimensional embedding vectors through the embedding layer, and the similarity of the features is compared in the high-dimensional space. The calculation expression is:

[0110] f = Conv(AvgPool(Target))

[0111] Among them, Conv is a 1×1×n convolutional layer, and AvgPool is average pooling.

[0112] Through the spatiotemporal matching module, the predicted candidate boxes and the prior candidate boxes in S2 are sequentially placed in the feature The corresponding mapping area features are used as input to calculate the average matching situation. The calculation formula is as follows:

[0113] Max(SPM(f1), SPM(f2))

[0114] Where f1 and f2 represent the mapping features of the prior candidate box and the predicted candidate box in the feature The mapping features on the target template are Max, which means the target corresponding to the target template with the highest average similarity is selected as the output result. SPM means the spatiotemporal matching module. The calculation expression is:

[0115] DisMax(AvgPool(Conv(f)))

[0116] Where Conv represents the convolution layer, AvgPool represents the global average pooling, DisMax represents the cosine distance matching module, and the calculation expression is:

[0117] Max(Dis(f1,P),Dis(f2,P))

[0118] Among them, Max represents the candidate box with the maximum value as the output result of the detection network, and Dis represents the cosine distance matching module.

[0119] In order to better measure and calculate similarity matching in high-dimensional space, this patent will use the cosine distance that can better highlight the angular relationship between vectors for similarity matching, and sequentially match the target feature embedding vector f with the target template. The correlation matching uses the cosine distance calculation, and the calculation expression is:

[0120]

[0121] Where f represents the embedding vector of the input feature, i represents the target template, and x i ,y i , z i Represents the embedding vector of the target in the i-th target template.

[0122] S32: Update the target template according to the final detection result of the detection network, and maintain the embedding vector in the template according to the update strategy.

[0123] The feature embedding vector corresponding to the prediction result is added to the target template set, and the number of embedding vectors in the target template corresponding to the prediction result is determined. When the number of embedding vectors in the target template set is greater than the specified threshold, the embedding vector with the highest matching degree is removed from the target template.

[0124] According to the matching situation of the matching network, the one with the highest average matching degree is selected as the output of the prior candidate box generation network, and prediction classification and bounding box regression are performed, and the network is used to obtain the detection result of the detection network. The vector set of the target template is maintained in the spatiotemporal matching network. The set retains the embedded vectors of different target features in the adjacent historical frames, and the embedded vectors in the target template are updated in time according to the update strategy, such as Figure 5 shown.

[0125] The overall steps of the update strategy are to add the feature embedding vector corresponding to the prediction result to the target template set, and at the same time determine the number of embedding vectors in the target template corresponding to the prediction result. When the number of embedding vectors in the target template set is greater than the specified threshold, the embedding vector with the highest matching degree is removed from the target template.

[0126] Specifically, each target template corresponds to a set of feature embedding vectors of different targets, which contains n sets of embedding vectors corresponding to the targets. The number of target templates is the same as the number of targets appearing in the historical frame. When the number of embedding vectors in the target template reaches a critical point, the update strategy will be used to update and maintain the embedding vector of the target in the current frame. The embedding vector of the current predicted feature is matched with the target embedding vector in turn, and the one with the highest average similarity is selected as the final detection result. When the target template needs to be updated, the current target feature embedding vector will be retained, and the embedding vector with the highest similarity in the current target template will be removed to ensure that the corresponding target template can better identify targets in different states, including the feature embedding vectors of the target in a variety of postures and scenes, which improves the detection effect under occlusion, deformation and other conditions.

[0127] S4, using the tracker to perform similarity matching between the target tracking detection result of the current frame and the prediction result of the detection network, and updating and maintaining the target state of the tracker according to the judgment detection result;

[0128] In this embodiment, in order to improve the detection accuracy of the algorithm and ensure continuous detection and tracking of aircraft in the scene, the detection network and the tracker will be used to monitor the targets in the scene together, such as Figure 6As shown. First, the detection network is used to detect the aircraft in the current scene to obtain the detection result set D. At the same time, the tracker is used to obtain the tracking target result set T. The corresponding features are respectively sent to the spatiotemporal similarity matching network to calculate their respective similarities. The matching results of the tracker and the detector are integrated to select the aircraft corresponding to the target template with the highest similarity as the final detection result set R, which further improves the detection effect. At the same time, in order to further improve the prediction trajectory accuracy of the tracker, the present invention will use IOU and similarity calculation to make a final association on the trajectory, calculate the similarity s through the tracking result and the final detection result, and use IOU as a supplementary evaluation indicator to calculate the IOU value r of the two target boxes. Establish a target tracking management mechanism to implement operations such as target initialization, update, and termination of tracking. The association calculation formula is as follows:

[0129] m=0.5×s+0.5×r

[0130] According to the association calculation formula, each trajectory T in the trajectory set T is i And each detection result D in the final detection result set j Perform correlation calculation to obtain the correlation matrix M. Use the Hungarian algorithm to optimize the correlation matrix M for the final correlation. When there is a test result D without correlation k When a target is associated with a new target, it is initialized as a new tracking target, and the status and information of the associated target are updated, further improving the detection effect of multi-target tracking.

[0131] S5. Perform continuous aircraft multi-target position tracking and matching based on the detection frame result of step S4 to determine whether the positioning information is missing. If so, reacquire the positioning information. If not, virtually label and display the aircraft based on the aircraft status and the detection frame result.

[0132] When the positioning information is missing, the sign information is displayed according to the positioning detection result, and the missing time of the positioning information is maintained. The credibility of the prediction result is proportionally reduced according to the missing time, and the positioning information can be obtained by flushing;

[0133] Using the tracking results obtained by S4, the tracking identifier and the state monitor are associated one-to-one through hash mapping, and the relevant information and matching status of the aircraft are updated, including the number of matches, matching time, motion trajectory, survival time, etc. When the aircraft enters the current frame, the survival time will begin to increase. When the aircraft disappears from the current frame, the hash mapping table still retains the relevant information of the aircraft and decays the survival time. The decay function expression is:

[0134]

[0135] Where n is the number of matches, t is the current time, and t i The last matching moment.

[0136] As the number of subsequent matches n decreases and the time t increases, the survival time T will gradually decay. When the survival time T decays to a negative number, the hash mapping table will delete the matching information of the aircraft, effectively avoiding the problem of repeated ID allocation caused by objects repeatedly entering and exiting the screen or detection errors, and improving the accuracy of the virtual listing.

[0137] In this embodiment, the detection results of the invented high-precision target detection algorithm are matched with the synchronous tracking of the multi-target tracker to complete the positioning and detection of the aircraft identification. According to whether the actual positioning information is missing, the positioning detection results are branched and displayed.

[0138] When the positioning information is missing, the sign information is displayed according to the positioning detection result, and the missing time of the positioning information is maintained. The credibility of the prediction result is proportionally reduced according to the missing time. When the positioning information is acquired again, the credibility will be re-determined and updated according to the current detection result to ensure the reliability of the detection result;

[0139] When positioning information exists, the current positioning information and the detection result are matched, and the historical positioning information and historical detection information of the aircraft are obtained from the status monitor for correlation calculation. If there is no matching object, the preset vacancy length is used for display; if it exists, it is weighted with the matching result of the current information to obtain the credibility of the current detection result, complete the update and correction of the detection result, and use the status monitor to maintain and update the current detection result.

[0140] The detection frame results of all aircraft are obtained through the detection network, and the target status monitor is used to continuously monitor the matching aircraft status, improve the continuous matching of subsequent detection results with the actual position targets, reduce the dependence on real positioning information, and improve the stability and reliability of the sign.

[0141] Finally, the virtual plaque information of the aircraft is displayed according to the survival status and detection results of each aircraft by the status monitor, such as Figure 7 shown.

[0142] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0143] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0145] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

[0146] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A method for virtual aircraft listing based on ADS-B information and time-space matching, characterized in that: The steps include: S1. Obtain and preprocess the panoramic surveillance video, the ground truth of the corresponding frame, and the ADS-B information, perform redundant segmentation on the high-resolution panoramic image, and convert the ground truth of the corresponding frame and the ADS-B coordinates accordingly, and name and save the preprocessed data; S2, taking the preprocessed data in S1 as input, using ResNet to extract features, combining the hybrid feature pyramid Mix-FPN to obtain multi-scale features, and generating prediction candidate boxes and prior candidate boxes through the prior region generation network Priori RPN, and using the region of interest pool ROI Pooling to obtain features for category determination and bounding box regression, specifically including the following steps: S21, using the preprocessed data obtained in S1 as input, extracting features from the input image through ResNet, and processing the features at different levels through Mix-FPN to obtain multi-scale features for prediction; S22, Priori RPN generates candidate boxes through the original candidate generation network RPN, and generates a priori target area probability map and corresponding label information through the ADS-B positioning information to generate additional candidate boxes; S23, the multi-scale features of the candidate box area are used as the region of interest as the prediction feature, and the target category judgment and detection box are obtained through Soft-NMS processing. During training, the complete intersection-over-union loss function CIOU Loss and the binary cross-entropy loss function are used as the loss function for border regression and category prediction. S3: Send the features obtained in S2 to the spatiotemporal matching network, match them with the historical target features in turn, match the results according to the matching strategy, use the prediction result with the highest average matching degree as the output of the detection network, and maintain and update the target template, including: S31, the original RPN and prior candidate box prediction results in the prior region generation network are combined in the multi-scale feature The features of the mapped area are sent to the spatiotemporal matching module in sequence, and the one with the highest average similarity is selected as the output result of the detection network; S32, updating the target template according to the final detection result of the detection network, and maintaining the embedding vector in the template according to the update strategy; S4, using the tracker to perform similarity matching between the target tracking detection result of the current frame and the prediction result of the detection network, and updating and maintaining the target state of the tracker according to the judgment detection result; S5. Perform continuous aircraft multi-target position tracking and matching based on the detection frame result of step S4 to determine whether the positioning information is missing. If so, reacquire the positioning information. If not, virtually label and display the aircraft based on the aircraft status and the detection frame result.

2. The method for virtual aircraft listing based on ADS-B information and time-space matching according to claim 1, characterized in that: The S1 specifically includes the following steps: S11, converting the coordinate information of ADS-B into the coordinates of the real world, and performing coordinate mapping through the homography matrix according to the arrangement position of the monitoring camera to obtain the corresponding coordinates of each aircraft in the high-resolution panoramic video image; S12, cutting the input high-resolution panoramic image and the corresponding bottom live image through a sliding window, and cutting an area with an overlapping ratio of 1 / 4 as the predicted input image data; S13, performing translation transformation on the corresponding ADS-B coordinate information using the cropped upper left corner coordinates to obtain the ADS-B coordinate information corresponding to the pre-processed image.

3. The method for virtual aircraft listing based on ADS-B information and time-space matching according to claim 1, characterized in that: The multi-scale features used for prediction in S21 are expressed as: Conv means convolution layer, Down means downsampling, and Up means upsampling. represents the feature concatenation of the channel dimension, represents the multi-scale features used for prediction, is the output feature of the FPN network, expressed as: ; To identify the convolutional layer, is the downsampling operation, are the output features of different levels in ResNet.

4. The method for virtual aircraft listing based on ADS-B information and time-space matching according to claim 2, characterized in that: In the S22, a probability map of the target area and corresponding marking information are generated according to the ADS-B information, an anchor point is added to the probability center, and anchor shapes of 1:1, 2:1, and 3:1 are used as the prior candidate boxes of the RPN network, respectively. The original candidate boxes and the prior candidate boxes are mapped to the multi-scale feature map through coordinates, and subsequent category prediction and frame regression are performed.

5. The method for virtual aircraft listing based on ADS-B information and time-space matching according to claim 1, characterized in that: The prediction result of the target detection network in S31 is matched with the features of the prior candidate box and the predicted candidate box through the spatiotemporal matching network and the target template through the spatiotemporal matching module, and the one with the highest average similarity is selected as the final detection result. The calculation expression is: in Respectively represent the mapping features of the prior candidate box and the predicted candidate box in the feature The mapping features on the target template are Max, which means the target corresponding to the target template with the highest average similarity is selected as the output result. SPM means the spatiotemporal matching module. The calculation expression is: Where Conv represents the convolution layer, AvgPool represents the global average pooling, DisMax represents the cosine distance matching module, and the calculation expression is: Among them, Max represents the candidate box with the maximum value as the output result of the detection network, Dis represents the cosine distance matching module, and the calculation expression is: Where f represents the embedding vector of the input feature, i represents the target template, represents the embedding vector of the target in the i-th target template.

6. The method for virtual aircraft listing based on ADS-B information and time-space matching according to claim 1, characterized in that: In S32, the prediction result obtained in S32 is used to modify and replace the embedded vector in the target template according to the update strategy; The feature embedding vector corresponding to the prediction result is added to the target template set, and the number of embedding vectors in the target template corresponding to the prediction result is determined. When the number of embedding vectors in the target template set is greater than the specified threshold, the embedding vector with the highest matching degree is removed from the target template.

7. The method for virtual aircraft listing based on ADS-B information and time-space matching according to claim 1, characterized in that: In S4, the detection network result and the tracker result are matched with the target template in turn through the spatiotemporal matching module, and the one with the highest average similarity is selected as the final detection result. The calculation expression is: Where D represents the feature corresponding to the detection result of the target detection network, and T represents the feature corresponding to the result of the tracker; Establish a target tracking management mechanism to implement operations such as target initialization, update, and termination of tracking. According to the final output results, combined with historical tracking data, use IOU and similarity calculation to make the final association of the trajectory, calculate the similarity s between the tracking result and the final detection result, and use IOU as a supplementary evaluation indicator to calculate the IOU value r of the two target boxes. Based on the calculation of the correlation between the two, the calculation expression is: in Respectively represent the current frame detection result and tracking target, Represents the IOU of the two, Dis represents the cosine distance; According to the association calculation formula, each trajectory in the trajectory set T is Each test result in the final test result set Perform correlation calculation to obtain the correlation matrix M, and use the Hungarian algorithm to optimize the correlation matrix M for the final correlation. When a target is associated with a new target, it is initialized as a new tracking target, and the status and information of the associated target are updated, further improving the detection effect of multi-target tracking.

8. The method for virtual aircraft listing based on ADS-B information and time-space matching according to claim 1, characterized in that: The S5 When the positioning information is missing, the sign information is displayed according to the positioning detection result, and the missing time of the positioning information is maintained. The credibility of the prediction result is proportionally reduced according to the missing time, and the positioning information can be obtained by flushing; When positioning information exists, the current positioning information and the detection result are matched, and the historical positioning information and historical detection information of the aircraft are obtained from the status monitor for correlation calculation. If there is no matching object, the preset vacancy length is used for display; if it exists, it is weighted with the matching result of the current information to obtain the credibility of the current detection result, complete the update and correction of the detection result, and use the status monitor to maintain and update the current detection result.

Citation Information

Patent Citations

  • Road traffic state detecting device based on omnibearing computer vision

    CN101710448A

  • Aircraft listing method based on combination of video analysis and positioning information

    CN108133028A