Airport simulation scene and video fusion method for abnormal flying object identification

By acquiring airport video and 3D scene data, performing frame sampling and abnormal flying object detection, the problems of dynamic object occlusion and lighting synchronization in traditional fusion methods are solved, thereby improving the accuracy of abnormal flying object recognition and aircraft safety.

CN121545104BActive Publication Date: 2026-04-24BEIJING JIRUIXIANG AVIATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING JIRUIXIANG AVIATION TECH CO LTD
Filing Date
2026-01-16
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In the fusion of traditional airport surveillance video and 3D scene, there are dynamic objects that can obstruct or penetrate the view, and the lighting and reflection properties are difficult to synchronize, which reduces the accuracy of abnormal flying object identification and affects aircraft flight safety.

Method used

By acquiring airport video sets and 3D scene data, frame sampling and abnormal flying object detection are performed to generate an initial abnormal flying object information group. Target association and fusion are then carried out to construct dynamic 3D scene data, and abnormal flying object information is added to improve recognition accuracy.

Benefits of technology

It improves the accuracy of anomalous flying object identification and aircraft flight safety, and enables full-trajectory, full-view tracking and visualization of anomalous flying objects, supporting bird deterrence or anti-drone strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545104B_ABST
    Figure CN121545104B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an airport simulation scene and video fusion method for abnormal flying object identification. A specific implementation of the method comprises: acquiring an airport video set and airport three-dimensional scene data; frame sampling each airport video in the airport video set to obtain a video frame image sequence set; performing abnormal flying object detection on each video frame image sequence in the video frame image sequence set to obtain an initial abnormal flying object information group set; performing target association on each initial abnormal flying object information group in the initial abnormal flying object information group set to obtain an abnormal flying object information set; fusing the airport three-dimensional scene data and each video frame image sequence in the video frame image sequence set to obtain dynamic three-dimensional scene data; and adding an abnormal flying object to the dynamic three-dimensional scene data according to the abnormal flying object information set to obtain abnormal flying object three-dimensional scene data. The implementation can improve the safety of aircraft flight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of digital twin technology, and more specifically to an airport simulation scene and video fusion method for identifying anomalous flying objects. Background Technology

[0002] Traditional airport surveillance videos suffer from limitations in scope and functionality, serving only to detect anomalies and lacking scene awareness and operational coordination capabilities. To overcome these limitations, overlaying and fusing live-action video with virtual simulation models can enhance the realism of virtual scenes, simulating and optimizing the identification process for unusual flying objects (such as birds and drones). Currently, the typical method for fusing 3D scenes with video involves manual adjustments or pre-defined rules.

[0003] However, when using the above method, a common technical problem is that unrealistic occlusion or penetration phenomena occur between dynamic objects and 3D models in the video, and physical properties such as lighting and reflection are difficult to synchronize in real time, resulting in deviations in the fused image, which reduces the accuracy of abnormal flying object recognition and thus reduces the safety of aircraft flight. Summary of the Invention

[0004] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0005] Some embodiments of this disclosure propose an airport simulation scene and video fusion method for identifying anomalous flying objects, in order to solve the technical problems mentioned in the background section above.

[0006] In a first aspect, some embodiments of this disclosure provide a method for airport simulation scene and video fusion for anomalous flying object identification. The method includes: acquiring an airport video set and airport 3D scene data of the airport to be simulated; performing frame sampling on each airport video in the aforementioned airport video set to generate a video frame image sequence, resulting in a video frame image sequence set, wherein each video frame image in the aforementioned video frame image sequence set corresponds to a frame label, the frame label including keyframes and non-keyframes; performing anomalous flying object detection on each video frame image sequence in the aforementioned video frame image sequence set to generate an initial anomalous flying object information group, resulting in an initial anomalous flying object information group set; performing target association on each initial anomalous flying object information group in the aforementioned initial anomalous flying object information group set, resulting in an anomalous flying object information set; fusing the aforementioned airport 3D scene data with each video frame image sequence in the aforementioned video frame image sequence set to obtain dynamic 3D scene data; and adding anomalous flying objects to the aforementioned dynamic 3D scene data according to the aforementioned anomalous flying object information set, resulting in 3D scene data containing anomalous flying objects.

[0007] Secondly, some embodiments of this disclosure provide an airport simulation scene and video fusion device for abnormal flying object identification. The device includes: an acquisition unit configured to acquire an airport video set and airport 3D scene data of the airport to be simulated; a frame sampling unit configured to perform frame sampling on each airport video in the aforementioned airport video set to generate a video frame image sequence, thereby obtaining a video frame image sequence set, wherein each video frame image included in the aforementioned video frame image sequence set corresponds to a frame label, and the frame label includes keyframes and non-keyframes; and an abnormal flying object detection unit configured to perform frame sampling on each video frame image in the aforementioned video frame image sequence set. The sequence performs abnormal flight object detection to generate an initial abnormal flight object information group, resulting in an initial abnormal flight object information group set; a target association unit is configured to perform target association on each initial abnormal flight object information group in the above-mentioned initial abnormal flight object information group set, resulting in an abnormal flight object information set; a fusion unit is configured to fuse the above-mentioned airport 3D scene data with each video frame image sequence in the above-mentioned video frame image sequence set, resulting in dynamic 3D scene data; an abnormal flight object addition unit is configured to add abnormal flight objects to the above-mentioned dynamic 3D scene data according to the above-mentioned abnormal flight object information set, resulting in 3D scene data containing abnormal flight objects.

[0008] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0009] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0010] The various embodiments of this disclosure have the following beneficial effects: The airport simulation scene and video fusion method for abnormal flying object identification according to some embodiments of this disclosure improves aircraft flight safety. Specifically, the reason for reduced aircraft flight safety is that unrealistic occlusion or penetration phenomena occur between dynamic objects and 3D models in the video, and physical properties such as illumination and reflection are difficult to synchronize in real time, leading to deviations in the fused image and thus reducing the accuracy of abnormal flying object identification. Based on this, the airport simulation scene and video fusion method for abnormal flying object identification according to some embodiments of this disclosure first acquires the airport video set and airport 3D scene data of the airport to be simulated. Then, frame sampling is performed on each airport video in the aforementioned airport video set to generate a video frame image sequence, resulting in a video frame image sequence set. Each video frame image included in the aforementioned video frame image sequence set corresponds to a frame label, which includes keyframes and non-keyframes. This reduces the amount of video data, and the division of frame labels can specifically improve the processing efficiency of subsequent anomaly detection. Next, abnormal flying object detection is performed on each video frame image sequence in the aforementioned video frame image sequence set to generate an initial abnormal flying object information group, resulting in an initial abnormal flying object information group set. By performing abnormal flying object detection on video frame images with frame labels, abnormal flying objects can be accurately identified while reducing computational complexity. Secondly, target association is performed on each initial abnormal flying object information group in the aforementioned initial abnormal flying object information group set to obtain an abnormal flying object information set. Through cross-view association, duplicate recording of abnormal flying objects can be avoided, integrating local abnormal information collected by each camera into global information, effectively improving scene perception capabilities and achieving full trajectory and full-view tracking of abnormal flying objects. Then, the aforementioned airport 3D scene data is fused with each video frame image sequence in the aforementioned video frame image sequence set to obtain dynamic 3D scene data. Thus, static 3D scenes can be combined with dynamic real-world video to construct a highly realistic dynamic 3D scene. Finally, based on the aforementioned abnormal flying object information set, abnormal flying objects are added to the aforementioned dynamic 3D scene data to obtain 3D scene data containing abnormal flying objects. Therefore, the visualized 3D scene can intuitively present the distribution and trajectory of abnormal flying objects, which helps in the formulation of bird deterrence or anti-drone strategies and ensures airport security. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0012] Figure 1 This is a schematic diagram of an application scenario of the airport simulation scene and video fusion method for identifying anomalous flying objects, which is one of the embodiments of this disclosure.

[0013] Figure 2 This is a flowchart of some embodiments of the airport simulation scene and video fusion method for identifying anomalous flying objects according to the present disclosure;

[0014] Figure 3 This is a framework diagram of an abnormal flying object detection model based on the airport simulation scene and video fusion method for abnormal flying object identification disclosed herein.

[0015] Figure 4 This is a structural schematic diagram of some embodiments of the airport simulation scene and video fusion apparatus for identifying anomalous flying objects according to the present disclosure;

[0016] Figure 5 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0020] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0021] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0022] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0023] Figure 1 This is a schematic diagram of an application scenario of the airport simulation scene and video fusion method for identifying anomalous flying objects, which is one of the embodiments of this disclosure.

[0024] exist Figure 1 In the application scenario, firstly, the computing device 101 can receive a video dataset 103 collected by a camera group 102 of the airport to be simulated. The camera group can include multiple cameras facing different angles towards the airport runway. Then, the computing device 101 can generate a video frame image sequence set 104 corresponding to the video dataset 103. Next, abnormal flying object detection is performed on the video frame image sequence set 104 to obtain an initial abnormal flying object information set 105. Afterwards, the computing device 101 can perform target association on each initial abnormal flying object information set in the initial abnormal flying object information set 105 to obtain an abnormal flying object information set 106. The computing device 101 can then fuse preset airport 3D scene data 107 with the video frame image sequence set 104 to obtain dynamic 3D scene data 108. Finally, the computing device 101 can add abnormal flying objects to the dynamic 3D scene data 108 based on the abnormal flying object information set 106 to obtain 3D scene data 109 containing abnormal flying objects.

[0025] It should be noted that the aforementioned computing device 101 can be either hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is software, it can be installed in the hardware devices listed above. It can be implemented as, for example, multiple software programs or software modules used to provide distributed services, or as a single software program or software module. No specific limitations are made here. It should be understood that... Figure 1 The number of computing devices in the system can be arbitrary, depending on the implementation requirements.

[0026] Continue to refer to Figure 2 The diagram illustrates a flow 200 of some embodiments of an airport simulation scene and video fusion method for anomalous flying object identification according to the present disclosure. This airport simulation scene and video fusion method for anomalous flying object identification includes the following steps:

[0027] Step 201: Obtain the airport video set and airport 3D scene data of the airport to be simulated.

[0028] In some embodiments, the entity executing the airport simulation scene and video fusion method for identifying anomalous flying objects (e.g., Figure 1 The computing device 101 shown can acquire a set of airport videos and 3D scene data of the airport to be simulated. The airport to be simulated may include a terminal building, runway, camera array, etc. Each airport video in the set can be a video captured by a camera in the camera array. Each camera in the camera array corresponds one-to-one with a video in the set. The set of airport videos can represent the environmental information around the airport runway over a period of time. The 3D scene data can be a simulation model obtained by modeling the airport to be simulated. The 3D scene data may include, but is not limited to, 3D geometric models of the terminal building and runway.

[0029] Step 202: Frame sampling is performed on each airport video in the above airport video set to generate a video frame image sequence, thus obtaining a video frame image sequence set.

[0030] In some embodiments, the execution entity may sample frames of each airport video in the aforementioned airport video set to generate a video frame image sequence, resulting in a video frame image sequence set. Each video frame image included in the video frame image sequence set corresponds to a frame label. The frame label includes keyframes and non-keyframes. In practice, firstly, each airport video in the aforementioned airport video set can be sampled according to a preset sampling rate to generate a video frame image sequence, resulting in a video frame image sequence set. For example, the sampling rate can be one frame per millisecond. Then, for each video frame image sequence in the aforementioned video frame image sequence set, a preset keyframe allocation algorithm can be used to determine the frame label corresponding to each video frame image in the sequence. The keyframe allocation algorithm can be a motion estimation-based keyframe selection algorithm or a time interval-based keyframe selection algorithm.

[0031] Step 203: Perform abnormal flying object detection on each video frame image sequence in the above video frame image sequence set to generate an initial abnormal flying object information group, thus obtaining the initial abnormal flying object information group set.

[0032] In some embodiments, the execution entity can perform abnormal flight object detection on each video frame image sequence in the video frame image sequence set to generate an initial abnormal flight object information group, thus obtaining an initial abnormal flight object information group set. Each initial abnormal flight object information group in the initial abnormal flight object information group set can represent a specific abnormal flight object in the corresponding airport video. Each initial abnormal flight object information group in the initial abnormal flight object information group set corresponds one-to-one with the airport videos in the airport video set. Each initial abnormal flight object information in the initial abnormal flight object information group can include a local target number, target category, and initial target trajectory. The local target number can be used to uniquely identify an abnormal flight object captured by a single camera. The target category can include birds, drones, etc. The initial target trajectory can be an ordered sequence of three-dimensional coordinates of the same abnormal flight object. In practice, a preset video target detection algorithm can be used to perform abnormal flight object detection on each video frame image sequence in the video frame image sequence set to generate an initial abnormal flight object information group, thus obtaining an initial abnormal flight object information group set. The video target detection algorithm can be DeepSORT or Video Swin Transformer.

[0033] Optionally, the execution entity performs abnormal flying object detection on each video frame image sequence in the video frame image sequence set to generate an initial abnormal flying object information group, which may include the following steps:

[0034] First, for each video frame image in the above video frame image sequence whose frame label is a keyframe, perform the following steps:

[0035] The first sub-step involves performing single-frame abnormal flying object detection on the aforementioned video frame images to generate a single-frame detection result set. Each single-frame detection result in the single-frame detection result set can characterize the category and location of the abnormal flying object in the aforementioned video frame image. Each single-frame detection result in the single-frame detection result set can include the target category, target bounding box, and target location. The target bounding box can be a rectangular box surrounding the abnormal flying object, represented by the coordinates of the four corners of the rectangle. The target location can be three-dimensional coordinates, specifically the coordinates representing the center point of the abnormal flying object. In practice, single-frame abnormal flying object detection can be performed on the aforementioned video frame images using a pre-defined abnormal flying object detection model that includes a single-frame detection layer to generate a single-frame detection result set. The abnormal flying object detection model can be a deep learning model used for abnormal flying object detection in airport videos. The abnormal flying object detection model can include a single-frame detection layer, a moving target detection layer, a target tracking layer, and a trajectory optimization layer. The framework diagram of the abnormal flying object detection model can be as follows: Figure 3 As shown.

[0036] The aforementioned single-frame detection layer can be a network layer that performs object detection on a single video frame image. This single-frame detection layer may include a region block segmentation network, a feature extraction network, an attention network, and a detection head network. The region block segmentation network may include a Region Proposal Network (RPN). The feature extraction network may include a Deep Residual Convolutional Network (ResNet) and a Pyramid Pooling Module (PPM). The attention network may include a CBAM attention network. The detection head network may include a Convolutional Neural Network (CNN), activation functions, and a Fully Connected Layer (FCN).

[0037] The aforementioned moving target detection layer can be a network layer used to identify targets that have undergone displacement in a scene. This moving target detection layer can be used to execute the following second, third, and fourth sub-steps to obtain a set of moving target detection results. The aforementioned moving target detection layer may include an optical flow information generation layer, an optical flow amplitude matrix generation layer, and an edge detection layer. The aforementioned optical flow information generation layer can generate forward and backward optical flow information based on the preceding video frame image corresponding to the aforementioned video frame image. The aforementioned optical flow amplitude matrix generation layer can generate an optical flow amplitude matrix based on the aforementioned forward and backward optical flow information. The aforementioned edge detection layer may include the Canny edge detection algorithm.

[0038] The aforementioned target tracking layer can be a network layer that transmits keyframe detection results from keyframes to non-keyframes. This target tracking layer may include a multi-target tracking network (Simple Online and Realtime Tracking, SORT).

[0039] The trajectory optimization layer described above can be a network layer that integrates detection results between keyframes and non-keyframes. This trajectory optimization layer may include a candidate region sequence set determination layer, a spatiotemporal feature extraction network, a spatiotemporal attention feature extraction network, and a head network. The candidate region sequence set determination layer can be used to generate a sequence of candidate regions for each frame's detection results in adjacent video frames. The spatiotemporal feature extraction network may include a 3D backbone network (e.g., Inflated-3D). The spatiotemporal attention feature extraction network may include a STAN network. The head network may include convolutional layers, activation functions, and fully connected layers.

[0040] Optionally, the aforementioned execution entity performs single-frame abnormal flying object detection on the aforementioned video frame images to generate a single-frame detection result group, which may include the following steps:

[0041] Sub-step one involves segmenting the aforementioned video frame image into region blocks to obtain a set of region blocks. Each region block in the set can be a portion of the aforementioned video frame image. In practice, firstly, a region block segmentation network can be used to determine the candidate bounding box set corresponding to the aforementioned video frame image. This network can be a Region Proposal Network (RPN) or a Selective Search algorithm. The width and height of each candidate bounding box in the set are smaller than the width and height of the aforementioned video frame image. Next, the area covered by each candidate bounding box in the set within the aforementioned video frame image can be defined as a region block, thus obtaining the region block set.

[0042] Sub-step two involves extracting features from each region block in the aforementioned region block set to generate region features, thus obtaining a region feature set. Each region feature in the aforementioned region representation set can represent the global and local texture features of the corresponding region block. In practice, the aforementioned feature extraction network can be used to extract features from each region block in the aforementioned region block set to generate region features and obtain a region feature set.

[0043] Sub-step three involves extracting regional attention features from each regional feature in the aforementioned regional feature set to generate regional attention features, thus obtaining a regional attention feature set. In practice, this can be achieved using the aforementioned attention network to extract regional attention features from each regional feature in the aforementioned regional feature set, thereby generating regional attention features and obtaining a regional attention feature set. The aforementioned attention network can be a CBAM attention network. Specifically, through the aforementioned attention network, channel attention extraction and spatial attention extraction can be performed sequentially on each regional feature to generate regional attention features that highlight key features.

[0044] Sub-step four involves determining the region detection result set corresponding to the aforementioned region attention feature set. Each region detection result in the aforementioned region detection result set can include the target category, target bounding box, target location, and confidence score. In practice, firstly, for each region attention feature in the aforementioned region attention feature set, the detection result set corresponding to the aforementioned region attention feature can be determined using the aforementioned detection head network. The aforementioned detection head network can include a classification branch, a bounding box regression branch, and a location regression branch. The classification branch can be used to generate multiple target categories based on the region attention features. The bounding box regression branch can generate multiple target bounding boxes based on the region attention features. The location regression branch can generate multiple 3D coordinates based on the region attention features. The classification branch, bounding box regression branch, and location regression branch of the aforementioned detection head network can all include convolutional layers (CNN), activation functions, and fully connected layers (FCN). Afterwards, the determined sets of detection results can be combined into a region detection result set.

[0045] Sub-step five: Based on the aforementioned region detection result set, generate a single-frame detection result group corresponding to the aforementioned video frame image. In practice, non-maximum suppression (NMS) can be used to filter the region detection results in the aforementioned region detection result set, resulting in a filtered region detection result set as the single-frame detection result group. Specifically, firstly, all region detection results can be sorted from high to low confidence. Secondly, the target bounding box of the region detection result with the highest confidence can be determined as the first bounding box, and the overlap area (Intersection over Union, IoU) between the first bounding box and all remaining target bounding boxes can be determined. If the overlap between a target bounding box and the first bounding box exceeds a preset threshold (e.g., 0.5), it is considered that the two region detection results detect the same object. To remove redundancy, region detection results with low confidence are deleted from the region detection result set. This process continues until all region detection results have been processed.

[0046] The second sub-step involves generating forward optical flow information and backward optical flow information based on the preceding video frame image corresponding to the aforementioned video frame image. The preceding video frame image can be the previous video frame image in the video frame image sequence. In practice, the optical flow information generation layer can be used to generate forward and backward optical flow information based on the preceding video frame image corresponding to the aforementioned video frame image. Specifically, firstly, the grayscale images of the aforementioned video frame image and the preceding video frame image can be determined as the current frame grayscale image and the preceding grayscale image, respectively. Next, a preset optical flow estimation algorithm can be used to determine the forward optical flow information from the preceding grayscale image to the current frame grayscale image. The forward optical flow information can characterize the motion offset of each pixel in the image from the preceding video frame image to the aforementioned video frame image. The optical flow estimation algorithm can be the Lucas-Kanade algorithm. Secondly, the optical flow estimation algorithm can be used to determine the backward optical flow information from the current frame grayscale image to the preceding grayscale image. The aforementioned reverse optical flow information can characterize the motion offset of each pixel in the image from the aforementioned video frame image to the aforementioned preceding video frame image. Both the aforementioned forward optical flow information and reverse optical flow information can be matrices of the same size as the video frame images, and the elements in the matrices can include horizontal and vertical displacements.

[0047] The third sub-step involves generating an optical flow amplitude matrix based on the aforementioned forward and reverse optical flow information. This optical flow amplitude matrix can be a matrix of the same size as the video frame image. The value of each element in the matrix can be the maximum value between the forward and reverse optical flow amplitudes. In practice, this optical flow amplitude matrix can be generated using the aforementioned optical flow amplitude matrix generation layer. Specifically, firstly, for each element included in the forward optical flow information, the L2 norm between the horizontal and vertical displacements of that element can be determined as the forward optical flow amplitude. Then, for each element included in the reverse optical flow information, the L2 norm between the horizontal and vertical displacements of that element can be determined as the reverse optical flow amplitude. Next, the maximum value between the forward and reverse optical flow amplitudes can be determined as the optical flow amplitude. Finally, the determined optical flow amplitudes can be merged into an optical flow amplitude matrix.

[0048] The fourth sub-step involves determining the moving target detection result group corresponding to the aforementioned optical flow amplitude matrix. Each moving target detection result in this group can include the target category, target bounding box, and target location. In practice, firstly, the aforementioned edge detection layer can be used to perform edge detection on the optical flow amplitude matrix to obtain the contours of each moving target, and the bounding box corresponding to each contour can be determined as the target bounding box, resulting in a target bounding box set. Secondly, target detection and monocular depth estimation can be performed on the image covered by each determined target bounding box to generate the corresponding target category and target location. Finally, each target bounding box and its corresponding target category and target location can be determined as the moving target detection result.

[0049] The fifth sub-step involves generating a keyframe detection result group based on the aforementioned single-frame detection result group and moving target detection result group. In practice, firstly, the single-frame detection result group and moving target detection result group can be merged to generate an initial detection result group. Then, the aforementioned non-maximum suppression technique can be used to filter out duplicate bounding boxes in the initial detection result group, resulting in a filtered initial detection result group, which serves as the keyframe detection result group. Specifically, the operation corresponding to sub-step five in the optional steps above can be used to filter out duplicate bounding boxes in the initial detection result group, which will not be elaborated further here.

[0050] The second step involves tracking and completing each video frame in the aforementioned video frame image sequence whose frame label is not a keyframe, based on the obtained keyframe detection result groups, to obtain a non-keyframe detection result group. Each non-keyframe detection result in this group can include the target category and the target bounding box. In practice, for each video frame in the sequence whose frame label is not a keyframe, the previous video frame labeled as a keyframe can first be identified as the target keyframe. Then, the target tracking layer, combined with the keyframe detection result group corresponding to the target keyframe, can track and complete the video frame to obtain the non-keyframe detection result group. Here, each non-keyframe detection result in the non-keyframe detection result group can share the same local target number as its corresponding keyframe detection result.

[0051] The third step is to generate an initial anomalous flying object information group based on the obtained keyframe detection result groups and non-keyframe detection result groups. In practice, the keyframe detection result groups and non-keyframe detection result groups can be combined into an initial anomalous flying object information group.

[0052] In practice, a common technical challenge in detecting anomalous flying objects from video data is that static information in a single frame is often affected by blurring or occlusion, leading to insufficient detection accuracy and consequently reducing the accuracy of target tracking. This is especially true for small targets and fast-moving objects, where detection robustness decreases significantly. Therefore, the following solution is proposed.

[0053] Optionally, the aforementioned execution entity generates an initial abnormal flying object information group based on the obtained keyframe detection result groups and non-keyframe detection result groups, which may include the following steps:

[0054] The first step is to merge the obtained keyframe detection result groups and non-keyframe detection result groups in chronological order to obtain a frame detection result set. Each frame detection result group in this set corresponds one-to-one with a video frame in the aforementioned video frame image sequence. In practice, the keyframe detection result groups and non-keyframe detection result groups can be sorted chronologically to obtain the frame detection result set.

[0055] The second step is to perform the following steps for each video frame in the above video frame image sequence:

[0056] The first sub-step involves determining the sequence of adjacent video frame images corresponding to the aforementioned video frame image. In practice, the combination of the aforementioned video frame image with each of the preceding and following n video frame images can be used to determine the sequence of adjacent video frame images. Here, n is a number; for example, n can be 3. The sequence of adjacent video frame images can include 2n+1 video frame images.

[0057] Specifically, when the aforementioned video frame image is a non-boundary frame, the first n video frame images, the current video frame image, and the last n video frame images can be combined into an adjacent video frame image sequence. When the aforementioned video frame image is located at the boundary of the aforementioned video frame image sequence, that is, when the aforementioned video frame image is the first n-1 frames or the last n-1 frames of the sequence, the first frame or the last frame can be repeated to complete the adjacent video frame image sequence, so that the length of the aforementioned adjacent video frame image sequence is 2n+1.

[0058] The second sub-step involves determining the frame detection result group corresponding to the aforementioned video frame image within the aforementioned frame detection result group set as the target frame detection result group. Each target frame detection result in the target frame detection result group can characterize an abnormal flying object contained within the aforementioned video frame image.

[0059] The third sub-step involves determining a candidate region sequence for each target frame detection result in the aforementioned target frame detection result group, corresponding to the adjacent video frame image sequence. This candidate region sequence characterizes the spatiotemporal features of the target frame detection result. Each candidate region in the candidate region sequence corresponds one-to-one with an adjacent video frame image in the adjacent video frame image sequence. Each candidate region in the candidate region sequence can be a region enclosed by a target bounding box. The candidate region sequence can be represented as a high-dimensional vector of w×h×t. w and h represent the width and height of each candidate region, i.e., the spatial dimension. t represents the number of candidate regions in the candidate region sequence, i.e., the temporal dimension.

[0060] In practice, the candidate region sequence corresponding to the target frame detection result can be determined by using the aforementioned candidate region sequence set to define the layer. Specifically, for each target frame detection result in the aforementioned target frame detection result group, the following steps can be performed: First, for each group of frame detection results in the aforementioned adjacent video frame image sequence, a feature matching algorithm can be used to determine the frame detection results in the aforementioned frame detection result group that correspond to the target frame detection result as associated detection results. The aforementioned feature matching algorithm can be an IoU matching algorithm or a feature similarity matching algorithm. Here, when there is no frame detection result corresponding to the target frame detection result in the adjacent video frame image, a preset zero matrix can be determined as the initial candidate region. The aforementioned zero matrix can be a matrix with element values ​​of 0. During the feature matching process, the same local target number can be assigned to the frame detection results of feature matching between different frames as a cross-frame identifier for the same anomalous flying object. Afterwards, the region enclosed by the target bounding box corresponding to each determined associated detection result can be determined as the initial candidate region, resulting in an initial candidate region sequence. Each initial candidate region in the above initial candidate region sequence can be represented by a matrix, the size of which is the same as the size of the target bounding box included in the frame detection result. Finally, each initial candidate region in the above initial candidate region sequence can be mapped to the same size to obtain a candidate region sequence.

[0061] The fourth sub-step involves extracting spatiotemporal features from each of the determined candidate region sequences to generate candidate spatiotemporal features, thus obtaining a candidate spatiotemporal feature group. Each candidate spatiotemporal feature in this group can characterize the information of the corresponding candidate region sequence in terms of both time dimension and spatial structure. In practice, the aforementioned spatiotemporal feature extraction network can be used to extract spatiotemporal features from each of the determined candidate region sequences to generate candidate spatiotemporal features, thereby obtaining a candidate spatiotemporal feature group.

[0062] The fifth sub-step involves extracting spatiotemporal attention features from each candidate spatiotemporal feature in the aforementioned candidate spatiotemporal feature group to generate candidate attention features, thus obtaining the candidate attention feature group. Each candidate attention feature in the aforementioned candidate attention feature group can characterize the key information of the corresponding candidate spatiotemporal feature in the temporal and spatial dimensions. In practice, the aforementioned spatiotemporal attention feature extraction network can be used to extract spatiotemporal attention features from each candidate spatiotemporal feature in the aforementioned candidate spatiotemporal feature group to generate candidate attention features.

[0063] The sixth sub-step involves determining the candidate detection result group corresponding to the aforementioned candidate attention feature group. Each candidate detection result in the aforementioned candidate detection result group may include: target category, local target number, target bounding box, and target location. In practice, the aforementioned head network can be used to determine the candidate detection result corresponding to each candidate attention feature in the aforementioned candidate attention feature group, thus obtaining the candidate detection result group. The aforementioned head network can be used for target classification, bounding box regression, and location regression of the candidate attention features.

[0064] As an example, the target classification described above can determine whether the target category corresponding to the candidate attention features is a bird or a drone. The bounding box regression described above can determine the coordinates of the four vertices of the bounding box corresponding to the candidate attention features (i.e., the position of the bounding box). The position regression described above can determine the three-dimensional coordinates of the target corresponding to the candidate attention features in the real scene.

[0065] The third step is to perform inter-frame consistency checks on the determined candidate detection result groups to obtain the initial anomalous flying object information group. In practice, firstly, the candidate detection result groups of all frames can be traversed, and candidate detection results that appear only in a single frame can be eliminated to obtain each checked detection result group. Next, the checked detection results corresponding to the same local target number in each of the above checked detection result groups can be merged to generate the initial anomalous flying object information, thus obtaining the initial anomalous flying object information group.

[0066] The first to third steps and related contents of the above-mentioned optional solutions constitute an inventive point of this disclosure, solving the aforementioned technical problem: "significantly reduced detection robustness." The reason for the significant decrease in detection robustness is often that single-frame static information is frequently affected by blurring or occlusion, leading to insufficient detection accuracy and thus reducing the accuracy of target tracking, especially for small targets and fast-moving objects, where detection robustness is significantly reduced. Solving these factors can improve the robustness of abnormal flying object detection. To achieve this effect, firstly, the detection results of key frames and non-key frames are unified to the same timeline. Then, a local time window is provided for each video frame image to improve the model's inference ability for brief occlusions and motion blur. Next, through feature matching, the detection results corresponding to the same target in adjacent frames are merged into a candidate region sequence. This greatly reduces background interference and improves target tracking performance. Secondly, the spatiotemporal features of the candidate region sequence can be extracted through a 3D backbone network. This allows for the simultaneous capture of spatial and temporal features, improving the model's robustness. Then, a spatiotemporal attention mechanism can be used to extract attention features from spatiotemporal features, improving the model's ability to learn key features. Finally, inter-frame consistency verification is used to eliminate false detections in single frames, resulting in the final stable trajectory information.

[0067] Step 204: Target association is performed on each initial abnormal flying object information group in the above initial abnormal flying object information group set to obtain an abnormal flying object information set.

[0068] In some embodiments, the aforementioned executing entity can perform target association on each initial abnormal flying object information group in the aforementioned initial abnormal flying object information set to obtain an abnormal flying object information set. Each abnormal flying object information in the aforementioned abnormal flying object information set may include a target number, target category, and target trajectory. In practice, a preset target association algorithm can be used to perform target association on each initial abnormal flying object information group in the aforementioned initial abnormal flying object information set to obtain a cross-view associated abnormal flying object information set. The aforementioned target association algorithm can be trajectory association, ReID, etc.

[0069] As an example, suppose the initial anomalous flying object information group detected by camera A includes one of the following initial anomalous flying object information: "Local target number: A1, target category: bird, initial target trajectory: {(x1_A1, y1_A1, z1_A1, t1), ..., (x n _A1, y n _A1, z n _A1,t nThe above x1_A1, y1_A1, z1_A1 represent the three-dimensional coordinates of the initial anomalous flying object (A1) at time t1. The initial anomalous flying object information group detected by camera B includes one initial anomalous flying object piece of information: "Local target number: B1, target category: bird, initial target trajectory: {(x...}". n _B1, y n _B1, z n _B1,t n ), ..., (x m _B1, y m _B1, z m _B1,t m )}".t m Greater than t n Using the target association algorithm described above, it is possible to determine whether two initial anomalous flying object information items correspond to the same target based on spatiotemporal similarity. For example, t n At time point x n _A1, y n _A1, z n _A1 and x n _B1, y n _B1, z n If the distance between A1 and B1 is within a preset distance threshold, then the initial anomalous flying object information corresponding to A1 and B1 corresponds to the same target. The aforementioned distance threshold can be 0.1m. Afterwards, a cross-view anomalous flying object information can be generated, represented as: "Target number: 1, Target category: Bird, Target trajectory: {(x1_A1, y1_A1, z1_A1, t1), ..., (x...}". n _A1, y n _A1, z n _A1,t n ), ..., (x m _B1, y m _B1, z m _B1,t m )}".

[0070] Optionally, the aforementioned executing entity performs target association on each initial abnormal flying object information group in the aforementioned initial abnormal flying object information group set to obtain an abnormal flying object information set, which may include the following steps:

[0071] The first step is to identify the initial anomalous flying object (FOO) information groups that meet the preset comparison criteria from the aforementioned initial FOO information group set as the comparison information group. Specifically, each initial FOO information group in the aforementioned initial FOO information group set corresponds one-to-one with a camera in the aforementioned camera group. The cameras in the aforementioned camera group may be numbered according to their installation location. The preset comparison criteria can be that the initial FOO information group corresponds to a target camera. The target camera's number can be 1. In practice, the initial FOO information groups in the aforementioned initial FOO information group set that correspond to the target camera can be identified as the comparison information group.

[0072] The second step involves performing the following association steps for each initial anomalous flying object information group in the aforementioned initial anomalous flying object information group set, excluding the aforementioned control information group:

[0073] The first sub-step involves determining the correlation degree between each initial anomalous flying object (FOO) information in the initial FOO information group and each control information in the control information group, thus obtaining a correlation degree set. Each correlation degree in the correlation degree set characterizes the degree to which the corresponding initial FOO information and control information represent the same target. In practice, for each initial FOO information in the initial FOO information group and each control information in the control information group, firstly, if the target category corresponding to the initial FOO information is different from the target category corresponding to the control information, the correlation degree can be set to 0. Secondly, if the target category corresponding to the initial FOO information is the same as the target category corresponding to the control information, a preset trajectory correlation degree algorithm can be used to determine the correlation degree between the target trajectory included in the initial FOO information and the target trajectory included in the control information. The trajectory correlation degree algorithm can be the Hausdorff distance.

[0074] The second sub-step involves, for each correlation degree in the aforementioned correlation set that satisfies a preset correlation requirement, merging the initial anomalous flying object information corresponding to that correlation degree with the control information to obtain merged initial anomalous flying object information; removing the initial anomalous flying object information from the initial anomalous flying object information group to obtain an updated initial anomalous flying object information group; and removing the control information from the control information group to obtain an updated control information group. The preset correlation requirement can be that the correlation degree meets a preset correlation threshold. This correlation threshold can be a numerical value and is not specifically limited here.

[0075] As an example, suppose the initial anomalous flying object information group includes "Initial Information 1, Initial Information 2, Initial Information 3", and the comparison information group includes "Comparison Information 1, Comparison Information 2". If the correlation between Initial Information 1 and Comparison Information 2 meets the above correlation requirements, then Initial Information 1 and Comparison Information 2 are merged. The merged initial anomalous flying object information can be: Target Number: 1, Target Category: Bird, Target Trajectory: (Initial Information 1 + Comparison Information 2). The updated initial anomalous flying object information group can include "Initial Information 2, Initial Information 3", and the updated comparison information group can include "Comparison Information 1".

[0076] The third sub-step involves generating an abnormal flight object information group based on the obtained merged initial abnormal flight object information groups, updated initial abnormal flight object information groups, and updated comparison information groups. In practice, the abnormal flight object information group can be determined by combining the above-mentioned merged initial abnormal flight object information groups, the updated initial abnormal flight object information groups within the above-mentioned updated initial abnormal flight object information groups, and the updated comparison information groups within the above-mentioned updated comparison information groups.

[0077] The fourth sub-step involves, in response to the determination that the initial abnormal flying object information group does not meet the preset iteration conditions, designating the abnormal flying object information group as a control information group and re-executing the aforementioned association steps. The iteration conditions can be that the initial abnormal flying object information group is the last initial abnormal flying object information group in the set of initial abnormal flying object information groups, meaning that each initial abnormal flying object information group in the set has undergone the aforementioned association steps. In practice, when it is determined that the initial abnormal flying object information group does not meet the aforementioned iteration conditions, the abnormal flying object information group is designated as a control information group, and the aforementioned association steps are executed again.

[0078] The third step involves determining that the initial abnormal flying object information group meets the preset iteration conditions, and then identifying the last generated abnormal flying object information group as the abnormal flying object information set. In practice, when each initial abnormal flying object information group in the above initial abnormal flying object information group set has undergone the above association steps, the last generated abnormal flying object information group can be identified as the abnormal flying object information set.

[0079] Step 205: The airport 3D scene data is fused with each video frame image sequence in the video frame image sequence set to obtain dynamic 3D scene data.

[0080] In some embodiments, the executing entity can fuse the airport 3D scene data with the video frame image sequences in the video frame image sequence set to obtain dynamic 3D scene data. The dynamic 3D scene data can be a virtual environment reflecting the actual airport status. The dynamic 3D scene data can be a collection of 3D scene data corresponding to various time points. In practice, the airport 3D scene data can be fused with the video frame image sequences in the video frame image sequence set based on a preset projection texture mapping method to obtain dynamic 3D scene data.

[0081] Optionally, the aforementioned executing entity fuses the aforementioned airport 3D scene data with the various video frame image sequences in the aforementioned video frame image sequence set to obtain dynamic 3D scene data, which may include the following steps:

[0082] The first step, for each video frame image corresponding to the same time in the above video frame image sequence set, is to perform the following steps:

[0083] The first sub-step involves projecting each of the aforementioned video frame images onto the aforementioned airport 3D scene data to obtain projected 3D scene data. This projected 3D scene data can be 3D scene data at a single point in time. In practice, video frame images captured by different cameras at the same time can be used as textures and projected onto the aforementioned 3D model to obtain projected 3D scene data. Specifically, for each video frame image captured by each camera in the aforementioned camera group, the video frame image can be projected onto the aforementioned 3D scene data according to the projection matrix between the video frame image and the aforementioned 3D scene data.

[0084] The second sub-step involves performing distortion optimization on the projected 3D scene data to obtain distortion-optimized 3D scene data. In practice, a preset distortion optimization algorithm can be used to optimize the distortion of the projected 3D scene data to obtain the distortion-optimized 3D scene data. This distortion optimization algorithm can be a Laplacian mesh deformation algorithm.

[0085] In practice, since objects in video images captured by cameras tend to appear larger when closer and smaller when farther away, the Laplacian mesh deformation algorithm can improve the misalignment and distortion during the fusion of video and 3D scene.

[0086] The third sub-step involves optimizing the overlapping regions of the projected 3D scene data to generate fused 3D scene data. This fused 3D scene data corresponds to specific time points, which are the same as the time points corresponding to the individual video frames. In practice, firstly, the ORB feature detection algorithm can be used to determine the overlapping regions corresponding to the projected 3D scene data. These overlapping regions can include images captured by multiple cameras. Then, a weighted fusion algorithm can be applied to weightedly fuse the multiple images included in each overlapping region, resulting in the fused overlapping regions. Finally, these fused overlapping regions can be mapped back to the projected 3D scene data to obtain the fused 3D scene data.

[0087] In practice, since there are overlapping areas between videos captured by multiple cameras, the overlapping areas can be optimized by adjusting the color transition of the overlapping areas to avoid seams and color differences, and ensure that the merged image is smooth and natural.

[0088] The second step is to generate dynamic 3D scene data based on the generated fused 3D scene data. This dynamic 3D scene data can be 3D spatial information that changes over time. At a single point in time, the dynamic 3D scene data can be represented as the fused 3D scene data at the corresponding timestamp. In practice, the fused 3D scene data at each timestamp can be rendered in chronological order to obtain the dynamic 3D scene data.

[0089] Step 206: Based on the above abnormal flying object information set, abnormal flying objects are added to the above dynamic three-dimensional scene data to obtain three-dimensional scene data containing abnormal flying objects.

[0090] In some embodiments, the aforementioned execution entity can add abnormal flying objects to the aforementioned dynamic 3D scene data based on the aforementioned abnormal flying object information set, thereby obtaining 3D scene data containing abnormal flying objects. In practice, for each abnormal flying object information in the aforementioned abnormal flying object information set, firstly, a corresponding virtual model can be rendered in the aforementioned dynamic 3D scene data according to the corresponding target category. Then, the target trajectory corresponding to the aforementioned abnormal flying object information can be used as the motion trajectory of the aforementioned virtual model in the aforementioned dynamic 3D scene data to generate 3D scene data containing abnormal flying objects.

[0091] In practice, in digital twin technology used for airports, flying objects are typically mapped as textures onto a 3D scene. The technical challenge is that the lack of precise location and intuitive spatial reference for anomalous flying objects leads to a lack of timeliness and accuracy in emergency response strategies, thus reducing the ability to address anomalous flying object threats in complex airspace environments. Therefore, the following solution is proposed.

[0092] Optionally, the aforementioned executing entity adds abnormal flying objects to the aforementioned dynamic 3D scene data based on the aforementioned abnormal flying object information set, thereby obtaining 3D scene data containing abnormal flying objects. This may include the following steps:

[0093] The first step is to determine the virtual model information of each anomalous flying object in the aforementioned anomalous flying object information set, based on the target category and target trajectory corresponding to the anomalous flying object information, within the aforementioned dynamic 3D scene data. The virtual model information includes: a virtual model and a virtual model trajectory. The virtual model can be a 3D model. The virtual model trajectory can be the path along which the virtual model moves within the aforementioned airport scene data. The virtual model trajectory can be a series of 3D coordinate points arranged in chronological order. In practice, firstly, a preset 3D model corresponding to the aforementioned target category can be determined as the virtual model. For example, when the target category is birds, the corresponding preset 3D model can be a bird 3D model. When the target category is a drone, the corresponding preset 3D model can be a drone 3D model. Secondly, the coordinates in the target trajectory can be transformed from the world coordinate system to the dynamic 3D scene coordinate system to obtain the virtual model trajectory. Optionally, the virtual model can also be scaled according to the physical dimensions of the anomalous flying object information to ensure that the size of the virtual model in the dynamic 3D scene matches the actual size. The physical dimensions can be calculated based on the target bounding box dimensions combined with camera intrinsic parameters.

[0094] The second step involves optimizing the textures of each determined virtual model based on a pre-defined generative adversarial model, resulting in a texture-optimized virtual model set. In practice, for each determined virtual model, texture optimization can be performed using a pre-defined generative adversarial model to obtain a texture-optimized virtual model.

[0095] Specifically, for each virtual model, firstly, the texture images covering the bounding boxes of each target corresponding to the abnormal flying object information of the virtual model can be obtained. The following information can be used as input to the generative adversarial model: the texture images covering the bounding boxes of each target; and the 3D mesh structure of the virtual model. The output of the generative adversarial model is a texture-mapped virtual model, achieving consistency between the texture and the actual target.

[0096] The aforementioned generative adversarial model can be a generative adversarial model (GAN) pre-trained based on real texture images and 3D model data of various abnormal flying objects (birds, drones) in airport scenarios.

[0097] Third, for each single frame of 3D scene data included in the above dynamic 3D scene data, perform the following steps:

[0098] The first sub-step involves constructing an initial virtual model group corresponding to the aforementioned single-frame 3D scene data, based on the determined information of each virtual model. In practice, for each virtual model trajectory, when the time point where a 3D coordinate point exists in the virtual model trajectory is the same as the time point corresponding to the aforementioned single-frame 3D scene data, the virtual model corresponding to the aforementioned virtual model trajectory is determined as the initial virtual model, and the aforementioned 3D coordinate point is determined as the coordinate of the initial virtual model.

[0099] The second sub-step involves performing collision detection on each of the initialized virtual models in the aforementioned initialized virtual model group to generate a collision detection result set. In practice, a preset 3D collision detection algorithm can be used to perform collision detection on any two initialized virtual models in the aforementioned initialized virtual model group, as well as equipment (e.g., an airport terminal) in a dynamic 3D scene, to generate a collision detection result set. The collision detection result can indicate whether a collision occurred or not.

[0100] The third sub-step involves modifying each initial virtual model in the initial virtual model group based on the aforementioned collision detection result set, resulting in a modified virtual model group. In practice, the initial virtual model corresponding to each collision detection result in the aforementioned collision detection result set can be translated to obtain the modified virtual model group.

[0101] As an example, when the collision detection result between two initial virtual models indicates a collision, each initial virtual model can be translated to obtain two translated virtual models. Collision detection is then performed again on these two translated virtual models. If neither of these translated virtual models collides with other models or devices, both translated virtual models are identified as corrected virtual models. If the two translated virtual models collide with other models or devices, they are translated again in different directions to avoid collision. When the collision detection result between the initial virtual model and a device in the dynamic 3D scene indicates a collision, the position of the virtual model can be adjusted by shifting it 5-10cm away from the device based on the surface normal vectors of the scene objects, resulting in a corrected virtual model.

[0102] The fourth sub-step involves rendering each of the corrected virtual models in the corrected virtual model group within the aforementioned single-frame 3D scene data to obtain a single-frame 3D scene data containing the anomalous flying object. In practice, pre-defined 3D scene processing software can be used to render each of the corrected virtual models in the aforementioned corrected virtual model group within the aforementioned single-frame 3D scene data to obtain a single-frame 3D scene data containing the anomalous flying object. This 3D scene processing software can be Unreal Engine.

[0103] The fourth step involves generating 3D scene data containing anomalous flying objects based on the obtained individual frames of 3D scene data. This 3D scene data can be temporally variable 3D spatial information. At a single point in time, this 3D scene data can be represented as a single frame of 3D scene data containing anomalous flying objects at the corresponding timestamp. In practice, the single frames of 3D scene data containing anomalous flying objects at each timestamp can be rendered sequentially to obtain the 3D scene data containing anomalous flying objects.

[0104] Steps one through four of the aforementioned optional solutions, along with their related content, constitute an inventive point of this disclosure, solving the aforementioned technical problem: "reduced ability to respond to the threat of anomalous flying objects in complex airspace environments." The low ability to respond to anomalous flying object threats is often due to a lack of precise location and intuitive spatial reference for the anomalous flying objects, resulting in a lack of timeliness and accuracy in emergency response strategies. Solving these factors can improve the ability to respond to anomalous flying object threats. To achieve this, firstly, a corresponding virtual model can be initialized in a 3D scene based on the category and trajectory of each anomalous flying object. Scale scaling ensures that each element in the dynamic 3D scene maintains a reasonable proportional relationship. Next, a generative adversarial model is used to extract texture features from the observed video data and map them to the 3D virtual model. This improves the realism of the virtual model. Secondly, the virtual model at each time point can be spatiotemporally aligned and fused with the 3D scene in chronological order to ensure inter-frame continuity and avoid motion trajectory distortion caused by timestamp misalignment or coordinate offset. Finally, collision detection and correction can be performed on each virtual model in the 3D scene at the same time to ensure that the virtual model of the anomalous flying object conforms to the actual airspace state of the 3D scene. Finally, the single-frame data is synthesized into a dynamic scene according to the timestamp, which fully reproduces the entire process of the abnormal flying object from appearance to disappearance. This makes it easier to trace the trajectory of the abnormal flying object and provides an intuitive and accurate spatial decision-making basis for formulating targeted interception, expulsion or early warning strategies.

[0105] Optionally, after step 206 above, the executing entity may also perform the following steps:

[0106] The first step is to generate a sample set of anomalous flying object information based on the aforementioned anomalous flying object information set. Each sample anomalous flying object information in this set can include a sample category and a sample trajectory. The sample category can include birds or drones. The sample trajectory can be an ordered sequence of flight coordinates of the sample anomalous flying object. In practice, different types of adjustments can be made to each anomalous flying object information in the aforementioned set to generate different sample anomalous flying object information. For example, for a single anomalous flying object information, the altitude of each coordinate point in the target trajectory can be increased by 5 meters to obtain the sample anomalous flying object information. Alternatively, the coordinate points contained in the target trajectory can be translated to obtain the sample anomalous flying object information.

[0107] In practice, the difficulty in accurately capturing the real-time location of anomalous flying objects in real-world airport scenarios leads to missing or inaccurate labeled data, directly impacting the training performance of anomalous flying object detection models. Digital twin models, by directly generating precisely labeled virtual data, not only increase the amount of data but also safely simulate various types of anomalous scenarios, providing data support for training high-precision, robust anomalous flying object detection models.

[0108] The second step involves optimizing the aforementioned 3D scene data containing anomalous flying objects based on the sample anomalous flying object information set, resulting in optimized 3D scene data. In practice, the operation in step 206 above can be used to add anomalous flying objects to the aforementioned dynamic 3D scene data based on the sample anomalous flying object information set, thus obtaining optimized 3D scene data.

[0109] The third step is to determine the sample airport video data corresponding to the optimized 3D scene data. In practice, video roaming technology can be used to determine the sample airport video data corresponding to the optimized 3D scene data. Specifically, multiple virtual cameras can be created in the optimized 3D scene data using 3D rendering software, and 3D scene roaming can be performed using these virtual cameras to obtain the sample airport video data. The coordinates of these virtual cameras can be the same as the coordinates of each camera in the actual camera group installed at the airport.

[0110] The fourth step involves detecting abnormal flying objects (FOOs) in the sample airport video data based on a pre-defined initial FOO detection model, thereby obtaining a predicted FOO information set. The initial FOO detection model can have the same structure as the FOO detection model mentioned above, but with different parameters. Each predicted FOO information in the predicted FOO information set can include a predicted category and a predicted trajectory. The predicted trajectory can be a sequence of coordinates of each predicted FOO.

[0111] The fifth step involves generating a loss value based on the predicted anomalous flying object information set and the sample anomalous flying object information set. In practice, for each sample anomalous flying object in the sample anomalous flying object information set, firstly, the classification loss between the predicted anomalous flying object information and the sample anomalous flying object information can be determined using the cross-entropy loss function. Then, the position regression loss between the predicted anomalous flying object information and the sample anomalous flying object information can be determined using the L1 loss function. This position regression loss may include coordinate loss. Finally, the weighted sum of the classification loss and the position regression loss can be used to determine the loss value.

[0112] Step 6: Based on the aforementioned loss value, train the initial anomalous flying object detection model to obtain a pre-trained anomalous flying object detection model. In practice, the initial anomalous flying object detection model can be subjected to parameter gradient descent based on the aforementioned loss value to obtain a pre-trained anomalous flying object detection model. The pre-trained anomalous flying object detection model can have the same structure as the initial anomalous flying object detection model but different parameters.

[0113] Step 7: Based on the pre-trained anomalous flight object detection model, perform anomalous flight object detection on the airport video data collected on-site to obtain an anomalous flight object information set. The airport video data can be the airport to be simulated. In practice, 3D scene data containing anomalous flight objects can be used for training, testing, and validation of the anomalous flight object detection model, further improving the airport's ability to identify anomalous flight objects.

[0114] Further reference Figure 4 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an airport simulation scene and video fusion device for identifying anomalous flying objects. These device embodiments are similar to... Figure 2 Corresponding to the method embodiments shown, the airport simulation scene and video fusion device for identifying anomalous flying objects can be specifically applied to various electronic devices.

[0115] like Figure 4As shown, an airport simulation scene and video fusion device 400 for abnormal flying object identification in some embodiments includes: an acquisition unit 401, a frame sampling unit 402, an abnormal flying object detection unit 403, a target association unit 404, a fusion unit 405, and an abnormal flying object addition unit 406. The acquisition unit 401 is configured to acquire a set of airport videos and airport 3D scene data of the airport to be simulated; the frame sampling unit 402 is configured to perform frame sampling on each airport video in the aforementioned airport video set to generate a video frame image sequence, resulting in a set of video frame image sequences, wherein each video frame image in the set of video frame image sequences corresponds to a frame label, and the frame label includes keyframes and non-keyframes; the abnormal flying object detection unit 403 is configured to perform abnormal flying object detection on each video frame image sequence in the aforementioned video frame image sequence set to generate initial abnormal flying object information. The system consists of three parts: an initial set of anomalous flying object information groups; a target association unit 404 configured to perform target association on each initial anomalous flying object information group in the initial set of anomalous flying object information groups to obtain an anomalous flying object information set; a fusion unit 405 configured to fuse the airport 3D scene data with each video frame image sequence in the video frame image sequence set to obtain dynamic 3D scene data; and an anomalous flying object addition unit 406 configured to add anomalous flying objects to the dynamic 3D scene data according to the anomalous flying object information set to obtain 3D scene data containing anomalous flying objects.

[0116] It is understandable that the units and references described in the airport simulation scene and video fusion device 400 for anomalous flying object identification are... Figure 2 The steps in the described method correspond accordingly. Therefore, the operations, features, and beneficial effects described above for the method are also applicable to the airport simulation scene and video fusion device 400 for anomalous flying object identification and the units contained therein, and will not be repeated here.

[0117] The following is for reference. Figure 5 It shows a schematic diagram of the structure of an electronic device 500 (e.g., a computing device) suitable for implementing some embodiments of the present disclosure. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0118] like Figure 5As shown, the electronic device 500 may include a processing unit 501 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0119] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 5 Each box shown can represent a device or multiple devices as needed.

[0120] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 509, or installed from storage device 508, or installed from ROM 502. When the computer program is executed by processing device 501, it performs the functions defined in the methods of some embodiments of this disclosure.

[0121] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0122] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0123] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a set of airport video images and airport 3D scene data of the airport to be simulated; perform frame sampling on each airport video in the aforementioned airport video set to generate a video frame image sequence, obtaining a set of video frame image sequences, wherein each video frame image included in the set of video frame image sequences corresponds to a frame label, the frame label including keyframes and non-keyframes; perform abnormal flight object detection on each video frame image sequence in the set of video frame image sequences to generate an initial abnormal flight object information group, obtaining an initial abnormal flight object information group set; perform target association on each initial abnormal flight object information group in the set of initial abnormal flight object information groups, obtaining an abnormal flight object information set; fuse the aforementioned airport 3D scene data with each video frame image sequence in the set of video frame image sequences to obtain dynamic 3D scene data; and add abnormal flight objects to the aforementioned dynamic 3D scene data according to the aforementioned abnormal flight object information set, obtaining 3D scene data containing abnormal flight objects.

[0124] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0126] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0127] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for fusing airport simulation scenes and videos for identifying anomalous flying objects, characterized in that, include: Acquire the airport video set and airport 3D scene data of the airport to be simulated; Each airport video in the airport video set is sampled frame by frame to generate a video frame image sequence, resulting in a video frame image sequence set. Each video frame image in the video frame image sequence set corresponds to a frame label, which includes keyframes and non-keyframes. Anomaly detection is performed on each video frame image sequence in the video frame image sequence set to generate an initial aomaly information group, thus obtaining an initial aomaly information group set; Target association is performed on each initial abnormal flying object information group in the initial abnormal flying object information group set to obtain an abnormal flying object information set; The airport 3D scene data is fused with each video frame image sequence in the video frame image sequence set to obtain dynamic 3D scene data; Based on the abnormal flying object information set, abnormal flying objects are added to the dynamic three-dimensional scene data to obtain three-dimensional scene data containing abnormal flying objects; The process of detecting abnormal flying objects in each video frame image sequence within the video frame image sequence set includes: For each video frame image in the video frame image sequence whose frame label is a keyframe, perform the following steps: Perform single-frame abnormal flying object detection on the video frame images to generate a single-frame detection result group; Based on the preceding video frame image corresponding to the video frame image, forward optical flow information and reverse optical flow information are generated; An optical flow amplitude matrix is ​​generated based on the forward optical flow information and the reverse optical flow information; Determine the moving target detection result group corresponding to the optical flow amplitude matrix; Based on the single-frame detection result group and the moving target detection result group, a keyframe detection result group is generated; For each video frame image in the video frame image sequence whose frame label is a non-key frame, the video frame image is tracked and completed according to the obtained key frame detection result group to obtain the non-key frame detection result group. Based on the obtained keyframe detection result groups and non-keyframe detection result groups, an initial abnormal flying object information group is generated.

2. The method according to claim 1, characterized in that, The step of performing single-frame abnormal flying object detection on the video frame image to generate a single-frame detection result group includes: The video frame image is segmented into region blocks to obtain a region block set; Feature extraction is performed on each region block in the region block set to generate region features, thus obtaining a region feature set; For each region feature in the region feature set, region attention features are extracted to generate region attention features, thus obtaining a region attention feature set; Determine the region detection result set corresponding to the region attention feature set; Based on the region detection result set, generate a single-frame detection result group corresponding to the video frame image.

3. The method according to claim 1, characterized in that, The step of associating targets among the initial anomalous flying object information groups in the initial anomalous flying object information group set to obtain the anomalous flying object information set includes: The initial abnormal flying object information groups that meet the preset comparison conditions are determined as the comparison information groups; For each initial anomalous flying object information group in the initial anomalous flying object information group set, excluding the control information group, the following association steps are performed: Determine the correlation degree between each initial abnormal flying object information in the initial abnormal flying object information group and each control information in the control information group to obtain a correlation degree set; For each correlation degree in the correlation degree set that meets the preset correlation requirements, the initial abnormal flying object information corresponding to the correlation degree is merged with the control information to obtain merged initial abnormal flying object information, and the initial abnormal flying object information is removed from the initial abnormal flying object information group to obtain updated initial abnormal flying object information group, and the control information is removed from the control information group to obtain updated control information group. Based on the obtained merged initial abnormal flying object information, updated initial abnormal flying object information group and updated comparison information group, an abnormal flying object information group is generated. In response to the determination that the initial abnormal flying object information group does not meet the preset iteration conditions, the abnormal flying object information group is determined as the control information group, and the association step is executed again; In response to the determination that the initial abnormal flying object information group meets the preset iteration conditions, the last abnormal flying object information group generated is determined as the abnormal flying object information set.

4. The method according to claim 1, characterized in that, The process of fusing the airport 3D scene data with the video frame image sequences in the video frame image sequence set to obtain dynamic 3D scene data includes: For each video frame image in the video frame image sequence set corresponding to the same time, perform the following steps: The images of each video frame are projected onto the airport 3D scene data to obtain the projected 3D scene data. The projected 3D scene data is subjected to distortion optimization to obtain distortion-optimized 3D scene data; The overlapping area of ​​the projected 3D scene data is optimized to generate fused 3D scene data; Based on the generated fused 3D scene data, dynamic 3D scene data is generated.

5. The method according to claim 1, characterized in that, The step of generating an initial abnormal flying object information group based on the obtained keyframe detection result groups and non-keyframe detection result groups includes: The obtained key frame detection result groups and non-key frame detection result groups are merged in chronological order to obtain the frame detection result group set. For each video frame in the video frame image sequence, perform the following steps: Determine the sequence of adjacent video frame images corresponding to the video frame image; The frame detection result group corresponding to the video frame image in the frame detection result group set is determined as the target frame detection result group; For each target frame detection result in the target frame detection result group, determine the candidate region sequence corresponding to the adjacent video frame image sequence; Spatiotemporal features are extracted from each of the determined candidate region sequences to generate candidate spatiotemporal features, thus obtaining a candidate spatiotemporal feature group. Spatiotemporal attention features are extracted from each candidate spatiotemporal feature in the candidate spatiotemporal feature group to generate candidate attention features, thus obtaining the candidate attention feature group; Determine the candidate detection result group corresponding to the candidate attention feature group; Inter-frame consistency verification is performed on each of the identified candidate detection result groups to obtain the initial abnormal flying object information group.

6. The method according to claim 1, characterized in that, The method further includes: Based on the aforementioned abnormal flying object information set, a sample abnormal flying object information set is generated; Based on the sample abnormal flying object information set, the three-dimensional scene data containing abnormal flying objects is optimized to obtain optimized three-dimensional scene data; Determine the sample airport video data corresponding to the optimized 3D scene data; Based on a preset initial abnormal flying object detection model, abnormal flying object detection is performed on the sample airport video data to obtain a predicted abnormal flying object information set; Based on the predicted abnormal flying object information set and the sample abnormal flying object information set, a loss value is generated; Based on the loss value, the initial abnormal flying object detection model is trained to obtain a pre-trained abnormal flying object detection model; Based on the pre-trained abnormal flight object detection model, abnormal flight objects are detected in the airport video data collected on site, and an information set of abnormal flight objects on site is obtained.

7. An airport simulation scene and video fusion device for identifying anomalous flying objects, characterized in that, include: The acquisition unit is configured to acquire the airport video set and airport 3D scene data of the airport to be simulated and tested; A frame sampling unit is configured to sample each airport video in the airport video set to generate a video frame image sequence, thereby obtaining a video frame image sequence set, wherein each video frame image included in the video frame image sequence set corresponds to a frame label, and the frame label includes keyframes and non-keyframes. An abnormal flying object detection unit is configured to perform abnormal flying object detection on each video frame image sequence in the video frame image sequence set to generate an initial abnormal flying object information group, thereby obtaining an initial abnormal flying object information group set. The process of detecting abnormal flying objects in each video frame image sequence within the video frame image sequence set includes: For each video frame image in the video frame image sequence whose frame label is a keyframe, perform the following steps: Perform single-frame abnormal flying object detection on the video frame images to generate a single-frame detection result group; Based on the preceding video frame image corresponding to the video frame image, forward optical flow information and reverse optical flow information are generated; An optical flow amplitude matrix is ​​generated based on the forward optical flow information and the reverse optical flow information; Determine the moving target detection result group corresponding to the optical flow amplitude matrix; Based on the single-frame detection result group and the moving target detection result group, a keyframe detection result group is generated; For each video frame image in the video frame image sequence whose frame label is a non-key frame, the video frame image is tracked and completed according to the obtained key frame detection result group to obtain the non-key frame detection result group. Based on the obtained key frame detection result groups and non-key frame detection result groups, an initial abnormal flying object information group is generated. The target association unit is configured to perform target association on each initial abnormal flying object information group in the initial abnormal flying object information group set to obtain an abnormal flying object information set; The fusion unit is configured to fuse the airport 3D scene data with each video frame image sequence in the video frame image sequence set to obtain dynamic 3D scene data. An anomalous flying object addition unit is configured to add anomalous flying objects to the dynamic three-dimensional scene data based on the anomalous flying object information set, thereby obtaining three-dimensional scene data containing anomalous flying objects.

8. An electronic device, characterized in that, include: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 6.

9. A computer-readable medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by a processor, implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Three-dimensional video fusion method and device, electronic equipment and storage medium

    CN117336459A

  • Airspace abnormal flyer intelligent analysis method combining FS-AFD model and combined agent analysis

    CN121191065A