Multi-modal data fusion-based anti-collision early warning method, device, equipment and medium for loader
The aircraft collision avoidance warning method based on multimodal data fusion utilizes the preprocessing and cross-modal filtering fusion of radar point clouds and photoelectric images to construct a threat assessment model. This solves the problems of misjudgment and low long-range detection accuracy in traditional aircraft collision avoidance systems in complex airspace, and achieves high-precision threat target detection and clear warning information presentation.
Patent Information
- Application Number
- CN202511090336.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-07
Smart Images

Figure CN120913459A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of aviation safety, and in particular to a multi-modal data fusion aircraft collision avoidance early warning method, device, equipment and medium. BACKGROUND
[0002] The aircraft collision avoidance system (such as TCAS (Traffic Collision Avoidance System) and ACAS (Airborne Collision Avoidance System)) is the core technology of aviation safety, which aims to prevent aircraft from colliding in the air within a few kilometers. Modern systems are based on secondary radar (such as S-mode transponder) and ADS-B (Automatic Dependent Surveillance-Broadcast) technology, which realizes target detection through radio signal exchange, can cover tens of kilometers of airspace, and calculates relative distance, height and heading, and provides traffic advisory (Traffic Advisory, TA) or resolution advisory (Resolution Advisory, RA). However, the traditional system is limited by the antenna angle measurement accuracy (the maximum error is 15°) and the body shielding, and in complex airspace or large maneuvering flight, the track may be interrupted or misjudged.
[0003] Laser radar detection includes single-mode detection and multi-modal detection, the current single-mode method has low cost, but each sensor has its own limitations, and cannot take into account the information, thereby increasing the risk of missed detection. And the current multi-modal method is mostly based on laser radar, or based on millimeter wave radar, the detection range of laser radar is limited within a few hundred meters, and due to the problems of high cost, long-distance signal attenuation, etc., it is difficult to directly adapt to long-distance detection scenarios.
[0004] In the multi-modal method, the current data superposition visualization only labels the abnormal position found by the radar on the image, lacks target level cognition, and cannot distinguish between interference: this scheme only identifies "abnormal points", but cannot distinguish whether the point is generated by a real target or by environmental interference or sensor noise. Therefore, false alarms are easy to occur, causing the problems of too close collision detection distance, low detection accuracy and untimely early warning. SUMMARY
[0005] Therefore, the present application provides a multi-modal data fusion aircraft collision avoidance early warning method, device, equipment and medium to solve the problem of short collision detection distance and low detection accuracy in the prior art.
[0006] In a first aspect, the present application provides a multi-modal data fusion aircraft collision avoidance early warning method, which comprises:
[0007] acquire multi-modal data, the multi-modal data comprising radar point cloud data and photoelectric image;
[0008] respectively pre-process the radar point cloud data and the photoelectric image to obtain radar effective point cloud, semantic mask and confidence;
[0009] perform cross-modal filter fusion based on the radar effective point cloud, the semantic mask and the confidence to obtain multi-dimensional feature point cloud;
[0010] construct a threat assessment model based on the multi-dimensional feature point cloud, and perform pre-warning for aircraft collision avoidance based on the threat assessment model.
[0011] The multi-modal data fusion aircraft collision avoidance pre-warning method provided by the application fuses the multi-modal data of radar point cloud and photoelectric image, realizes the complementary advantages of multi-modal data, greatly improves the robustness of flight environment perception, respectively pre-processes the radar point cloud data and the photoelectric image to obtain radar effective point cloud, semantic mask and confidence, performs cross-modal filter fusion based on the radar effective point cloud, the semantic mask and the confidence to obtain multi-dimensional feature point cloud, eliminates radar invalid point cloud and background interference, realizes cross-modal filter fusion, and the double-factor weighted threat assessment model constructed based on the multi-dimensional feature point cloud realizes automatic detection of threat targets of the aircraft in flight, has a longer measurement distance and higher measurement accuracy, and solves the problems of short detection distance and low detection accuracy in the prior art.
[0012] In an optional implementation, respectively pre-processing the radar point cloud data and the photoelectric image to obtain radar effective point cloud and semantic mask comprises:
[0013] performing field of view filtering and distance threshold filtering on the radar point cloud data to obtain the radar effective point cloud;
[0014] performing image semantic segmentation on the photoelectric image by using a preset neural network to generate the semantic mask and the confidence.
[0015] The multi-modal data fusion aircraft collision avoidance pre-warning method provided by the application greatly reduces data redundancy, improves processing efficiency, and enhances the physical effectiveness of point cloud data by combining field of view filtering and distance threshold filtering, thereby laying a foundation for cross-modal fusion.
[0016] In an optional implementation, the neural network comprises a multi-layer encoder and a multi-layer decoder, and the image semantic segmentation on the photoelectric image by using the preset neural network to generate the semantic mask and the confidence comprises:
[0017] extracting multi-scale features of the photoelectric image by using the multi-layer encoder, and fusing the multi-scale features by using the multi-layer decoder to obtain a semantic prediction result;
[0018] Generate a pixel-level class probability distribution based on the semantic prediction result, and generate a semantic mask and a confidence based on the pixel-level class probability distribution.
[0019] The application provides a multi-modal data fusion aircraft collision warning method, a multi-layer encoder can capture multi-scale features of photoelectric images through hierarchical processing, a shallow layer encoder can extract detailed information of the images, a multi-layer decoder fuses the multi-scale features, can complement the advantages of features at different levels, generates a pixel-level class probability distribution, can provide a quantitative basis for the possibility of each pixel belonging to different classes, and provides an important reference for subsequent cross-modal filtering and fusion, can filter point clouds corresponding to low confidence regions, reduces false detection, and improves the accuracy of target screening.
[0020] In an optional embodiment, the method further comprises:
[0021] The radar effective point cloud and the semantic mask are time-aligned by using the near-sequence frame extraction method.
[0022] A coordinate projection model is established, and the radar effective point cloud and the semantic mask are spatially aligned based on the coordinate projection model.
[0023] The application provides a multi-modal data fusion aircraft collision warning method, the near-sequence frame extraction method matches the radar point cloud and the photoelectric image frame with the closest time stamp, effectively solves the problem that the sampling frequencies of the radar and the photoelectric sensor are not synchronized, realizes time alignment, the coordinate projection model accurately converts the three-dimensional coordinates of the radar effective point cloud into two-dimensional pixel coordinates of the image plane, achieves spatial position matching of the radar point cloud and the semantic mask, realizes spatial alignment, and through the cooperation of time alignment and spatial alignment, cross-modal data consistent in time and space is constructed.
[0024] In an optional embodiment, cross-modal filtering and fusion are performed based on the radar effective point cloud, the semantic mask and the confidence, and multi-dimensional feature point clouds are obtained, including:
[0025] The radar effective point cloud is transmitted to the image plane through the coordinate projection model to obtain projection pixel coordinates corresponding to each point;
[0026] It is judged whether the projection pixel coordinates fall within the effective area of the semantic mask, and the point cloud with a confidence lower than a preset threshold in the effective area is deleted to obtain a point cloud projection area;
[0027] The point cloud projection area is subjected to semantic-guided point cloud filtering to obtain multi-dimensional feature point clouds.
[0028] The application provides a carrier anti-collision early warning method based on multi-modal data fusion, which projects radar effective point clouds to an image plane through a coordinate projection model to obtain corresponding projection pixel coordinates of each point, establishes a direct association between radar three-dimensional space information and image two-dimensional semantic information, realizes cross-modal data binding in the spatial dimension, judges whether the projection pixel coordinates are in the effective area of a semantic mask, can preliminarily screen out point clouds related to a target semantic, eliminates invalid point clouds falling in a background area, reduces the interference of irrelevant data, and semantic-guided point cloud filtering gives the point clouds rich attributes, so that the obtained multi-dimensional feature point clouds contain not only accurate spatial position information but also clear semantic categories and physical attribute information, thereby providing comprehensive and accurate input data for subsequent threat assessment.
[0029] In an optional implementation, the point cloud filtering of the point cloud projection area is guided by semantics to obtain multi-dimensional feature point clouds, including:
[0030] An image semantic category corresponding to the point cloud projection area is acquired, and the image semantic category is given to corresponding point cloud points to obtain category labels of each point cloud point, and the category labels of each point cloud point are compared with a target category set to delete point cloud points not belonging to the target category set to obtain category attributes of each point cloud point.
[0031] A ground plane equation is estimated through a plane fitting algorithm, and a vertical distance of each point cloud point relative to the ground is calculated based on the ground plane equation to obtain a height attribute of each point cloud point.
[0032] A straight-line distance between each point cloud and the carrier is calculated based on the projection pixel coordinates of each point cloud point to obtain a distance attribute of each point cloud point.
[0033] Multi-dimensional feature point clouds are generated based on the category attributes, the height attributes and the distance attributes.
[0034] The carrier anti-collision early warning method based on multi-modal data fusion provided by the application acquires an image semantic category corresponding to a point cloud projection area and gives the semantic category to point cloud points, and then compares the point cloud points with a target category set to delete point cloud points not belonging to the target category set, so that a threat target of interest can be accurately locked, a ground plane equation is estimated through a plane fitting algorithm, and then a vertical distance of each point cloud point relative to the ground is calculated to obtain a height attribute, thereby providing key spatial dimension information for threat assessment, a straight-line distance between each point cloud and the carrier is calculated based on the projection pixel coordinates of each point cloud point to obtain a distance attribute, so that the spatial distance between the target and the carrier can be intuitively reflected, and multi-dimensional feature point clouds generated based on the category attributes, the height attributes and the distance attributes integrate semantic information, spatial height information and distance information of the target, thereby providing comprehensive and rich input data for threat assessment.
[0035] In an optional implementation, a threat assessment model is constructed based on the multi-dimensional feature point cloud, and a pre-warning for aircraft collision avoidance is performed based on the threat assessment model, comprising:
[0036] Within the category attribute, a distance threat function is constructed based on the distance attribute using an exponential decay function, and a height threat function is constructed based on the height attribute using a preset function;
[0037] A weighted two-factor threat assessment model is constructed based on the distance threat function, a preset distance weight, the height threat function and a preset height weight;
[0038] A threat value of the threat target is calculated based on the weighted two-factor threat assessment model, a threat level is divided based on the threat value, and a pre-warning for aircraft collision avoidance is performed based on the threat level.
[0039] The multi-modal data fusion aircraft collision pre-warning method provided by the application constructs threat functions in category attributes, realizes differentiated assessment of the same type of targets, avoids threat confusion of different types of targets, uses an exponential decay function for the distance threat function to accurately depict the rule that the closer the distance, the more the threat increases non-linearly, uses a preset function for the height threat function to highlight the threat sudden increase effect of height difference in the danger threshold interval, realizes dynamic adaptation of threat factors by preset distance weight and height weight in the weighted two-factor threat assessment model, and solves the one-sidedness of single-factor assessment. The threat value is divided into levels to realize graded response, accurate intervention, and balance the effectiveness of pre-warning and the cognitive load of pilots.
[0040] In an optional implementation, the method further comprises:
[0041] The threat targets are sorted in descending order of threat level, and a structured threat target list and an optoelectronic image marking the threat target are output in real time; each threat target comprises a threat ID, a threat type, a threat value and a relative height.
[0042] The multi-modal data fusion aircraft collision pre-warning method provided by the application can intuitively present the urgency of the threat in descending order of threat level, help pilots quickly focus on high-priority targets, present the core information (threat ID, type, threat value and relative height) of each target in a standardized format to realize information clarification and standardization, and combine abstract threat information with intuitive visual images to realize associated presentation of data and images, so that the pilot's perception of the threat is more three-dimensional, the actual situation of the target can be more accurately judged by combining the texture, color and other details of the image, and the overall cognition of the threat scene is improved.
[0043] In a second aspect, the present application provides a multi-modal data fusion aircraft collision warning device, which comprises:
[0044] a multi-modal data acquisition module, configured to acquire multi-modal data, the multi-modal data comprising radar point cloud data and photoelectric image data;
[0045] a preprocessing module, configured to preprocess the radar point cloud data and the photoelectric image data respectively to obtain radar effective point cloud, semantic mask and confidence;
[0046] a cross-modal filtering fusion module, configured to perform cross-modal filtering fusion based on the radar effective point cloud, the semantic mask and the confidence to obtain multi-dimensional feature point cloud;
[0047] a threat assessment and collision warning module, configured to construct a threat assessment model based on the multi-dimensional feature point cloud and to perform collision warning for the aircraft based on the threat assessment model.
[0048] In a third aspect, the present application provides a computer device, comprising a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the multi-modal data fusion aircraft collision warning method of the first aspect or any of the corresponding embodiments thereof.
[0049] In a fourth aspect, the present application provides a computer readable storage medium, which stores computer instructions, and the computer instructions are used to make the computer execute the multi-modal data fusion aircraft collision warning method of the first aspect or any of the corresponding embodiments thereof.
[0050] In a fifth aspect, the present application provides a computer program product, which comprises computer instructions, and the computer instructions are used to make the computer execute the multi-modal data fusion aircraft collision warning method of the first aspect or any of the corresponding embodiments thereof. BRIEF DESCRIPTION OF DRAWINGS
[0051] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed in the specific embodiments or the prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0052] Figure 1 is a flowchart of the multi-modal data fusion aircraft collision warning method according to the embodiments of the present application;
[0053] Figure 2is a flowchart of another multi-modal data fusion aircraft collision warning method according to an embodiment of the present application;
[0054] Figure 3 is a flowchart of still another multi-modal data fusion aircraft collision warning method according to an embodiment of the present application;
[0055] Figure 4 is a flowchart of yet another multi-modal data fusion aircraft collision warning method according to an embodiment of the present application;
[0056] Figure 5 is a structural block diagram of a multi-modal data fusion aircraft collision warning device according to an embodiment of the present application;
[0057] Figure 6 is a hardware structure schematic diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION
[0058] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0059] At present, for the related detection of aircraft collision, laser radar detection is mostly used, single modal target detection mainly includes radar detection and photoelectric detection. The radar detection has high distance resolution but low angle resolution and no semantic information. The photoelectric detection contains semantic information and has high angle resolution but depends on light and has no direct depth information. The single modal method has low cost, but each sensor has its own limitations and cannot consider information, thereby increasing the risk of missed detection.
[0060] Multi-modal detection combines the advantages of radar and photoelectric sensors and performs fusion at three levels of data level, feature level and decision level:
[0061] Data level fusion directly integrates the original data of different sensors, which can retain the most abundant information, but has high calculation cost and strict requirements for space-time synchronization. Small deviations will lead to fusion errors.
[0062] Feature level fusion extracts intermediate features of each modal for fusion, which needs to rely on feature alignment algorithm and is still limited by the space-time registration accuracy between sensors.
[0063] In decision level fusion, each sensor independently processes and outputs results, and then the results are fused by algorithm, but the bottom layer data interaction is lacking and the global understanding of complex scenes is weak.
[0064] Current multi-modal methods are mostly based on laser radar or millimeter wave radar. Laser radar detection range is limited within a few hundred meters, and due to high cost, long distance signal attenuation and other problems, it is difficult to directly adapt to long distance detection scene. In the long distance scene, the point cloud data becomes extremely sparse, resulting in loss of target details, especially the detection ability of small objects is significantly reduced. In addition, point cloud lacks color and texture information, making it difficult to distinguish different objects with similar geometric shapes. While the method based on pure photoelectric image is richer in texture and semantic information, but limited by the principle of passive imaging, the depth estimation accuracy is insufficient in long distance scene, the target scale is reduced and the texture is blurred, and it is easily affected by light conditions (such as strong light, low illumination) and weather (such as fog, rain), resulting in target missed detection or false detection.
[0065] In the long distance (such as 1000 meters away) scene, multi-modal fusion technology faces significant challenges: first, radar point cloud is extremely sparse due to increased distance, and photoelectric image is affected by atmospheric scattering and diffraction, resulting in reduced resolution, making the noise of semantic segmentation on which traditional fusion methods rely increase, reducing the reliability of fusion. Second, the time and space synchronization error is amplified due to the difference in sensor sampling frequency and the offset of the viewing angle, making it difficult to align cross-modal features, especially in the BEV space, the point cloud and image scale difference of the long distance target is significant. In addition, the computational complexity of existing algorithms increases non-linearly with the detection distance, making it difficult to meet the real-time requirements, and the multi-modal complementarity is easily ineffective in extreme environments such as rain, fog, and strong light.
[0066] The applicability of the above various methods is different, but when facing long distance environment, the above methods all have the problem of balancing performance and accuracy. There is an urgent need for an aircraft collision avoidance warning method that improves the robustness of the system in complex environments (such as low visibility, high dynamic airspace) while meeting the collision avoidance needs of different application scenarios.
[0067] For long distance high precision collision avoidance needs, the fusion technology of radar point cloud and photoelectric image has become a research hotspot. Laser radar (LiDAR) provides centimeter-level precision 3D structure information, but is limited by sparsity and short distance coverage (usually <300 meters); while multi-view vision system can achieve target recognition and ranging up to 1600 meters through deep learning, but relies on light conditions. Multi-modal fusion cooperates at data level, feature level and decision level, making up for the limitations of single sensor, and realizing dynamic threat level division. The present embodiment provides a multi-modal data fusion aircraft collision avoidance warning method, which realizes the automatic detection of threat targets of the aircraft in flight by multi-modal data fusion based on millimeter wave radar and photoelectric image, optimizes the data fusion strategy, improves the cross-modal alignment accuracy and enhances the robustness of the algorithm, realizes the automatic detection of threat targets of the aircraft in flight, measures a longer distance, and has higher measurement accuracy, solving the problem of short detection distance and low detection accuracy in the prior art.
[0068] According to the embodiment of the present application, a multi-modal data fusion aircraft collision warning method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.
[0069] In the present embodiment, a multi-modal data fusion aircraft collision warning method is provided, which can be used in the above-mentioned computer device, wherein the aircraft can be an airplane, other aircraft, etc. Figure 1 The flowchart of the multi-modal data fusion aircraft collision warning method according to the embodiment of the present application is shown in FIG. 1, which includes the following steps: Figure 1
[0070] In step S101, multi-modal data is obtained, including radar point cloud data and photoelectric image.
[0071] Specifically, the environmental data is collected in real time by a multi-modal sensor. For example, a millimeter wave radar sensor is used as a data collection device, which is usually installed at the nose, wing or belly of the aircraft, etc., to ensure that the detection field of view covers the key airspace in front of and around the aircraft. The radar point cloud data is collected by a frequency-modulated continuous wave radar or a pulsed Doppler radar, and each frame of radar point cloud contains a large number of three-dimensional points. The original information of each point cloud point includes three-dimensional coordinates in the radar coordinate system, reflection intensity (reflecting the reflection characteristics of the target surface material, such as the reflection intensity of a metal target being higher than that of a non-metal), and time stamp, etc.
[0072] An RGB camera and an infrared camera are provided to form a multi-spectral photoelectric collection system. The RGB camera is used for daytime or well-lit environments to capture details such as the color and texture of the target; the infrared camera is used for night, fog or low-light environments to image the temperature difference between the target and the background. The installation positions of the two cameras are kept overlapping (overlapping rate ≥ 80%) to ensure consistent collection range. The RGB image is stored in RGB three-channel pixel values (0-255) to reflect the visual appearance of the target, such as the color of the aircraft body and the shape of the wing; the infrared image is stored in grayscale values (0-255), and the higher the brightness, the higher the temperature of the target, which can distinguish the temperature difference between the engine and the body.
[0073] The radar point cloud provides three-dimensional position, speed and reflection intensity information of the target, covering long distances and all-weather scenes; the photoelectric image (including RGB image and infrared image) supplements rich texture and temperature characteristics, which is particularly suitable for close-range fine identification. The two types of data are synchronized by hardware or time stamp to ensure time consistency, laying a foundation for subsequent fusion. For example, the millimeter wave radar can detect point clouds at a distance of thousands of meters, while the infrared camera can still image clearly at night or in fog.
[0074] Step S102, respectively, pre-process the radar point cloud data and the photoelectric image to obtain radar effective point cloud, semantic mask and confidence.
[0075] Specifically, in the radar point cloud preprocessing process, field of view filtering is the first key step. The main purpose of this step is to quickly eliminate invalid point cloud data outside the field of view according to the actual detection range of the radar sensor, thereby significantly reducing the computational load of subsequent processing. The detection capability of the radar is usually limited by the horizontal and vertical angles, which together define the effective detection space of the radar. For the near distance interval with more sidelobes and the false alarm point cloud points that are too far away, distance threshold filtering is also performed to eliminate invalid point cloud points outside the effective detection space.
[0076] The photoelectric image preprocessing uses a preset neural network to generate a pixel-level semantic mask, which labels the semantic category and confidence.
[0077] Step S103, cross-modal filtering fusion based on radar effective point cloud, semantic mask and confidence to obtain multi-dimensional feature point cloud.
[0078] Specifically, regional-level spatial filtering is used to perform cross-modal filtering fusion based on radar effective point cloud, semantic mask and confidence, and only 3D point cloud points in the overlapping area of the point cloud and the semantic mask (such as vehicle / pedestrian category area) are retained. According to the segmentation confidence, low-reliability point cloud points are removed to reduce false detection. Attribute enhancement is used to assign each point cloud point with multiple attributes such as class label, relative ground height (calculated by ground fitting) and Euclidean distance, and generate multi-dimensional feature point cloud.
[0079] Step S104, constructing a threat assessment model based on the multi-dimensional feature point cloud, and warning the aircraft against collision based on the threat assessment model.
[0080] Specifically, in the target category set, a threat assessment model is constructed by weighted fusion of distance and height double factors, and the threat value of the threat target is calculated based on the threat assessment model, and the aircraft is warned against collision through the warning level of the threat value.
[0081] The multi-modal data fusion aircraft collision warning method provided in the embodiment realizes the complementary advantages of multi-modal data, greatly improves the robustness of flight environment perception, respectively pre-processes radar point cloud data and photoelectric image, obtains radar effective point cloud, semantic mask and confidence, performs cross-modal filtering fusion based on the radar effective point cloud, the semantic mask and the confidence, obtains multi-dimensional feature point cloud, eliminates invalid radar point cloud and background interference, realizes cross-modal filtering fusion, and the double-factor weighted threat evaluation model constructed based on the multi-dimensional feature point cloud realizes automatic detection of the threat target of the aircraft in flight, the measuring distance is farther, and the measuring precision is higher, thereby solving the problems of short collision detection distance and low detection precision in the prior art.
[0082] In the embodiment, a multi-modal data fusion aircraft collision warning method is provided, which can be used for the computer device described above, Figure 2 is a flowchart of the multi-modal data fusion aircraft collision warning method according to the embodiment of the application, as Figure 2 shown, the flowchart includes the following steps:
[0083] Step S201, multi-modal data is obtained, and the multi-modal data includes radar point cloud data and photoelectric image. For details, refer to step S101 of the embodiment shown in Figure 1 , which will not be repeated here.
[0084] Step S202, respectively pre-process the radar point cloud data and the photoelectric image to obtain radar effective point cloud, semantic mask and confidence.
[0085] Specifically, the above step S202 includes:
[0086] Step S2021, field of view filtering and distance threshold filtering are performed on the radar point cloud data to obtain radar effective point cloud.
[0087] Specifically, the field of view filtering refers to spatial range screening of the original radar point cloud data according to the physical detection field of view range of the radar sensor to eliminate invalid point cloud outside the field of view.
[0088] In the radar point cloud preprocessing flowchart, the field of view filtering is the first key step. The main purpose of this step is to quickly eliminate invalid point cloud data outside the field of view according to the actual detection range of the radar sensor, thereby significantly reducing the calculation amount of subsequent processing. The detection capability of the radar is usually limited by the horizontal and vertical angles, and these limitations jointly define the effective detection space of the radar.
[0089] The detection field of view of the radar sensor is determined by its hardware design, including the horizontal detection angle (e.g. ±90°) and the vertical detection angle (e.g. ±30°). The field of view filtering retains the point cloud whose azimuth angle (the angle between the horizontal direction and the central axis of the radar) and elevation angle (the angle between the vertical direction and the central axis of the radar) are within the radar field of view, and deletes the point cloud whose azimuth angle and elevation angle are out of the angle range (e.g. the point cloud with an azimuth angle of 100° is out of the ±90° field of view).
[0090] In a specific implementation, for each radar point [θ, φ, r], the angle values are compared with the preset field of view range [θ_min, θ_max] and [φ_min, φ_max] to determine whether the point is located in the effective detection area. In addition, distance threshold filtering is also performed on the near distance range with many sidelobes and the false alarm points that are too far away, to remove the invalid points outside [r_min, r_max]. Wherein, θ, φ, and r respectively represent the azimuth angle, the elevation angle, and the distance, θ_min and θ_max respectively represent the minimum azimuth angle and the maximum azimuth angle, φ_min and φ_max respectively represent the minimum elevation angle and the maximum elevation angle, and r_min and r_max respectively represent the minimum distance and the maximum distance.
[0091] The point cloud data retained after the field of view filtering not only greatly reduces the data size, but also ensures that the subsequent processing process only focuses on the truly meaningful detection area, laying a good foundation for subsequent semantic analysis and target recognition. This filtering method based on geometric constraints has the advantages of high computational efficiency and clear physical meaning, and is an indispensable important link in the preprocessing of radar point cloud data.
[0092] In step S2022, a preset neural network is used for image semantic segmentation of the photoelectric image to generate a semantic mask and a confidence.
[0093] Specifically, the preset neural network selects a lightweight SegFormer network. The lightweight SegFormer is a semantic segmentation framework combining a hierarchical Transformer encoder and a lightweight multilayer perceptron (MLP) decoder.
[0094] In an optional implementation, the neural network includes a multi-layer encoder and a multi-layer decoder, and the step S2022 includes:
[0095] In step a1, the multi-layer encoder is used to extract multi-scale features of the photoelectric image, and the multi-layer decoder is used to fuse the multi-scale features to obtain a semantic prediction result.
[0096] The photoelectric image preprocessing generates a pixel-level semantic mask using a lightweight SegFormer network, labels the class and confidence. The lightweight SegFormer network extracts multi-scale features through a hierarchical Transformer encoder, where the output of the ith layer encoder can be expressed as:
[0097] F i = Encoder i (I), i ∈ {1, 2, 3, 4} (1);
[0098] where I represents the input original photoelectric image; Encoder i (I) represents the output of the ith layer hierarchical encoder of the lightweight SegFormer network, and in this embodiment, a four-layer encoder is selected, which receives the image I as input and outputs a multi-scale feature map. With the deepening of the level, the resolution of this multi-scale feature map gradually decreases, but the semantic information it contains is more rich.
[0099] The resolution gradually decreases with the depth of the network. These multi-scale features are fused through a lightweight multi-layer MLP decoder to generate the final semantic prediction result:
[0100]
[0101] where Decoder represents the lightweight MLP decoder in the lightweight SegFormer network, which receives multiple different scale feature maps from each level of the encoder, effectively fuses them, and finally generates a pixel-level semantic prediction result. The specific details include: first, multiple feature maps from different levels of the encoder are processed through independent lightweight multi-layer MLP decoders to unify the channel dimension; then, all the processed are up-sampled to make their resolutions reach a common size, and all the up-sampled are spliced in the channel dimension; finally, it is input into another layer of MLP decoder to generate the final semantic prediction result.
[0102] Step a2, generating a pixel-level class probability distribution based on the semantic prediction result, and generating a semantic mask and a confidence based on the pixel-level class probability distribution.
[0103] Specifically, in order to obtain the class probability distribution of each pixel point, the softmax function (normalized exponential function) is used to normalize to obtain the pixel-level class probability distribution, and the formula is as follows:
[0104]
[0105] Wherein, u, v are pixel coordinates, u is the horizontal coordinate (column), v is the vertical coordinate (row), c is the semantic class, M is the semantic prediction result fused by the lightweight multi-layer MLP decoder, and is the original result before the softmax function.
[0106] Based on this, the semantic mask and the confidence are generated:
[0107]
[0108] Wherein, S(u, v) represents the semantic mask, and conf(u, v) represents the confidence.
[0109] In step S203, the near-sequence frame extraction method is used to time-align the radar effective point cloud and the semantic mask; a coordinate projection model is established, and the radar effective point cloud and the semantic mask are spatially aligned based on the coordinate projection model.
[0110] Specifically, the spatio-temporal alignment is the core link for realizing the fusion of the radar point cloud and the image semantic information, which is divided into two parts of time alignment and spatial coordinate conversion, and finally realizes the accurate mapping of each radar point and image pixel.
[0111] Wherein, the details of the time alignment are as follows:
[0112] 1. Time stamp collection and storage: the radar point cloud frame and the photoelectric image frame both carry high-precision time stamps (provided by the aircraft synchronous clock, with a precision of microseconds). For example, the time stamp of the radar point cloud is Tr=[t r1 ,t r2 ,...,t rn ] (unit: ms), corresponding to the generation time of each frame of point cloud; the time stamp of the photoelectric image is Ti=[t i1 ,t i2 ,...,t im ], corresponding to the acquisition time of each image.
[0113] 2. For the current image frame to be processed (time stamp t i ), the time stamp sequence of the radar point cloud is traversed, and the time difference△t k =|t i -t rk | is calculated. The radar point cloud frame with the smallest△t k is selected as the matching frame, so that the time deviation is controlled within a preset threshold (such as≤10ms). If the minimum deviation exceeds the threshold (such as due to sensor failure leading to asynchronization), the system will trigger a warning and use the nearest valid frame to replace, so as to avoid that the time misplacement is too large to affect the spatial alignment accuracy.
[0114] The spatial alignment details are described as follows: the core of the spatial alignment is to convert the 3D point (xi, yi, zi) in the radar coordinate system into the 2D image pixel coordinate (ui, vi) in the image plane through the coordinate projection model, that is, to convert the radar coordinate system to the image plane through the calibration file.
[0115] 1. Conversion of the radar coordinate system to the camera coordinate system:
[0116] Radar coordinate system: taking the center of the radar antenna as the origin Or, the xr axis points to the front of the aircraft, the yr axis points to the right of the aircraft, and the zr axis is vertically upward (right-hand coordinate system).
[0117] Camera coordinate system: taking the camera optical center as the origin Oc, the xc axis is parallel to the image plane and horizontally to the right, the yc axis is parallel to the image plane and vertically downward, and the zc axis is along the camera optical axis and forward (in the same direction as the radar coordinate system).
[0118] For each 3D point (xi, yi, zi) in the radar coordinate system, the corresponding image pixel coordinate can be calculated through the camera intrinsic matrix K and the transformation matrix [R|t] from the radar to the camera, and the formula is as follows:
[0119]
[0120] Where, u i , v i represent the 2D image pixel coordinates of the image plane.
[0121] The multi-modal data fusion aircraft collision warning method provided in the embodiment effectively solves the problem of different sampling frequencies of the radar and the photoelectric sensor by selecting the radar point cloud and the photoelectric image frame with the closest time stamp for matching, realizes time alignment, the coordinate projection model accurately converts the three-dimensional coordinates of the effective radar point cloud into the two-dimensional pixel coordinates of the image plane, achieves spatial position matching of the radar point cloud and the semantic mask, realizes spatial alignment, and through the cooperation of time alignment and spatial alignment, constructs the cross-modal data consistent in time and space.
[0122] Step S204: performing cross-modal filtering fusion based on the effective radar point cloud, the semantic mask and the confidence to obtain a multi-dimensional feature point cloud. For details, please refer to the step S103 of the embodiment shown in Figure 1 The step S103 of the embodiment shown in
[0123] Step S205: constructing a threat assessment model based on the multi-dimensional feature point cloud, and warning the aircraft collision based on the threat assessment model. For details, please refer to the step S104 of the embodiment shown in Figure 1 The step S104 of the embodiment shown in
[0124] The multi-modal data fusion aircraft collision warning method provided in the embodiment can greatly reduce data redundancy, improve processing efficiency, and enhance the physical effectiveness of point cloud data by combining field of view filtering and distance threshold filtering, thereby laying a foundation for cross-modal fusion. The multi-layer encoder can capture multi-scale features of the photoelectric image through hierarchical processing, the shallow layer encoder can extract detailed information of the image, the multi-layer decoder can fuse the multi-scale features, and the multi-layer decoder can complement the advantages of features at different levels to generate a pixel-level class probability distribution, thereby providing a quantitative basis for the possibility of each pixel point belonging to different categories and providing an important reference for subsequent cross-modal filtering and fusion. The cross-modal filtering and fusion can filter point clouds corresponding to a low confidence area, reduce false detection, and improve the accuracy of target screening.
[0125] In the embodiment, a multi-modal data fusion aircraft collision warning method is provided, which can be used for the computer device described above, Figure 3 is a flowchart of the multi-modal data fusion aircraft collision warning method according to the embodiment of the present application, as shown in the figure, the flowchart includes the following steps: Figure 3
[0126] In step S301, multi-modal data is acquired, and the multi-modal data includes radar point cloud data and photoelectric image. For details, refer to step S201 of the embodiment shown in Figure 2 , which will not be repeated here.
[0127] In step S302, the radar point cloud data and the photoelectric image are respectively preprocessed to obtain radar effective point cloud, semantic mask, and confidence. For details, refer to step S202 of the embodiment shown in Figure 2 , which will not be repeated here.
[0128] In step S303, cross-modal filtering and fusion are performed based on the radar effective point cloud, the semantic mask, and the confidence to obtain multi-dimensional feature point cloud.
[0129] Specifically, the above step S303 includes:
[0130] In step S3031, an image semantic category corresponding to a point cloud projection area is acquired, and the image semantic category is assigned to a corresponding point cloud point to obtain a class label of each point cloud point. The class label of each point cloud point is compared with a target class set, and point cloud points not belonging to the target class set are deleted to obtain a class attribute of each point cloud point.
[0131] Specifically, after semantic segmentation of the photoelectric image, a pixel-level semantic mask is generated (stored in the form of a two-dimensional array, the size is consistent with the image, and each element is the class label of the corresponding pixel, such as 0 = background, 1 = aircraft, 2 = bird, and 3 = large obstacle). For the radar point cloud that has completed spatial alignment, the projected pixel coordinates (u, v) can directly index the semantic mask to obtain the corresponding class label. For example, a certain radar point is projected to the image (u = 500, v = 300), and the label at this position in the semantic mask is 1, so the initial class label of the point cloud is aircraft.
[0132] The target class set is preset by the system according to the anti-collision requirements, for example (corresponding to aircraft and large obstacle). The class label of each point cloud point is judged: if the label belongs to , the label is converted into a structured description (such as 1→aircraft); if it does not belong to (such as label 2 = bird), the point cloud is directly deleted. For example, the point cloud with label 2 is excluded because it is not in the target set, and the final retained point cloud only contains two types of aircraft and large obstacle, achieving preliminary filtering of background interference.
[0133] In step S3032, the ground plane equation is estimated by a plane fitting algorithm, and the vertical distance of each point cloud point relative to the ground is calculated based on the ground plane equation to obtain the height attribute of each point cloud point.
[0134] Specifically, the ground candidate points are selected from the radar effective point cloud: in combination with the point cloud corresponding to the ground class (label 0) in the semantic mask, or by a prior height threshold (such as points with a relative aircraft height lower than -5m, assuming that the aircraft height from the ground is >10m) to preliminarily judge as ground points, forming a ground point cloud subset P ground .
[0135] The ground plane equation is estimated by a plane fitting algorithm (such as a random sample consensus algorithm), and the steps are as follows:
[0136] Random sampling: 3 non-collinear points are randomly selected from P ground , and an initial plane ax+by+cz+d=0 (the plane equation is standardized as a 2 +b 2 +c 2 =1) is fitted.
[0137] Inlier judgment: the distances di=|axi+byi+czi+d| of all points in P ground to the plane are calculated, and points with a distance less than a threshold (such as 0.5m) are regarded as inliers.
[0138] Iterative optimization: repeat the sampling, fitting, and judging process (set the number of iterations to 50-100 times), and keep the plane with the most inliers as the final ground plane equation Ax+By+Cz+D=0.
[0139] For each point cloud point (xi, yi, zi), substitute the ground plane equation to calculate the vertical distance.
[0140] Step S3033, calculate the straight-line distance between each point cloud and the carrier based on the projection pixel coordinates of each point cloud point to obtain the distance attribute of each point cloud point.
[0141] Specifically, the carrier is set as the origin (0, 0, 0) in the radar coordinate system, and the three-dimensional coordinates of each radar point cloud are (xi, yi, zi) (after preprocessing, the effective point cloud), and the straight-line distance between the carrier is the Euclidean distance:
[0142] Step S3034, generate a multi-dimensional feature point cloud based on the category attribute, height attribute, and distance attribute.
[0143] Specifically, for the projection result, perform semantic-guided point cloud filtering, and the screening condition is:
[0144]
[0145] wherein, represents the target category set, P filtered represents the final radar point cloud set reserved after filtering under the guidance of the semantic mask, S(u i ,v i represents the semantic category label of the pixel after the radar point is projected, i.e., the semantic mask, p i represents the i-th point cloud, conf(u i ,v i represents the confidence score predicted by the model after the pixel, i.e., the probability that the pixel is of this category (semantic mask), τ is the confidence threshold, according to the confidence, there may be pixels that are difficult to distinguish (such as a point at the junction 55% sky and 45% mountain), and the point cloud on such difficult-to-distinguish pixels is filtered out.
[0146] This step effectively eliminates background interference, and finally obtains the fused point cloud, such as a point cloud labeled as "aircraft-distance 120m-height difference 15m", forming a multi-dimensional feature point cloud. The final features of a certain effective point cloud are: (50.2, 10.3, 5.2, aircraft, 15m, 120m).
[0147] Step S304, construct a threat assessment model based on the multi-dimensional feature point cloud, and perform pre-warning for the carrier collision avoidance based on the threat assessment model.
[0148] Specifically, the step S304 includes:
[0149] Step S3041, within the category attribute, a distance threat function is constructed based on the distance attribute using an exponential decay function, and a height threat function is constructed based on the height attribute using a preset function.
[0150] Specifically, the distance threat is modeled using an exponential decay function, reflecting the characteristic that the threat decreases with the increase of distance, and the formula is as follows:
[0151]
[0152] where T d is the distance threat function, λ d is the distance threat decay coefficient, which is a normal number, used to control the speed of the distance threat function decay, d is the target distance, d max is the maximum distance.
[0153] The height threat uses a Sigmoid function (activation function) to evaluate the threshold effect, and the height threat function formula is as follows:
[0154]
[0155] where T h is the height threat function, λ h is the height threat coefficient, h is the absolute value of the height difference, h0 is the preset height difference threat threshold, which is the critical height difference for the threat level to change significantly.
[0156] Step S3042, a weighted two-factor threat assessment model is constructed based on the distance threat function, the preset distance weight, the height threat function and the preset height weight.
[0157] Specifically, the weighted two-factor threat assessment model formula is as follows:
[0158] T = w d ·T d (d) + w h ·T h (h) (9);
[0159] where T is the threat value, w d is the distance weight, w h is the height weight, the setting of the distance weight and the height weight is according to the flight scene and environmental characteristics of the carrier (such as the flight height interval, airspace complexity), the task and performance parameters of the carrier (such as the task type and maneuvering performance), the threat characteristics of the target type (such as large aircraft, small unmanned aerial vehicle and fixed obstacle), which is not limited here.
[0160] Step S3043, the threat value of the threat target is calculated based on the weighted two-factor threat evaluation model, and the threat level is divided based on the threat value, and the warning of the carrier collision is prewarned based on the threat level.
[0161] Specifically, when the carrier detects a threat target, the threat value of the threat target is calculated according to formula (8), the warning is divided into three levels according to the threat value range, and the threshold can be dynamically adjusted according to the type of the carrier (such as a civil aviation passenger plane or a helicopter):
[0162] Low threat warning (T≤0.3): The target threat degree is low, and it does not constitute a collision risk. For example, the target with a distance of more than 500m from the carrier and a height difference of more than 100m, the threat value is usually in this interval.
[0163] Medium threat warning (0.3<T≤0.7): The target has potential collision risk and needs to be tracked continuously. For example, the target with a distance of 200-500m and a height difference within the dangerous reference ±30m, the threat value falls into this range.
[0164] High threat warning (T>0.7): The target has approached the collision danger threshold and needs to take immediate evasive measures. For example, the target with a distance of less than 200m and a height difference within the dangerous reference ±10m, the threat value will exceed 0.7.
[0165] Step S305, the threat targets are sorted in descending order according to the threat level, and a structured threat target list and an optoelectronic image marked with the threat target are output in real time; each threat target includes threat ID, threat type, threat value and relative height.
[0166] Specifically, the structured threat target list is output in real time through multi-sensor fusion, and is displayed in descending order according to the threat level, each threat target contains ID, type, threat value, relative height and other key information, and the threat target is marked on the output image. In the visualization layer, the system highlights the individual with the highest threat value among the same objects, providing pilots with decision assistance that combines spatial perception and situation assessment.
[0167] The multi-modal data fusion aircraft collision warning method provided by the embodiment can accurately lock the threat target of interest, estimate the ground plane equation through a plane fitting algorithm, calculate the vertical distance of each point cloud point relative to the ground to obtain the height attribute, and provide key spatial dimension information for threat assessment, calculate the linear distance of the projection pixel coordinates of the point cloud point and the aircraft to obtain the distance attribute, which can intuitively reflect the spatial distance between the target and the aircraft, and generate multi-dimensional feature point clouds based on the category attribute, the height attribute and the distance attribute, which integrates the semantic information, spatial height information and distance information of the target, and provides comprehensive and rich input data for threat assessment. The threat function is constructed in the category attribute, the differential assessment of the same type of target is realized, the threat confusion of different types of targets is avoided, the distance threat function adopts an exponential decay function, the non-linear surge rule of the threat with the closer distance is accurately described, the height threat function adopts a preset function, the threat sudden increase effect of the height difference in the dangerous threshold interval can be highlighted, the dynamic adaptation of the threat factors is realized through the preset distance weight and height weight of the weighted double-factor threat assessment model, and the one-sidedness of single-factor assessment is solved. Based on the threat value division level and the linkage warning, the hierarchical response and accurate intervention are realized, and the effectiveness of the warning and the flight crew's cognitive load are balanced.
[0168] As one or more specific application embodiments of the present application, the present application is described in conjunction with Figure 4 The multi-modal data fusion aircraft collision warning method provided by the present application is further described in detail as follows. Figure 4 As shown in the figure, the specific process is as follows:
[0169] In order to be able to carry out collision warning at a long distance, the present application provides an aircraft collision warning method based on millimeter wave radar and photoelectric image. The method can carry out multi-modal fusion detection, optimize the data fusion strategy, improve the cross-modal alignment accuracy, and enhance the robustness of the algorithm.
[0170] 1. Data input:
[0171] Radar point cloud: input the original point cloud data.
[0172] Photoelectric image: input the synchronously collected RGB image or infrared image.
[0173] Specifically, the environmental data is collected in real time by multi-modal sensors. Radar point cloud provides three-dimensional position, velocity and reflection intensity information of the target, covering long-distance, all-weather scenes; photoelectric image (RGB image / infrared image) supplements rich texture and temperature features, especially suitable for close-range fine identification. The two kinds of data are synchronized by hardware or timestamp alignment to ensure temporal consistency and lay a foundation for subsequent fusion. For example, millimeter wave radar can detect point cloud at a distance of thousands of meters, while infrared camera can still image clearly at night or in foggy conditions.
[0174] 2. Preprocessing and alignment:
[0175] 1) Radar preprocessing: filtering noise points, region segmentation.
[0176] Specifically, in the radar point cloud preprocessing process, field of view filtering is the first key step. The main purpose of this step is to quickly eliminate invalid point cloud data outside the field of view according to the actual detection range of the radar sensor, thereby significantly reducing the computational load of subsequent processing. The detection capability of the radar is usually limited by the horizontal and vertical angles of view, which together define the effective detection space of the radar.
[0177] In specific implementation, for each radar point [θ, φ, r], the angle values are compared with the preset field of view range [θ_min, θ_max] and [φ_min, φ_max] to determine whether the point is located within the effective detection area. Secondly, for the near distance interval with many sidelobes and the false alarm points that are too far away, distance threshold filtering is also performed to eliminate invalid points outside [r_min, r_max]. Wherein, θ, φ, r represent azimuth angle, elevation angle, distance respectively, θ_min and θ_max represent minimum azimuth angle and maximum azimuth angle respectively, φ_min and φ_max represent minimum elevation angle and maximum elevation angle respectively, r_min and r_max represent minimum distance and maximum distance respectively.
[0178] After field of view filtering, the retained point cloud data not only greatly reduces the data size, but also ensures that the subsequent processing process only focuses on the truly meaningful detection area, laying a good foundation for subsequent semantic analysis and target identification. This filtering method based on geometric constraints has the advantages of high computational efficiency and clear physical meaning, and is an indispensable important link in radar point cloud data preprocessing.
[0179] 2) Image preprocessing: generating segmented semantic masks and corresponding confidence (including class labels) through a semantic segmentation network.
[0180] The photoelectric image preprocessing generates a pixel-level semantic mask using a lightweight SegFormer network, labels the class and confidence. The lightweight SegFormer network extracts multi-scale features through a hierarchical Transformer encoder, where the output of the ith layer encoder can be expressed as:
[0181] F i = Encoder i (I), i e {1, 2, 3, 4} (1);
[0182] where I represents the input original photoelectric image; Encoder i (I) represents the output of the ith layer hierarchical encoder of the lightweight SegFormer network, and in this embodiment, a four-layer encoder is selected, which receives the image I as input and outputs a multi-scale feature map. With the deepening of the level, the resolution of this multi-scale feature map gradually decreases, but the semantic information it contains is more rich.
[0183] The resolution gradually decreases with the depth of the network. These multi-scale features are fused through a lightweight multi-layer MLP decoder to generate the final semantic prediction result:
[0184]
[0185] where Decoder represents a lightweight MLP decoder in the lightweight SegFormer network, which receives multiple different scale feature maps from each level of the encoder, effectively fuses them, and finally generates a pixel-level semantic prediction result. The specific details include: first, the multiple feature maps from different levels of the encoder are processed through independent lightweight multi-layer MLP decoders to unify the channel dimension; then, all the processed are up-sampled to make their resolutions reach a common size, and all the up-sampled are spliced in the channel dimension, and finally, it is input into another layer of MLP decoder to generate the final semantic prediction result.
[0186] In order to obtain the class probability distribution of each pixel point, the pixel-level class probability distribution is obtained by normalizing through the softmax function (normalized exponential function), and the formula is as follows:
[0187]
[0188] where u and v are pixel coordinates, u is the horizontal coordinate (column), v is the vertical coordinate (row), c is the semantic class class, and M is the semantic prediction result after fusion through the lightweight multi-layer MLP decoder, which is the original result before the softmax function.
[0189] Based on this, the semantic mask and confidence are generated:
[0190]
[0191] where S(u, v) denotes semantic mask, conf(u, v) denotes confidence.
[0192] 3) Spatio-temporal alignment:
[0193] Temporal alignment: Interpolation or matching based on timestamps.
[0194] To achieve accurate alignment of radar point cloud and image semantic information, the time is aligned using the near-sequence frame extraction method, and the coordinate projection model is used to align the space. For specific details, see step S2023 described above, which will not be repeated here.
[0195] Spatial alignment: Project point cloud to image plane through calibration parameters.
[0196] The core of spatial alignment is to convert 3D points (xi, yi, zi) in the radar coordinate system to 2D image pixel coordinates (ui, vi) in the image plane through the coordinate projection model, that is, to convert the radar coordinate system to the image plane through the calibration file. For each 3D point (xi, yi, zi) in the radar coordinate system, the corresponding image pixel coordinates can be calculated through the camera intrinsic matrix K and the transformation matrix [R|t] of the radar to the camera, as follows:
[0197]
[0198] where u i , v i denote 2D image pixel coordinates in the image plane.
[0199] 3. Cross-modal fusion and filtering:
[0200] 1) Region-level spatial filtering: Only keep the points in the point cloud that overlap with the image segmentation mask (such as vehicle / pedestrian category regions), and remove low-reliability points according to the segmentation confidence.
[0201] Specifically, the system performs region-level spatial filtering: only keep the 3D points in the point cloud that overlap with the image semantic mask, and filter low-confidence regions to reduce false positives.
[0202] 2) Attribute enhancement: Assign class, relative height, and distance information to the fused point cloud.
[0203] The attribute enhancement stage assigns each point cloud point with a class label, a relative ground height (calculated by ground fitting), and a Euclidean distance. For specific details, see step S303 described above, which will not be repeated here.
[0204] For the projection result, perform semantic-guided point cloud filtering with the following screening conditions:
[0205]
[0206] wherein, represents the target category set, which effectively eliminates background interference, and finally obtains the fused point cloud, such as a point cloud marked as "aircraft-distance 120 m-height difference 15 m", forming a multi-dimensional feature point cloud. The final features of a certain valid point cloud are: (50.2, 10.3, 5.2, aircraft, 15 m, 120 m).
[0207] 4. Threat assessment:
[0208] Feature calculation:
[0209] Dynamic threat: based on distance.
[0210] Static threat: based on height, category.
[0211] Threat level: weighted fusion of the above static threat and dynamic threat features.
[0212] Specifically, the distance threat is modeled using an exponential decay function, reflecting the characteristic that the threat decreases with increasing distance, as follows:
[0213]
[0214] wherein, T d is the distance threat function, λ d is the distance threat decay coefficient, which is a normal number, used to control the speed of distance threat function decay, d is the target distance, d max is the maximum distance.
[0215] The height threat adopts a Sigmoid function (activation function) to evaluate the threshold effect, and the height threat function formula is as follows:
[0216]
[0217] wherein, T h is the height threat function, λ h is the height threat coefficient, h is the absolute value of height difference, h0 is the preset height difference threat threshold, which is the critical height difference for the threat level to change significantly.
[0218] The weighted two-factor threat assessment model formula is as follows:
[0219] T=w d ·T d (d)+w h ·T h (h) (9);
[0220] wherein, T is the threat value, wd is a distance weight, w h is a height weight, and the settings of the distance weight and the height weight are set according to the flight scene and environmental characteristics (such as the flight height interval, airspace complexity) of the carrier aircraft, the task and performance parameters (such as the task type and maneuvering performance) of the carrier aircraft, and the threat characteristics (such as large aircrafts, small unmanned aerial vehicles, and fixed obstacles) of the target type, which are not specifically limited herein.
[0221] 5. Output and visualization:
[0222] Sorted output: The target points are arranged in descending order of threat level.
[0223] System alarm: The high-threat targets (such as location, category, and threat value) are output.
[0224] Through multi-sensor fusion, a structured threat target list is output in real time, and each threat target contains ID, type, threat value, relative height, and other key information, and the threat target is marked on the output image. In the visualization layer, the system highlights the individual with the highest threat value among the same objects, providing pilots with decision assistance that combines spatial perception and situation assessment.
[0225] The multi-modal data fusion carrier collision warning method provided by the embodiment has the following innovative points:
[0226] 1) Multi-modal data fusion strategy: High-precision cross-modal target detection is achieved through time and space synchronization calibration and regional-level spatial filtering of millimeter wave radar and photoelectric images.
[0227] 2) Optimized data processing and feature enhancement: Noise filtering and semantic segmentation are used to enhance the attributes (category, height, distance) of point clouds, improving target feature expression.
[0228] 3) Dynamic and static threat assessment model: The threat level is calculated based on distance, category, and height, improving the accuracy of the warning.
[0229] The embodiment of the present application optimizes 1-5 km detection: The existing technology focuses on short and medium distance detection for autonomous driving, but the present application uses distance gate filtering and image semantic recognition to detect targets (mountains, etc.) above 1 km, and can fuse radar point clouds and photoelectric images in a long distance (1000 meters away) scene for multi-modal fusion detection. Through optimization of the data fusion strategy, improvement of the cross-modal alignment accuracy, and enhancement of the algorithm robustness, the key problems of sparse radar point clouds and small image target pixels in a long distance scene are solved.
[0230] The embodiment of the present application fuses millimeter wave radar (precise ranging) and photoelectric image (high angular resolution + semantic information recognition), combines time-space alignment and semantic filtering, uses a distance + height threat model, and realizes more intelligent aircraft collision avoidance warning decision through weighted fusion.
[0231] In the embodiment, a multi-modal data fusion aircraft collision avoidance warning device is also provided, which is used to realize the above-mentioned embodiments and preferred embodiments, and will not be described again. As used below, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the device described in the following embodiments is preferably realized in software, hardware, or a combination of software and hardware is also possible and is conceived.
[0232] The embodiment provides a multi-modal data fusion aircraft collision avoidance warning device, as shown in Figure 5 , comprising:
[0233] A multi-modal data acquisition module 501 is configured to acquire multi-modal data, and the multi-modal data includes radar point cloud data and photoelectric image.
[0234] A preprocessing module 502 is configured to respectively preprocess the radar point cloud data and the photoelectric image to obtain radar effective point cloud, semantic mask and confidence.
[0235] A cross-modal filtering and fusion module 503 is configured to perform cross-modal filtering and fusion based on the radar effective point cloud, the semantic mask and the confidence to obtain multi-dimensional feature point cloud.
[0236] A threat assessment and collision avoidance warning module 504 is configured to construct a threat assessment model based on the multi-dimensional feature point cloud, and to perform collision avoidance warning for an aircraft based on the threat assessment model.
[0237] In some optional embodiments, the preprocessing module 502 comprises:
[0238] A filtering unit is configured to perform field of view filtering and distance threshold filtering on the radar point cloud data to obtain radar effective point cloud.
[0239] A semantic segmentation unit is configured to perform image semantic segmentation on the photoelectric image using a preset neural network to generate a semantic mask and a confidence. In some optional embodiments, the neural network comprises a multi-layer encoder and a multi-layer decoder, and the semantic segmentation unit comprises:
[0240] A multi-scale feature extraction and fusion subunit is configured to extract multi-scale features of the photoelectric image using the multi-layer encoder, and to fuse the multi-scale features using the multi-layer decoder to obtain a semantic prediction result.
[0241] The semantic mask and confidence generation subunit is configured to generate a pixel-level class probability distribution based on the semantic prediction result, and generate a semantic mask and confidence based on the pixel-level class probability distribution.
[0242] In some optional embodiments, the multi-modal data fusion airborne collision avoidance warning device further comprises:
[0243] The space-time alignment module is configured to perform time alignment on the radar effective point cloud and the semantic mask by using a near-sequence frame extraction method; establish a coordinate projection model, and perform space alignment on the radar effective point cloud and the semantic mask based on the coordinate projection model.
[0244] In some optional embodiments, the cross-modal filtering fusion module 503 comprises:
[0245] The projection unit is configured to project the radar effective point cloud to an image plane through the coordinate projection model to obtain a projection pixel coordinate corresponding to each point.
[0246] The screening unit is configured to determine whether the projection pixel coordinate falls within an effective area of the semantic mask, and delete a point cloud with a confidence lower than a preset threshold in the effective area to obtain a point cloud projection area.
[0247] The point cloud filtering unit is configured to perform semantic-guided point cloud filtering on the point cloud projection area to obtain a multi-dimensional feature point cloud.
[0248] In some optional embodiments, the point cloud filtering unit comprises:
[0249] The class attribute determination subunit is configured to obtain an image semantic class corresponding to the point cloud projection area, and assign the image semantic class to a corresponding point cloud point to obtain a class label of each point cloud point, and compare the class label of each point cloud point with a target class set, delete a point cloud point not belonging to the target class set, and obtain a class attribute of each point cloud point.
[0250] The height attribute determination subunit is configured to estimate a ground plane equation by a plane fitting algorithm, and calculate a vertical distance of each point cloud point relative to the ground based on the ground plane equation to obtain a height attribute of each point cloud point.
[0251] The distance attribute determination subunit is configured to calculate a straight-line distance between each point cloud and the aircraft based on the projection pixel coordinate of each point cloud point to obtain a distance attribute of each point cloud point.
[0252] The multi-dimensional feature point cloud generation subunit is configured to generate a multi-dimensional feature point cloud based on the class attribute, the height attribute, and the distance attribute.
[0253] In some optional embodiments, the threat assessment and collision avoidance warning module 504 comprises:
[0254] The threat function construction unit is configured to construct a distance threat function based on an exponential decay function of the distance attribute and construct a height threat function based on a preset function of the height attribute within the category attribute.
[0255] The threat evaluation model construction unit is configured to construct a weighted two-factor threat evaluation model based on the distance threat function, a preset distance weight, the height threat function and a preset height weight.
[0256] The collision avoidance warning unit is configured to calculate a threat value of the threat target based on the weighted two-factor threat evaluation model, divide a threat level based on the threat value, and warn the aircraft collision avoidance based on the threat level.
[0257] In some optional embodiments, the multi-modal data fusion aircraft collision avoidance warning device further comprises:
[0258] The output and visualization module is configured to sort the threat targets in descending order of the threat level, and output a structured threat target list and an optoelectronic image marking the threat targets in real time, wherein each threat target comprises a threat ID, a threat type, a threat value and a relative height.
[0259] Further function descriptions of the above-mentioned various modules and units are the same as those of the above-mentioned corresponding embodiments, and will not be described here again.
[0260] The multi-modal data fusion aircraft collision avoidance warning device in the embodiment is presented in the form of functional units, wherein the units refer to ASIC (Application Specific Integrated Circuit, Application Specific Integrated Circuit) circuits, processors and memories executing one or more software or fixed programs, and / or other devices that can provide the above-mentioned functions.
[0261] The embodiment of the application further provides a computer device having the above-mentioned Figure 5 multi-modal data fusion aircraft collision avoidance warning device.
[0262] Please refer to Figure 6 , Figure 6 is a structural schematic diagram of a computer device provided by an optional embodiment of the application, as Figure 6As shown, the computer device includes one or more processors 10, memory 20, and interfaces 30 for external devices such as a keyboard and a mouse and a disk drive. One or more of the interfaces 30 enable a user to interact with the computer device. In some embodiments, the interface 30 also includes an input device, such as a microphone, or output device, such as a speaker. Figure 6 The processor 10 is used in the description as an example.
[0263] The processor 10 can be a central processing unit, a network processor, or a combination thereof. The processor 10 can further include a hardware chip. The hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device can be a complex programmable logic device, a field programmable logic device, a general array logic, or any combination thereof.
[0264] The memory 20 stores instructions that can be executed by the at least one processor 10 to cause the at least one processor 10 to perform the methods described in the above embodiments.
[0265] The memory 20 can include a program storage area and a data storage area. The program storage area can store an operating system, application programs, and the like for use by the at least one processor 10. The data storage area can store data created by the computer device, etc. Additionally, the memory 20 can include a volatile memory, such as a random access memory, and a non-volatile memory, such as at least one disk storage device, flash memory, or other non-volatile solid state storage device. In some embodiments, the memory 20 can optionally include a memory that is remote from the processor 10, such as a network storage device connected to the computer device through a network. Examples of the network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and a combination thereof.
[0266] The memory 20 can include a volatile memory, such as a random access memory, and a non-volatile memory, such as at least one disk storage device, flash memory, or other non-volatile solid state storage device.
[0267] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 can be connected by a bus or other means, Figure 6 The bus connection is taken as an example.
[0268] The input device 30 can receive inputted digital or character information, and generate key signal input related to user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 can include a display device, an auxiliary lighting device (e.g., an LED), a tactile feedback device (e.g., a vibration motor), etc. The display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0269] The embodiments of the present application also provide a computer readable storage medium, and the method according to the embodiments of the present application can be implemented in hardware, firmware, or recorded in a storage medium, or stored in a remote storage medium or a non-transitory machine readable storage medium downloaded through a network and stored in a local storage medium, so that the method described herein can be processed by such software on a storage medium using a general purpose computer, a special purpose processor, or programmable or special hardware. The storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk or a solid state disk, etc. Further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that the computer, the processor, the microprocessor controller or the programmable hardware include storage components that can store or receive software or computer code, which, when accessed and executed by the computer, the processor or the hardware, implements the method shown in the above embodiments.
[0270] Part of the present application can be applied as a computer program product, for example, computer program instructions, when executed by a computer, the operation of the computer can invoke or provide the method and / or technical solutions according to the present application. Those skilled in the art should understand that the form of computer program instructions in computer readable medium includes but is not limited to source file, executable file, installation package file, etc. Correspondingly, the way of computer program instructions executed by computer includes but is not limited to: the computer directly executes the instructions, or the computer executes the corresponding compiled program after compiling the instructions, or the computer reads and executes the instructions, or the computer executes the corresponding installed program after reading and installing the instructions. Here, the computer readable medium can be any available computer readable storage medium or communication medium accessible to the computer.
[0271] While embodiments of the application have been described in connection with the preferred embodiments of the various figures, those of ordinary skill in the art will appreciate that various modifications and changes can be made without departing from the spirit and scope of the application, and that such modifications and changes fall within the scope of the appended claims.
Claims
1. A multi-modal data fusion based aircraft collision warning method, characterized in that, The method comprises: acquiring multi-modal data, the multi-modal data comprising radar point cloud data and photoelectric image data; respectively pre-processing the radar point cloud data and the photoelectric image data to obtain radar effective point cloud, semantic mask and confidence; performing cross-modal filtering fusion based on the radar effective point cloud, the semantic mask and the confidence to obtain multi-dimensional feature point cloud; constructing a threat assessment model based on the multi-dimensional feature point cloud, and performing pre-warning for aircraft collision avoidance based on the threat assessment model.
2. The method of claim 1, wherein, The respective pre-processing of the radar point cloud data and the photoelectric image data to obtain radar effective point cloud and semantic mask comprises: performing field of view filtering and distance threshold filtering on the radar point cloud data to obtain radar effective point cloud; performing image semantic segmentation on the photoelectric image data using a preset neural network to generate a semantic mask and confidence.
3. The method of claim 2, wherein, The neural network comprises a multi-layer encoder and a multi-layer decoder, and the image semantic segmentation performed on the photoelectric image data using the preset neural network to generate a semantic mask and confidence comprises: extracting multi-scale features of the photoelectric image data using the multi-layer encoder, and fusing the multi-scale features using the multi-layer decoder to obtain a semantic prediction result; generating a pixel-level class probability distribution based on the semantic prediction result, and generating a semantic mask and confidence based on the pixel-level class probability distribution.
4. The method of claim 1, wherein, The method further comprises: performing time alignment on the radar effective point cloud and the semantic mask using a near-sequence frame extraction method; establishing a coordinate projection model, and performing spatial alignment on the radar effective point cloud and the semantic mask based on the coordinate projection model.
5. The method of claim 4, wherein, The cross-modal filtering fusion performed based on the radar effective point cloud, the semantic mask and the confidence to obtain multi-dimensional feature point cloud comprises: transmitting the radar effective point cloud to an image plane through the coordinate projection model to obtain a projection pixel coordinate corresponding to each point; judging whether the projection pixel coordinate falls within an effective area of the semantic mask, and deleting point cloud with confidence lower than a preset threshold within the effective area to obtain a point cloud projection area; performing semantic-guided point cloud filtering on the point cloud projection area to obtain multi-dimensional feature point cloud.
6. The method of claim 5, wherein, The semantic-guided point cloud filtering performed on the point cloud projection area to obtain multi-dimensional feature point cloud comprises: acquiring an image semantic class corresponding to the point cloud projection area, and assigning the image semantic class to a corresponding point cloud point to obtain a class label of each point cloud point, and comparing the class label of each point cloud point with a target class set to delete point cloud points not belonging to the target class set to obtain a class attribute of each point cloud point; estimating a ground plane equation through a plane fitting algorithm, and calculating a vertical distance of each point cloud point relative to the ground plane based on the ground plane equation to obtain a height attribute of each point cloud point; calculating a straight-line distance between each point cloud and an aircraft based on a projection pixel coordinate of each point cloud point to obtain a distance attribute of each point cloud point; generating multi-dimensional feature point cloud based on the class attribute, the height attribute and the distance attribute.
7. The method of claim 6, wherein, The construction of a threat assessment model based on the multi-dimensional feature point cloud, and the pre-warning for aircraft collision avoidance based on the threat assessment model comprises: In the category attribute, a distance threat function is constructed based on the distance attribute using an exponential decay function, and a height threat function is constructed based on the height attribute using a preset function; A weighted two-factor threat assessment model is constructed based on the distance threat function, a preset distance weight, the height threat function, and a preset height weight; A threat value of a threat target is calculated based on the weighted two-factor threat assessment model, a threat level is divided based on the threat value, and a pre-warning for aircraft collision avoidance is performed based on the threat level.
8. The method of claim 7, wherein, The method further includes: The threat targets are sorted in descending order of the threat level, and a structured threat target list and an optoelectronic image marking the threat targets are output in real time; each threat target includes a threat ID, a threat type, a threat value, and a relative height.
9. A multi-modal data fusion based aircraft collision warning device, characterized in that, The device includes: A multi-modal data acquisition module for acquiring multi-modal data, the multi-modal data including radar point cloud data and optoelectronic images; A preprocessing module for preprocessing the radar point cloud data and the optoelectronic images respectively to obtain radar effective point clouds, semantic masks, and confidence levels; A cross-modal filtering and fusion module for performing cross-modal filtering and fusion based on the radar effective point clouds, the semantic masks, and the confidence levels to obtain multi-dimensional feature point clouds; A threat assessment and collision avoidance pre-warning module for constructing a threat assessment model based on the multi-dimensional feature point clouds and performing a pre-warning for aircraft collision avoidance based on the threat assessment model.
10. A computer device, comprising: It includes: A memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the multi-modal data fusion aircraft collision pre-warning method of any one of claims 1 to 8.
Citation Information
Patent Citations
Air traffic collision avoidance method based on state prediction
CN106548661A
Civil aviation clearance safety risk assessment method and device, computer equipment and storage medium
CN113177719A
Three-dimensional target detection method based on monocular vision and radar pseudo image fusion
CN115082924A
Anti-collision radar track threat sorting method based on fuzzy set
CN116894608A
Method and device for dynamically sensing target threat area of unmanned aerial vehicle
CN119397141A
Cited By
Airborne synthetic visual dynamic threat intelligent identification system based on multi-source fusion
CN121325159A
A low-altitude flying object threat identification method and system based on dynamic density clustering
CN122469335A
A method and system for identifying low-altitude flying object threats based on dynamic density clustering
CN122469335B