Airport multi-mode sensing scene monitoring system and method based on 5GAeroMACS

Through the airport multimodal sensing scene monitoring system based on 5GAeroMACS, the information fusion is performed using signal transmission and reception nodes and deep learning networks to solve the limitations of the existing system in monitoring accuracy and data processing speed, and efficient and accurate airport scene monitoring is achieved, and safety and efficiency are improved.

CN120580894APending Publication Date: 2025-09-02BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510622506.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing airport scene monitoring system has limitations in monitoring accuracy, data processing speed and real-time performance, and it is difficult to meet the needs of large-scale data processing, affecting the safety management of airport scenes and air transportation efficiency.

Method used

The airport multimodal sensing scene monitoring system based on 5GAeroMACS is adopted, including signal transmission nodes, signal reception nodes, working mode switching modules, spectral estimation modules and fusion detection modules. 5G AeroMACS technology is used to integrate information with high bandwidth and low latency data transmission and deep learning networks to achieve efficient and accurate monitoring of airport scenes.

Benefits of technology

It improves the accuracy and accuracy of airport scene monitoring, ensures the real-time and completeness of data transmission, improves monitoring efficiency and security, and provides smart services for airport management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580894A_ABST
    Figure CN120580894A_ABST
Patent Text Reader

Abstract

The invention relates to an airport multi-modal sensing scene monitoring system and method based on 5GAeroMACS, and belongs to the technical field of airport scene monitoring, and the airport multi-modal sensing scene monitoring system and method provided by the invention realize efficient and accurate monitoring of an airport scene through an innovative technical means. The design of the system fully considers the whole process of data acquisition, transmission, processing and application, and aims to improve the monitoring capability and safety of the airport scene through an intelligent means. The airport multi-mode sensing scene monitoring system based on the 5GAeroMACS comprises a signal transmitting node, a signal receiving node, an airport scene communication sensing working mode switching module, a spectrogram estimation module and a fusion detection module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of airport scene monitoring, and in particular relates to an airport multimodal perception scene monitoring system and method based on 5G AeroMACS. Background Art

[0002] With the rapid development of the air transport industry, the use of radar, cameras, and sensors to monitor and manage airport safety (e.g., the movement of aircraft, vehicles, and personnel on runways and aprons) has become increasingly important. While traditional monitoring systems, such as radar, cameras, and sensors, provide essential monitoring data to a certain extent, they have significant limitations in terms of monitoring accuracy, data processing speed, and real-time performance. These limitations not only impact airport safety management but also restrict the efficiency and safety of air transport.

[0003] In the current technological landscape, emerging technologies such as 5G communications, artificial intelligence, and the Internet of Things (IoT) are bringing new opportunities to airport surface communication and perception. These advances offer the potential for higher-precision and more efficient monitoring, but they also present new challenges. For example, effectively integrating these technologies into existing airport monitoring systems, ensuring real-time transmission and processing of monitoring data, and improving target recognition accuracy are all pressing challenges.

[0004] Furthermore, the complexity of airport surface communication and perception systems is constantly increasing. As airports expand and the number of flights increases, the amount of data that monitoring systems must process increases dramatically, placing higher demands on data transmission and processing systems. Traditional monitoring systems often experience delays and errors when processing large amounts of data, which not only affects the accuracy of monitoring results but also increases safety management risks. Summary of the Invention

[0005] In response to these challenges, this paper proposes a multimodal airport surface monitoring system and method based on 5G AeroMACS (Aviation 5G Airport Surface Broadband Mobile Communication System). This system utilizes innovative technologies to achieve efficient and accurate monitoring of airport surfaces. The system's design fully considers the entire process of data collection, transmission, processing, and application, aiming to enhance airport surface monitoring capabilities and safety through intelligent means.

[0006] The present invention provides an airport multimodal perception scene monitoring system based on 5G AeroMACS, comprising a signal transmitting node, a signal receiving node, a working mode switching module for airport scene communication perception, a spectrum estimation module, and a fusion detection module;

[0007] A signal transmitting node, including a 5G AeroMACS base station and / or a 5G AeroMACS repeater station, is configured to transmit a 5G AeroMACS signal when performing perception spectrum estimation on a target object;

[0008] A signal receiving node, including a 5G AeroMACS base station, a 5G AeroMACS repeater station, and / or a 5G AeroMACS mobile station, is configured to receive a 5G AeroMACS signal when performing perception spectrum estimation on a target object;

[0009] A working mode switching module is used to select the working mode of airport surface communication perception according to the airport surface;

[0010] A spectrogram estimation module is used to generate a multi-dimensional perception spectrogram based on the 5G AeroMACS signals transmitted between the signal transmitting node, the signal receiving node and the target object;

[0011] The fusion detection module includes an airport scene target object fusion detection model, which is used to detect the distance, speed and angle values ​​of the target object from the perception spectrum, give the confidence of the target object, and determine the category of the target object.

[0012] Optionally, a 5G AeroMACS communication transmission module is further included, which is used to transmit the multi-dimensional perception spectrum of the 5G AeroMACS forwarding station and / or the 5G AeroMACS mobile station back to the 5G AeroMACS base station.

[0013] Optionally, a network training module is also included for training a multimodal perception airport scene target object fusion detection model using a loss function and a three-dimensional detection frame and a label frame of a target object obtained by the airport scene target object fusion detection model.

[0014] Optionally, the 5G AeroMACS base station and the 5G AeroMACS forwarding station are provided with antennas to receive signals through the antennas; the mobile station is a vehicle; and the 5G AeroMACS base station is a 5G AeroMACS radar.

[0015] Optionally, the working modes of airport surface communication perception include 5G AeroMACS base station perception mode, 5G AeroMACS forwarding station perception mode and 5G AeroMACS mobile station perception mode; in the 5G AeroMACS base station perception mode, the 5G AeroMACS base station transmits signals and serves as a signal transmitting node, the 5G AeroMACS base station serves as a signal receiving node, and the received signals realize perception of the target object; in the 5G AeroMACS forwarding station perception mode, the 5G AeroMACS base station transmits signals and serves as a signal transmitting node, the 5G AeroMACS forwarding station serves as a signal receiving node, and the received signals realize perception of the target object; in the 5G AeroMACS mobile station perception mode, the 5G AeroMACS base station or forwarding station transmits signals and serves as a signal transmitting node, the 5G AeroMACS mobile station serves as a signal receiving node, and the received signals realize perception of the target object.

[0016] Another aspect of the present invention provides an airport multimodal perception scene monitoring method based on 5G AeroMACS, which comprises the following steps:

[0017] Step 1: Under 5G AeroMACS communication, according to the working mode of airport surface communication perception, obtain the receiving signal delay at the receiving node based on the target object, the Doppler frequency shift of the signal between the signal transmitting node and the signal receiving node, and the receiving angle;

[0018] Step 2: According to the working mode of airport surface communication perception, the transmission signal of the signal transmitting node under 5G AeroMACS communication is obtained;

[0019] Step 3: According to the working mode of airport surface communication perception, based on the received signal delay at the receiving node in step 1, the Doppler frequency shift and receiving angle of the signal between the signal transmitting node and the signal receiving node, and the transmitted signal obtained in step 2, the received signal of the signal receiving node under 5G AeroMACS communication is obtained;

[0020] Step 4: According to the working mode of the airport surface communication perception, the transmission signal of the transmitting node in the received signal obtained in step 3 is processed to obtain a multi-dimensional perception spectrum;

[0021] Step 5: Construct an airport scene target object fusion detection model that includes a fusion detection network, a spatial channel attention model, a shared multi-layer perceptron, a spatial attention module, a module for generating a spatial attention map, and an attention mechanism module. Use the airport scene target object fusion detection model to fuse the multidimensional perception spectrum obtained in step 4 to obtain a three-dimensional detection frame of the target object.

[0022] Step 6: Based on the target object's label frame and the final 3D detection frame obtained in step 5, a fusion detection model for airport scene target objects based on multimodal perception is trained.

[0023] Step 7: Use the trained multimodal perception airport scene target object fusion detection model to obtain the target object perception results.

[0024] Optionally, the working modes of airport surface communication perception include 5G AeroMACS base station perception mode, 5G AeroMACS forwarding station perception mode and 5G AeroMACS mobile station perception mode; in 5G AeroMACS base station perception mode, the 5G AeroMACS base station transmits signals and serves as a signal transmitting node, the 5G AeroMACS base station serves as a signal receiving node, and the received signals realize perception of the target object; in 5G AeroMACS forwarding station perception mode, the 5G AeroMACS base station transmits signals and serves as a signal transmitting node, the 5G AeroMACS forwarding station serves as a signal receiving node, and the received signals realize perception of the target object; in 5G AeroMACS mobile station perception mode, the 5G AeroMACS base station or forwarding station transmits signals and serves as a signal transmitting node, the 5G AeroMACS mobile station serves as a signal receiving node, and the received signals realize perception of the target object.

[0025] Optionally, the multi-dimensional perception spectrum in step 4 includes a distance-speed-angle perception spectrum of the target object.

[0026] Optionally, the target object perception result in step 7 is a detection frame of the target object, including the target object position, category and confidence level.

[0027] The third aspect of the present invention provides a computer-readable storage medium, which stores computer instructions. When the computer reads the computer instructions in the storage medium, the computer executes the aforementioned airport multimodal perception scene monitoring method based on 5G AeroMACS.

[0028] Compared with the prior art, the present invention has at least the following beneficial effects:

[0029] (1) The airport multimodal perception scene monitoring system of the present invention performs perception in three working modes for 5G AeroMACS base stations, 5G AeroMACS forwarding stations and 5G AeroMACS mobile stations, and extracts three-dimensional information of the distance, speed and angle of airport scene targets from the target reflection echoes. At the same time, based on the deep learning network enhanced by the point product module and the space-channel attention module, it simultaneously detects and classifies three types of targets: aircraft, vehicles and people in three-dimensional space, realizing multimodal monitoring of multi-base station mode collaborative perception, multi-type target classification and multi-dimensional space detection, thereby improving the accuracy and precision of airport scene monitoring.

[0030] (2) The airport multimodal sensing scene monitoring system of the present invention is also specially designed with single-station active sensing mode, dual-station active sensing mode and dual-station passive sensing mode to adapt to different monitoring needs and environmental conditions.

[0031] (3) The present invention’s airport multimodal sensing scene surveillance system utilizes the advanced 5G AeroMACS technology’s high-bandwidth and low-latency data transmission capabilities, as well as its extensive connectivity, enabling it to connect a large number of sensors and devices. This enables the system to achieve extensive coverage and detailed monitoring of the airport scene, while ensuring the real-time and integrity of data transmission.

[0032] (4) The airport multimodal sensing scene surveillance system of the present invention uses multi-antenna reception technology and Doppler frequency shift technology to accurately capture the reflected signals of target objects. These signals are transmitted through 5G AeroMACS technology, ensuring that the data reaches the processing center quickly and accurately.

[0033] (5) The present invention's multimodal airport perception scene monitoring method employs a convolutional neural network based on a channel attention mechanism to fuse information from multidimensional perception spectra. This method can simultaneously determine the distance, speed, and angle of three types of targets at the airport: aircraft, vehicles, and people. It can effectively extract key features from the perception data to achieve accurate identification and location of target objects. The present invention's deep learning model also employs a multi-task learning framework, integrating multiple related tasks to enhance the model's generalization and robustness.

[0034] (6) The application layer of the present invention's airport multimodal perception scene surveillance system converts processed data into actual monitoring results and decision-making support information, providing diversified and customized intelligent services to airport management departments, airlines, airports, and other stakeholders. This intelligent airport scene communication perception system not only improves the accuracy and efficiency of monitoring but also provides strong technical support for airport safety management, with significant social and economic value. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Schematic diagram of the working mode of the airport multimodal perception scene monitoring system based on 5G AeroMACS of the present invention.

[0036] Figure 2 It is a schematic diagram of 5G AeroMACS signal transmission and reception in the airport multimodal perception scene monitoring system based on 5G AeroMACS of the present invention.

[0037] Figure 3 A multi-level spectrum graph processing framework diagram based on a channel attention mechanism convolutional neural network in the 5GAeroMACS-based airport multimodal perception scene monitoring system of the present invention.

[0038] Figure 4 Schematic diagram of neural network labels in the airport multimodal perception scene monitoring system based on 5G AeroMACS of the present invention.

[0039] Figure 5 Schematic diagram of the backbone network in the airport multimodal perception scene monitoring system based on 5G AeroMACS of the present invention.

[0040] Figure 6 Schematic diagram of the neck network in the airport multimodal perception scene monitoring system based on 5G AeroMACS of the present invention.

[0041] Figure 7 Schematic diagram of the space-channel attention module in the airport multimodal perception scene monitoring system based on 5G AeroMACS of the present invention.

[0042] Figure 8 It is a schematic diagram of the distance-speed profile of the three-dimensional detection target in the airport multimodal perception scene monitoring method based on 5G AeroMACS of the present invention.

[0043] Figure 9 It is a schematic diagram of the distance-angle profile of the three-dimensional detection target in the airport multimodal perception scene monitoring method based on 5G AeroMACS of the present invention. DETAILED DESCRIPTION

[0044] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. In addition, the present invention can also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited by the specific embodiments disclosed below.

[0045] A specific embodiment of the present invention, as Figures 1-9, discloses an airport multimodal perception scene monitoring system based on 5GAeroMACS, including a signal transmitting node, a signal receiving node, a working mode switching module for airport scene communication perception, a spectrum estimation module and a fusion detection module;

[0046] A signal transmitting node, including a 5G AeroMACS base station and / or a 5G AeroMACS repeater station, is configured to transmit a 5G AeroMACS signal when performing perception spectrum estimation on a target object;

[0047] A signal receiving node, including a 5G AeroMACS base station, a 5G AeroMACS repeater station, and / or a 5G AeroMACS mobile station, is configured to receive a 5G AeroMACS signal when performing perception spectrum estimation on a target object;

[0048] A working mode switching module is used to select the working mode of airport surface communication perception according to the airport surface;

[0049] A spectrogram estimation module is used to generate a multi-dimensional perception spectrogram based on the 5G AeroMACS signals transmitted between the signal transmitting node, the signal receiving node and the target object;

[0050] The fusion detection module is embedded with an airport scene target object fusion detection model, which is used to simultaneously detect the distance, speed and angle values ​​of the target object from the perception spectrum, provide the confidence level of the target object, and determine the category of the target object;

[0051] Furthermore, the categories of target objects are aircraft, vehicles and people; the measured distance of the target object is the sum of the distances between the target object and the signal transmitting node and the signal receiving node; the transmitting and receiving signals of the present invention are AeroMACS signals, which are used to provide data communication services for air traffic control, airport operation command, airline operation management and services, and on-site unit operation management.

[0052] It can be understood that the spectrum is a distance-speed-angle spectrum of the target, which is a three-dimensional feature map.

[0053] The 5G AeroMACS communication transmission module is used to transmit the multi-dimensional sensing spectrum of the 5G AeroMACS forwarding station and / or 5G AeroMACS mobile station back to the 5G AeroMACS base station.

[0054] Furthermore, it also includes a network training module for training a multimodal perception airport scene target object fusion detection model through a loss function, a final three-dimensional detection frame and a label frame of the target object.

[0055] Furthermore, the 5G AeroMACS base station and the 5G AeroMACS forwarding station are equipped with antennas to receive signals through the antennas; the mobile station is a vehicle; and the 5G AeroMACS base station is a 5G AeroMACS radar.

[0056] Furthermore, the working modes of airport surface communication perception of the airport multimodal perception surface surveillance system based on 5G AeroMACS include 5G AeroMACS base station perception mode, 5G AeroMACS forwarding station perception mode and 5G AeroMACS mobile station perception mode; in the 5G AeroMACS base station perception mode, the 5G AeroMACS base station transmits signals and serves as a signal transmitting node, while the 5G AeroMACS base station serves as a signal receiving node, and receives signals to realize perception of the target object; in the 5G AeroMACS forwarding station perception mode, the 5G AeroMACS base station transmits signals and serves as a signal transmitting node, while the 5G AeroMACS forwarding station serves as a signal receiving node, and receives signals to realize perception of the target object; in the 5G AeroMACS mobile station perception mode, the 5G AeroMACS base station or forwarding station transmits signals and serves as a signal transmitting node, while the 5G AeroMACS mobile station serves as a signal receiving node, and receives signals to realize perception of the target object.

[0057] Another specific embodiment of the present invention further discloses a method for monitoring an airport multimodal perception scene based on 5G AeroMACS. Using the aforementioned airport multimodal perception scene monitoring system based on 5G AeroMACS, the specific steps are as follows:

[0058] Step 1: Under 5G AeroMACS communication, according to the working mode of airport surface communication perception, obtain the receiving signal delay at the receiving node based on the target object, the Doppler frequency shift of the signal between the signal transmitting node and the signal receiving node, and the receiving angle.

[0059] Furthermore, (1) when the working mode of airport surface communication perception is 5G AeroMACS base station perception mode, the 5G AeroMACS base station is a signal transmitting node and a signal receiving node, which is used to transmit signals and receive echo signals through reflection from target objects. In this perception working mode, since the signal transmitting node and the signal receiving node are both performed at the same node, the signal receiving node can obtain complete transmission waveform information.

[0060] Specifically, in the 5G AeroMACS base station sensing mode, the time delay of receiving the signal at the signal receiving node and the Doppler frequency shift of the signal between the signal transmitting node and the signal receiving node are expressed as follows:

[0061] τ k,basestation=2τ 1,k,basestation ,

[0062] f d,k,basestation =2f d,1,k,basestation

[0063]

[0064] Among them, τ k,basestation represents the time delay of the signal receiving node receiving the reflected signal from the kth target object, k = 1, 2, ..., K, K represents the total number of target objects on the airport surface; τ 1,k,basestation represents the time it takes for the signal transmitting node to transmit the signal to the kth target object; f d,k,basestation represents the Doppler frequency shift of the signal passing through the kth target object and back to the signal transmitting node and the signal receiving node; f d,1,k,basestation represents the Doppler frequency shift of the signal from the signal transmitting node to the kth target object; R 1,k,basestation is the distance from the signal transmitting node to the kth target; c is the speed of light; θ 1,k,basestation is the angle between the moving direction of the kth target object and the direction of the signal transmitted by the signal transmitting node; f c is the carrier frequency; v k,basestation represents the velocity of the kth target object relative to the 5G AeroMACS base station.

[0065] The receiving angle is the angle of the target object relative to the 5G AeroMACS base station.

[0066] Specifically, (2) when the working mode of airport surface communication perception is the 5G AeroMACS forwarding station perception mode, the 5G AeroMACS base station is the signal transmitting node, and the 5G AeroMACS forwarding station is the signal receiving node. The transmitting node and the receiving node are separate and different devices. The receiving node receives the signal sent by the transmitting node and demodulates, decodes, and reconstructs it for perception of the spatial area to be perceived.

[0067] Specifically, in the 5G AeroMACS forwarding station sensing mode, the delay of receiving the signal at the signal receiving node and the Doppler shift of the signal between the transmitting node and the receiving node are expressed as follows:

[0068] τ k,relaystation =τ 1,k,relaystation +τ 2,k,relaystation ;

[0069] f d,k,relaystation =f d,1,k,relaystation +f d,2,k,relaystation ;

[0070]

[0071] Among them, τ k,relaystation represents the time delay of the signal receiving node receiving the reflected signal from the kth target object; τ 1,k,relaystation represents the time it takes for the signal transmitting node to propagate the signal to the kth target object; τ 2,k,relaystation represents the time it takes for the kth target object to transmit the signal to the signal receiving node; f d,k,relaystation represents the Doppler frequency shift of the signal passing through the kth target object and returning to the signal transmitting node and the signal receiving node; f d,1,k,relaystation represents the Doppler frequency shift of the signal from the signal transmitting node to the kth target object; R 1,k,relaystation is the distance from the signal transmitting node to the kth target object; R 2,k,relaystation is the distance from the signal receiving node to the kth target object; c is the speed of light;

[0072] θ 1,k,relaystation is the angle between the moving direction of the kth target object and the direction of the signal transmitted by the signal transmitting node; θ 2,k,relaystation is the angle between the moving direction of the kth target object and the receiving signal direction of the signal receiving node; f c is the carrier frequency; v k,relaystation represents the speed of the kth target object relative to the 5GAeroMACS forwarding station.

[0073] The receiving angle is the angle of the target object relative to the 5G AeroMACS repeater station.

[0074] Furthermore, (3) when the working mode of airport surface communication perception is the 5G AeroMACS mobile station perception mode, the 5G AeroMACS base station or 5G AeroMACS forwarding station is the signal transmitting node, and the rotating station is the signal receiving node. The signal transmitting node and the signal receiving node are separate and distinct devices. The signal receiving node has two channels: a reference channel and a monitoring channel. The perception of the spatial area is achieved by comparing the signal correlation of these two channels.

[0075] Specifically, in the 5G AeroMACS mobile station sensing mode, the time delay of receiving the signal at the signal receiving node and the Doppler shift of the signal between the transmitting node and the receiving node are expressed as follows:

[0076] τ k,movestation =τ 1,k,movestation +τ 2,k,movestation ;

[0077] f d,k,movestation =f d,1,k,movestation +f d,2,k,movestation ;

[0078]

[0079] Among them, τ k,movestation represents the time delay of the signal receiving node receiving the reflected signal from the kth target object; τ 1,k,movestation represents the time it takes for the signal transmitting node to propagate the signal to the kth target object; τ 2,k,movestation represents the time it takes for the kth target object to transmit the signal to the signal receiving node; f d,k,movestation represents the Doppler frequency shift of the signal passing through the kth target object and returning to the signal transmitting node and the signal receiving node; f d,1,k,movestation represents the Doppler frequency shift of the signal from the signal transmitting node to the kth target object; R 1,k,movestation is the distance from the signal transmitting node to the kth target object;

[0080] R 2,k,movestation is the distance from the signal receiving node to the kth target object; c is the speed of light;

[0081] θ 1,k,movestation is the angle between the moving direction of the kth target object and the direction of the signal transmitted by the signal transmitting node; θ 2,k,movestation is the angle between the moving direction of the kth target object and the receiving signal direction of the signal receiving node; f c is the carrier frequency.

[0082] The receiving angle is the angle of the target object relative to the sensing pattern of the 5G AeroMACS mobile station.

[0083] Step 2: According to the working mode of airport surface communication perception, obtain the transmission signal of the signal transmitting node under 5G AeroMACS communication, and the expression is:

[0084]

[0085] Where x(t) represents the transmitted signal of the signal transmitting node at time t; N sym represents the number of symbols of the signal transmitted by the signal transmitting node; μ represents the μth symbol of the signal transmitted by the signal transmitting node; a(.) represents the modulation data of 5GAeroMACS; j represents an imaginary number; π represents the ratio of pi; T OFDM represents the symbol interval of the Orthogonal Frequency Shift Keying (OFDM) symbol of the 5G AeroMACS signal; f n Indicates the frequency of the nth signal subcarrier; N c Indicates the total number of signal subcarriers; exp(.) indicates exponential calculation; rect(.) indicates the time rectangular window function.

[0086] Step 3: According to the working mode of airport surface communication perception, based on the received signal delay at the receiving node in step 1, the Doppler frequency shift and receiving angle of the signal between the signal transmitting node and the signal receiving node, and the transmitted signal obtained in step 2, the received signal of the signal receiving node under 5G AeroMACS communication is obtained.

[0087] Furthermore, in 5G AeroMACS base station sensing mode, R k =cτ k,basestation ;5GAeroMACS forwarding station perception mode, R k =cτ k,relaystation ; In 5G AeroMACS mobile station perception mode, R k =cτ k,movestation ; R k represents the distance between a single antenna of the signal receiving node and the kth target object. Furthermore, in the 5G AeroMACS base station sensing mode, In 5G AeroMACS forwarding station perception mode, In 5G AeroMACS mobile station perception mode, Where λ represents the wavelength of the signal, v k represents the velocity of the kth target object.

[0088] Furthermore, multiple antennas are provided on the 5G AeroMACS base station, 5G AeroMACS forwarding station and mobile station.

[0089] Specifically, for a single antenna on a signal receiving node, the received signal y k (t) is expressed as:

[0090]

[0091] Among them, y k (t) represents the received signal of a single antenna of the k-th target object at time t; R k represents the distance between a single antenna of the signal receiving node and the kth target object; v k represents the speed of the kth target object; Δf represents the subcarrier spacing.

[0092] Specifically, for multiple antennas on a signal receiving node, the received signal expression of the multiple antennas is:

[0093]

[0094] Where y(t) represents the multi-antenna array receiving signal of the signal receiving node at time t; a(θ k ) represents the steering vector of the multi-antenna array; y k(t) represents the received signal of a single antenna of the k-th target object at time t; θ k Represents the angular direction of the kth target object relative to the receiving node, k = 1, 2, ..., K, and K represents the total number of target objects.

[0095] Furthermore, the steering vector a(θ k ) is:

[0096]

[0097] Wherein, e represents a natural constant; d represents the spacing between adjacent antennas in the multi-antenna array; and M represents the total number of antennas in the multi-antenna array.

[0098] Furthermore, the received signal y of a single antenna k (t) is sampled to obtain the received signal matrix D at time t Rx , the received signal matrix D at time t Rx It consists of the nth subcarrier signal receiving signal of the μth symbol of the signal receiving node.

[0099] The expression of the nth subcarrier signal received by the signal receiving node of the μth symbol is:

[0100]

[0101]

[0102] Among them, (D Rx ) μ,n represents the nth subcarrier signal receiving signal of μth symbol at the signal receiving node; (D Tx ) μ,n represents the transmitted signal of the nth subcarrier of the μth symbol at the receiving node; A(μ,n) represents the complex amplitude of the reflected signal of the nth subcarrier of the μth symbol at the signal receiving node; Represents the distance dimension steering vector between the signal receiving node and the kth target object; represents the velocity-dimensional steering vector of the k-th target.

[0103] Step 4: According to the working mode of airport surface communication perception, the transmission signal of the transmitting node in the receiving signal obtained in step 3 is processed to obtain a multi-dimensional perception spectrum.

[0104] Preferably, the multi-dimensional perception spectrum includes a distance-speed-angle perception spectrum of the target object.

[0105] Specifically, (1) in the 5G AeroMACS base station sensing mode, the transmission signal is eliminated by element-wise division to obtain the 5G AeroMACS sensing matrix D after the transmission signal is eliminated.div , 5G AeroMACS perception matrix D after eliminating the transmitted signal div It is composed of 5GAeroMACS perception of the nth subcarrier signal of the μth symbol after removing the transmitted signal, and the expression is:

[0106]

[0107] Among them, (D div ) μ,n represents the 5G AeroMACS perception of the n-th subcarrier signal of the μ-th symbol after the transmission signal is eliminated; represents the Kronecker product.

[0108] Estimate the target object’s distance-speed perceptual spectrum Ω based on inverse Fourier transform and Fourier transform:

[0109] Ω=G H D div F

[0110] Where G and F represent Fourier matrices of different dimensions; H represents the complex conjugate transpose;

[0111]

[0112] Among them, N sym Indicates the number of symbols of the signal transmitted by the signal transmitting node, N c Indicates the total number of signal subcarriers;

[0113] Specifically, (2) in the 5G AeroMACS forwarding station sensing mode, the signal receiving node receives the signal sent by the signal transmitting node and demodulates, decodes and reconstructs it to obtain the receiving signal matrix based on the single antenna on the signal receiving node. The received signal matrix D at time t Tx It is used to eliminate the transmission signal in the single-station active sensing mode and obtain the 5G AeroMACS sensing matrix D after eliminating the transmission signal. dib , 5G AeroMACS perception matrix D after eliminating the transmission signal div It consists of the 5G AeroMACS perception of the nth subcarrier signal of the μth symbol after the transmission signal is removed, and is expressed as:

[0114]

[0115] Among them, (D div ) μ,n It represents the 5G AeroMACS perception of the n-th subcarrier signal of the μ-th symbol after the transmission signal is eliminated.

[0116] The perception spectrum Ω of the target object's range-velocity is estimated based on the inverse Fourier transform and the Fourier transform. The expression is the same as that in the 5G AeroMACS base station perception mode and will not be repeated here.

[0117] (3) In the 5G AeroMACS mobile station sensing mode, the transmitted signal is processed by the mutual ambiguity function of the reference channel and the monitoring channel of the signal receiving node to obtain the range-velocity spectrum Ω of the target object. The range-velocity spectrum Ω of the target object is composed of the range-velocity spectrum of the n-th subcarrier signal of the μ-th symbol Ω(μ,n), which is expressed as:

[0118]

[0119] Among them, y(t) represents the signal of the monitoring channel at time t, Represents the signal of the reference channel at time t.

[0120] Furthermore, in the three working modes of airport surface communication perception, multi-antenna reception is used to realize the angle estimation of the target object, and the angle spectrum P of the target object is obtained through Fourier transform. DFT (θ), the expression is:

[0121] P DFT (θ)=|a H (θ)y(t)| 2 (12)

[0122] Where θ is the discrete value of the angle of the target object relative to the receiving node; a H (θ) represents the complex conjugate transpose of the array steering vector of the multi-antenna array;

[0123] In this invention, 5G AeroMACS (Aviation 5G Airport Surface Broadband Mobile Communication System) is simultaneously used to achieve efficient transmission and processing of airport surface awareness data. 5G AeroMACS also supports extensive connectivity, enabling the connection of a large number of sensors and devices. This is a significant advantage for airport surface awareness systems, as it allows the integration of more sensors and monitoring devices, achieving wider coverage and more detailed monitoring. This extensive connectivity also enables the system to be flexibly expanded to adapt to changing airport surface awareness needs.

[0124] The high bandwidth and reliability of 5G AeroMACS allow for the rapid transmission of large amounts of data. In airport surface awareness systems, this means real-time transmission of high-resolution spectrograms and extensive sensor data without data loss or delays due to bandwidth limitations. This high-bandwidth capability is crucial for handling complex spectrum analysis and data fusion tasks, ensuring data integrity and real-time performance. The spectrograms sensed by each sensing station are quantized, compressed, and encoded into dedicated sensing data packets. Advanced network management and error control mechanisms reduce errors and loss during data transmission, and the sensing information from each station is aggregated to a central processing base station using 5G channels.

[0125] Furthermore, the low latency of 5G AeroMACS is crucial for real-time data processing. Rapid response is crucial for airport scene perception. 5G AeroMACS can achieve millisecond-level latency, making the entire process from data acquisition to processing and decision support faster and more efficient. This is crucial for rapid response and decision-making in emergency situations. An adaptive scheduling algorithm coordinates network access and routing, ensuring timely transmission of perception data packets.

[0126] Step 5: Figure 3 As shown in the figure, an airport scene target object fusion detection model is constructed, which includes a fusion detection network, a spatial channel attention model, a shared multi-layer perceptron, a spatial attention (SA) module, a spatial attention map generation and an attention mechanism module. The airport scene target object fusion detection model is used to fuse the multi-dimensional perception spectrum obtained in step 4 to obtain a three-dimensional detection frame of the target object.

[0127] Step 51: normalize the multi-dimensional perceptual spectrum obtained in step 4 to obtain a normalized spectrum. The specific steps are: subtract the arithmetic mean from the multi-dimensional perceptual spectrum and divide it by its standard deviation to obtain the normalized spectrum.

[0128] It can be understood that the multidimensional perception spectrum includes the distance-speed spectrum of the target object (which can be disassembled into a distance spectrum and a speed spectrum) and an angle spectrum, and can be stacked into a multidimensional perception spectrum.

[0129] Specifically, the multi-dimensional perception spectrum is obtained by combining the distance spectrum and the speed spectrum and superimposing them along the angle dimension of the angle spectrum.

[0130] Step 52: Process the standardized atlas based on the fusion detection network to obtain a three-dimensional detection frame of the scene target.

[0131] Specifically, the fusion detection network includes a 3D backbone network, a 3D neck network and a 3D detection head.

[0132] Step 521: Input the standardized atlas into the three-dimensional backbone network of the fusion detection network to obtain the aggregated feature map output by the three-dimensional backbone network.

[0133] Specifically, in the forward propagation of the 3D backbone network, the generated feature map is Figure 5 The data is transmitted in the direction of the arrow shown to obtain the aggregated feature map output by the three-dimensional backbone network.

[0134] Specifically, if Figure 5 As shown, the 3D backbone network consists of network modules 0 through 9, which are used to extract high-level features from the input normalized graph and reduce redundant information in the normalized graph to facilitate in-depth processing by subsequent networks. The 10 network modules in the 3D backbone network have sequential inputs and outputs. For example, the output of network module 0 is the input of network module 1, and the output of network module 1 is the input of network module 2. Furthermore, network modules 5, 7, and 9 have duplicate output ports, corresponding to output ports 15, 12, and 10, respectively.

[0135] Preferably, in the third network module, the three-dimensional backbone network is composed of alternating PointMultiply and RepNACPELAN4 network structures. The present invention uses the PointMultiply module for initial feature extraction and the minimum feature map in the three-dimensional backbone to improve the ability to capture global and small target information.

[0136] Furthermore, feature enhancement is achieved using a point-wise multiplication network. This operation can map the input into a high-dimensional nonlinear feature space without increasing the network width. Within a single layer of the neural network structure of the fusion detection network, high-dimensional mapping of features is achieved through element-by-element multiplication.

[0137] By employing this approach, the present invention enhances the model's representational capabilities without changing the network width. This is because element-by-element multiplication essentially creates all possible interactions between features, increasing the feature dimensionality without increasing the number of parameters. Stacking multiple such layers exponentially increases the implicit dimensionality of features, theoretically achieving an infinite-dimensional feature space. This provides the model with powerful capabilities to learn and represent complex data patterns.

[0138] Preferably, RepNACPELAN4 in the 5th and 7th network modules includes a CBS module, a RepNBottleNeck module and a RepNCSP module.

[0139] Preferably, the CBS module includes two-dimensional convolution, two-dimensional batch normalization and SiLU activation function modules.

[0140] Preferably, the RepNBottleneck module includes a RepNBottleneck[CV1] convolutional layer and a RepNBottleneck[CV2] convolutional layer, and the RepNBottleneck[CV1] convolutional layer and

[0141] The RepNBottleneck[CV2] convolutional layers are set up in parallel; the RepNBottleneck[CV1] convolutional layer includes 1 CBS module, and the RepNBottleneck[CV2] convolutional layer includes 2 CBS modules. The outputs of the RepNBottleneck[CV1] convolutional layer and the RepNBottleneck[CV2] convolutional layer are added to obtain the output feature map of the RepNBottleneck module.

[0142] Optionally, the RepNCSP module is a sequential structure model including multiple submodules, each of which includes a CBS module and a RepNBottleneck module. In each submodule, the features of the feature maps obtained by the CBS module and the RepNBottleneck module are concatenated to obtain aggregated features, and the CBS module is used to perform further feature extraction to obtain an aggregated feature map.

[0143] Furthermore, the forward propagation expression of the RepNCSP module is:

[0144] f RepNCSP (x) = f(f(x))

[0145] Among them, f RepNCSP (x) represents the output of the RepNCSP model; f(x) represents the sub-module of RepNCSP; x represents the feature map obtained by the CBS module and RepNBottleneck module input to the RepNCSP module.

[0146] Furthermore, f(x)=CVout,3(torch.cat((m(CVout,1(x)),CVout,2(x)),dim=1))

[0147] Among them, CVout,3(x) represents the final convolutional layer of the RepNCSP module; torch.cat(.) represents the three-dimensional feature map splicing operation; m(.) represents the feature map processing using the RepNBottleneck module; CVout,1(x) represents the first CBS module; CVout,2(x) represents the output of the second CBS module; dim represents the first dimension along the three-dimensional feature map.

[0148] Finally, the RepNCSPELAN4 module is a superposition of the RepNCSP module and the CBS module, which enhances the feature extraction and integration mechanism. It first consists of a top-level convolutional layer CV up,1 (x), the initial feature extraction is performed on the input tensor. Subsequently, the feature map undergoes a Chunk operation, which bifurcates the channels and guides them into the processing paths of the subsequent convolutional layers CVup,2(x) and CVup,3(x), each of which is a component of RepNCSP. After the feature concatenation, the final CV up,4 (x) Further extract features to obtain the output feature map of the RepNCSPELAN4 module.

[0149] Step 522: Use the 3D neck network of the fusion detection network to process the aggregated feature map output by the 3D backbone network to obtain the output F″ of the spatial attention mechanism module.

[0150] Furthermore, the 3D neck network uses a feature pyramid network (FPN) to perform secondary processing on the high-level features (aggregated feature maps) extracted by the 3D backbone network to fully integrate features of different scales, thereby enhancing the multi-scale detection performance of the detector.

[0151] Further, if Figure 6 As shown, the three-dimensional neck network includes network modules from 10 to 25. The input of the 10th network module 10 is the output of the replica port of the 9th network module 9, the input of the 11th network module 11 is the output of the network module 10, the input of the 12th network module 12 is the output of the 7th network module 7 and the output of the 11th network module 11, the input of the 13th network module 13 is the output of the 12th network module 12, the input of the 14th network module 14 is the output of the 13th network module 13, the input of the 15th network module 15 is the output of the 15th network module 5 and the output of the 14th network module 14, the input of the 16th network module 16 is the output of the 15th network module 15, the input of the 17th network module 17 is the output of the 16th network module 16, and the input of the 18th network module 18 is the output of the 19th network module 20. The input of network module 18 is the output of the 17th network module 17, the 19th network module 19 is spliced ​​with the output of the 13th network module 13 and the output of the 18th network module 18, the input of the 20th network module 20 is the output of the 19th network module 19, the input of the 21st network module 21 is the output of the 20th network module 20, the input of the 22nd network module 22 is the output of the 21st network module 21, the 23rd network module 23 is spliced ​​with the output of the 10th network module 10 and the output of the 22nd network module 22, the input of the 24th network module 24 is the output of the 23rd network module 23, and the input of the 25th network module 25 is the output of the 24th network module 24.

[0152] Preferably, in the present invention, the structures of the 16th, 20th and 24th network modules introduce a spatial channel attention (SCAM) module composed of a spatial attention module and a channel attention module. The SCAM module is designed to enhance the model's ability to capture key information through fine feature reconstruction, thereby improving the accuracy of target object detection. Three spatial channel attention modules SCAMBlocks are added before the RepNACPELAN4 blocks close to the three detection heads. By utilizing the channel and spatial attention mechanism of SCAMBlocks, the three-dimensional neck network can better distinguish and focus on the most informative features, thereby improving the accuracy and robustness of the target object detection task.

[0153] Specifically, the structure of the spatial channel attention module is designed to aggregate the input feature map Input spatial channel attention module, use average pooling AvgPool and maximum pooling MaxPool to process the compressed spatial dimension of the aggregated feature map F, and extract the channel average value and channel maximum The expression is:

[0154]

[0155] Average the channels and channel maximum Pass it to the shared multi-layer perceptron (MLP) to obtain the feature map F′ after the channel attention operation. The specific steps are:

[0156] First, the channel average and channel maximum The channel attention map M is formed by element-wise summing and merging c (F), the expression is:

[0157]

[0158] Among them, σ represents the sigmoid activation function, which ensures that the value of the channel attention map is between 0 and 1 and adjusts the channels of the input feature map.

[0159] Furthermore, a multilayer perceptron (MLP) consists of hidden layers.

[0160] Then, the channel attention mechanism is implemented by element-by-element multiplication to obtain the feature map F′ after the channel attention operation, which is expressed as:

[0161] F′=M c (F)e F

[0162] Among them, F represents the aggregated feature map of the input; M c(F) represents the channel attention map; F′ represents the feature map after the channel attention operation.

[0163] Furthermore, based on the feature map F′ after the channel attention operation, the average spatial attention feature map is obtained by performing average pooling and maximum pooling operations on the channel axis in the spatial attention (SA) module. and the maximum spatial attention feature map Concatenated average spatial attention feature map and the maximum spatial attention feature map Get a valid spatial feature descriptor F s .

[0164] Furthermore,

[0165] Furthermore, the average spatial attention feature map is concatenated and the maximum spatial attention feature map Get a valid spatial feature descriptor F s The expression is:

[0166]

[0167] Furthermore, a 7×7 convolutional layer f is used. 7×7 Generate spatial attention map M s Processing effective spatial feature descriptors F s Get the spatial attention map M s (F′), the expression is:

[0168] M s (F′)=σ(f 7×7 (F s )).

[0169] Furthermore, based on the feature map F′ after the channel attention operation and the spatial attention map M s (F′), the spatial attention operation is performed through the attention mechanism module to obtain the output of the channel attention mechanism module, which is expressed as;

[0170] F″=M s (F′)e F′

[0171] Among them, F′ represents the output of the channel attention module; F″ represents the output of the spatial attention mechanism module; M s (F′) represents the spatial attention map.

[0172] Step 523: Use the 3D detection head of the fusion detection network to extract the category and spatial position of the target object from the output of the 3D neck network, and output the final 3D detection frame.

[0173] Furthermore, the label box of each target object is represented by a 10-parameter vector. The final 3D detection box includes 6 parameters of the target object position, 3 parameters describing the target object category, and 1 parameter describing the target object confidence.

[0174] It can be understood that the dimensions of the detection box and the label box are the same. The detection box is the predicted value of the target object, and the label box is the true value of the target object, which is used for network training.

[0175] Specifically, the six parameters describing the position of the target object include target average distance, target distance dimension expansion width, target average speed, target speed dimension expansion width, target average angle, and target angle dimension expansion width. Figure 8 and Figure 9 In the experiment, a total of 3 targets were detected, with average target distances of 200m, 600m, and 600m, respectively; target distance dimension expansion widths of 64m, 46m, and 75m, respectively; target average speeds of 5m / s, 20m / s, and -15m / s, respectively; target speed dimension expansion widths of 7m / s, 6m / s, and 8m / s, respectively; target average angles of 30°, 0°, and -45°, respectively; and target angle dimension expansion widths of 23°, 22°, and 26°, respectively.

[0176] Specifically, target object categories include airplanes, vehicles, and people. For example, the parameter [1, 0, 0] is used to represent the airplane category, the parameter [0, 1, 0] is used to represent the vehicle category, and the parameter [0, 0, 1] is used to represent the person category.

[0177] Furthermore, the target confidence range is 0-1. The closer the value is to 1, the higher the detection accuracy of the target. Figure 8 and Figure 9 The confidence levels are 0.92, 0.81, and 0.86 respectively.

[0178] Specifically, for the category information of the target object, the softmax function is used to convert the predicted values ​​of the three target object categories into probability values ​​whose sum is 1, and the target object category with the largest probability value is used as the predicted category of the target object. The expression is:

[0179]

[0180] Among them, l represents the category number, y l represents the predicted category of the target object; x l Represents the target object category prediction value output by the detection head network; Represents the probability value of the lth target object category based on the natural constant e.

[0181] For example, the target object and label box on the multi-dimensional perception spectrum are shown in Figure 4 The label box and detection box are represented by solid lines and dotted lines respectively. Figure 4 In the figure, the width and height of the target object's range-velocity spectrum are normalized, and all line segments are scaled proportionally.

[0182] In step 6, based on the target object's label frame and the final 3D detection frame obtained in step 5, a fusion detection model for airport scene target objects based on multimodal perception is trained.

[0183] First, the initialization network parameters of the airport scene target object fusion detection model based on multimodal perception are obtained.

[0184] Secondly, the final 3D detection frame is obtained by forward propagation through the airport scene target object fusion detection model based on multimodal perception;

[0185] Then, the loss between the final 3D detection frame and the label frame of the target object (ie, the real label) is obtained through the loss function, which includes classification loss, positioning loss and confidence loss.

[0186] Then, the back propagation algorithm is used to obtain the network parameter gradient, and the verification set is used to determine whether the airport scene target object fusion detection model based on multimodal perception has reached performance convergence. If the performance converges, the trained airport scene target object fusion detection model based on multimodal perception is obtained. If the performance convergence is not reached, the initial weights and learning rate during network training are modified.

[0187] Furthermore, the classification loss describes the category of the target object on the airport scene, and the category loss L class The expression is:

[0188]

[0189] Among them, l represents the category number, y l Represents the predicted category of the target object;

[0190] Furthermore, the target object position loss L local The expression is:

[0191]

[0192] Among them, w R represents the distance weight of the target object, R represents the distance value output by the airport scene target object fusion detection model based on multimodal perception, Represents the target distance label value, w δrepresents the distance dimension expansion width weight, δ represents the output distance dimension expansion width of the airport scene target object fusion detection model based on multimodal perception, Represents the target object distance width label value, w V represents the speed weight of the target object, V represents the output speed value of the airport scene target object fusion detection model based on multimodal perception, Represents the target speed label value, w υ represents the speed width weight, υ represents the network output speed dimension expansion width, Represents the target speed width label value, w A represents the angle weight of the target object, A represents the output angle value of the airport scene target object fusion detection model based on multimodal perception, Represents the target angle label value, w α represents the angle width weight, α represents the output angle dimension expansion width of the airport scene target object fusion detection model based on multimodal perception, Represents the target angle width label value.

[0193] Furthermore, the target object confidence loss L gr The expression is:

[0194] L gr =1-p gr

[0195] Among them, p gr Output confidence for the target.

[0196] Step 7: Use the trained multimodal perception airport scene target object fusion detection model to obtain the target object perception results.

[0197] Specifically, the target object perception result is the detection frame of the target object, including the target object position, category and confidence.

[0198] The third aspect of the present invention further discloses a computer-readable storage medium, which stores computer instructions. When the computer reads the computer instructions in the storage medium, the computer executes the aforementioned airport multimodal perception scene monitoring method based on 5G AeroMACS.

[0199] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.

Claims

1. An airport multimodal perception scene monitoring system based on 5G AeroMACS, characterized by: It includes signal transmitting node, signal receiving node, working mode switching module of airport scene communication perception, spectrum estimation module and fusion detection module; A signal transmitting node, including a 5G AeroMACS base station and / or a 5G AeroMACS repeater station, is configured to transmit a 5G AeroMACS signal when performing perception spectrum estimation on a target object; A signal receiving node, including a 5G AeroMACS base station, a 5G AeroMACS repeater station, and / or a 5G AeroMACS mobile station, is configured to receive a 5G AeroMACS signal when performing perception spectrum estimation on a target object; A working mode switching module is used to select the working mode of airport surface communication perception according to the airport surface; A spectrogram estimation module is used to generate a multi-dimensional perception spectrogram based on the 5G AeroMACS signals transmitted between the signal transmitting node, the signal receiving node and the target object; The fusion detection module includes an airport scene target object fusion detection model, which is used to detect the distance, speed and angle values ​​of the target object from the perception spectrum, give the confidence of the target object, and determine the category of the target object.

2. The airport multimodal perception scene monitoring system based on 5G AeroMACS according to claim 1 is characterized in that: It also includes a 5G AeroMACS communication transmission module for transmitting the multi-dimensional sensing spectrum of the 5G AeroMACS forwarding station and / or the 5G AeroMACS mobile station back to the 5G AeroMACS base station.

3. The airport multimodal perception scene monitoring system based on 5G AeroMACS according to claim 1 is characterized in that: It also includes a network training module for training a multimodal perception airport scene target object fusion detection model using a loss function and the three-dimensional detection frame and label frame of the target object obtained by the airport scene target object fusion detection model.

4. The airport multimodal perception scene monitoring system based on 5G AeroMACS according to any one of claims 1 to 3, characterized in that: The 5G AeroMACS base station and 5G AeroMACS repeater station are equipped with antennas to receive signals through the antennas; the mobile station is a vehicle; and the 5G AeroMACS base station is a 5G AeroMACS radar.

5. The airport multimodal perception scene monitoring system based on 5G AeroMACS according to any one of claims 1 to 3, characterized in that: The operating modes of airport surface communication perception include 5G AeroMACS base station perception mode, 5G AeroMACS forwarding station perception mode, and 5G AeroMACS mobile station perception mode. In 5G AeroMACS base station perception mode, the 5G AeroMACS base station transmits signals and acts as a signal transmitting node, while the 5G AeroMACS base station acts as a signal receiving node, receiving signals to perceive the target object. In the 5G AeroMACS forwarding station perception mode, the 5G AeroMACS base station transmits signals and serves as a signal transmitting node, and the 5G AeroMACS forwarding station serves as a signal receiving node. The received signals enable perception of the target object. In the 5G AeroMACS mobile station perception mode, the 5G AeroMACS base station or forwarding station transmits signals and serves as a signal transmitting node. The 5G AeroMACS mobile station serves as a signal receiving node. The received signals enable perception of the target object.

6. A multimodal sensing scene monitoring method for airports based on 5G AeroMACS, characterized in that: The specific steps are as follows: Step 1: Under 5G AeroMACS communication, according to the working mode of airport surface communication perception, obtain the receiving signal delay at the receiving node based on the target object, the Doppler frequency shift of the signal between the signal transmitting node and the signal receiving node, and the receiving angle; Step 2: According to the working mode of airport surface communication perception, the transmission signal of the signal transmitting node under 5G AeroMACS communication is obtained; Step 3: According to the working mode of airport surface communication perception, based on the received signal delay at the receiving node in step 1, the Doppler frequency shift and receiving angle of the signal between the signal transmitting node and the signal receiving node, and the transmitted signal obtained in step 2, the received signal of the signal receiving node under 5G AeroMACS communication is obtained; Step 4: According to the working mode of the airport surface communication perception, the transmission signal of the transmitting node in the received signal obtained in step 3 is processed to obtain a multi-dimensional perception spectrum; Step 5: Construct an airport scene target object fusion detection model that includes a fusion detection network, a spatial channel attention model, a shared multi-layer perceptron, a spatial attention module, a module for generating a spatial attention map, and an attention mechanism module. Use the airport scene target object fusion detection model to fuse the multidimensional perception spectrum obtained in step 4 to obtain a three-dimensional detection frame of the target object. Step 6: Based on the target object's label frame and the final 3D detection frame obtained in step 5, a fusion detection model for airport scene target objects based on multimodal perception is trained. Step 7: Use the trained multimodal perception airport scene target object fusion detection model to obtain the target object perception results.

7. The airport multimodal perception scene monitoring method based on 5G AeroMACS according to claim 6 is characterized in that: The working modes of airport surface communication perception include 5G AeroMACS base station perception mode, 5G AeroMACS forwarding station perception mode and 5G AeroMACS mobile station perception mode; in 5G AeroMACS base station perception mode, the 5G AeroMACS base station transmits signals and serves as a signal transmitting node, and the 5G AeroMACS base station serves as a signal receiving node, and the received signals realize perception of the target object; in 5G AeroMACS forwarding station perception mode, the 5G AeroMACS base station transmits signals and serves as a signal transmitting node, and the 5G AeroMACS forwarding station serves as a signal receiving node, and the received signals realize perception of the target object; in 5G AeroMACS mobile station perception mode, the 5G AeroMACS base station or forwarding station transmits signals and serves as a signal transmitting node, and the 5G AeroMACS mobile station serves as a signal receiving node, and the received signals realize perception of the target object.

8. The airport multimodal perception scene monitoring method based on 5G AeroMACS according to claim 6 is characterized in that: The multi-dimensional perception spectrum in step 4 includes a perception spectrum of distance-speed-angle of the target object.

9. The airport multimodal perception scene monitoring method based on 5G AeroMACS according to claim 6 is characterized in that: The target object perception result in step 7 is the detection frame of the target object, including the target object position, category and confidence.

10. A computer-readable storage medium, characterized in that The storage medium stores computer instructions. When the computer reads the computer instructions in the storage medium, the computer executes the airport multimodal perception scene monitoring method based on 5G AeroMACS as described in any one of claims 6 to 9.