Camera Relocalization Method for Fixed-Wing Aircraft Approach and Landing
The approach landing image is processed through the onboard forward-view camera and relocation network, and the inaccurate positioning problem caused by signal loss during the approach landing phase of fixed-wing aircraft is solved, and precise aircraft positioning is achieved in complex environments.
Patent Information
- Application Number
- CN202510299636.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The existing navigation methods are problem of inaccurate positioning due to signal loss during the approach and landing of fixed-wing aircraft.
The on-board forward-view camera is used to collect approach landing images, acquire the aircraft's posture information, convert it into camera posture information through the coordinate conversion module, and build a camera repositioning network for feature extraction and fusion, and output the repositioned camera posture information.
Provide accurate vehicle positioning in case of signal limitations, improving safety and accuracy of fixed-wing aircraft approach landings.
Smart Images

Figure CN119832459B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a camera relocalization method for fixed-wing aircraft approach and landing. Background Art
[0002] During the approach and landing phase of a fixed-wing aircraft, it usually relies on satellite navigation systems (GPS) and instrument landing systems (ILS) for positioning. However, these traditional navigation methods have obvious defects in complex environments. For example, satellite signals are prone to loss or reduced accuracy under adverse weather, terrain occlusion, or interference from high-rise buildings around the airport, resulting in inaccurate pose estimation of the aircraft. In addition, the instrument landing system relies on ground signals and may not be able to provide stable positioning information under equipment failures or environmental factors. Summary of the Invention
[0003] This application provides a camera relocalization method for fixed-wing aircraft approach and landing, which is used to solve the technical problem that the existing navigation methods may lead to inaccurate positioning due to signal loss during the approach and landing phase.
[0004] This application provides a camera relocalization method for fixed-wing aircraft approach and landing. The method includes: using an on-board forward-looking camera to collect images of the fixed-wing aircraft approach and landing; obtaining the aircraft pose information of the fixed-wing aircraft, where the aircraft pose information includes the position information and attitude information of the aircraft; performing conversion on the aircraft pose information through a coordinate conversion module; performing conversion on the position and attitude information of the fixed-wing aircraft through the coordinate conversion module to obtain camera pose information, where the camera pose information includes the position information and attitude information of the forward-looking camera in the world coordinate system; constructing a camera relocalization network, inputting the image and the corresponding camera pose information into the camera relocalization network for feature extraction, outputting a fused feature map, and performing camera relocalization based on the fused feature map to output camera pose information, where the camera pose information is the relocalized camera pose information.
[0005] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0006] The camera relocalization method for fixed-wing aircraft approach and landing provided by this application relates to the technical field of data processing. It collects approach and landing images through an airborne forward-looking camera, obtains the position and attitude information of the fixed-wing aircraft, and obtains the camera pose information through a coordinate conversion module. It constructs a camera relocalization network, extracts features from the input images and pose information, outputs a fused feature map, and performs relocalization based on this feature map, and finally outputs the relocalized camera pose information. It solves the technical problem that existing navigation methods may have inaccurate positioning due to signal loss during the approach and landing phase, and realizes the technical effect of providing accurate aircraft positioning through an airborne forward-looking camera and a relocalization network when the signal is limited. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0008] Figure 1 Schematic flowchart of the camera relocalization method for fixed-wing aircraft approach and landing provided by the embodiments of this application;
[0009] Figure 2 Schematic flowchart of generating the second camera pose information in the camera relocalization method for fixed-wing aircraft approach and landing provided by the embodiments of this application;
[0010] Figure 3 Schematic diagram of the structure of the global interaction module of the feature interaction module in the camera relocalization method for fixed-wing aircraft approach and landing provided by the embodiments of this application;
[0011] Figure 4 Schematic diagram of the structure of the local interaction module of the feature interaction module in the camera relocalization method for fixed-wing aircraft approach and landing provided by the embodiments of this application;
[0012] Figure 5 Schematic diagram of the coordinate system in the camera relocalization method for fixed-wing aircraft approach and landing provided by the embodiments of this application;
[0013] Figure 6 Schematic diagram of the structure of the proximity relocalization network LandNet in the camera relocalization method for fixed-wing aircraft approach and landing provided by the embodiments of this application;
[0014] Figure 7 Schematic diagram of the bottleneck block in the CNN branch in the camera relocalization method for fixed-wing aircraft approach and landing provided by the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] This application provides a camera relocalization method for the approach and landing of fixed-wing aircraft, which is used to solve the technical problem that the existing navigation methods may lead to inaccurate positioning due to signal loss during the approach and landing phase.
[0016] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the protection scope of this application.
[0017] It should be noted that the terms "first", "second", etc. in the description of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0018] Embodiment 1, as Figure 1 shown, this application provides a camera relocalization method for the approach and landing of fixed-wing aircraft, and this method includes:
[0019] P10: Use an airborne forward-looking camera to collect images of the fixed-wing aircraft during approach and landing.
[0020] Specifically, first use an airborne forward-looking camera to collect images during the approach and landing process of a fixed-wing aircraft. The airborne forward-looking camera is a sensor installed in the front of the aircraft, mainly used to capture visual information of the environment in front of the aircraft. These environmental image data are crucial for the navigation of the aircraft during the approach and landing phase, because they can provide real-time environmental information for the pilot or the autopilot system to help with tasks such as path planning, obstacle detection, and landing point confirmation.
[0021] The approach and landing phase is the most risky part of the aircraft flight process. The airborne forward-looking camera captures real-time image information of the runway, ground facilities, meteorological conditions and other key factors by shooting the area passed by the aircraft during flight. Through the analysis of these images, the system can update the position and attitude information of the aircraft in real time, so that the aircraft can land safely according to the predetermined flight path.
[0022] To ensure that the captured images have sufficient quality and accuracy, airborne forward-looking cameras are usually equipped with high-resolution imaging sensors, such as CMOS or CCD sensors, which can provide clear images under different lighting conditions. In addition, forward-looking cameras usually also adopt optical zoom or have high dynamic range (HDR) capabilities to cope with complex scenarios that may be encountered during flight, such as strong lighting contrast or night flight.
[0023] P20: Obtain the aircraft pose information of the fixed-wing aircraft, where the aircraft pose information includes the position information and attitude information of the aircraft.
[0024] Optionally, obtain the pose information of the fixed-wing aircraft. The pose information consists of two key parts: the position information and attitude information of the aircraft. The position information is usually represented by three-dimensional coordinates (X, Y, Z), where X, Y, and Z respectively correspond to the position information of the aircraft in the world coordinate system. The attitude information describes the orientation and posture of the aircraft relative to itself, usually represented by three angles (roll angle, pitch angle, yaw angle), which are used to define the spatial orientation of the aircraft relative to the ground and the flight path.
[0025] Exemplarily, the acquisition of aircraft pose information can rely on multiple sensors and navigation systems. The most commonly used systems include the Global Positioning System (GPS), Inertial Measurement Unit (IMU), and barometric altimeter. The GPS system can provide the real-time position of the aircraft globally. Through satellite signal positioning, the longitude, latitude, and altitude of the aircraft are obtained. The IMU measures the acceleration and angular velocity of the aircraft through accelerometers and gyroscopes, and thus infers the attitude changes of the aircraft, especially the roll angle, pitch angle, and yaw angle. The barometric altimeter estimates the altitude information of the aircraft based on changes in atmospheric pressure and is usually used in combination with the GPS to provide a more accurate flight altitude.
[0026] The aircraft pose information is the basis for coordinate transformation and camera repositioning in subsequent steps. It can provide the current position and attitude of the aircraft. Combining with the image data collected by the airborne forward-looking camera, through feature extraction and fusion, it can further help the repositioning system accurately estimate the attitude and position of the aircraft, thereby improving the safety and accuracy of landing.
[0027] P30: As Figure 5 shown, transform the aircraft pose information through a coordinate transformation module to obtain the first camera pose information, where the first camera pose information includes the position information and attitude information of the forward-looking camera.
[0028] Furthermore, step P30 of the embodiment of the present application further includes:
[0029] P31: Define multiple coordinate systems, where the multiple coordinate systems include the Earth-Centered Earth-Fixed (ECEF) coordinate system, the navigation coordinate system, the vehicle coordinate system, and the camera coordinate system; P32: Perform coordinate unified transformation on the multiple coordinate systems, output a coordinate transformation matrix, and analyze the aircraft pose information according to the coordinate transformation matrix to obtain the first camera pose information.
[0030] Among them, the expression for performing coordinate unified transformation on the multiple coordinate systems is:
[0031] ; where is the coordinate transformation for converting the geographical location to the three-dimensional coordinates in the camera coordinate system, including converting the longitude, latitude, and altitude in the WGS84 coordinate system to the three-dimensional coordinates of the Earth-Centered Earth-Fixed (ECEF) coordinate system, and converting the coordinates of the Earth-Centered Earth-Fixed (ECEF) coordinate system to the world coordinate system ; including the transformation from the navigation coordinate system to the vehicle coordinate system , and defining the transformation from the vehicle coordinate system to the camera coordinate system according to the camera installation relationship.
[0032] It should be understood that the aircraft pose needs to undergo a series of coordinate transformations through a coordinate transformation module to generate the camera pose information. Specifically, this process can be divided into two main stages: coordinate system definition and transformation, and coordinate unified transformation.
[0033] First of all, it is necessary to define multiple coordinate systems to ensure that all spatial data can be processed uniformly. The multiple coordinate systems include: the Earth-Centered Earth-Fixed (ECEF) coordinate system, which is a global coordinate system with the Earth's center of mass as the origin and is applicable to data processing in geographic information systems such as the Global Positioning System (GPS). The navigation coordinate system (NED), which is a coordinate system corresponding to the local geographical coordinate system of the aircraft's current position and provides the relative position information of the aircraft during flight in particular. The vehicle coordinate system (Body), which is the coordinate system of the aircraft itself and is usually aligned with the physical structure of the aircraft, such as the "front-right-up" direction of the fuselage. The camera coordinate system (Camera), which is a coordinate system related to the camera hardware and is used to represent the camera's perspective and the direction of capturing images.
[0034] Then, perform unified transformation on these coordinate systems. The goal of coordinate transformation is to convert the position and attitude information of the fixed-wing aircraft to the camera coordinate system. The transformation process can rely on the relationships between multiple coordinate systems, represented as a series of transformation matrices. Among them, each transformation matrix represents the relative position and attitude relationship between different coordinate systems, ensuring the conversion of the aircraft pose information from the global coordinate system to the camera coordinate system.
[0035] Exemplarily, the geodetic coordinate system WGS84 is converted to the Earth-Centered Earth-Fixed coordinate system ECEF. A point on the runway plane, which are longitude, latitude, and altitude respectively, is projected onto a point in the Earth-Centered Earth-Fixed coordinate system due to the flattening of the Earth ellipsoid. The conversion relationship is as follows:
[0036] ;
[0037] Where, represents the coordinate point in the Earth-Centered Earth-Fixed coordinate system, the three-dimensional coordinates after conversion, is the normal radius of the Earth (i.e., the radius of the Earth's surface at a certain latitude), and its value changes with latitude, is the altitude, that is, the height of a certain point on the ground, Lon is the longitude, Lat is the latitude, is the square of the eccentricity e of the Earth, is the ellipsoid radius of a certain point on the ground, and T is the transpose.
[0038] The Earth-Centered Earth-Fixed coordinate system ECEF is converted to the world coordinate system . For a point in the Earth-Centered Earth-Fixed coordinate system, it is converted to the ENU coordinate system with point A on the runway as the origin, and this coordinate system is used as the world coordinate system. The conversion relationship is as follows:
[0039] ;
[0040] Navigation coordinate system is converted to the vehicle coordinate system . By rotating the navigation coordinate system three times around the Z-axis, Y-axis, and X-axis, the vehicle coordinate system can be obtained, and its conversion relationship is as follows:
[0041] ;
[0042] Vehicle coordinate system is converted to the camera coordinate system . The aircraft fuselage is rigidly connected to the camera. The conversion model between the two coordinate systems includes the rotation matrix and the translation matrix . The coordinate conversion relationship from the vehicle coordinate system to the camera coordinate system is:
[0043] ;
[0044] Through the above conversions, the position and attitude of the camera in the world coordinate system can be obtained, and the image data captured by the forward-looking camera can be further used to complete the pose regression of the camera, thereby providing accurate visual data support for the autonomous navigation and approach landing of the aircraft.
[0045] P40: As Figure 6 shown, a camera relocalization network is constructed. The camera relocalization network includes an initial feature extraction module, a branch feature extraction module, a feature interaction module, and a pose regressor. The image and the camera pose information are input into the camera relocalization network for feature extraction, and a fused feature map is output. Camera relocalization is performed based on the fused feature map, and second camera pose information is output, where the camera pose information is the camera pose information after relocalization.
[0046] Further, as Figure 2 shown, step P40 of the embodiment of the present application further includes:
[0047] P41: Obtain an initial feature map according to the initial feature extraction module; P43: Input the initial feature map into the branch feature extraction module for enhanced feature extraction to obtain a local feature map and a global feature map; P44: Input the local feature map and the global feature map into the feature interaction module for feature interaction to output the fused feature map; P45: The pose regressor performs regression analysis on the fused feature map to output second camera pose information, including the position information and the attitude information after relocalization.
[0048] It should be understood that the camera pose is predicted by constructing a camera relocalization network. This process mainly includes inputting an image and first camera pose information, performing feature extraction and fusion through a series of network modules, and finally outputting second camera pose information, that is, the camera pose information after relocalization.
[0049] First, a camera relocalization network is constructed. This network consists of multiple modules, including an initial feature extraction module, a branch feature extraction module, a feature interaction module, and a pose regressor. The role of these modules is to extract, interact, and fuse features from the input image and camera pose information in order to more accurately predict the camera pose.
[0050] Specifically, the main task of the initial feature extraction module is to perform preliminary processing on the input image and extract basic image features from it. Usually, the initial feature extraction module uses two layers of convolution to obtain an initial feature map. These initially extracted features provide a basis for subsequent feature extraction and pose estimation.
[0051] The branch feature extraction module is used to perform more in-depth feature processing and enhancement on the initial feature map. This module includes two branches, which are respectively responsible for extracting local features and global features in the image. Local features usually focus on the detailed parts in the image. The extraction of local and global features helps to capture information at different scales, thereby improving the accuracy of camera relocalization.
[0052] Next, the local feature map and the global feature map after branch feature extraction are fed into the feature interaction module. The function of this module is to fuse the local features and the global features. Through feature interaction, the system can better understand the overall situation and details of the image, thereby enhancing the accuracy of pose prediction.
[0053] Finally, the fused feature map is fed into the pose regressor for regression analysis. The pose regressor is a multi-layer perceptron (MLP) or a fully connected network, and its task is to output the pose information of the camera based on the fused feature map. Specifically, the pose regressor will predict the position information of the camera (including the coordinates in the X, Y, and Z directions) and the attitude information (including the roll angle, pitch angle, and yaw angle). Through this regression process, the spatial information in the image is converted into the accurate camera pose, thus realizing the repositioning of the camera.
[0054] Throughout the process, the accuracy of the feature extraction module and the pose regressor is crucial because they directly affect the accuracy of camera repositioning. The deep learning-based feature extraction method can automatically learn useful spatial features from the data, not only reducing the workload of manually designing features but also being able to handle complex environmental changes such as illumination changes and perspective changes.
[0055] Furthermore, step P43 of the embodiment of the present application further includes:
[0056] P43-1: Input the initial feature map into the branch feature extraction module for enhanced feature extraction to obtain a local feature map and a global feature map, where the branch feature extraction module includes a CNN branch and a Transformer branch; P43-2: Among them, as Figure 7 shown, the CNN branch includes four CNN convolutional structural blocks, and the convolution sizes of each convolutional structural block are different. The Transformer branch includes four Transformer structural blocks, and the splitting granularities of each Transformer structural block are different; P43-3: Perform enhanced feature extraction on the initial feature map according to the CNN branch and output a local feature map; P43-4: Perform enhanced feature extraction on the initial feature map according to the Transformer branch and output a global feature map.
[0057] In a possible embodiment of the present application, the initial feature map is subjected to enhanced feature extraction by the branch feature extraction module to obtain a local feature map and a global feature map.
[0058] First, the initial feature map is further processed by being input into the branch feature extraction module. The branch feature extraction module consists of a CNN branch and a Transformer branch. The CNN branch is mainly used to extract local features of the image, while the Transformer branch is used to extract global features of the image. The purpose of doing this is to capture feature information at different levels from the image through two different network structures, thereby enhancing the understanding ability of complex scenes in the image.
[0059] Specifically, the CNN branch includes four CNN convolutional structural blocks, and the convolutional kernel sizes of each convolutional structural block are different, aiming to capture local features at different scales in the image. The change in the convolutional kernel size enables each convolutional structural block to perform feature extraction at different scales, thereby enhancing the recognition ability of local details. For example, small convolutional kernels can capture fine features in the image (such as edges, corners, etc.), while large convolutional kernels are helpful for capturing larger local structures (such as textures, shapes of objects, etc.). The detailed structure of the CNN branch is shown in Table 1:
[0060] Table 1 Detailed Structure Table of the CNN Branch
[0061]
[0062]
[0063] On the other hand, the Transformer branch consists of four Transformer structural blocks, and the segmentation granularity of each structural block is different. The segmentation granularity of the Transformer structure usually refers to the input image being divided into multiple non-overlapping image patches, and each patch is processed by the multi-head attention mechanism of the Transformer. Different segmentation granularities enable the Transformer to capture global features of the image at different scales and levels of detail. Especially when dealing with long-distance or complex image content, the Transformer architecture shows stronger global dependency modeling ability than traditional CNNs.
[0064] Exemplarily, according to the CNN branch, the initial feature map will undergo multiple convolutional processes to extract local features. Through convolutional operations, the network can automatically extract key information from the image, such as the edges, shapes, textures of objects, etc. The convolutional layer scans the image through a sliding window and extracts the features of each local region, thereby forming a local feature map. In this way, the CNN can effectively capture low-level and mid-level features in the image, providing necessary detailed information for subsequent pose and position prediction.
[0065] Meanwhile, the Transformer branch performs global feature extraction on the initial feature map. The self-attention mechanism of the Transformer allows the model to not only focus on local regions but also capture the context information of the entire image when processing the image. By dividing the image into multiple blocks and calculating the relationships between them, the Transformer can understand the mutual influence between various parts of the image, thereby generating a global feature map.
[0066] Exemplarily, when inputting the image obtained by the front-view camera into the Transformer branch, an image embedding operation needs to be performed first, that is, expanding the input image into a series of non-overlapping image patches . Among them, represents the size of the input image, represents the number of non-overlapping patches, and P represents the resolution of the image patch. represents the dimension after the image embedding operation. After the image embedding operation, the input is mapped to three matrices through a linear transformation: the query matrix , the key matrix , and the value matrix . This linear transformation is as shown in the formula:
[0067] ; among them, the values of the linear mapping matrices , , and need to be determined through learning. After obtaining the query matrix , the key matrix , and the value matrix , on this basis, the multi-head attention mechanism is calculated, and the calculation process is as shown in the formula:
[0068] ; among them, represents the output of the multi-head attention mechanism, represents the query matrix, K represents the key matrix, represents the transpose of the key matrix, is the dimension of the key, represents the dimension of the K vector.
[0069] The global feature map is obtained through calculation. The global feature map is crucial for the overall understanding of the scene. Especially when dealing with complex backgrounds or long-range dependencies, the Transformer shows stronger capabilities than traditional CNNs.
[0070] By combining these two network architectures, local and global information can be fully utilized, and thus richer and more accurate input features can be provided for subsequent feature fusion and pose regression.
[0071] Furthermore, step P44 of the embodiment of the present application further includes:
[0072] P44-1: Input the local feature map and the global feature map into the feature interaction module for feature interaction. As Figure 2 , Figure 3 shown, the feature interaction module includes a global interaction module and a local interaction module; P44-2: The global interaction module is used to upsample the local feature map output by the CNN branch, perform feature concatenation, dimensionality reduction, and activation with the global feature map output by the Transformer branch, and output an optimized global feature map; P44-3: The local interaction module is used to downsample the global feature map output by the Transformer branch, perform feature concatenation, dimensionality reduction, and activation with the local feature map output by the CNN branch, and output an optimized local feature map; P44-4: Output a fused feature map according to the optimized global feature map and the optimized local feature map.
[0073] Optionally, the local feature map and the global feature map are further fused and optimized through the feature interaction module to provide more accurate input features for subsequent pose regression. This process mainly includes the collaborative work of the global interaction module and the local interaction module, and the functions of these two modules are to optimize the feature map and perform information fusion.
[0074] First, the local feature map and the global feature map obtained from the CNN branch and the Transformer branch respectively are input into the feature interaction module.
[0075] In the global interaction module, first, the local feature map output by the CNN branch is upsampled. The upsampling operation is usually used to increase the resolution of the feature map so that it can match the global feature map. Then, the local feature map is concatenated with the global feature map output by the Transformer branch, that is, these two different types of feature maps are combined in the channel dimension to form a new comprehensive feature map. The concatenated feature map will also be dimensionally reduced, and the number of channels is reduced through a convolutional or fully connected layer to retain important feature information. Finally, after being processed by the activation function, an optimized global feature map is output. The role of the activation function is to introduce a non-linear transformation so that the model can capture more complex features.
[0076] Similar to the global interaction module, the local interaction module downsamples the global feature map output by the Transformer branch. Downsampling operations are typically used to reduce the resolution of the feature map to match the local feature map. Then, the optimized local feature map and the global feature map are feature concatenated. In this way, the information of the local feature map and the global feature map is complementary, facilitating subsequent feature processing. The concatenated feature map will undergo dimensionality reduction to further reduce redundant information and be activated to output the optimized local feature map. In this way, local information is strengthened and effectively combined with global information.
[0077] Finally, based on the optimized global feature map and the optimized local feature map, the two are fused to output the final fused feature map. This fused feature map contains the global information and local details in the image and is the input for the next pose regression. Through this multi-level and multi-dimensional information fusion, the spatial structure in the image can be understood more precisely, thereby improving the accuracy of camera pose repositioning.
[0078] Furthermore, step P45 of the embodiment of the present application further includes:
[0079] P45-1: Among them, the pose regressor is a multi-layer perceptron, including two fully connected layers and an activation function; P45-2: Input the fused feature map into the pose regressor, concatenate the obtained three-dimensional position feature vector and three-dimensional attitude feature vector, and output a six-dimensional pose feature vector; P45-3: Output the second camera pose information according to the six-dimensional pose feature vector.
[0080] Specifically, regression analysis is performed on the input fused feature map through a pose regressor to predict the pose information of the camera. Specifically, this process mainly includes the extraction of three-dimensional position and attitude features, feature concatenation, and the final pose output.
[0081] Among them, the pose regressor is a multi-layer perceptron (MLP), composed of fully connected layers and an activation function. The fully connected layer is used to convert the input feature map into the required pose information. The fully connected layer is a common network layer in deep learning, which maps the input data to the output space through weighted sum and bias, and then generates the final prediction result. In this step, the role of the fully connected layer is to convert the fused feature map that has passed through the feature interaction module and optimization into a specific pose vector. The activation function is usually applied after each fully connected layer to introduce non-linearity, enabling the model to capture complex patterns. Common activation functions include ReLU, Sigmoid or Tanh. Here, they help the neural network learn and adjust complex mapping relationships, enhancing the expression ability of the model.
[0082] First, the optimized fused feature map is input into the pose regressor for processing. The pose regressor extracts a three-dimensional position feature vector and a three-dimensional attitude feature vector from it. The three-dimensional position feature vector contains the X, Y, and Z coordinate information of the camera in the air, indicating the position of the camera in the world coordinate system; while the three-dimensional attitude feature vector contains the roll angle, pitch angle, and yaw angle of the camera, describing the orientation and spatial attitude of the camera. After being processed by the regressor, these two vectors respectively represent the position information and attitude information of the camera. Next, these two feature vectors are concatenated to form a six-dimensional pose feature vector, which simultaneously contains the three-dimensional position and three-dimensional attitude information of the camera. The role of feature concatenation is to combine the position and attitude information into a complete feature vector that contains all the necessary pose information.
[0083] Finally, according to the concatenated six-dimensional pose feature vector, the second camera pose information is output. This output is the final predicted pose of the camera, which contains the three-dimensional position and three-dimensional attitude of the camera, that is, the accurate coordinates and orientation of the camera. Through this regression process, the pose of the camera in the approach and landing phase can be accurately predicted, providing key data for subsequent navigation and control.
[0084] Furthermore, the camera relocalization network in the embodiments of the present application is obtained through supervised learning training by introducing a loss function. The loss function includes:
[0085] ; where is used to measure the distance error between the real position and the predicted position, is the dynamic weight factor of the position error, is the weight for dynamically adjusting the position error, is used to measure the normalized difference between the real attitude and the predicted attitude, is the dynamic weight factor of the attitude error, is the weight for dynamically adjusting the attitude error; , is the actual camera position information, is the predicted camera position information, is the Euclidean norm used to calculate the Euclidean distance between two position vectors; , is the actual camera attitude information, is the predicted camera attitude information, is the normalization process of the predicted camera attitude information.
[0086] It should be understood that in this step, a loss function is introduced for supervised learning training to optimize the performance of the camera relocalization network. The loss function is an essential part of supervised learning, whose purpose is to measure the difference between the predicted value and the true value, and update the weights of the model through this difference to improve the accuracy of the model. In the embodiments of the present application, the loss function includes two main components: position error and attitude error.
[0087] Among them, the position error ( ) calculates the Euclidean distance between the true position of the aircraft ( ) and the predicted position ( ), and the attitude error ( ) is used to calculate the error between the true attitude of the aircraft ( ) and the predicted attitude ( ). Since the attitude is usually represented in the form of unit quaternions, the validity of the rotation transformation is ensured. The total loss function is composed of the weighted sum of the position error and the attitude error. The introduction of dynamic weight factors (such as and ) is to balance the position change rate and the attitude change rate during the training process to ensure that both the position and attitude errors can be effectively optimized.
[0088] By introducing such a loss function with dynamic adjustment, during the approach and landing process of the aircraft, the different change rates of position and attitude can be comprehensively considered, the accuracy of camera relocalization can be optimized, the adaptability of the loss function is ensured, and it helps the camera relocalization network to better fit the actual flight path and attitude changes.
[0089] Exemplarily, as shown in Table 2, the comparison of the solution accuracy between the LandNet disclosed in the present invention and the existing camera relocalization methods:
[0090] Table 2 Comparison of the solution accuracy between the LandNet of the present invention and the existing camera relocalization methods
[0091]
[0092] As shown in Table 2, the position errors of the camera relocalization method disclosed in the present invention in the X, Y, and Z directions are all optimal. The average error in the X direction is 6.4 meters, the average error in the Y direction is 4.1 meters, and the average error in the height direction is 0.67 meters. The position accuracy meets the requirements during aircraft landing; in terms of the angular error, the LandNet disclosed in the present invention achieves a roll angle error of 0.011 degrees, a pitch angle error of 0.003 degrees, and a yaw angle error of 0.011 degrees. The attitude accuracy also meets the navigation requirements of the aircraft during the approach and landing process.
[0093] In summary, the embodiments of the present application at least have the following technical effects:
[0094] This application uses an airborne forward-looking camera to collect images of a fixed-wing aircraft during approach and landing, and obtains the pose information (including position and attitude) of the aircraft. Through a coordinate conversion module, the position and attitude information of the fixed-wing aircraft is converted into the first camera pose information. Then, a camera relocalization network is constructed, and the image and the first camera pose information are input into the network for feature extraction, and a fused feature map is output. Camera relocalization is performed based on the fused feature map, and finally the second camera pose information, that is, the position and attitude of the relocalized camera, is output.
[0095] It achieves the technical effect of providing accurate aircraft positioning through an airborne forward-looking camera and a relocalization network when the signal is limited.
[0096] It should be noted that the above sequence of embodiments of this application is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of this specification have been described. In addition, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0097] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.
[0098] This specification and the drawings are only exemplary descriptions of this application and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of this application. Obviously, those skilled in the art can make various changes and modifications to this application without departing from the scope of this application. Thus, if these modifications and variations of this application fall within the scope of this application and its equivalent technologies, this application is intended to include these changes and modifications.
Claims
1. A camera relocalization method for fixed-wing aircraft approach and landing, characterized in that The method includes: Collecting images of a fixed-wing aircraft approaching and landing using an on-board forward-looking camera; Obtaining the aircraft pose information of the fixed-wing aircraft, where the aircraft pose information includes the position information and attitude information of the aircraft; Converting the aircraft pose information through a coordinate conversion module to obtain first camera pose information, where the first camera pose information includes the position information and attitude information of the forward-looking camera in the world coordinate system; Constructing a camera relocalization network, where the camera relocalization network includes an initial feature extraction module, a branch feature extraction module, a feature interaction module, and a pose regressor; Inputting the image and the first camera pose information into the camera relocalization network for feature extraction, outputting a fused feature map, and performing camera relocalization based on the fused feature map to output second camera pose information, where the second camera pose information is the relocalized camera pose information; The method for outputting the second camera pose information includes: Obtaining an initial feature map according to the initial feature extraction module; Inputting the initial feature map into the branch feature extraction module for enhanced feature extraction to obtain a local feature map and a global feature map; Inputting the local feature map and the global feature map into the feature interaction module for feature interaction to output the fused feature map; The pose regressor performs regression analysis on the fused feature map and outputs camera pose information, that is, the relocalized position information and attitude information; The branch feature extraction module includes a CNN branch and a Transformer branch; Among them, the CNN branch includes four CNN convolutional structural blocks, and the convolution sizes of each convolutional structural block are different. The Transformer branch includes four Transformer structural blocks, and the segmentation granularities of each Transformer structural block are different; Performing enhanced feature extraction on the initial feature map according to the CNN branch and outputting a local feature map; Performing enhanced feature extraction on the initial feature map according to the Transformer branch and outputting a global feature map; The feature interaction module includes a global interaction module and a local interaction module; The global interaction module is used to upsample the local feature map output by the CNN branch, perform feature splicing, dimensionality reduction, and activation with the global feature map output by the Transformer branch, and output an optimized global feature map; The local interaction module is used to downsample the global feature map output by the Transformer branch, perform feature splicing, dimensionality reduction, and activation with the local feature map output by the CNN branch, and output an optimized local feature map; Outputting a fused feature map according to the optimized global feature map and the optimized local feature map.
2. The method according to claim 1, wherein The pose regressor performs regression analysis on the fused feature map and outputs the second camera pose information. The method includes: Among them, the pose regressor is a multi-layer perceptron, including two fully connected layers and an activation function; Input the fused feature map into the pose regressor, concatenate the obtained three-dimensional position feature vector and three-dimensional attitude feature vector, and output a six-dimensional pose feature vector; Output the second camera pose information according to the six-dimensional pose feature vector.
3. The method according to claim 1, characterized in that, The camera relocalization network is obtained through supervised learning training by introducing a loss function, and the loss function includes: ; Among them, is used to measure the distance error between the true position and the predicted position, is the dynamic weight factor of the position error, is to dynamically adjust the weight of the position error, is used to measure the normalized difference between the true attitude and the predicted attitude, is the dynamic weight factor of the attitude error, is to dynamically adjust the weight of the attitude error; , is the actual camera position information, is the predicted camera position information, is the Euclidean norm used to calculate the Euclidean distance between two position vectors; , is the actual camera pose information, is the predicted camera pose information, is the normalization of the predicted camera pose information.
4. The method according to claim 1, wherein The method of the coordinate transformation module includes: Define multiple coordinate systems, and the multiple coordinate systems include the Earth-centered Earth-fixed coordinate system, the navigation coordinate system, the vehicle coordinate system, the camera coordinate system, and the world coordinate system; Perform coordinate unified transformation on the multiple coordinate systems, output a coordinate transformation matrix, and analyze the aircraft pose information according to the coordinate transformation matrix to obtain the pose information of the camera in the world coordinates.
5. The method according to claim 4, wherein The expression for performing coordinate unified transformation on the multiple coordinate systems: ; Among them, is a coordinate transformation for converting a geographical location into three-dimensional coordinates in the camera coordinate system, including converting longitude, latitude, and altitude in the WGS84 coordinate system into three-dimensional coordinates of the Earth-Centered Earth-Fixed coordinate system ECEF, and converting the coordinates of the Earth-Centered Earth-Fixed coordinate system ECEF into the world coordinate system ; including the transformation from the navigation coordinate system to the vehicle coordinate system , and the transformation from the vehicle coordinate system defined according to the camera installation relationship to the camera coordinate system .
Citation Information
Patent Citations
Flight guidance system for approaching stage of aircraft, verification method for image acquisition module of flight guidance system and flight guidance method
CN115183780A
Camera repositioning system based on lightweight Transform model
CN117726676A