Method and apparatus for positioning an aircraft, aircraft, medium and program product
By employing omnidirectional fisheye vision and adaptive weighting methods, and utilizing multi-eye fisheye cameras and a unified coordinate system transformation matrix, the visual marker positioning technology was flexibly applied and achieved high-precision positioning in complex environments. This solved the problem of limited deployment of visual marker positioning technology in complex scenarios, ensuring the stability and accurate positioning of UAVs in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-08
- Publication Date
- 2026-03-31
AI Technical Summary
Existing visual marker localization technology requires all markers to be placed on the same plane, which cannot adapt to complex environments with multiple angles and directions. This limits its application in complex scenarios. Furthermore, single-camera viewpoint recognition and localization is unstable, prone to field of view occlusion and localization failure, resulting in low localization accuracy.
A method based on omnidirectional fisheye vision and adaptive weighting is adopted. Images are acquired by multi-eye fisheye cameras, and arbitrarily arranged visual marker objects are identified. By combining the transformation matrix of a unified coordinate system and the adaptive weighting mechanism, an objective function is constructed to solve the pose, thereby realizing the information fusion of multiple visual marker objects.
It improves the flexibility and positioning accuracy of visual marker positioning technology in complex scenarios, ensuring continuous, stable and high-precision positioning of UAVs in complex environments, adapting to multi-angle and multi-directional deployment, and avoiding the problem of field of view obstruction.
Smart Images

Figure CN121475158B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of flight positioning technology, and more specifically, to a positioning method, device, aircraft, medium, and program product for an aircraft. Background Technology
[0002] With the rapid development of drones in related fields, higher demands are being placed on their high-precision autonomous positioning in complex environments. Traditional solutions mainly rely on Global Positioning System (GNSS) or Visual-Inertial Odometry (VIO) for drone positioning. However, in indoor environments and areas with strong magnetic interference, such as substations, GNSS positioning suffers from signal obstruction, making accurate positioning impossible. VIO positioning methods are easily affected by factors such as changes in lighting, texture loss, and cumulative drift, resulting in insufficient positioning stability for drones in long-term flight scenarios.
[0003] Currently, in order to solve the problems of signal occlusion and long-term unstable positioning in traditional solutions, visual marker-based positioning technology is usually used instead. This technology is based on the known coordinates of each marker in the three-dimensional world. By detecting their projection positions in a two-dimensional image, the pose information of the camera relative to the marker can be calculated according to the intrinsic and extrinsic parameters calibrated by the camera.
[0004] However, existing visual marker positioning technology requires all markers to be placed on the same plane (such as the ground) to achieve accurate positioning. This limits the flexibility of marker placement and cannot adapt to complex environments with multi-angle and multi-directional distribution, thus limiting the application of this positioning technology in complex scenarios. Summary of the Invention
[0005] The purpose of this application is to provide a method, device, aircraft, medium, and program product for locating an aircraft, so as to improve the application flexibility of visual marker positioning technology in complex scenarios.
[0006] In a first aspect, embodiments of this application provide a method for locating an aircraft, including:
[0007] Acquire images to be identified from cameras on the aircraft, and identify at least one visually marked object in the images to be identified;
[0008] Based on the identification information of each visual marker object, the two-dimensional coordinate information corresponding to each visual marker object is matched and associated with the actual three-dimensional coordinate information; wherein, each actual three-dimensional coordinate information is obtained based on a preset coordinate transformation rule and represented by a unified transformation matrix parameter;
[0009] Construct and solve an objective function for the two-dimensional coordinate information, the actual three-dimensional coordinate information, and the pose of the aircraft to be solved. Obtain the aircraft pose information based on the solution.
[0010] In this embodiment, the position information of each marker is represented by a transformation matrix based on a unified coordinate system, thereby allowing different markers to be arbitrarily placed in different planes, effectively improving the application flexibility of visual marker positioning technology in complex scenarios.
[0011] In some embodiments, the step of constructing and solving an objective function for two-dimensional coordinate information, actual three-dimensional coordinate information, and the pose of the aircraft to be solved, and obtaining the aircraft pose information based on the solution results, includes:
[0012] The theoretical two-dimensional coordinate information corresponding to each of the actual three-dimensional coordinate information is determined according to the preset projection parameters; wherein, the projection parameters are constructed by the preset camera external factors to solve the pose of the aircraft.
[0013] The aircraft pose is solved by minimizing the sum of the differences between each of the two-dimensional coordinate information and the corresponding theoretical two-dimensional coordinate information, and the aircraft pose information is obtained based on the solution results.
[0014] In this embodiment, the three-dimensional coordinates of the detection result are projected according to preset projection parameters, and an objective function is constructed based on the theoretical two-dimensional coordinates obtained by projection and the two-dimensional coordinates detected by the image for solution, thereby further improving the accuracy of pose solution.
[0015] In some embodiments, the step of constructing and solving an objective function for two-dimensional coordinate information, actual three-dimensional coordinate information, and the pose of the aircraft to be solved, and obtaining the aircraft pose information based on the solution results, includes:
[0016] By combining the weight information corresponding to each visual marker object, an objective function is constructed and solved for the two-dimensional coordinate information, the actual three-dimensional coordinate information, and the pose of the aircraft to be solved. The aircraft pose information is obtained based on the solution result. The weight information is determined based on the graphic size of each visual marker object in the image to be identified.
[0017] In this embodiment, the weights of each visual marker object are determined based on the size of the graphic in the image, and the weights of each visual marker object are combined in the pose solving process for weighted solution, thereby further improving the accuracy of aircraft pose solving.
[0018] In some embodiments, the method for determining the weight information includes:
[0019] The side length parameters corresponding to each visually labeled object are determined based on a dynamically adjusted weight index factor.
[0020] Determine the sum of the side length parameters of all visually marked objects in the image to be identified;
[0021] The weight information of the visual marker object is determined based on the ratio of the side length parameter corresponding to the visual marker object to the sum of the side length parameters.
[0022] The weighting index factor is obtained by dynamically adjusting a preset base index based on the side length difference index of each visual marker object in the image to be identified.
[0023] In this embodiment of the application, the weight index used to calculate the weight value is dynamically adjusted according to the difference in side length corresponding to each visual marker in the image, thereby further improving the accuracy of weight allocation.
[0024] In some embodiments, the actual three-dimensional coordinate information of each visual marker object is obtained in the following ways:
[0025] Obtain the geographic location information of the four corner points of each of the aforementioned visual marker objects; wherein, the geographic location information includes latitude, longitude, and altitude;
[0026] Based on the geographical location information of each corner point in the visual marker object, the local coordinate system pose information of the visual marker object relative to the Earth coordinate system is determined; wherein, the local coordinate system pose information is represented by the first transformation matrix parameter corresponding to the Earth coordinate system;
[0027] A unified coordinate system is constructed based on the aircraft's takeoff position as the origin and the pose information of the origin relative to the Earth coordinate system; wherein, the pose information includes the aircraft's yaw angle and latitude and longitude information at the takeoff position;
[0028] The pose information of each visual marker object in the local coordinate system is transformed into the unified coordinate system to obtain the actual three-dimensional coordinate information of each visual marker object; wherein the actual three-dimensional coordinate information is represented by the second transformation matrix parameter corresponding to the unified coordinate system.
[0029] In this embodiment of the application, by establishing a local coordinate system corresponding to each visual marker object based on the Earth coordinate system, and then uniformly representing the local coordinate system corresponding to each visual marker object based on the unified coordinate system with the aircraft take-off position as the origin, the application flexibility of visual marker positioning technology in complex scenarios is effectively improved.
[0030] In some embodiments, acquiring the image to be identified captured by a camera on the aircraft includes:
[0031] Acquire multiple images to be identified from the multi-view fisheye camera on the aircraft.
[0032] In this embodiment, the reliability of pose determination is further improved by using a multi-eye fisheye camera to acquire multiple images for coordinate information fusion and pose determination.
[0033] In some embodiments, before solving for the aircraft pose, the method further includes:
[0034] Based on a preset fisheye camera calibration model, the original two-dimensional coordinate information corresponding to each visual marker object is corrected to obtain the corrected two-dimensional coordinate information.
[0035] In this embodiment of the application, the accuracy of pose calculation is further improved by first performing distortion correction processing on the image before pose calculation.
[0036] In some embodiments, the two-dimensional coordinate information and the actual three-dimensional coordinate information both include the coordinates of multiple corner points corresponding to the visual marker object.
[0037] In the embodiments of this application, the reliability of pose solving is further improved by using multiple corner coordinates to represent the two-dimensional and three-dimensional coordinates corresponding to each marker.
[0038] Secondly, embodiments of this application provide a positioning device for an aircraft, comprising:
[0039] The marker recognition module is used to acquire images to be recognized captured by cameras on the aircraft and to identify at least one visual marker object in the images to be recognized;
[0040] The matching and association module is used to match and associate the two-dimensional coordinate information corresponding to each visual marker object with the actual three-dimensional coordinate information based on the identification information of each visual marker object; wherein, each actual three-dimensional coordinate information is obtained based on a preset coordinate transformation rule and represented by a unified transformation matrix parameter;
[0041] The pose solving module is used to construct and solve the objective function of the two-dimensional coordinate information, the actual three-dimensional coordinate information, and the pose of the aircraft to be solved, and obtain the pose information of the aircraft based on the solution results.
[0042] Thirdly, embodiments of this application provide an aircraft including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, can implement the method described in any embodiment of the first aspect.
[0043] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the method described in any embodiment of the first aspect.
[0044] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program, wherein when the computer program is executed by a processor, it can implement the method described in any embodiment of the first aspect. Attached Figure Description
[0045] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 A flowchart illustrating a method for locating an aircraft, as provided in an embodiment of this application;
[0047] Figure 2 This is a schematic diagram of the structure of a positioning device for an aircraft provided in an embodiment of this application;
[0048] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0049] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.
[0050] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0051] It should be noted that among existing positioning methods based on visual tags (such as QR codes), commonly used visual tags include QRCode, ArUco, and AprilTag. Among them, AprilTag is increasingly widely used in the fields of UAV positioning and robot navigation due to its advantages such as strong uniqueness, high noise resistance, fast detection speed, and high corner positioning accuracy.
[0052] However, existing AprilTag-based localization methods have at least one of the following problems:
[0053] 1. Limited Deployment Flexibility: Existing AprilTag positioning methods typically require all QR codes to be placed on the same plane (e.g., the ground) and identified by a camera with a top-down view. This method limits the angle and position of the QR codes, making it unsuitable for complex environments with multi-angle and multi-directional distribution, thus limiting its application in actual drone inspections or indoor scenarios.
[0054] 2. Using a single camera view for identification and positioning cannot cover the entire area. This can easily lead to situations where the drone cannot detect the AprilTag in the field of view at a certain location, resulting in positioning failure and affecting flight continuity and mission stability.
[0055] 3. The accuracy is unstable when the positioning is based on only a single AprilTag in each recognition: When a single AprilTag is located at the edge of the image or far away from the camera, the detection accuracy of its position will decrease significantly due to factors such as reduced resolution and projection distortion, which will lead to increased pose calculation error and insufficient stability.
[0056] 4. If a fisheye lens is used for image acquisition, the positioning accuracy will be affected by image distortion: UAVs often use fisheye or wide-angle lenses to obtain images with a large field of view, but severe distortion will affect the corner accuracy of AprilTag.
[0057] 5. Insufficient utilization of multiple AprilTag information: In scenarios where a single frame image contains multiple AprilTags, traditional methods often only solve the pose based on each independent QR code, failing to effectively integrate multiple AprilTag information in different coordinate systems, resulting in insufficient overall positioning accuracy and robustness.
[0058] To address at least one of the problems existing in the prior art, this application provides a drone positioning method based on omnidirectional fisheye vision and adaptive weighted arbitrary layout QR codes.
[0059] like Figure 1 As shown in the figure, this application provides a method for locating an aircraft, which may include the following steps:
[0060] S1. Acquire the image to be identified captured by the camera on the aircraft, and identify at least one visually marked object in the image to be identified.
[0061] It should be noted that the main objective of this application embodiment is to use visual marking technology to assist in achieving continuous, stable, and high-precision positioning when GNSS signals are blocked or interfered with, or when the cumulative error of VIO visual positioning is too large, during the process of UAV performing tasks such as indoor and outdoor inspections and substation equipment inspections.
[0062] For example, multiple visual marker objects (such as AprilTag QR codes) can be deployed in a preset environment, and the actual coordinate information corresponding to each QR code (e.g., latitude and longitude coordinates and altitude of corner and center points) can be recorded. The three-dimensional coordinate information of each QR code is configured according to its actual installation location and size, serving as a reference for subsequent pose calculation. It is understood that each QR code can be configured independently, and its number can be flexibly increased or decreased according to task requirements. By increasing the QR code deployment density in the environment, the positioning coverage and accuracy can be further expanded to meet the autonomous inspection and positioning needs of UAVs in scenarios of different scales.
[0063] For example, when an aircraft or drone needs to locate itself during flight, it can use a camera (camera) on the aircraft to capture images of the surrounding environment (identify the image to be identified) and identify visual markers in the image. Based on this, it can calculate the drone's pose information (including latitude and longitude, altitude and direction information) at its current position (the position where the image was taken).
[0064] For example, the same image to be identified may contain one or more visually labeled objects.
[0065] S2. Based on the identification information of each visual marker object, match and associate the two-dimensional coordinate information corresponding to each visual marker object with the actual three-dimensional coordinate information; wherein, each actual three-dimensional coordinate information is obtained based on a preset coordinate transformation rule and represented by a unified transformation matrix parameter.
[0066] It should be noted that when multiple visual marker objects are deployed in the preset environment, each AprilTag (visual marker object) has a unique identifier (such as ID information), which is generated through a specific encoding method.
[0067] By identifying each visual marker object in the image to be identified, the identification information corresponding to each visual marker object in the image can be determined. Then, based on the identification information, the visual marker objects in the image can be matched one-to-one with the actual visual marker objects, thereby matching and associating the two-dimensional coordinate information (i.e. the position of the visual marker object in the image) of each visual marker object with the actual three-dimensional coordinate information.
[0068] For example, each visual marker object can correspond to the two-dimensional coordinate information of a point or a set of points (such as the four corner points of a QR code, or specific key points). Accordingly, the two-dimensional coordinate information of a point corresponds to an actual three-dimensional coordinate information.
[0069] It should be noted that the actual three-dimensional coordinate information of each visual marker object is represented by a unified transformation matrix parameter. In other words, each visual marker object can be set in any direction and any position, such as on the ground, on a wall, or on a column, which greatly improves the flexibility of visual marker placement and is suitable for various complex indoor and outdoor scenes.
[0070] By establishing a matching relationship between the detected visual marker objects (corner points) in two dimensions and their corresponding three-dimensional spatial coordinates, two-dimensional-three-dimensional matching point pairs are formed. The coordinates of these point pairs are used as input to the pose calculation algorithm, providing constraints for the pose calculation of the UAV camera.
[0071] S3. Construct and solve the objective function for the two-dimensional coordinate information, the actual three-dimensional coordinate information, and the attitude of the aircraft to be solved. Obtain the attitude information of the aircraft based on the solution results.
[0072] Finally, using the matched coordinate information pairs (two-dimensional coordinate information paired with actual three-dimensional coordinate information one-to-one) as constraints, the attitude information of the UAV is jointly solved using a preset Bundle Adjustment (BA) optimization algorithm. Specifically, an objective function is constructed based on the two-dimensional coordinate information, the actual three-dimensional coordinate information, and the pose of the aircraft to be solved, and then solved to obtain the current pose information of the aircraft.
[0073] Based on this, the positional information of each marker is represented by a transformation matrix based on a unified coordinate system, which allows different markers to be arbitrarily placed in different planes, effectively improving the application flexibility of visual marker positioning technology in complex scenarios.
[0074] In some embodiments, step S3, which involves constructing and solving an objective function for the two-dimensional coordinate information, the actual three-dimensional coordinate information, and the aircraft pose to be solved, and obtaining the aircraft pose information based on the solution result, may include:
[0075] S301. Determine the theoretical two-dimensional coordinate information corresponding to each actual three-dimensional coordinate information according to the preset projection parameters; wherein, the projection parameters are constructed by the preset camera external factors to the pose of the aircraft to be solved.
[0076] S302. The goal is to minimize the sum of the differences between each two-dimensional coordinate information and the corresponding theoretical two-dimensional coordinate information. The attitude of the aircraft to be solved is then calculated, and the attitude information of the aircraft is obtained based on the solution results.
[0077] Specifically, in the objective function, for each actual three-dimensional coordinate information, a relational expression for projecting the actual three-dimensional coordinate information onto the image plane using preset projection parameters can be constructed based on a preset reprojection calculation function. Here, the projection parameters refer to the relational expression constructed based on preset camera extrinsic factors to solve for the aircraft pose.
[0078] Meanwhile, in the objective function, the goal is to minimize the sum of the differences between each two-dimensional coordinate information and its corresponding theoretical two-dimensional coordinate information. This is used to solve for the aircraft pose, and the aircraft pose information can be obtained from the solution. In other words, given the actual two-dimensional coordinate information and the corresponding actual three-dimensional coordinate information of each point (one or more points of each visual marker object) in the image detection results, and also knowing the camera's extrinsic parameters (the projection parameters of the UAV in the same coordinate system can be calculated based on the preset camera extrinsic parameters and the UAV pose), the actual three-dimensional coordinate information can be converted into the corresponding theoretical two-dimensional coordinate information according to the projection parameters. Finally, by minimizing the sum of the errors between the actual and theoretical two-dimensional coordinate information of each point, the aircraft pose can be obtained.
[0079] Based on this, the three-dimensional coordinates of the detection results are projected according to preset projection parameters, and the objective function is constructed and solved based on the theoretical two-dimensional coordinates obtained by projection and the two-dimensional coordinates detected by the image, thereby further improving the accuracy of pose solving.
[0080] In some embodiments, an objective function is constructed and solved for two-dimensional coordinate information, actual three-dimensional coordinate information, and the pose of the aircraft to be solved. Based on the solution, the aircraft pose information is obtained, including:
[0081] By combining the weight information corresponding to each visual marker object, an objective function is constructed and solved for the two-dimensional coordinate information, the actual three-dimensional coordinate information, and the pose of the aircraft to be solved. The aircraft pose information is obtained based on the solution results. The weight information is determined based on the graphic size of each visual marker object in the image to be recognized.
[0082] It should be noted that when multiple visual markers exist in an image, the closer the visual marker is to the camera, the higher its detection accuracy. Therefore, in the process of calculating the drone position, more weight can be assigned to the coordinate information of the visual markers that are closer to the camera, so as to dynamically adjust the contribution of different visual markers in pose calculation and improve the overall robustness and accuracy.
[0083] Specifically, the weight value of each visual marker object can be determined based on its graphic size in the image to be recognized. For example, the side length of the visual marker object in the image can be used as a measure of its size. The longer the side length, the closer the visual marker object is to the camera, and therefore a larger weight is assigned to it; the shorter the side length of the visual marker object in the image, the farther the visual marker object is from the camera, and therefore a smaller weight is assigned to it.
[0084] For example, due to factors such as shooting angle, the actual square visual marker object may appear as an irregular shape in the image to be recognized, such as a quadrilateral with four unequal sides. Therefore, the sum (or average) of the lengths of the four sides can be used as the "side length" of the visual marker object, and this can be used as the basis for determining the graphic size of each visual marker object in the image to be recognized.
[0085] It should be noted that during the pose calculation process, different visual marker objects in the image have different weights; however, if the pose calculation is performed using the coordinate information of multiple points of each visual marker object, then for the same visual marker object, its multiple points (such as the four corner points) have the same weight.
[0086] Based on this, the weights of each visual marker object in the image are determined by the size of the object, and the weights of each visual marker object are combined in the pose solving process for weighted solution, thereby further improving the accuracy of aircraft pose solving.
[0087] In some embodiments, the method for determining weight information includes:
[0088] The side length parameters corresponding to each visually labeled object are determined based on a dynamically adjusted weight index factor.
[0089] Determine the sum of the side length parameters of all visually labeled objects in the image to be identified;
[0090] The weight information of a visually labeled object is determined based on the ratio of its corresponding side length parameter to the sum of its side length parameters.
[0091] The weighting index factor is obtained by dynamically adjusting the preset base index based on the side length difference index of each visual marker object in the image to be identified.
[0092] It should be noted that the distribution ratio between various weights can be adjusted by dynamically adjusting the weight index factor to achieve adaptive weight allocation based on the side length of QR code detection. Specifically, when the side lengths of the visual marker objects differ significantly, the weight index factor is increased to widen the distribution gap between weights; conversely, when the side lengths of the visual marker objects are similar, the weight index factor is decreased to reduce the distribution gap between weights and balance the contribution of each weight to pose calculation.
[0093] For example, assuming the current frame detects N AprilTags (visual marker objects), the average detection side length of the i-th QR code in the image (e.g., the average of the four sides of an AprilTag) is: Then its weight can be defined. for:
[0094] ;
[0095] in For dynamically adjusted weighting index factors ( >0), used to control the non-linear relationship between side length and weight, which can be dynamically adjusted according to the side length distribution (the side length difference index of each visually labeled object):
[0096] ;
[0097] in, The preset base index can be initially set to 1.0; k is the adjustment coefficient, which is usually set to an empirical value of 1.0. This represents the index indicating the difference in side lengths of each visually labeled object in the image to be identified, where The side length of all visually marked objects in this frame The standard deviation of the side lengths of all visually labeled objects. The mean value. When the side lengths of different visually labeled objects differ significantly, the side length difference index increases accordingly.
[0098] Based on this, by introducing a weighting factor determined by the side length of the QR code into an adaptive weighting mechanism, the QR code, which is closer and has higher detection accuracy, can play a greater role in pose estimation, thereby further improving the overall solution accuracy.
[0099] In some embodiments, the actual three-dimensional coordinate information of each visual marker object is obtained in the following ways:
[0100] Obtain the geographic location information of the four corner points of each visually marked object; the geographic location information includes latitude, longitude and altitude.
[0101] The local coordinate system pose information of the visually marked object relative to the Earth coordinate system is determined based on the geographical location information of each corner point in the visually marked object; wherein, the local coordinate system pose information is represented by the parameters of the first transformation matrix corresponding to the Earth coordinate system;
[0102] A unified coordinate system is constructed based on the aircraft's takeoff position as the origin and the attitude information of the origin relative to the Earth coordinate system; the attitude information includes the aircraft's yaw angle and latitude and longitude information at the takeoff position.
[0103] The pose information of each visual marker object in the local coordinate system is transformed into a unified coordinate system to obtain the actual three-dimensional coordinate information of each visual marker object; wherein, the actual three-dimensional coordinate information is represented by the second transformation matrix parameter corresponding to the unified coordinate system.
[0104] It should be noted that the actual three-dimensional coordinate information of each visual marker object can be obtained offline, that is, coordinate system transformation is performed in advance after the visual marker objects are installed and placed and before the drone takes off.
[0105] For example, the actual three-dimensional coordinate transformation of this application embodiment is mainly divided into two steps: 1. Representing the local coordinate information corresponding to each visual marker object based on the transformation matrix parameters of the global coordinate system; 2. Transforming the local coordinate information corresponding to each visual marker object to a unified coordinate system with the UAV as the origin, that is, using the transformation matrix parameters of the unified coordinate system to represent the local coordinate information corresponding to each visual marker object.
[0106] For example, in step 1, by reading the latitude and longitude information (including latitude and longitude position and altitude) of the four corner points of each AprilTag, the system can automatically calculate the planar orientation, normal direction, and scale of the tag in the Earth coordinate system, thereby automatically generating its three-dimensional local coordinate system and attitude parameters. This method removes the limitations of position, angle, or number of AprilTags, and the system can automatically adapt to the coordinate system under different installation methods (ground, wall, column, etc.), greatly improving the flexibility and automation of the layout, and effectively improving the versatility and deployment efficiency of the system.
[0107] Specifically, for each AprilTag (visual tag object), the geographical locations of its four corner points are recorded in the form of latitude, longitude, and altitude. The geographical location information of the four corner points is recorded as follows: , i =1, 2, 3, 4; where the three parameters represent longitude, latitude, and altitude, respectively. First, the coordinates of the above geographical location information are converted into three-dimensional coordinates in the Earth-Centered Earth-Fixed (ECEF) coordinate system (i.e., the Earth coordinate system). Then, centered on AprilTag Construct a local coordinate system of East-North-Sky (ENU) with the origin as the origin. The transformation matrix is determined by the coordinates of the center point:
[0108] ;
[0109] At this point, all four corner points are mapped to the local coordinate system of the AprilTag center. Finally, based on the positional relationship of the four corner points in the local ENU coordinate system, the plane and orientation of the AprilTag are defined: Let the first and second corner points form the x-axis direction vector:
[0110] ;
[0111] Let the first and third corner points determine the z-axis of the plane normal vector:
[0112] ;
[0113] Calculate the y-axis direction using a right-handed coordinate system: Finally, the rotation matrix of the AprilTag local coordinate system is obtained:
[0114] ;
[0115] Its translation vector is the coordinate of the center point of AprilTag. .
[0116] At this point, the parameters of the first transformation matrix (rotation matrix and translation vector) can be determined as follows: Therefore, the pose of each AprilTag in the local ENU coordinate system, that is, the pose information of the visually labeled object relative to the Earth coordinate system in the local coordinate system, can be represented as:
[0117] ;
[0118] For example, in step 2, when the camera detects multiple visual marker objects simultaneously, to facilitate pose calculation, the local coordinate information of each visual marker object can be transformed to a unified coordinate system so that information between different QR codes can be fused. This coordinate system transformation is based on the known installation position of the QR code in Earth coordinates, achieving global alignment of camera observation information.
[0119] Specifically, the takeoff location of the drone can be converted from latitude and longitude to Earth coordinates. Then, this point is set as the origin in a unified coordinate system. Assuming the yaw angle of the UAV at the takeoff position is yaw, the rotation matrix of the origin attitude relative to the Earth coordinate system is:
[0120] ;
[0121] At this point, the rotation of each AprilTag from the drone's origin can be obtained as follows: .
[0122] Therefore, the position of each AprilTag relative to the drone's origin is: .
[0123] Therefore, the parameters of the second transformation matrix (rotation matrix and translation vector) can be determined as follows: Therefore, the pose of each AprilTag in a unified coordinate system, i.e., the actual 3D coordinate information of each visually labeled object, can be represented as:
[0124] ;
[0125] Based on this, by establishing a local coordinate system corresponding to each visual marker object based on the Earth coordinate system, and then uniformly representing the local coordinate system corresponding to each visual marker object based on a unified coordinate system with the aircraft's takeoff position as the origin, visual markers can be placed at any location and angle without the need for a unified plane or a specific viewpoint. The system can automatically establish a unified coordinate system based on the geographic coordinates and attitude of each visual marker label, realizing automatic registration and fusion of multiple labels, greatly simplifying the on-site deployment and calibration process, and improving the system's adaptability and deployment efficiency.
[0126] In some embodiments, acquiring the image to be identified captured by a camera on the aircraft includes:
[0127] Acquire multiple images to be identified from the multi-view fisheye camera on the aircraft.
[0128] It should be noted that, in order to solve the problem that a single camera's viewpoint cannot cover the entire field of view and may fail to detect the AprilTag, thus causing positioning failure, a multi-view fisheye camera can be used for image acquisition.
[0129] For example, fisheye cameras can be installed in the front, back, left, and right directions of the drone to form an omnidirectional vision system, achieving 360° detection without blind spots and effectively compensating for blind spots and occlusion problems existing in a single viewpoint. Even if the drone's attitude changes or it is in a complex environment, the system can still continuously observe multiple QR code targets, thereby effectively avoiding positioning interruption problems caused by field of view occlusion.
[0130] Based on this, the reliability of pose determination is further improved by using a multi-eye fisheye camera to acquire multiple images for coordinate information fusion and pose determination.
[0131] In some embodiments, before solving for the aircraft pose, the method further includes:
[0132] Based on the preset fisheye camera calibration model, the original two-dimensional coordinate information corresponding to each visual marker object is corrected to obtain the corrected two-dimensional coordinate information.
[0133] It should be noted that, due to the inherent distortion effect of fisheye cameras, this embodiment performs distortion correction processing on the original two-dimensional coordinate information corresponding to each detected visual marker object before solving the pose, obtaining corrected two-dimensional coordinate information, which is then used as input data for pose solving. By employing a preset fisheye camera calibration model to perform geometric correction on corner points, the accuracy and consistency of point pair matching are improved.
[0134] For example, the original two-dimensional coordinate information corresponding to each visually marked object can be distorted based on the DS (Double Sphere) fisheye distortion model, and the correction formula is as follows:
[0135] ;
[0136] ;
[0137] ;
[0138] ;
[0139] ;
[0140] ;
[0141] The input parameters (known pixel coordinates of each point and camera intrinsic parameters) include:
[0142] (u, v) represents the pixel coordinates on the original fisheye image, which are the distortion points that need to be corrected;
[0143] f x , f y These represent the camera's focal length in the x and y directions, respectively, and describe the scaling ratio of the camera sensor.
[0144] c x , c y This indicates the principal point coordinates of the camera, which is usually close to the center of the image and is the intersection of the optical axis and the imaging plane.
[0145] The parameters of the DS distortion model include:
[0146] alpha, represents the second spherical scaling parameter, which usually ranges between (0, 1). It controls the continuous transition of the projection model between the pinhole model and the orthographic projection model and is one of the key parameters that determine the distortion shape and the effective field of view (FOV).
[0147] x i The first spherical offset parameter defines the offset of the center of the first virtual sphere relative to the origin of the camera coordinate system. alpha Together, they can accurately model field of view angles exceeding 180°;
[0148] Intermediate calculation variables include:
[0149] and The pixel coordinates are normalized. By subtracting the principal point and dividing by the focal length, the pixel coordinates are transformed into physical imaging plane coordinates in meters. However, the z coordinate is normalized to 1 (corresponding to the ideal pinhole model).
[0150] r² It is the sum of squares of the normalized coordinates of a pixel, that is, the square of the distance from that point to the origin of the normalized plane;
[0151] As an intermediate quantity in the model, its calculation formula is the core of the DS model. Essentially, it represents, in the double-sphere refraction path model, the point after the action of the second sphere, and its relationship with the normalized plane point (…). m x , m y The associated depth correlation quantity directly affects the geometric relationship of the final back projection.
[0152] The output includes:
[0153] and For the point after distortion removal, the X and Y coordinate components on the unit sphere.
[0154] Based on this, the accuracy of pose calculation is further improved by first performing distortion correction on the image before pose calculation.
[0155] In some embodiments, both the two-dimensional coordinate information and the actual three-dimensional coordinate information include the coordinates of multiple corner points corresponding to the visually marked object.
[0156] It should be noted that during the pose determination process, the coordinate information of each visual marker object can be represented by the coordinates of its multiple corner points. For example, each visual marker object can be represented and used in the calculation by the coordinates of its four corner points. It can be understood that the two-dimensional coordinate information of each point corresponds to an actual three-dimensional coordinate information, and their pairing relationship can be constructed through step S2.
[0157] Based on this, by using multiple corner coordinates to represent the two-dimensional and three-dimensional coordinates corresponding to each marker, the reliability of pose solving is further improved.
[0158] It should be noted that, through the introduction of adaptive weights, the system in this application embodiment can achieve robust fusion localization in the presence of multiple visually labeled objects, and finally output the precise pose of the UAV in a unified coordinate system.
[0159] For example, assuming a drone using four fisheye cameras for localization, when N AprilTags are detected in a frame of the four fisheye camera images, the pose can be solved using the following objective function:
[0160] ;
[0161] in, and This indicates the aircraft pose to be solved. The projection parameters of the j-th camera in a unified coordinate system are calculated from the extrinsic parameters of the j-th camera (c represents camera in the formula, with a value from 1 to 4);
[0162] For the j-th camera, the first... i Adaptive weights for each visually labeled object;
[0163] For the j-th camera i The two-dimensional coordinates of the kth corner point of a QR code in the image;
[0164] For this corner point (and) The known spatial coordinates (actual three-dimensional coordinate information) of the corresponding corner points in a unified coordinate system.
[0165] Let be the reprojection calculation function, representing the rotation matrix of the j-th camera used to calculate the corner points. Translation vector (Projection parameters) Projecting three-dimensional spatial points onto the image plane yields theoretical two-dimensional coordinates.
[0166] Finally, by minimizing the above objective function using a nonlinear least squares algorithm (such as the Levenberg-Marquardt algorithm), the precise pose of the UAV in a unified coordinate system is obtained:
[0167] ;
[0168] It should be noted that this embodiment of the application achieves high-precision, continuous, and stable positioning of UAVs in complex environments by adaptively weighting arbitrarily arranged AprilTags and combining them with an omnidirectional fisheye vision system. This method not only possesses strong robustness and high precision but also has significant advantages such as flexible arrangement, strong scalability, and ease of deployment. It can be widely applied to scenarios such as indoor and outdoor UAV inspection, substation equipment testing, and autonomous navigation of unmanned systems, demonstrating good engineering application value and promising prospects for widespread adoption.
[0169] Please refer to Figure 2 , Figure 2 A block diagram illustrating the composition of a positioning device for an aircraft according to some embodiments of this application is shown. It should be understood that the positioning device for this aircraft is similar to that described above. Figure 1 Corresponding to the method embodiments, it is able to perform each step involved in the above method embodiments. The specific functions of the positioning device of the aircraft can be found in the description above. To avoid repetition, detailed descriptions are appropriately omitted here.
[0170] Figure 2 The positioning device for the aircraft includes at least one software function module that can be stored in a memory or embedded in the positioning device of the aircraft in the form of software or firmware. The positioning device for the aircraft includes:
[0171] The marker recognition module 210 is used to acquire the image to be recognized captured by the camera on the aircraft, and to identify at least one visual marker object in the image to be recognized;
[0172] The matching and association module 220 is used to match and associate the two-dimensional coordinate information corresponding to each visual marker object with the actual three-dimensional coordinate information based on the identification information of each visual marker object; wherein, each actual three-dimensional coordinate information is obtained based on a preset coordinate transformation rule and represented by a unified transformation matrix parameter;
[0173] The pose solving module 230 is used to construct and solve the objective function of the two-dimensional coordinate information, the actual three-dimensional coordinate information and the pose of the aircraft to be solved, and obtain the pose information of the aircraft based on the solution results.
[0174] It is understood that the above-described device embodiments correspond to the method embodiments of the present invention. The aircraft positioning device provided by the embodiments of the present invention can implement the aircraft positioning method provided by any one of the method embodiments of the present invention.
[0175] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the aforementioned method, and will not be elaborated further here.
[0176] like Figure 3 As shown, some embodiments of this application provide an electronic device 300, which includes a memory 310, a processor 320, and a computer program stored in the memory 310 and executable on the processor 320. When the processor 320 reads the program from the memory 310 via a bus 330 and executes the program, it can implement any of the methods included in the above-described aircraft positioning method.
[0177] Processor 320 can process digital signals and may include various computing architectures. For example, it may be a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements multiple instruction set combinations. In some examples, processor 320 may be a microprocessor.
[0178] The memory 310 can be used to store instructions executed by the processor 320 or data related to the execution of instructions. These instructions and / or data may include code used to implement some or all of the functions of one or more modules described in the embodiments of this application. The processor 320 of this disclosure embodiment can be used to execute the instructions in the memory 310 to implement the methods shown above. The memory 310 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memories well known to those skilled in the art.
[0179] Some embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, describes the method described in the method embodiments.
[0180] Some embodiments of this application also provide a computer program product that, when run on a computer, causes the computer to perform the method described in the method embodiments.
[0181] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0182] It should be understood, in the several embodiments provided in this application, that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0183] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0184] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0185] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0186] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0187] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A method of positioning an aircraft, characterized in that, The method comprises: acquiring an image to be identified collected by a camera on an aircraft, and identifying at least one visual marker object in the image to be identified; based on the identification information of each visual marker object, matching and associating the two-dimensional coordinate information corresponding to each visual marker object with the actual three-dimensional coordinate information, wherein each actual three-dimensional coordinate information is obtained based on a preset coordinate conversion rule and is represented by a unified transformation matrix parameter; constructing a target function about the two-dimensional coordinate information, the actual three-dimensional coordinate information, and the aircraft pose to be solved, and solving, and obtaining the aircraft pose information according to the solving result; the method of constructing a target function about the two-dimensional coordinate information, the actual three-dimensional coordinate information, and the aircraft pose to be solved, and solving, and obtaining the aircraft pose information according to the solving result comprises: determining the theoretical two-dimensional coordinate information corresponding to each actual three-dimensional coordinate information according to a preset projection parameter, wherein the projection parameter is constructed by a preset camera extrinsic parameter and the aircraft pose to be solved; minimizing the sum of the differences between each two-dimensional coordinate information and the corresponding theoretical two-dimensional coordinate information as the target, solving the aircraft pose to be solved, and obtaining the aircraft pose information according to the solving result.
2. The method of positioning an aircraft according to claim 1, wherein, the method of constructing a target function about the two-dimensional coordinate information, the actual three-dimensional coordinate information, and the aircraft pose to be solved, and solving, and obtaining the aircraft pose information according to the solving result comprises: combining the weight information corresponding to each visual marker object, constructing a target function about the two-dimensional coordinate information, the actual three-dimensional coordinate information, and the aircraft pose to be solved, and solving, and obtaining the aircraft pose information according to the solving result; wherein the weight information is determined based on the graphic size of each visual marker object in the image to be identified.
3. The method of positioning an aircraft according to claim 2, wherein, the determination method of the weight information comprises: determining the side length parameter corresponding to each visual marker object based on a dynamically adjusted weight exponential factor; determining the sum of the side length parameters of all visual marker objects in the image to be identified; determining the weight information of the visual marker object based on the ratio of the side length parameter corresponding to the visual marker object to the sum of the side length parameters; wherein the weight exponential factor is obtained by dynamically adjusting a preset basic exponential based on the side length difference index of each visual marker object in the image to be identified.
4. The method of positioning an aircraft of claim 1, wherein, the acquisition method of the actual three-dimensional coordinate information of each visual marker object comprises: acquiring the geographic position information of four corner points in each visual marker object; wherein the geographic position information includes latitude, longitude, and altitude; determining the local coordinate system pose information of the visual marker object relative to the earth coordinate system based on the geographic position information of each corner point in the visual marker object; wherein the local coordinate system pose information is represented by a first transformation matrix parameter corresponding to the earth coordinate system; taking the takeoff position of the aircraft as the origin, constructing a unified coordinate system based on the pose information of the origin relative to the earth coordinate system; wherein the pose information includes the yaw angle and the latitude and longitude information of the aircraft at the takeoff position. The local coordinate system pose information corresponding to each visual marker object is converted into the unified coordinate system to obtain actual three-dimensional coordinate information of each visual marker object, wherein the actual three-dimensional coordinate information is represented by second transformation matrix parameters corresponding to the unified coordinate system.
5. The method of positioning an aircraft of claim 1, wherein, The method comprises: The method comprises:
6. The method of positioning an aircraft according to claim 5, wherein, Before solving the pose of the aircraft, the method further comprises: The original two-dimensional coordinate information corresponding to each visual marker object is corrected based on a preset fish-eye camera calibration model to obtain corrected two-dimensional coordinate information.
7. The method of positioning an aircraft according to any one of claims 1 to 6, characterized in that, The two-dimensional coordinate information and the actual three-dimensional coordinate information each comprise a plurality of corner point coordinates corresponding to the visual marker object.
8. A positioning device for an aircraft, characterized in that The method comprises: The marker identification module is configured to acquire a to-be-identified image captured by a camera on the aircraft and identify at least one visual marker object in the to-be-identified image; The matching and correlating module is configured to match and correlate the two-dimensional coordinate information and the actual three-dimensional coordinate information corresponding to each visual marker object based on the identification information of each visual marker object, wherein each actual three-dimensional coordinate information is obtained based on a preset coordinate conversion rule and represented by a unified transformation matrix parameter; The pose solving module is configured to construct a target function about the two-dimensional coordinate information, the actual three-dimensional coordinate information, and the to-be-solved aircraft pose and solve the target function to obtain the aircraft pose information according to a solving result; The pose solving module is specifically configured to: determine theoretical two-dimensional coordinate information corresponding to each actual three-dimensional coordinate information according to a preset projection parameter, wherein the projection parameter is constructed from a preset camera external parameter and the to-be-solved aircraft pose; solve the to-be-solved aircraft pose by taking the sum of differences between each two-dimensional coordinate information and corresponding theoretical two-dimensional coordinate information as a target, and obtain the aircraft pose information according to a solving result.
9. An aircraft, characterized in that The computer readable storage medium stores a computer program, and the computer program is run on the processor to implement the positioning method of the aircraft according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is run on the processor to implement the positioning method of the aircraft according to any one of claims 1-7.
11. A computer program product, characterised in that, The computer program product comprises a computer program, and the computer program is run on the processor to implement the positioning method of the aircraft according to any one of claims 1-7.
Citation Information
Patent Citations
Pose estimation method and device and electronic equipment
CN120125648A