High-robust visual positioning and action prediction method and device based on adaptive prediction

By employing a central LED transmissive double-layer target, global optimal stereo matching, and dynamic process noise Kalman filtering algorithm, the problems of target recognition error and state estimation lag in visual guidance schemes under high dynamic environments were solved, achieving high-precision, all-weather UAV positioning and motion prediction, and ensuring safe docking under high-speed relative motion conditions.

CN122265391APending Publication Date: 2026-06-23XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2026-03-11
Publication Date
2026-06-23

Smart Images

  • Figure CN122265391A_ABST
    Figure CN122265391A_ABST
Patent Text Reader

Abstract

This invention discloses a robust visual localization and motion prediction method and device based on adaptive prediction. It acquires two images of a target with a central LED-transmitted double-layer target at each moment and extracts the target center to obtain a target point set. Based on the target point set of the two images at each moment, a three-dimensional observation point set is determined. Using this point set, along with the target model point set and the target's position in the world coordinate system, the three-dimensional position of the target in the camera coordinate system at each moment is estimated, and filtered using an improved Kalman filter algorithm. Based on the velocity, acceleration, and filtered three-dimensional position at each moment, as well as the filtered three-dimensional position, velocity, and trajectory curvature at multiple historical moments, along with a constant acceleration model and a deep nonlinear residual network based on LSTM units, the three-dimensional position of the target in the future is predicted. This invention maintains high accuracy and robustness even under complex lighting, strong interference, and long time-delay environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of visual positioning technology, specifically relating to a highly robust visual positioning and motion prediction method and device based on adaptive prediction. Background Technology

[0002] With the rapid development of modern aerospace engineering, especially in high-precision missions such as UAV aerial handover, automatic carrier landing, and rendezvous and docking of spacecraft, the demand for high-precision relative positioning and control between aircraft is becoming increasingly urgent. In such missions, UAVs need to achieve continuous tracking, stable locking, and safe docking with the mother aircraft with centimeter-level precision under high-speed relative motion conditions. This poses extremely stringent challenges to the system's visual perception and closed-loop control capabilities. On the one hand, the flight environment is filled with uncontrollable unstructured interference, and lighting conditions fluctuate drastically with cloud cover, direct sunlight, and aircraft shadows, easily creating high dynamic range scenes that lead to local overexposure or blackout of visual images. On the other hand, there is a cumulative time delay of 900 milliseconds to 1 second in the entire closed-loop control loop from sensor exposure, algorithm processing, data transmission to actuator response. For high-speed maneuvering aircraft, this means that the control system often has to respond based on severely lagging "past" states. Without high-precision real-time perception and forward-looking prediction methods, it is easy to cause system oscillations, overshoot, or even catastrophic collisions.

[0003] To address the aforementioned challenges, existing visual guidance solutions primarily rely on passive planar target technology, such as ArUco QR codes, AprilTags, or ordinary circular targets. While these solutions are cost-effective, they exhibit significant limitations in practical engineering applications. First, passive targets depend entirely on ambient light reflection imaging. In nighttime, overcast skies, strong backlighting, or at long distances, image contrast drops sharply, and corner features are easily obscured by noise or blurred due to insufficient resolution, leading to a significant decrease in recognition rate and accuracy. More seriously, planar circular targets have an inherent geometric defect: "perspective projection eccentricity error." According to the principles of projective geometry, although the image of a spatial circle under perspective projection is an ellipse, the geometric center of this ellipse is not equivalent to the perspective projection point of the physical circle's center. When there is a large angle (oblique view) between the camera's optical axis and the target plane, this eccentricity error increases non-linearly with the angle, resulting in systematic deviations in pose calculations that cannot be eliminated through simple calibration.

[0004] At the algorithmic processing level, existing stereo matching and state estimation techniques are also insufficient to meet the demands of highly dynamic environments. Traditional stereo matching algorithms often rely on local searches based on epipolar constraints. In aviation contexts, interference such as cloud edges and aircraft textures can easily induce mismatches, leading to outliers in depth calculations and compromising positioning stability. Furthermore, the widely used standard Extended Kalman Filter (EKF) employs a fixed process noise covariance matrix, which presents a dilemma between "smoothness" and "response speed": if the parameter settings favor smoothness, the filter is sluggish in responding to sudden maneuvers, resulting in severe tracking lag; if the parameter settings favor responsiveness, it cannot effectively filter out sensor noise, leading to trajectory jitter. In addition, to compensate for system latency, existing technologies often use linear motion models such as constant velocity or constant acceleration to extrapolate the target's future position. However, within long-latency prediction windows, UAVs are affected by aerodynamic drag and attitude adjustments, resulting in highly nonlinear trajectories (such as S-shaped maneuvers). Linear models cannot describe this variable acceleration motion, causing the predicted trajectory to diverge along the tangential direction, ultimately leading to docking failure due to significant prediction errors. In summary, the targets used in the existing technology have eccentricity errors, and the related stereo matching and state estimation techniques are difficult to meet the needs of highly dynamic environments. Summary of the Invention

[0005] To address the aforementioned problems in the prior art, this invention provides a robust visual localization and motion prediction method and device based on adaptive prediction.

[0006] The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides a robust visual localization and action prediction method based on adaptive prediction, comprising: Two images of a target with multiple targets are acquired at the current moment, each image containing the image regions of the multiple targets; the outer layer of each target is a planar ring, the inner layer is a light-transmitting hole with the same center as the outer layer, and an actively emitting LED light source is set behind the light-transmitting hole; Image processing and target center extraction are performed on each image to obtain a target point set for each image, wherein the target point set contains multiple target center points; Global optimal stereo matching is performed on the target point set based on the two images at the current time, and the observed three-dimensional point set is determined. Then, using the observed three-dimensional point set and the target model point set in the world coordinate system and the position of the target, the three-dimensional position of the target in the camera coordinate system at the current time is determined based on the alternating optimization rigid body pose estimation method. An improved Kalman filter algorithm is used to filter the target's three-dimensional position at the current moment to obtain the filtered three-dimensional position of the target at the current moment. Based on the target's velocity, acceleration, and filtered 3D position at the current moment, the filtered 3D position, velocity, and trajectory curvature of the target at multiple historical moments at the current moment, as well as the constant acceleration model and a trained deep nonlinear residual network based on LSTM units, the 3D position of the target at future times is predicted.

[0007] The present invention also provides a highly robust visual localization and action prediction device based on adaptive prediction, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is used to store computer programs; When the processor executes the program stored in the memory, it implements the steps of the above-described highly robust visual localization and action prediction method based on adaptive prediction.

[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: 1) The target designed in this invention utilizes the imaging characteristics of the LED light source formed by the inner layer micropores to restrict the light source, and combined with the short exposure acquisition strategy, it completely eliminates the perspective projection eccentricity error of traditional planar targets at non-frontal viewing angles from a physical level. At the same time, the double-layer structure combines the advantages of the outer passive visual anchor point and the inner active anti-interference light source, effectively solving the recognition failure problem caused by strong light reflection, shadows and long-distance blurring, and greatly improving the positioning accuracy (error <1cm) and all-weather robustness of the system in extreme lighting environments.

[0009] 2) The pose calculation method based on global optimal association and alternating iteration designed in this invention fundamentally overcomes the shortcomings of traditional local matching algorithms, which are prone to mismatches in complex backgrounds (such as cloud edges) and partial occlusion. Furthermore, it achieves rapid convergence and high-precision output of pose calculation, ensuring the data purity of the system in noisy environments.

[0010] 3) The improved Kalman filter algorithm proposed in this invention mainly addresses the difficulty of balancing smoothness and response speed in the standard Kalman filter by designing an adaptive adjustment mechanism based on the Mahalanobis distance of the observation residuals. This mechanism can perceive the target's maneuvering intentions (such as sharp turns or accelerations) in real time. By nonlinearly adjusting the process noise covariance matrix, it reduces the trust in the old model during sudden maneuvers to achieve a zero-delay response, and enhances the model trust during stable flight to maintain trajectory smoothness, significantly improving the system's tracking stability for highly dynamic targets.

[0011] 4) This invention constructs a dual-stream fusion architecture of "physical dynamics baseline (i.e., constant acceleration model, abbreviated as CA model) + deep nonlinear residual network," and in particular, introduces trajectory curvature as a high-order dynamic feature input. This design uses the physical baseline to prevent prediction drift and uses the residual network to fit complex maneuvering errors. It specifically solves the problem that traditional linear models cannot predict nonlinear trajectories such as "S"-shaped maneuvers or right-angle turns under system control loop delays of up to 900ms-1000ms, effectively compensating for time lag and ensuring the safety of docking missions.

[0012] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Attached Figure Description

[0013] Figure 1 This is a comparison diagram of the target design provided in this embodiment of the invention with several other existing targets; Figure 2 This is a flowchart illustrating the highly robust visual localization and action prediction method based on adaptive prediction provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the system implementation scenario and binocular imaging geometry provided in the embodiments of the present invention; Figure 4 This is a schematic diagram of the triangulation principle provided in an embodiment of the present invention; Figure 5 This is an exemplary comparison diagram of the estimated trajectory and the actual trajectory provided in an embodiment of the present invention; Figure 6 This is another exemplary comparison diagram of the estimated trajectory and the actual trajectory provided by an embodiment of the present invention. Detailed Implementation

[0014] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0015] Existing related technologies have the following drawbacks: (1) Disadvantages of existing targets: They rely entirely on ambient reflected light for imaging, and are prone to failure due to overexposure, shadows, or blurring at long distances under complex lighting conditions such as aviation docking. More seriously, planar circular targets have an inherent "perspective projection eccentricity error", that is, when shooting from an oblique view, the center of the ellipse in the image is not the same as the perspective projection point of the physical center of the circle. This geometric error increases nonlinearly with the increase of the shooting angle (usually resulting in a positioning error greater than 30mm), and cannot be eliminated by conventional calibration.

[0016] (2) Disadvantages of existing matching and filtering algorithms: Traditional stereo matching algorithms are mostly based on local epipolar constraints, which are prone to mismatches when the background texture is messy (such as cloud edges) or there is occlusion, resulting in outliers in depth calculation. At the same time, the widely used standard extended Kalman filter (EKF) adopts a process noise model with fixed parameters, which makes it difficult to balance smoothness and response speed. If the parameter settings are biased towards smoothness, the filter will be sluggish in response to the sudden maneuvers of the UAV (such as sharp turns), resulting in serious tracking lag; if the settings are biased towards response, it will not be able to effectively filter out noise.

[0017] (3) Existing shortcomings in trajectory prediction and time delay compensation: In the context of cumulative time delays of 900ms to 1 second in the airborne control loop, existing linear prediction models such as constant velocity (CV) or constant acceleration (CA) cannot describe the complex dynamic characteristics of UAVs under the influence of aerodynamic drag and attitude adjustment. Within the long-delay prediction window, the extrapolated trajectory of the linear model often diverges along the tangential direction and cannot fit nonlinear trajectories such as "S"-shaped maneuvers or spiral ascents, causing the prediction error to accumulate and fail over time.

[0018] To address the aforementioned shortcomings, this invention proposes a central LED transmissive dual-layer target structure. Utilizing the high penetration of an active light source, it solves the environmental adaptability problem. Based on the imaging characteristics of a "transmissive point light source," it completely eliminates perspective eccentricity errors at the physical level, achieving centimeter-level high-precision positioning (error <1cm) under all weather and all-angle conditions. This invention proposes a global optimal matching strategy based on the Hungarian algorithm and Dynamic Process Noise Kalman Filtering (DPN-KF). It solves the mismatch problem through global optimization and uses observation residuals to adaptively adjust model confidence, achieving zero-delay response and stable tracking of highly maneuverable targets. This invention proposes a constrained perception neural prediction model (N-TDMNet), constructing a dual-layer architecture of physical baseline + deep residual network. By introducing trajectory curvature features, it accurately learns and predicts the nonlinear maneuvering intentions of UAVs, effectively compensating for perception lag caused by long system latency and ensuring the safety of docking missions.

[0019] First, in order to solve the problems of traditional passive targets failing under strong light and "perspective eccentricity error" at non-frontal viewing angles, this invention designs a dual-layer composite target that combines active and passive approaches. It can be used with an efficient recognition algorithm based on morphological and geometric constraints to achieve all-weather, high-precision feature extraction.

[0020] The target proposed in this invention adopts a concentric double-layer structure of "outer passive color ring + inner active light source". Specifically, the outer layer is designed as a planar circular ring of a high-saturation single color (such as red), mainly used to provide a stable visual anchor point and region of interest (ROI) in a wide field of view. The inner layer is a light-transmitting hole located at the center of the outer ring, and a high-brightness LED active light source is placed behind the light-transmitting hole. The mechanism of the target anti-eccentricity imaging proposed in this invention is as follows: using a pinhole camera imaging model, the inner LED, through the micro-hole, is optically approximated as a physical point light source. Regardless of the angle between the camera optical axis and the target normal, the point light source always strictly corresponds to the perspective projection point of the physical center of the image plane. This design physically eliminates the geometric center deviation (eccentricity error) caused when a planar circular target is imaged as an ellipse, ensuring sub-pixel-level positioning reference. For example, Figure 1 This is a comparison diagram of the target design of this invention with several other existing target designs. Specifically, Figure 1 In the diagram, a, b, and c represent the simulated and actual shapes of a single circular target, a large circular target enclosing a small circular target, and a central LED-transmitting double-layer target (the target proposed in this invention), respectively. Figure 1 The shapes shown above a, b, and c represent simulated shapes of a single circular target, a large circular target containing a small circular target, and a double-layered target with a central LED transmission, respectively. Figure 1 The figures shown below a, b, and c represent the actual shapes of a single circular target, a large circular target containing a small circular target, and a double-layered target with a central LED transmission, respectively.

[0021] Based on the target designed above, the present invention also provides a robust visual localization and action prediction method based on adaptive prediction. Figure 2 This is a flowchart illustrating the highly robust visual localization and action prediction method based on adaptive prediction provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the method includes: S101. Acquire two images of a target with multiple targets at the current time. Each image contains image regions of multiple targets. The outer layer of each target is a planar ring, and the inner layer is a light-transmitting hole with the same center as the outer layer. An actively emitting LED light source is set behind the light-transmitting hole.

[0022] It should be noted that the target can be a moving target in space, such as an aircraft. Multiple targets are present at different locations on the target. The device (referred to as the tracking device) implementing the above-described highly robust visual localization and motion prediction method based on adaptive prediction acquires images of the target at each time step using its binocular camera, thus obtaining two images at each time step, and each image contains these target patterns. For example, Figure 3This is a schematic diagram of the system implementation scenario and binocular imaging geometry of the aforementioned tracking device. To clearly illustrate the target image, Figure 3 The aircraft in the image only shows a target pattern, such as Figure 3 As shown, the target pattern has an outer passive color ring and a central LED.

[0023] S102. Perform image processing and target center extraction on each image to obtain a target point set for each image, which contains multiple target center points.

[0024] Specifically, S102 is achieved through steps S1021 to S1024: S1021. After converting each image to the HSV color space, candidate regions are extracted using a preset HSV threshold and mask operation to obtain multiple candidate regions for each image. Each candidate region contains an image region of a target.

[0025] Here, the preset HSV threshold can be set according to actual needs. The number of candidate regions extracted from each image is determined by the number of targets contained in that image. Furthermore, due to the use of masking operations, each candidate region is a binary image. It should be noted that this candidate region extraction step utilizes the stability of chromaticity information to effectively reduce the impact of changes in ambient light intensity on the initial recognition.

[0026] S1022. Perform morphological noise suppression and contour restoration on each candidate region of each image to obtain multiple contour-restored candidate regions for each image.

[0027] Here, for the noise and contour breaks present in each image, a "dilation-erosion" morphological processing is performed. That is, for each candidate region, a dilation operation is first performed to fill the breaks or small holes at the edge of the target and enhance the connectivity of the contour. Then, an erosion operation is performed to remove isolated background noise points and restore the true contour shape of the target.

[0028] S1023. Calculate the ratio between the area of ​​the circumcircle of the contour of each candidate region after contour repair and the actual area, and filter out the candidate regions after contour repair whose ratio does not meet the preset conditions.

[0029] Here, for each candidate region after contour repair, the area of ​​the circumcircle of the contour of that candidate region is first calculated. And the actual area of ​​the contour of the candidate region after contour repair. Next, the candidate regions after contour repair are calculated. and The ratio between ,Right now When the candidate region after contour repair Meet the preset conditions (i.e.) When (considering tolerance range), if the contour of the candidate region after contour repair is determined to be a circular structure, then the candidate region after contour repair is retained; otherwise, when the contour of the candidate region after contour repair is not... If the preset conditions are not met, the shape of the candidate region after contour repair is determined to be abnormal; therefore, the candidate region after contour repair is discarded. It should be noted that if... If the difference between 1 and 1 is within the set tolerance range, then it indicates that... .

[0030] S1024. In each remaining candidate region after contour restoration, extract the bright LED region. When the center of the extracted bright LED region coincides with the center of the candidate region after contour restoration, determine a target center point based on the pixel coordinates of the extracted bright LED region. After traversing each remaining candidate region after contour restoration, obtain the target point set for each image.

[0031] To avoid numerous false positives and improve robustness, a dual validity determination is introduced. Specifically, for each retained candidate region after contour restoration, the bright LED region is first segmented using an adaptive threshold. If the geometric center of the segmented bright LED region coincides with the geometric center of the candidate region after contour restoration, it indicates that the bright LED region is located inside the candidate region, thus eliminating interference from background reflections and stray light sources. Next, the average value of all pixel coordinates in the bright LED region is calculated, and this average value is used as the center point of a target corresponding to the candidate region after contour restoration to alleviate positioning errors caused by uneven local brightness. Thus, for each image, after calculating the center point of a target corresponding to each candidate region after contour restoration, these target center points constitute the target point set for the image.

[0032] S103. Perform global optimal stereo matching based on the target point set of the two images at the current time, determine the observed three-dimensional point set, and use the observed three-dimensional point set and the target model point set in the world coordinate system to determine the target's three-dimensional position in the camera coordinate system at the current time based on the alternating optimization rigid body pose estimation method.

[0033] Specifically, S103 is achieved through steps S1031 to S1034: S1031. Based on the point-to-point matching algorithm and the outlier removal algorithm, perform global optimal stereo matching on the target point sets of the two images at the current time to obtain multiple pairs of matching points.

[0034] Let the target point sets of the two images at the current time be respectively and The purpose of optimal stereo matching is to solve... and The optimal pair of matching points between them, such that any pair of matching points The following dual constraints must be met: a. Epipolar constraint: According to the geometry of a standard binocular system, the difference between corresponding points on the image's vertical coordinate (y-axis) should approach zero; b. Disparity consistency constraint: Considering that the target is fixed on a rigid surface and its depth changes continuously, the disparity (x-axis coordinate difference) between matching point pairs should have statistical consistency.

[0035] Based on this, first utilize and Construct the cost matrix , of which elements Indicates the matching cost: ,in, , They are respectively The first in points and The first in The ordinates of the points; for The first in points and The first in The disparity of each point, and , They are respectively The first in points and The first in The x-coordinates of the points; Let be the median or cluster center of the disparity set, where the disparity set is... The points in The set of all parallaxes between all points in the array; , These are weighting coefficients, which can be set according to actual needs. This indicates taking the absolute value. Next, the Hungarian Algorithm is used to solve for the cost matrix. The problem of global minimum weight matching is solved by combining the Z-score statistical method to remove outliers with abnormal distances, thus ensuring the robustness of point pair association and finally obtaining multiple pairs of matching points.

[0036] S1032. Based on multiple pairs of matching points and the camera's intrinsic and extrinsic parameters, determine the observed three-dimensional point set.

[0037] Here, the camera's intrinsic parameters include the camera's focal length. The camera's extrinsic parameters include the camera's baseline length. For each pair of matching points, combined with the camera's focal length... and baseline length ,use Figure 4 The triangulation principle shown can be used to recover the coordinates of a 3D point P corresponding to this pair of matching points. Ideally, the coordinates of the 3D point P... The calculation formulas are as follows: ,in, , and They are and , and These represent the x-coordinates of the pair of matching points. and These represent the pairs of matching points that belong to the target point set. The x and y coordinates of the points in the graph. After obtaining the coordinates of a 3D point corresponding to each matched point, a set of points consisting of these 3D coordinates can be obtained, that is, the observed 3D point set. .

[0038] S1033. When iteratively solving the pose at the current moment, the pose at the previous moment is used as the initial pose at the current moment. Each time, the target model point set in the world coordinate system is projected using the pose obtained from the previous iteration. Then, a point pair matching algorithm is used to find matching point pairs between the projected point set and the observed 3D point set. Based on the matching point pairs and the Kabsch algorithm, the current optimal pose is solved to obtain the pose obtained in this iteration. If the change in the pose obtained in this iteration is less than or equal to the preset pose threshold, the pose obtained in this iteration is used as the pose at the current moment. If the change in the pose obtained in this iteration is greater than the preset pose threshold, the iterative solution of the pose continues until the pose at the current moment is obtained.

[0039] Here, a target model point set in the world coordinate system is pre-constructed. It should be noted that this target model point set... It is the set of three-dimensional coordinates of the center points of each target within the target array, in the world coordinate system. Therefore, the observed three-dimensional point set... Compared with the known set of real target model points in the world coordinate system Registration is a typical rigid body registration problem, and the purpose of rigid body registration is to solve for the pose (i.e., the rotation matrix). Translation vector Minimize the mean square error of reprojection: ,in, This indicates the calculation of the L2 norm. Considering the inefficiency of brute-force enumeration and the continuous motion of the target, this invention proposes a "prediction-correction" alternating optimization algorithm. This algorithm is used to solve for the pose at the current moment. The principle of this "prediction-correction" alternating optimization is as follows: a. Initialization: Take the pose obtained from the previous time step. Used as the initial pose when solving for the pose at the current moment.

[0040] b. Projection Matching: Obtain the pose required for this step, and use the obtained pose to match the model point set. Projecting is performed based on pose, and the Hungarian algorithm is used to find the projected point set and the observed 3D point set. The matching point pairs between them, wherein the pose required in this step is the initial pose in step a or the current optimal pose solved in step c.

[0041] c. Pose Update: Based on the matching point pairs obtained in step b, the current optimal pose is solved using the Kabsch algorithm (the covariance matrix after SVD decomposition).

[0042] d. Determine whether the change in the latest optimal pose obtained in step c compared to the previous optimal pose obtained in step c is less than the threshold. If not, return to step b above to continue iterating. If yes, use the latest optimal pose obtained in step c as the final high-precision pose obtained at the current moment.

[0043] S1034. Using the pose at the current moment, project the target's position in the world coordinate system onto the camera coordinate system to obtain the target's three-dimensional position at the current moment.

[0044] Here, in the world coordinate system, in addition to the pre-built target model point set... In addition, it also has the coordinates of a point representing the target's position in the world coordinate system. This point is a point on the target, and this point is called the target's position in the world coordinate system. Therefore, after obtaining the pose at the current moment, the target's position in the world coordinate system can be projected onto the camera coordinate system using this pose to obtain the target's three-dimensional position at the current moment.

[0045] S104. Use the improved Kalman filter algorithm to filter the target's three-dimensional position at the current time to obtain the filtered three-dimensional position of the target at the current time.

[0046] To address the motion characteristics of UAVs, which are primarily translational with occasional abrupt changes, this invention proposes a dynamic process noise Kalman filter algorithm (i.e., an improved Kalman filter algorithm). To overcome the lag problem of traditional Kalman filtering, the improved algorithm uses the Mahalanobis distance of the observation residuals at each time step as a maneuver detection index. This index is used to dynamically adjust the process noise covariance matrix at each time step, and the adjusted covariance matrix is ​​then used for filtering.

[0047] For example, in the improved Kalman filter algorithm, for the current time... Motor detection indicators and the adjusted process noise covariance matrix The calculation formulas are as follows: ; ; in, Indicates the current time Observation residuals Mahalanobis distance, It is a preset factor with a value greater than 1. Indicates the preset detection threshold. This represents the original process noise covariance matrix (i.e., the current time calculated using the improved Kalman filter algorithm). (process noise covariance matrix). Represents the innovation covariance matrix. express The inverse matrix. According to From the formula, we can see that when Greater than the threshold At that time, it was determined to be a rate mutation, through factor... Increase Forced filter to trust observations, enabling rapid response maneuvers. When Less than or equal to the threshold At that time, keep a small This invention utilizes the smoothing properties of filtering to suppress noise. The present invention references this dynamic filtering algorithm, which can further suppress sensor noise and handle sudden maneuvers.

[0048] S105. Based on the target's velocity, acceleration, and filtered 3D position at the current moment, the target's filtered 3D position, velocity, and trajectory curvature at multiple historical moments at the current moment, as well as the constant acceleration model and the trained deep nonlinear residual network based on LSTM units, predict the target's 3D position at future times.

[0049] Specifically, the above S105 is achieved through steps S1051 to S1053: S1051. Using the target's velocity, acceleration, and filtered three-dimensional position at the current moment, as well as the time interval between two adjacent moments, predict the target's position components in the future based on a constant acceleration model. .

[0050] Specifically, The expression is as follows: ; in, , , These represent the target at the current moment. The velocity, acceleration, and filtered 3D position. Indicates the current time Compared to the previous moment The time interval between them.

[0051] It should be noted that future time refers to the period between 900ms and 1000ms in the future.

[0052] S1052. Using the target's filtered 3D position, velocity, and trajectory curvature at the current moment, as well as the target's velocity, trajectory curvature, and filtered 3D position at multiple historical moments, a historical state sequence is constructed. Based on the historical state sequence and a trained deep nonlinear residual network based on LSTM units, the nonlinear residual of the target at future times is predicted. .

[0053] Here, the curvature of the target's trajectory at the current moment. The expression is: , This represents the magnitude of the vector.

[0054] Here, the historical state sequence is the current moment. of , and and the current moment The former The state sequence is composed of the filtered three-dimensional position, velocity, and trajectory curvature at each consecutive moment, and this state sequence is a 7-dimensional feature vector, which can be represented as: ,in, for , for The rest follows the same principle. It should be noted that the historical state sequence is input into a pre-trained deep nonlinear residual network based on LSTM units, which then predicts the nonlinear residuals of the target in future times. .

[0055] This invention presents a deep nonlinear residual network based on LSTM units, specifically designed to process flight trajectory data with long-term temporal dependencies. It aims to solve the gradient vanishing problem inherent in traditional recurrent neural networks (RNNs), thereby accurately capturing nonlinear positional deviations caused by UAV maneuvers (such as S-turns and spiral ascents) within a time window of time delay τ. The deep nonlinear residual network based on LSTM units comprises: an input layer, at least two LSTM layers, a fully connected layer, and an output layer. Specifically, the input to the input layer is a three-dimensional tensor with dimension τ. .in, Represents the time step, set to (i.e., the above) ), representing the historical length of the network "review". The representative feature dimension, as mentioned above. The dimensions of motion are defined by three-dimensional position and velocity, which provide the basic motion state, while trajectory curvature... As a high-order geometric constraint feature, it enables the network to sensitively perceive abrupt changes in the degree of trajectory curvature. For example, after the input layer, there are 2-3 stacked LSTM units (i.e., 2-3 stacked LSTM layers). The first LSTM layer is responsible for extracting shallow, short-term fluctuation features from the original high-dimensional data, outputting sequence features to the next layer; the second / third LSTM layers are responsible for integrating long-term temporal dependencies, remembering the long-term trend of trajectory evolution (such as the start and reverse phases of an "S"-shaped maneuver), and using the Dropout mechanism (dropout rate set to 0.2) to prevent overfitting and enhance the model's generalization ability. The hidden states of the LSTM layers at the last moment are fed into a fully connected layer. The fully connected layer typically contains two linear layers with a ReLU activation function in between, used to map the extracted high-dimensional temporal features back to three-dimensional Euclidean space, regressing the final nonlinear error vector. The output layer outputs a three-dimensional vector. This refers to the amount of correction to the predicted value of the physical baseline at the current moment.

[0056] It should be noted that, in order to effectively memorize motion features under long time delays, the LSTM unit used in this invention introduces a unique "cell state." Through the fine gating mechanism of this LSTM unit, the network can "lock" the feature in the cell state when the input sequence undergoes a sharp curvature change (such as the start of a turn), and release it through the output gate at the prediction time τ, thereby achieving accurate compensation for large-lag maneuvers. Since the introduction of a unique "cell state" LSTM unit is an existing technology, the specific working principle of such LSTM units will not be elaborated here. It should be noted that the loss function used when training the above-mentioned deep nonlinear residual network based on LSTM units is the mean squared error (MSE) loss function.

[0057] S1053, Based on the trajectory curvature and position components of the target at the current moment. Nonlinear residuals Based on historical state sequences, determine the target's three-dimensional location in the future. .

[0058] Specifically, The expression is: .

[0059] In this invention, the aforementioned deep nonlinear residual network based on LSTM units and the computational principle of S1051 described above are used to calculate... The modules constitute an N-TDMNet model for compensating for the long time delay τ caused by the overall perception and execution links of the system. The deep nonlinear residual network serves as the residual correction layer of the model, and the module serves as the physical baseline layer. This model utilizes trajectory curvature to keenly capture steering intentions, effectively correcting the tangential error of the linear model under "S"-shaped maneuvers.

[0060] The present invention also provides a robust visual localization and motion prediction device based on adaptive prediction, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When the processor executes the program stored in the memory, it implements the steps of the above-described highly robust visual localization and action prediction method based on adaptive prediction.

[0061] Compared with existing related technologies, the present invention has the following advantages: 1) This invention proposes a central LED transmissive dual-layer target structure: mainly a composite structure design of "outer layer high-saturation color ring + inner layer transmissive point light source". This invention utilizes the imaging characteristics of the inner layer micro-pores to confine the LED light source to form a physical point light source, combined with a short exposure acquisition strategy, to completely eliminate the perspective projection eccentricity error of traditional planar targets at non-frontal viewing angles from a physical perspective; at the same time, the dual-layer structure combines the advantages of the outer layer passive visual anchor point and the inner layer active anti-interference light source, effectively solving the recognition failure problems caused by strong light reflection, shadows and long-distance blurring, and significantly improving the positioning accuracy (error <1cm) and all-weather robustness of the system in extreme lighting environments.

[0062] 2) This invention proposes a pose calculation strategy based on global optimal association and alternating iteration: This invention designs a bipartite graph matching model that integrates epipolar constraints and disparity consistency constraints, and uses the Hungarian algorithm to solve for the global optimal solution, fundamentally overcoming the shortcomings of traditional local matching algorithms that are prone to mismatches in complex backgrounds (such as cloud edges) and partial occlusion. Furthermore, the proposed rigid body pose calculation method based on the Kabsch algorithm of the previous frame prior and alternating iteration of point-pair updates achieves fast convergence and high-precision output of pose calculation, ensuring the data purity of the system in noisy environments.

[0063] 3) This invention proposes a Dynamic Process Noise Kalman Filter (DPN-KF) ​​mechanism: This invention primarily addresses the challenge of balancing smoothness and response speed in standard Kalman filtering by designing an adaptive adjustment mechanism based on the Mahalanobis distance of the observation residuals. This mechanism can perceive the target's maneuvering intentions (such as sharp turns or accelerations) in real time. By nonlinearly adjusting the process noise covariance matrix, it reduces the trust in the old model during sudden maneuvers to achieve a zero-delay response, and enhances the model trust during stable flight to maintain trajectory smoothness, significantly improving the system's tracking stability for highly dynamic targets.

[0064] 4) This invention proposes a constraint-aware neural prediction model (N-TDMNet) architecture: This invention constructs a dual-stream fusion architecture of "physical dynamics baseline (CA model) + deep nonlinear residual network (LSTM / GRU)," specifically introducing trajectory curvature as a high-order dynamic feature input. This design utilizes the physical baseline to prevent prediction drift and uses the residual network to fit complex maneuvering errors. It specifically addresses the problem that traditional linear models cannot predict nonlinear trajectories such as "S"-shaped maneuvers or right-angle turns under system control loop delays of up to 900ms-1000ms, effectively compensating for time lag and ensuring the safety of the docking mission.

[0065] The technical effects of this invention are illustrated below with some experimental data: 1) Regarding positioning accuracy, this invention completely eliminates perspective eccentricity error, achieving millimeter-level high-precision measurement. Existing technologies commonly use single-planar circular targets, which inherently suffer from perspective projection eccentricity at non-frontal shooting angles. Experimental data shows that their average positioning error is as high as 36.677 mm, severely impacting docking safety. In contrast, the central LED transmissive double-layer target designed in this invention utilizes micro-holes in the inner layer to confine the light source, forming a physical point light source, ensuring consistency between the imaging center and the physical center at any tilted viewing angle. Simulation and physical experiments verify that this invention significantly reduces the average positioning error to 9.836 mm, improving accuracy by approximately 73.2%, successfully converging the measurement error to within 1 cm, providing a reliable sensing benchmark for high-precision aerospace docking.

[0066] 2) In terms of environmental adaptability, this invention demonstrates strong robustness, effectively solving the problem of target loss under complex lighting conditions. Traditional vision solutions (such as ordinary color rings or concentric circle targets) often suffer from low recognition rates, with an average precision of only 0.73, due to interference such as strong light reflection, shadow variations, and partial occlusion common in aviation scenarios. This invention, through a dual-constraint design of an "outer layer high-saturation color ring + inner layer active light source," combined with a short-exposure acquisition strategy, effectively suppresses background clutter. In a harsh environment test set containing strong light interference and partial occlusion, the system exhibits excellent stability in recognizing the central LED transmissive double-layer target, with both average precision and recall reaching 1.00, achieving "zero missed detections" in all weather conditions, significantly outperforming existing technologies. Table 1 below compares the average precision and average recall for different types of targets, showing that both are higher than the other two target types.

[0067]

[0068] Table 1 Comparison of average precision and average recall for different types of targets. 3) In terms of dynamic tracking, this invention eliminates phase lag under high-speed maneuvers, achieving precise tracking with zero latency. Traditional Extended Kalman Filter (EKF) is limited by fixed process noise parameters, often exhibiting significant lag in tracking trajectories when facing sharp turns or sudden speed changes by UAVs. The Dynamic Process Noise Kalman Filter (DPN-KF) ​​proposed in this invention introduces the Mahalanobis distance of the observation residual as a maneuver detection index, enabling real-time adaptive adjustment of the model confidence. As shown in Table 2, experimental comparisons demonstrate that DPN-KF maintains trajectory smoothness while rapidly responding to target maneuver changes. Its average dynamic tracking error is superior to traditional filtering schemes on all three experimental trajectories, effectively resolving the contradiction between "smoothness" and "response speed." Figure 5 shows the difference between the estimated trajectory (denoted by pred) and the actual trajectory (denoted by gt) for a certain dimension of motion on trajectory 3. It can be seen that the difference between the two is very small under translational and directional changes, reflecting the accuracy and low latency of the method proposed in this invention.

[0069]

[0070] Table 2 Comparison of the performance of each module in the pose optimization algorithm for dynamic process noise Kalman filtering 4) In terms of prediction and time delay compensation, this invention possesses superior capabilities in handling long-delay nonlinear trajectories, ensuring the safety of the control loop. Faced with long time delays of 900ms-1000ms in aviation control loops, traditional constant acceleration (CA) models cannot fit complex nonlinear maneuvers and are prone to generating significant overshoot errors at turns. The N-TDMNet model proposed in this invention introduces trajectory curvature features and utilizes a deep residual network to accurately fit nonlinear dynamic errors. As experimental visualization results confirm, when the UAV performs an "S" maneuver or a right-angle turn, the predicted trajectory of N-TDMNet closely matches the actual path, significantly reducing the root mean square error (RMSE), effectively compensating for system time delays, and giving the system the ability to "predict" the target's motion intentions. Figure 6 shows the difference between the predicted trajectory (denoted by pred) and the actual trajectory (denoted by gt) on a certain trajectory. It can be seen that the prediction algorithm is very stable during translation, but at the change of direction, the prediction lags behind the actual trajectory by about 900ms. This demonstrates that once the system detects a change of direction in the estimated trajectory, it immediately responds and changes the predicted direction. This rapid and accurate response capability is the advantage of the method proposed in this invention.

[0071] This invention has the following application prospects: With the rapid development of aerospace, unmanned systems, and intelligent manufacturing, the demand for high-precision perception and prediction of high-speed moving targets in unstructured environments is becoming increasingly urgent. The method proposed in this invention, leveraging the strong light resistance and zero eccentricity of its central LED transmissive double-layer target, and the accurate prediction capability of the N-TDMNet model for long-delay nonlinear trajectories, overcomes the technical bottlenecks of traditional visual navigation systems in harsh environments and highly maneuverable scenarios. This invention has extremely high engineering application value and broad market prospects, mainly reflected in the following areas: 1) Autonomous Aerial Docking and Formation Coordination for Unmanned Aerial Vehicles: Autonomous aerial docking is hailed as the "crown jewel" of aviation technology. Its core challenge lies in achieving centimeter-level precision physical connection between the docking aircraft and the target cone under high-speed airflow disturbances. Existing vision-based solutions are prone to target loss under strong sunlight reflection or cloud cover and cannot effectively handle the random swaying of the docking mechanism. The active luminescent target of this invention can penetrate strong light and thin fog, ensuring stable locking in all weather conditions. Simultaneously, the DPN-KF and N-TDMNet algorithms can predict the nonlinear swaying trajectory of the target cone one second in advance, effectively compensating for communication delays in the control system and significantly improving the success rate and safety of automatic aerial docking. This is a key supporting technology for enhancing the aerial recovery and mission expansion capabilities of future unmanned wingmen and strategic UAVs.

[0072] 2) Autonomous Landing and Recovery of Shipborne UAVs in High Sea States: When recovering UAVs on mobile maritime platforms (such as destroyers and aircraft carriers), the deck experiences severe six-degree-of-freedom swaying with the waves, and sea surface moisture, salt spray, and variable natural light significantly interfere with visual sensors. This invention effectively eliminates sea surface reflections and clutter interference through a globally optimal matching strategy using binocular vision; utilizing a nonlinear prediction model, the system can accurately predict the deck's heave and sway trend, guiding the UAV to land during a relatively stable window of opportunity. This is of great significance for improving the fully autonomous sortie and recovery capabilities of naval UAVs in high sea states.

[0073] 3) Spacecraft rendezvous and docking and on-orbit servicing: The space environment is characterized by a deep black background and extremely high contrast (high dynamic range) due to direct sunlight. Traditional planar targets are prone to docking failure due to "eccentricity error." The transmissive target designed in this invention utilizes the principle of physical light spot imaging to maintain the consistency of the geometric center projection at any oblique viewing angle, eliminating systematic errors. This technology can be widely applied in satellite on-orbit maintenance, space station cargo transfer, and non-cooperative target acquisition missions, providing reliable visual guidance for spacecraft precision operations under extreme lighting conditions.

[0074] 4) High-speed industrial automation and flexible manufacturing: In high-end manufacturing, workpiece gripping on high-speed production lines or dynamic assembly by collaborative robots often faces problems of motion fuzziness and system latency. The short-exposure imaging and dynamic process noise filtering technology proposed in this invention can adapt to sudden changes in conveyor belt speed or high-speed maneuvers of robotic arms, achieving zero-latency tracking and gripping of fast-moving workpieces. Compared to expensive LiDAR solutions, the pure vision solution of this invention is lower in cost, smaller in size, and easier to integrate into the end effector of industrial robotic arms, resulting in significant economic benefits.

[0075] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0076] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0077] In this specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. While different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce a good effect.

[0078] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A robust visual localization and action prediction method based on adaptive prediction, characterized in that, include: Two images of a target with multiple targets are acquired at the current moment, each image containing the image regions of the multiple targets; the outer layer of each target is a planar ring, the inner layer is a light-transmitting hole with the same center as the outer layer, and an actively emitting LED light source is set behind the light-transmitting hole; Image processing and target center extraction are performed on each image to obtain a target point set for each image, wherein the target point set contains multiple target center points; Global optimal stereo matching is performed on the target point set based on the two images at the current time, and the observed three-dimensional point set is determined. Then, using the observed three-dimensional point set and the target model point set in the world coordinate system and the position of the target, the three-dimensional position of the target in the camera coordinate system at the current time is determined based on the alternating optimization rigid body pose estimation method. An improved Kalman filter algorithm is used to filter the target's three-dimensional position at the current moment to obtain the filtered three-dimensional position of the target at the current moment. Based on the target's velocity, acceleration, and filtered 3D position at the current moment, the filtered 3D position, velocity, and trajectory curvature of the target at multiple historical moments at the current moment, as well as the constant acceleration model and a trained deep nonlinear residual network based on LSTM units, the 3D position of the target at future times is predicted.

2. The robust visual localization and action prediction method based on adaptive prediction according to claim 1, characterized in that, The method of determining the predicted future position of the target based on the target's velocity, acceleration, and filtered 3D position at the current moment, the filtered 3D position, velocity, and trajectory curvature of the target at multiple historical moments at the current moment, and a trained deep nonlinear residual network based on LSTM units includes: Using the target's velocity, acceleration, and filtered three-dimensional position at the current moment, as well as the time interval between two adjacent moments, the position components of the target in future times are predicted based on the constant acceleration model. ; By utilizing the filtered 3D position, velocity, and trajectory curvature of the target at the current moment, as well as the velocity, trajectory curvature, and filtered 3D position of the target at multiple historical moments, a historical state sequence is constructed. Based on this historical state sequence and a trained deep nonlinear residual network based on LSTM units, the nonlinear residual of the target at future times is predicted. ; Based on the trajectory curvature of the target at the current moment and the position components The nonlinear residual Based on the historical state sequence, determine the three-dimensional position of the target at future time. .

3. The robust visual localization and action prediction method based on adaptive prediction according to claim 2, characterized in that, The target's position component in future time The expression is as follows: ; in, , , These respectively represent the target at the current time. The velocity, acceleration, and filtered 3D position. Indicates the current time Compared to the previous moment The time interval between them.

4. The robust visual localization and action prediction method based on adaptive prediction according to claim 2, characterized in that, The target's three-dimensional position in future time The expression is as follows: ; in, This represents the curvature of the target's trajectory at the current moment. This represents the historical state sequence.

5. The robust visual localization and action prediction method based on adaptive prediction according to claim 1, characterized in that, The process of image processing and target center extraction for each image to obtain a target point set for each image includes: After converting each image to the HSV color space, candidate regions are extracted using a preset HSV threshold and masking operation to obtain multiple candidate regions for each image. Each candidate region contains an image region of a target. Morphological noise suppression and contour restoration are performed on each candidate region of each image to obtain multiple contour-restored candidate regions for each image. Calculate the ratio between the area of ​​the circumcircle of the contour of each candidate region after contour repair and the actual area, and filter out the candidate regions after contour repair whose ratio does not meet the preset conditions. In each of the remaining candidate regions after contour restoration, a highlight LED region is extracted. When the extracted highlight LED region coincides with the geometric center of the candidate region after contour restoration, a target center point is determined based on the pixel coordinates of the extracted highlight LED region. After traversing each of the remaining candidate regions after contour restoration, a target point set for each image is obtained.

6. The robust visual localization and action prediction method based on adaptive prediction according to claim 1, characterized in that, The process involves performing global optimal stereo matching on the target point set based on the two images at the current moment, determining the observed 3D point set, and using the observed 3D point set and the target model point set in the world coordinate system with the target's position, determining the target's 3D position in camera coordinates at the current moment based on an alternating optimization rigid body pose estimation method, including: Based on the point-to-point matching algorithm and the outlier removal algorithm, the target point set of the two images at the current time is subjected to global optimal stereo matching to obtain multiple pairs of matching points; Based on the multiple pairs of matching points and the camera's intrinsic and extrinsic parameters, the observed three-dimensional point set is determined; When iteratively solving for the pose at the current moment, the pose at the previous moment is used as the initial pose at the current moment. Each time, the target model point set in the world coordinate system is projected using the pose obtained from the previous iteration. Then, a point pair matching algorithm is used to find matching point pairs between the projected point set and the observed 3D point set. Based on the matching point pairs and the Kabsch algorithm, the current optimal pose is solved to obtain the pose obtained in this iteration. If the change in the pose obtained in this iteration is less than or equal to a preset pose threshold, the pose obtained in this iteration is used as the pose at the current moment. If the change in the pose obtained in this iteration is greater than the preset pose threshold, the iterative solution of the pose continues until the pose at the current moment is obtained. Using the current pose, the target's position in the world coordinate system is projected onto the camera coordinate system to obtain the target's three-dimensional position at the current moment.

7. The robust visual localization and action prediction method based on adaptive prediction according to claim 1, characterized in that, The improved Kalman filter algorithm uses the Mahalanobis distance of the observation residuals at the current time as a maneuver detection index, dynamically adjusts the process noise covariance matrix at the current time using the maneuver detection index, and uses the adjusted process noise covariance matrix for filtering.

8. The robust visual localization and action prediction method based on adaptive prediction according to claim 7, characterized in that, In the improved Kalman filter algorithm, the dynamically adjusted process noise covariance matrix at the current time step The expression is as follows: ; in, Indicates the current time Mahalanobis distance of the observation residuals It is a preset factor with a value greater than 1. Indicates the preset detection threshold. This represents the original process noise covariance matrix.

9. The robust visual localization and action prediction method based on adaptive prediction according to claim 2, characterized in that, The deep nonlinear residual network based on LSTM units comprises, in sequence: an input layer, at least two LSTM layers, a fully connected layer, and an output layer.

10. A robust visual localization and motion prediction device based on adaptive prediction, comprising a processor, a communication interface, a memory, and a communication bus, characterized in that, The processor, the communication interface, and the memory communicate with each other via the communication bus; The memory is used to store computer programs; When the processor executes a program stored in the memory, it implements the steps of the method according to any one of claims 1-9.