Vision-strain fusion based pose estimation and online self-correction system for aerial flexible arms

The vision-strain fusion aerial flexible arm pose estimation system, which combines a depth camera and a resistance strain gauge, solves the problem of unstable accuracy in pose estimation of the aerial flexible robotic arm end effector in complex environments, and achieves high-precision and stable pose perception and control.

CN121521156BActive Publication Date: 2026-03-24SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-15
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing methods for estimating the pose of the end effector of a flexible aerial robotic arm are not accurate enough in complex environments and are easily affected by changes in lighting, vibration, and texture. Furthermore, they are limited by computation and resources, making it difficult to achieve efficient and stable pose perception and control.

Method used

A vision-strain fusion aerial flexible arm pose estimation system is adopted, which combines a depth camera and a resistance strain gauge. The system uses the VINS-Fusion algorithm and a deep neural network for information fusion and online self-correction to achieve high-precision pose estimation.

Benefits of technology

High-precision and stable pose estimation was achieved in complex environments, which improved the robustness and continuous operation capability of the system and ensured the operational stability and control accuracy under all weather and working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121521156B_ABST
    Figure CN121521156B_ABST
Patent Text Reader

Abstract

The application relates to an aerial flexible arm pose estimation and online self-correction system based on vision-strain fusion, and relates to the technical field of pose estimation. The application comprises a physical system framework and an information processing framework. The application has the beneficial effects that: the application not only realizes multi-sensor fusion and online updating of the end pose of an aerial flexible arm, but also realizes autonomous evolution and knowledge migration of the perception ability, greatly improves the robustness and continuous operation ability of the system, realizes optimal fusion and dynamic timing enhancement of information, and constructs a self-correcting perception-control integrated closed loop.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pose estimation technology, specifically to a pose estimation and online self-correction system for a flexible aerial arm based on vision-strain fusion. Background Technology

[0002] As the low-altitude economy becomes a global strategic emerging industry, the demand for efficient aerial operations in complex scenarios continues to rise. As a carrier for the deep integration of UAV platforms and continuum robot technology, the research of aerial flexible robotic arms is of great strategic significance and opens up a new application dimension for the low-altitude economy.

[0003] Aerial manipulation refers to the use of unmanned aerial vehicles (UAVs) with robotic arms to perform tasks such as grasping, sampling, light assembly, and valve operation.

[0004] The operating mechanisms of aerial flexible robotic arms are mainly divided into two categories: grippers and robotic arms. Grippers are usually mounted on the bottom of the drone, close to the center of gravity, and have little impact on flight dynamics, but are only suitable for simple grasping tasks. In contrast, robotic arms have multiple degrees of freedom, a larger workspace, and more diverse functions, and can adapt to complex operational needs.

[0005] Drone robotic arms can be broadly categorized into rigid and flexible robotic arms. Rigid robotic arms are typically small articulated arms with 2–6 degrees of freedom, mounted on quadcopter / hexacopter fuselages. Their advantages include mature control and high end-effector stiffness, but disadvantages include significant weight, substantial impact on the drone platform, and extremely noticeable disturbances to flight attitude. Flexible aerial robotic arm systems, on the other hand, have gained attention due to their high flexibility, low weight, and good environmental adaptability. These systems possess high compliance and unlimited-dimensional deformation, making them suitable for safe interaction in confined spaces, near walls, or complex environments.

[0006] The flexible aerial robotic arm system, through its biomimetic structure and multi-degree-of-freedom deformation design, performs better in constrained scenarios such as pipeline inspection and navigation in curved spaces. Simulating a pipeline environment using a prototype flexible robotic arm reveals that, compared to traditional rigid robotic arms, the flexible robotic arm offers more efficient task execution capabilities in complex and dynamic environments, although modeling and perception are more challenging.

[0007] The typical working conditions and challenges in the application scenarios of aerial flexible robotic arms are as follows:

[0008] (1) Outdoor weak texture / strong light / backlight / night: visual reliability fluctuates, and is prone to loss of lock or posture change;

[0009] (2) High dynamic / vibration environment: The vibration of the thruster is coupled with the micro-vibration of the flexible arm, which affects both the IMU and the camera;

[0010] (3) Long-term operation: Temperature drift, strain gauge bonding aging, and mechanical wear cause calibration drift;

[0011] (4) Limited onboard resources: Computation, power consumption and weight constraints make it unsuitable for heavy external positioning systems or highly complex algorithms.

[0012] Existing end-effector pose estimation and flexible arm bending and yaw angle estimation schemes mainly include methods based on inertial measurement units (IMU), vision-based methods—using external or airborne cameras for visual positioning, methods based on motion capture systems, and methods based on traditional tendon length models in kinematics.

[0013] The inertial measurement unit (IMU)-based method involves mounting an IMU sensor at the end of a flexible arm to obtain pose information through integrated acceleration and angular velocity. However, this method suffers from significant cumulative errors (especially in aerial flexible robotic arm scenarios), a sharp decline in positioning accuracy after prolonged operation, and susceptibility to magnetic field interference. Furthermore, the non-rigid structure of the flexible arm generates complex vibration modes, and these high-frequency vibrations severely contaminate the IMU's accelerometer readings, leading to attitude calculation divergence.

[0014] Vision-based methods utilize external or airborne cameras for visual localization. Based on multi-view geometry and nonlinear optimization theory, these methods can provide high-precision metric-level pose under ideal conditions (such as rich texture and stable lighting). However, their robustness bottleneck lies in their strong dependence on the quality of visual features. When ambient lighting changes drastically, the background is a large area of ​​weak texture (such as a white wall or sky), or the image is blurred due to rapid drone maneuvers, feature extraction and matching will fail extensively, leading to a sharp drop in visual localization accuracy or even complete loss of accuracy.

[0015] Motion capture systems severely limit the working range of drones to expensive and complex indoor calibration fields, making them unsuitable for deployment in real outdoor or wilderness scenarios. Furthermore, their reliance on unobstructed line of sight makes them highly susceptible to failure due to arm self-occlusion or environmental obstacles.

[0016] Traditional kinematic tendon length models rely on the constant curvature (CC) assumption. They calculate the bending angle and direction of the robotic arm by measuring the length of the driving tendon, thereby inferring the end-effector pose. However, a significant drawback of this model is that the accuracy of the calculated shape and pose deteriorates severely under external loads due to nonlinear and uneven force distribution.

[0017] Therefore, real-time, stable, and continuous estimation of the end-effector pose (including at least the bending angle θ and yaw angle φ) is a prerequisite for achieving high-quality closed-loop control and safe operation. However, factors such as flight vibration and wind disturbance, changes in lighting and texture conditions, load coupling, and structural nonlinearity can lead to sensor degradation, model mismatch, and control instability. Therefore, establishing a robust, low-latency, and long-term operational end-effector pose perception and control scheme has become a core technical challenge in this field. Summary of the Invention

[0018] The purpose of this invention is to address the shortcomings and defects of the existing technology by providing a vision-strain fusion-based system for aerial flexible arm pose estimation and online self-correction.

[0019] To achieve the above objectives, the present invention adopts the following technical solution: a vision-strain fusion-based aerial flexible arm pose estimation and online self-correction system, which includes a physical system framework and an information processing framework. The physical system framework includes a quadcopter drone 1 serving as a mobile base, a flexible continuum manipulator 2 mounted on the quadcopter drone 1 as an execution carrier, a resistance strain gauge 3 mounted on the base of the flexible continuum manipulator 2, a depth camera module 4 mounted at the end of the flexible continuum manipulator 2, and a drive unit for driving the flexible continuum manipulator 2. A real-time data processing unit is installed between the quadcopter drone 1 and the flexible continuum manipulator 2. According to the embedded computing platform 5, the information processing framework is a layered, modular software framework responsible for transforming the raw sensing data of the physical system into precise control commands. It includes a perception layer, a vision-strain pose fusion and state estimation layer, and a decision and control layer. Specifically: the perception layer is responsible for processing the raw data from the heterogeneous sensing subsystems and transforming it into meaningful physical quantities; the vision-strain pose fusion and state estimation layer is responsible for intelligently processing and fusing the information from the perception layer to obtain the optimal estimate of the current state of the system; and the decision and control layer is responsible for generating and executing control commands based on the optimal state estimate and the task objective.

[0020] Furthermore, the physical system framework specifically comprises: a quadcopter UAV 1: serving as a mobile base, providing a wide range of movement and hovering capabilities within three-dimensional space; the UAV's body coordinate system is denoted as... Its origin is located at the center of mass of the aircraft; Flexible continuum manipulator 2: Composed of multiple connected, tendon-driven continuum structures, containing a semi-rigid central skeleton and equidistant disks, covered by an elastic cladding. Flexible continuum manipulator 2 is lightweight, highly compliant, and inherently safe. The base coordinate system of flexible continuum manipulator 2 is denoted as... ,and Rigid connection, with only fixed vertical displacement. The attitude of the flexible continuum manipulator 2 is provided by the inertial measurement unit of the quadcopter drone 1; the drive unit consists of four high-precision digital servos mounted at the top of the flexible continuum manipulator 2, which control the length of four low-elasticity tendons via a reel. The tendon passes through and is fixed to the distal disc at 0°, 90°, 180°, and 270° along the circumference of the arm, changing the position of the tendon. This allows the arm to bend and extend in any direction, with the driving space denoted as... Depth camera module 4, also known as the external perception-visual inertial unit, is installed on the end of the flexible continuum robotic arm 2. Its visual and inertial information jointly drive the VINS-Fusion algorithm to output high-precision 6-DOF pose. and real-time confidence level The depth camera has an RGB resolution of 1920×1080@30fps and a depth resolution of 1280×720@90fps, with a built-in IMU sampling at 400Hz. The resistance strain gauge 3, also known as the internal sensing-body strain unit, is installed on the base of the flexible continuous robotic arm 2. Four resistance strain gauges 3 are symmetrically attached along the circumference of the flexible continuous robotic arm 2 in orthogonal directions of 0°, 90°, 180°, and 270°. The resistance strain gauge 3 can accurately measure the minute strain generated on the surface of the flexible continuous robotic arm 2 due to overall bending at its root. This orthogonal layout can decouple the deformation information of the arm body on the two main bending planes to the greatest extent, providing structured and information-rich raw input for subsequent machine learning models. It adopts a full-bridge circuit configuration and is sampled by a 24-bit high-precision analog-to-digital converter at a frequency of 1000Hz.

[0021] Furthermore, the perception layer includes a VINS module, a strain-visual model module, and a training strategy and online self-calibration module. Specifically, the VINS module receives images and IMU data streams from the depth camera module 4, runs the VINS-Fusion algorithm, and outputs the 6-DOF pose of the flexible arm's end effector in real time. The arm flexion angle and yaw angle were extracted by pose. And a confidence level that quantifies the quality of the current estimate. Strain-Vision Module: Receives four strain signal vectors from resistance strain gauges. It is then fed into a pre-trained deep neural network model, which outputs the predicted pose of the arm's end effector in real time. And the uncertainty of the model in this prediction Training strategy and online self-calibration module: Based on a robust model foundation of offline course-based pre-training, it is key to achieving the system's adaptive capability. It adopts a closed-loop mechanism of "sample selection - experience playback - safety rollback": the system continuously monitors the confidence level of the VINS module. ,when When the values ​​exceed a preset threshold, it indicates that the visual positioning is extremely reliable. At this point, the "VINS pose and strain readings" data are considered high-quality samples, and a hybrid training set is constructed using an empirical replay buffer. Heteroscedasticity probability loss and geometric consistency constraints are employed to perform online, incremental parameter fine-tuning of the "strain-pose model." Simultaneously, exponential moving averages and validation monitoring are introduced. Once a performance degradation is detected, automatic rollback is triggered to ensure absolute system stability while continuously resisting sensor aging and temperature drift. (The angle...) All are defined in the base coordinate system Below, the angles of arm bending and direction are indicated.

[0022] Furthermore, the vision-strain pose fusion and state estimation layer includes a confidence-weighted fusion module and a dynamic filtering module, specifically: the confidence-weighted fusion module: this module will, based on... The adaptive weighting approach first unifies the visual confidence score and the prediction variance of the strain network into an observation variance, and then uses an inverse variance weighting strategy to calculate the intermediate fused pose. The dynamic filtering module, based on a unified uncertainty quantization mechanism for heterogeneous information sources, maps the confidence score of the visual system to an equivalent observation variance with the same dimensions as the prediction variance of the strain network, and constructs an adaptive observation noise matrix accordingly. Subsequently, the pose is smoothed and differentiated using a kinematic model, ultimately outputting the system's optimal state vector containing angles and angular velocities. .

[0023] Furthermore, the decision and control layer includes an inverse kinematics module and a closed-loop controller module, specifically: the inverse kinematics module receives the desired pose. It is analyzed as the expected change in length of the four tendons in the driving space. Closed-loop controller module: Receives the output of the inverse kinematics module as a feedforward signal, and simultaneously receives the optimal pose and angular velocity estimation from the fusion layer as a full-state feedback signal. Subsequently, it can execute a two-layer closed-loop control law, that is, use angular velocity for feedforward compensation and D-term damping control to calculate the precise command to drive the four servo motors, thereby completing the precise tracking of the pose of the flexible arm end effector.

[0024] Furthermore, the VINS module includes the following steps: Visual-inertial tight coupling theoretical framework: The VINS-Fusion algorithm is used as a high-precision benchmark for the pose estimation of the flexible robotic arm's end effector. The VINS-Fusion algorithm achieves optimal information fusion from multiple sensors within a sliding window optimization framework by tightly coupling high-frequency dynamic information from the IMU with low-frequency global information from the vision. The sliding window contains... One keyframe, and The observed landmarks represent the complete state vector of the system to be estimated. The definition is as follows: in, Representing the The IMU navigation state of a frame includes its position in the world coordinate system. The lower position ,speed and quaternions representing attitude ,also, It also includes accelerometer zero bias. and gyroscope zero bias These two quantities drift slowly over time; online estimation of them can effectively improve the accuracy of long-term operation. This represents the external parameters between the camera and the IMU. For the first Inverse depth of feature points; IMU pre-integration model: To reduce the computational burden of repeatedly integrating massive amounts of raw IMU measurement data due to state updates in sliding window optimization, IMU pre-integration technology is adopted. This technology preprocesses the IMU measurements between two frames into relative motion constraints, decoupling them from absolute pose. In keyframes... arrive Within the time interval, the relative rotation, velocity, and displacement pre-calculated based on IMU measurements are respectively Constructing IMU pre-integrated residuals as follows: ,in, For the rotation matrix from the IMU to the world system; This is the vector of gravitational acceleration; It is a logarithmic mapping function that maps rotational errors from Lie group space to vector space. By minimizing this residual, high-frequency inertial information can be used to constrain the short-term violent motion of the UAV. Visual reprojection error: In order to eliminate the drift of the IMU accumulated over time, the system uses visual observation to introduce global geometric constraints. For the first... Let the nth landmark be the nth road sign. The observation coordinates on the frame image are Its predicted coordinates are based on the projection of the current estimated pose onto the image plane. Visual reprojection residual Defined as the deviation between the observed value and the predicted value: , ,in, These are the three-dimensional coordinates of the landmark points in the camera coordinate system, and the function... This is a projection model of a pinhole camera. Focal length Using the principal point coordinates, by minimizing this geometric error on the image plane, the camera pose and feature point depth can be accurately corrected; sliding window joint optimization: the marginalization prior, IMU pre-integration residual, and visual reprojection residual are unified into a nonlinear least squares objective function for joint solution: The objective function Denotes the square of the Mahalanobis distance, where This is the inverse of the information matrix covariance, used to automatically assign weights based on the noise characteristics of different sensors. To marginalize prior residuals, historical constraint information of the sliding window's moveout frames was preserved. The Huber robust kernel function is used to suppress the impact of visual mismatches on optimization, and the Levenberg-Marquardt algorithm is used to iteratively update the state increment. Until convergence; Pose parameterization and angle extraction: VINS outputs the end pose. To serve the control of the flexible continuous body robotic arm, it is necessary to... Extraction from the system and Coordinate transformation: ; bending angle : ; Yaw angle : Singular value handling: When like At that time, it was believed Unstable, take the previous moment Alternatively, its gain can be reduced to avoid jumps in near-straightened states; dynamic filtering and temporal smoothing: to suppress high-frequency jitter in visual positioning and eliminate outliers, the angle sequence output by VINS is smoothed and preprocessed. By constructing a Kalman filter with a constant velocity model, visual observation noise is suppressed before the data is sent to the fusion layer, ensuring the continuity of the input to the subsequent fusion module. During the update, the angle difference is normalized around: ,use A smooth filter can be obtained by completing the filter update. and its angular velocity.

[0025] Furthermore, the VINS module includes the following steps: pose parameterization and angle extraction: VINS outputs the end-effector pose. To serve the control of the flexible continuous body robotic arm, it is necessary to... Extraction from the system and Coordinate transformation: ; bending angle : ; Yaw angle : Singular value handling: When like At that time, it was believed Unstable, take the previous moment Alternatively, its gain can be reduced to avoid jumps in near-straightened states; dynamic filtering and temporal smoothing: to suppress high-frequency jitter in visual positioning and eliminate outliers, the angle sequence output by VINS is smoothed and preprocessed. By constructing a Kalman filter with a constant velocity model, visual observation noise is suppressed before the data is sent to the fusion layer, ensuring the continuity of the input to the subsequent fusion module. During the update, the angle difference is normalized around the perimeter: .use A smooth filter can be obtained by completing the filter update. and its angular velocity.

[0026] Furthermore, the bending angle The predicted branch is specifically: bending angle. Range of non-periodic variables However, in weak bending Due to poor observability under certain operating conditions, a hybrid architecture of "triangular encoding + direct regression + gating fusion" was added: a hybrid output head design, which includes three parallel output heads: Trig head: outputs a unit circular vector. Geometric constraints are used to ensure the continuity of predicted values; Direct head: directly outputs scalar angles. Captures local linear relationships; Uncertainty header: outputs the log-variance of the predicted values. Gating weight generation: To adaptively select the optimal prediction strategy under different bending intensities, a gating network is designed based on the amplitude... Generate weights : .in, Using the Sigmoid activation function, the network is trained to trust the Direct head more under strong bending conditions and the Trig head more under weak bending conditions. The final output of the branch is: first, the Trig head is back-projected into an angle. Then, a weighted fusion is performed to obtain the final predicted value. : Simultaneously, the variance is directly output for subsequent confidence calculation: .in, It is determined by the gating weight The adjusted fusion coefficient; the yaw angle φ prediction branch specifically refers to: yaw angle It is a periodic variable Directly returning to the face The jump problem includes the following: First stage: Geometric baseline calculation: utilizing the aforementioned principal axis difference feature. By using a pre-fitted polynomial design matrix Perform ridge regression and calculate the geometric baseline vector. : ,in, For the regression coefficient matrix fitted offline, the geometric baseline vector model provides a stable but biased initial estimate; the second stage: expert residual learning and uncertainty estimation: the neural network aims to learn the residual angles relative to the baseline. The network adopts a "hybrid expert model" structure, which will Divided into 8 sectors, the gated network is based on features Calculate the activation probability of experts in each sector and output the comprehensive residual vector. And uncertainty: ,in, For the first The third stage: adaptive fusion output based on amplitude prediction of local residual vectors; Adaptive weights The baseline vector is fused with the residual vector predicted by the neural network: Finally, the predicted yaw angle is obtained through back projection. and its variance: The yaw angle φ prediction branch ensures that the system automatically reverts to a stable geometric baseline during weak curves, while compensating for strong curves using the high-precision residuals of the neural network. It can accurately reflect the credibility of current predictions.

[0027] Furthermore, the training strategy and online self-calibration module include the following steps: Offline training strategy: Based on uncertain supervised learning, offline training aims to establish an initial mapping from strain features to pose and endow the model with self-evaluation capabilities; Data construction and "strong bending priority" course learning: First, define the strong bending dataset. :, .in, This is the bending amplitude threshold. During the initial training phase, only... The model is iterated on to ensure it quickly learns the geometric features of the principal bending plane; then the full dataset is gradually added. Fine-tuning is performed, and a gating mechanism is used to automatically process noise in weak bending data; bending angle The probability loss function: For the curved angle branch, the training objective is to simultaneously optimize prediction accuracy and uncertainty estimation. Heteroscedastic negative log-likelihood loss is used as the core objective, supplemented by geometric constraints, and is defined as follows: Total channel loss : Geometric consistency loss Combining triangular head cosine similarity with direct head Huber loss ensures the geometric continuity of predicted values ​​on the unit circle; uncertainty loss. Forced model learning predicts variance : The above formula forces the model to operate with larger errors. Large regions actively predict larger variances To reduce losses; yaw angle The residual probability loss function: For the yaw angle branch, the network learns the residuals relative to the geometric baseline. Heteroscedastic negative log-likelihood loss is used to supervise the network output. Defined... Channel loss : This loss function ensures that the expert network can not only correct the deviation of the geometric baseline, but also accurately assess the reliability of the current sector prediction; self-supervision of the gated network: introducing amplitude-based... Soft tag supervision: ,in for The quantiles guide the network to automatically increase the weights of the direct regression head during strong bending; online self-calibration mechanism; anti-drift continuous learning: during operation, high-confidence segments of VINS are used as "pseudo-true values" to continuously fine-tune the "strain-visual model"; high-confidence sample filter, the system monitors the VINS confidence level in real time. The data at the current time is valid only if the following conditions are met. Only then will it be collected: Furthermore, the IMU / reprojection error remains stable. This ensures the purity of the signal itself, preventing the VINS positioning errors from being propagated to the strain model; Experience replay buffer: To prevent "catastrophic forgetting" in online learning, a fixed-capacity first-in-first-out queue is constructed. : When sampling batches for training, an amplitude-based approach is used. The bucket sampling strategy ensures a balanced distribution of training data across different curvatures, preventing the model from overfitting to a specific pose; incremental parameter fine-tuning employs a hybrid batch strategy for parameter updates, with each iteration using a different training set. From the new data currently collected and historical data in the buffer pool Composition, online optimization target for: ,in, With an extremely small learning rate, this incremental, rapid learning approach allows the model to correct zero-point drift caused by strain gauge creep or temperature changes in real time; safety rollback and EMA protection are also included. To prevent divergence during online training, the system maintains an exponential moving average model for actual inference and monitors validation errors. Once a continuous increase in verification error is detected, or If the system remains sluggish for an extended period, it will immediately discard the current update and roll back to the last stable parameter checkpoint to ensure the absolute safety of flight operations.

[0028] Furthermore, the vision-strain pose fusion and state estimation layer specifically involves: unified quantification of the uncertainty of heterogeneous information sources. To achieve effective fusion of the vision system and the strain system, it is first necessary to uniformly map the confidence indices of the vision system and the strain system to statistical observation variances, using this as the sole standard for measuring the "credibility weight" of their respective data. The variance mapping of the strain-vision model: During the inference phase, the strain-vision model can directly output the variance of the predicted values ​​based on the distribution of input features. For bending angles... and yaw angle Define its observation variance as follows: The physical meaning of this formula is: the variance of the network output. It directly quantifies the model's "self-doubt" regarding the current prediction results. Equivalent variance modeling of the VINS system: Confidence level of the VINS system output. This mainly reflects the number of feature points, tracking stability, and reprojection error. To align with the variance of the strain model, an inverse proportional mapping model from confidence level to variance needs to be established, defining the equivalent observation variance of VINS. as follows: ,in, To adjust the scaling factor between the dimensions of the vision and strain systems; To prevent tiny constants with a denominator of zero; As the penalty factor in the near-extended state, it can be seen that when the visual environment deteriorates, its equivalent variance will increase exponentially until it approaches infinity; Static fusion: based on maximum likelihood inverse variance weighting: bending angle Scalar fusion: For non-periodic bending angles, the reciprocals of the two differences are used as weights for linear combination to calculate the fused bending angle. : The above formula ensures that the fusion result always favors the side with the smaller variance. Yaw angle Vector domain fusion: To address the periodicity of the yaw angle and avoid clipping errors caused by arithmetic averaging, it is mapped to a unit circle space for vector weighting, defining a weighted fusion vector. : The yaw angle after fusion is then extracted using inverse trigonometric functions. : This ensured that... The fused trajectory in the vicinity maintains geometric continuity, eliminating the impact of angle jumps on the control system; Dynamic fusion: Observation noise adaptive EKF: State prediction and kinematic modeling: Selecting a state vector that includes the angle and its first derivative. A constant angular velocity model is used for time updates; the state prior prediction equation is as follows: The current pose is predicted using the angular velocity from the previous moment, and a smoothing constraint in the time dimension is introduced, which can effectively suppress instantaneous glitch noise from the sensor. Real-time construction of the observation noise matrix: Static fusion: The fused pose is calculated based on the inverse variance weighted by maximum likelihood. As observed values, their corresponding synthetic variances are used to dynamically fill the observation noise covariance matrix. Construct the real-time observation noise matrix : The magnitude of this matrix value reflects the reliability of the current fusion result in real time, especially when both visual and strain methods are unreliable. The surge in Kalman gain leads to a decrease in the filter's accuracy, causing it to automatically refuse updates and instead rely on prior predictions. This prevents output divergence; posterior state update: utilizing the calculated Kalman gain. The prior state is corrected to obtain the final optimal estimate. The state update equation is as follows: The final output It not only contains the fused high-precision pose information, but also the smoothed angular velocity estimate, which is directly used as the feedback signal for the subsequent closed-loop control system.

[0029] By adopting the above technical solution, the present invention has at least the following advantages:

[0030] 1. Achieved "autonomous evolution and knowledge transfer" of perception capabilities: Not only did the low-cost strain gauge system gain high-precision pose perception capabilities through visual guidance, but it also achieved continuous knowledge transfer and adaptive optimization of the model through an online self-correction mechanism.

[0031] 2. Significantly improves the robustness and continuous operation capability of the system: When the vision system temporarily fails or its confidence level decreases due to low light, occlusion, or motion blur, the continuously evolving and drift-free strain model dominates or provides stable estimates, ensuring the system's continuous operation capability under all weather and working conditions.

[0032] 3. Achieved optimal information fusion and dynamic temporal enhancement: Based on the smooth fusion and adaptive filtering strategy of bidirectional confidence, it not only achieves Pareto optimal utilization of the two information channels, but also accurately estimates the motion speed of the flexible arm while suppressing high-frequency noise by introducing kinematic constraints, ensuring that the system can output the pose result with the best accuracy and stability at any time.

[0033] 4. A self-correcting integrated closed loop of "perception-control" has been constructed: The innovative online self-correction mechanism enables the entire "perception-control" link to have the ability to adaptively optimize. The continuous improvement of pose estimation accuracy is directly converted into the improvement of control performance, and finally a higher level of autonomous operation capability is achieved. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a schematic diagram of the structure of the present invention.

[0036] Figure 2 This is a schematic diagram of the installation of the resistance strain gauge and the depth camera module in this invention.

[0037] Figure labeling: 1. Quadcopter UAV; 2. Flexible continuum robotic arm; 3. Resistance strain gauge; 4. Depth camera module; 5. Embedded computing platform. Detailed Implementation

[0038] The system consists of a four-tendon driven flexible arm, an end-effector depth camera (running VINS), and strain gauges on all four sides of the arm. It constructs a heterogeneous sensing system comprised of a depth camera and a set of low-cost strain gauges. A robust pose estimation is achieved through an innovative, dynamically learning-enabled integrated "learning-correction-fusion" framework, providing accurate pose data for subsequent flexible arm motion control. The core of this framework is:

[0039] First, the Vins-Fusion algorithm is used with a depth camera and an onboard NUC to obtain the quaternions at the end of the flexible robotic arm. The bending and yaw angles are then calculated using vector calculations. A basic model is established through "visual-guided learning" to supervise the training of a deep neural network model. In the initial stage, the high-precision visual inertial odometry (VINS) pose output by the depth camera under good operating conditions is used as the "teacher" signal. Interpretable geometric baselines (global and regional robust mappings) are established using physical features such as strain differentials and resultant quantities to obtain initial yaw estimates. A lightweight learning model is then used to regionally compensate for the geometric residuals, while a small-scale angle network learns the bending angles, making both strong and weak bending conditions observable and thus endowing the simple strain gauge system with preliminary, independent pose perception capabilities. This model aims to learn and construct a complex nonlinear mapping from the resistance changes (AD values) of the four strain gauges to the high-dimensional pose (bending angle θ and yaw angle φ) of the flexible arm's end effector.

[0040] Secondly, the model achieves "online self-calibration" for continuous evolution: In actual operation, this invention does not statically use the pre-trained model, but establishes an online self-calibration mechanism. Based on quality indicators such as the number of feature points, reprojection error, and covariance of the VINS, the system obtains visual confidence. It continuously monitors the confidence of the VINS output. When the VINS is determined to be in a high-confidence state (e.g., rich and stable visual features), it selectively uses the current high-quality "VINS pose-strain reading" data pair as a new, effective training sample, triggering online incremental learning to perform small-step self-calibration on the strain residual layer. Sliding window verification and rollback mechanisms suppress divergence, providing long-term resistance against temperature drift and assembly aging. This process endows the strain pose model with the ability to continuously learn, self-calibrate, and adapt to environmental changes (such as arm material aging and temperature changes) during tasks.

[0041] Finally, a "spatiotemporal adaptive fusion" is executed to output the optimal estimate and achieve closed-loop control: at any given time, the system runs VINS and a continuously self-calibrating strain-pose model in parallel. This invention proposes a dynamic weighted fusion strategy based on uncertainty quantification. First, by evaluating the quality of the VINS output and the uncertainty of the neural network model in real time, weights are adaptively assigned to the two pose signals, and fusion is performed in the vector domain to obtain intermediate observations. Subsequently, an adaptive extended Kalman filter (Adaptive EKF) is introduced to adapt the observation noise, and the kinematic model is used to perform temporal smoothing and state prediction on the fusion result, ultimately outputting the optimal posterior estimate containing high-precision angles and angular velocities. This continuously optimized, highly robust full-state estimate is directly used as the core feedback signal of the closed-loop control system. The subsequent execution layer can use analytical-learning hybrid inverse kinematics to obtain the changes in the four tendons, thereby controlling the flexible arm end to reach the desired bending and yaw angles.

[0042] See Figures 1-2 As shown, the technical solution adopted in this specific embodiment includes a physical system framework and an information processing framework.

[0043] The physical system framework includes a quadcopter drone 1 serving as a mobile base, a flexible continuum robotic arm 2 mounted on the quadcopter drone 1 as an execution carrier, a resistance strain gauge 3 mounted on the base of the flexible continuum robotic arm 2, a depth camera module 4 mounted at the end of the flexible continuum robotic arm 2, and a drive unit for driving the flexible continuum robotic arm 2. An embedded computing platform 5 for real-time data processing is installed between the quadcopter drone 1 and the flexible continuum robotic arm 2.

[0044] The physical system framework is specifically as follows:

[0045] Quadcopter UAV 1: Serving as a mobile base, it provides a wide range of movement and hovering capabilities in three-dimensional space. The UAV's coordinate system is denoted as... Its origin is located at the center of mass of the spacecraft;

[0046] Flexible continuum robotic arm 2: Composed of multiple interconnected, tendon-driven continuum structures, it contains a semi-rigid central skeleton and equidistant disks, covered by an elastic cladding. Flexible continuum robotic arm 2 is lightweight, highly compliant, and inherently safe. The base coordinate system of flexible continuum robotic arm 2 is denoted as... ,and Rigid connection, with only fixed vertical displacement. The attitude of the flexible continuum robotic arm 2 is provided by the inertial measurement unit of the quadcopter drone 1;

[0047] Drive unit: Four high-precision digital servos are mounted on the top of the flexible continuous robotic arm 2, and the length of four low-elasticity tendons is controlled by a reel. The tendon passes through and is fixed to the distal disc at 0°, 90°, 180°, and 270° along the circumference of the arm, changing the position of the tendon. This allows the arm to bend and extend in any direction, with the driving space denoted as... ;

[0048] Depth camera module 4, also known as the external perception-visual inertial unit, is installed on the end effector of the flexible continuum robotic arm 2. Its visual and inertial information jointly drive the VINS-Fusion algorithm to output high-precision 6-DOF pose. and real-time confidence level The depth camera has an RGB resolution of 1920×1080@30fps and a depth resolution of 1280×720@90fps, with a built-in IMU sampling at 400Hz.

[0049] Resistance strain gauge 3: also known as internal sensing-body strain unit, is installed on the base of the flexible continuous robot arm 2. Four resistance strain gauges 3 are symmetrically attached along the circumference of the flexible continuous robot arm 2 in orthogonal directions of 0°, 90°, 180° and 270°. The resistance strain gauge 3 can accurately measure the small strain generated on the surface of the flexible continuous robot arm 2 due to the overall bending. This orthogonal layout can decouple the deformation information of the arm body on the two main bending planes to the greatest extent, providing a structured and information-rich raw input for the subsequent machine learning model. It adopts a full-bridge circuit configuration and is sampled by a 24-bit high-precision analog-to-digital converter at a frequency of 1000Hz.

[0050] The information processing framework is a layered, modular software framework responsible for transforming raw sensor data from the physical system into precise control commands. It includes a perception layer, a vision-strain pose fusion and state estimation layer, and a decision and control layer, specifically:

[0051] The perception layer is responsible for processing raw data from heterogeneous sensing subsystems and transforming it into meaningful physical quantities. The perception layer includes a VINS module, a strain-visual model module, and a training strategy and online self-calibration module, specifically:

[0052] VINS module: Receives image and IMU data streams from depth camera module 4, runs the VINS-Fusion algorithm, and outputs the 6-DOF pose of the flexible arm's end effector in real time. The arm flexion angle and yaw angle were extracted by pose. And a confidence level that quantifies the quality of the current estimate. ;

[0053] Strain-Vision Module: Receives four strain signal vectors from resistance strain gauges. It is then fed into a pre-trained deep neural network model, which outputs the predicted pose of the arm's end effector in real time. And the model's uncertainty (variance) regarding this prediction. ;

[0054] Training strategy and online self-calibration module: Based on a robust model foundation built through offline course-based pre-training, this is key to achieving the system's adaptive capabilities. It employs a closed-loop mechanism of "sample selection - experience replay - safety rollback": the system continuously monitors the confidence level of the VINS module. ,when When the values ​​exceed a preset threshold, it indicates that the visual positioning is extremely reliable. At this point, the "VINS pose and strain readings" data are considered high-quality samples, and a hybrid training set is constructed using an empirical replay buffer. Heteroscedasticity probability loss and geometric consistency constraints are employed to perform online, incremental parameter fine-tuning of the "strain-pose model." Simultaneously, exponential moving average (EMA) and validation monitoring are introduced. Once a performance degradation is detected, automatic rollback is triggered to ensure absolute system stability while continuously resisting sensor aging and temperature drift. The angle... All are defined in the base coordinate system Below, the angles of arm bending and direction are indicated.

[0055] The visual-strain pose fusion and state estimation layer is responsible for intelligently processing and fusing information from the perception layer to obtain the optimal estimate of the system's current state. This layer includes a confidence-weighted fusion module and a dynamic filtering module, specifically:

[0056] Confidence-weighted fusion module: This module will be based on From the adaptive weight allocation, the visual confidence and the prediction variance of the strain network are first quantified into the observation variance, and the intermediate fused pose is calculated using the inverse variance weighting strategy.

[0057] The Adaptive Filtering Module (EKF) is based on a unified uncertainty quantification mechanism for heterogeneous information sources. It maps the confidence score of the vision system to an equivalent observation variance with the same dimensions as the prediction variance of the strain network, and constructs an adaptive observation noise matrix accordingly. Subsequently, it uses a kinematic model to perform Kalman filtering smoothing and differentiation on the pose, and finally outputs the optimal state vector of the system containing angles and angular velocities. .

[0058] Decision and Control Layer: Responsible for generating and executing control commands based on optimal state estimation and task objectives. This layer includes an inverse kinematics module and a closed-loop controller module, specifically:

[0059] Inverse Kinematics (IK) Module: Receives the desired pose It is analyzed as the expected change in length of the four tendons in the driving space. Closed-loop controller module: Receives the output of the inverse kinematics module as a feedforward signal, and simultaneously receives the optimal pose and angular velocity estimation from the fusion layer as a full-state feedback signal. Subsequently, it can execute a two-layer closed-loop control law, that is, use angular velocity for feedforward compensation and D-term damping control to calculate the precise command to drive the four servo motors, thereby completing the precise tracking of the pose of the flexible arm end effector.

[0060] More specifically, the VINS module includes the following steps:

[0061] Step 1, Visual-Inertial Tight Coupling Theoretical Framework:

[0062] The VINS-Fusion algorithm is used as a high-precision benchmark for pose estimation of the flexible robotic arm's end effector. The VINS-Fusion algorithm achieves optimal information fusion from multiple sensors within a sliding window optimization framework by tightly coupling high-frequency dynamic information from the IMU (400Hz) and low-frequency global information from the vision system (30Hz). The sliding window contains... One keyframe, and The observed landmark points (feature points) and the complete state vector of the system to be estimated. The definition is as follows: in, Representing the The IMU navigation state of a frame includes its position in the world coordinate system. The lower position ,speed and quaternions representing attitude (corresponding rotation matrix) ),also, It also includes accelerometer zero bias. and gyroscope zero bias These two quantities drift slowly over time; online estimation of them can effectively improve the accuracy of long-term operation. This represents the external parameters between the camera and the IMU. For the first Inverse depth of each feature point;

[0063] Step 2, IMU pre-integration model:

[0064] To reduce the computational burden of repeatedly integrating massive amounts of raw IMU measurement data (400Hz) due to state updates in sliding window optimization, an IMU pre-integration technique is employed. This technique preprocesses the IMU measurements between two frames into relative motion constraints, decoupling them from absolute pose. In keyframes... arrive Within the time interval, the relative rotation, velocity, and displacement pre-calculated based on IMU measurements are respectively Constructing IMU pre-integrated residuals as follows: The physical meaning of the above formula is to ensure that the "relative motion calculated based on the state to be optimized" remains consistent with the "relative motion actually measured by the sensor," where, For the rotation matrix from the IMU to the world system; This is the vector of gravitational acceleration; It is a logarithmic mapping function that maps rotation errors from Lie group space to vector space. By minimizing this residual, high-frequency inertial information can be used to constrain the short-term violent motion of the UAV; Step 3, visual reprojection error: In order to eliminate the drift of the IMU accumulated over time, the system introduces global geometric constraints using visual observations. Let the nth landmark be the nth road sign. The observation coordinates on the frame image are Its predicted coordinates are based on the projection of the current estimated pose onto the image plane. Visual reprojection residual Defined as the deviation between the observed value and the predicted value: in, These are the three-dimensional coordinates of the landmark points in the camera coordinate system, and the function... This is a projection model of a pinhole camera. Focal length Using the principal point coordinates, by minimizing this geometric error on the image plane, the camera pose and feature point depth can be accurately corrected;

[0065] Step 4, joint optimization of sliding windows:

[0066] To ensure both accuracy and real-time performance, this invention unifies the marginalization prior, IMU pre-integration residual, and visual reprojection residual into a single nonlinear least squares objective function for joint solution. The objective function Denotes the square of the Mahalanobis distance, where This is the information matrix (the inverse of the covariance), used to automatically assign weights based on the noise characteristics of different sensors. To marginalize prior residuals, historical constraint information of the sliding window's moveout frames was preserved. The Huber robust kernel function is used to suppress the impact of visual mismatches (outside points) on the optimization, and the Levenberg-Marquardt (LM) algorithm is used to iteratively update the state increment. until convergence;

[0067] Step 5, Pose parameterization and angle extraction:

[0068] VINS outputs end pose To serve the control of the flexible continuous body robotic arm, it is necessary to... Extraction from the system and ;

[0069] Coordinate transformation: ;

[0070] Bending angle (end) Shaft and base (Axis angle, numerically stable): ;

[0071] Yaw angle (The curved plane is in) Tie Orientation of the plane: ;

[0072] Singular value handling: When (like When ), it is believed Unstable, take the previous moment Or reduce its gain to avoid jumps in the near-extended state;

[0073] Step 6, Dynamic Filtering and Timing Smoothing

[0074] To suppress high-frequency jitter in visual localization and eliminate outliers, the angle sequence output by VINS is smoothed and preprocessed. A Kalman filter with a constant velocity model is constructed to suppress visual observation noise before the data is fed into the fusion layer, ensuring the continuity of the input to the subsequent fusion module. During the update, the angle difference is normalized around the perimeter: use A smooth filter can be obtained by completing the filter update. and its angular velocity.

[0075] Through the above steps, this invention constructs a theoretically complete and robust VINS pose calculation module. This step, as the system's 'pre-signal conditioning' stage, aims to provide smooth and continuous visual observation input for subsequent heterogeneous fusion. It not only provides high-precision bending angles and yaws, but also provides quantified and theoretically reasonable uncertainties for their estimation, providing crucial input for subsequent visual-strain confidence fusion.

[0076] More specifically, the strain-vision module includes the following steps:

[0077] The strain-vision module details the implementation of the "strain-vision" deep neural network model. This module aims to establish a high-precision, high-dimensional model for the pose (bending angle) of the flexible arm end effector, ranging from low-cost, easily disturbed strain gauge electrical signals to low-cost, high-dimensional strain gauge electrical signals. With yaw angle The strain-vision module not only outputs high-precision static angles, but its high-frequency and continuous inference energy (>100Hz) also provides rich time series information for the subsequent fusion layer to accurately solve angular velocities.

[0078] Step 1: Feature Vector Construction and Physical Preprocessing

[0079] To improve the learning efficiency and generalization ability of neural networks, instead of directly using the original voltage values, feature vectors are constructed based on physical interpretability and rotation invariance.

[0080] 1) Baseline correction and orthogonal decoupling:

[0081] First, baseline correction needs to be performed on the original voltage signal (to eliminate the initial zero-point drift caused by the strain gauge bonding process), and the corrected voltage signal is defined. as follows: in, This is the static bridge baseline voltage. To acquire voltage in real time, a principal axis differential feature reflecting the deformation of the two principal bending planes is constructed by using an orthogonal arrangement (0°, 90°, 180°, 270°) of resistance strain gauges 3 on the circumference of the flexible continuous robotic arm 2. and : in, The bending component approximately corresponds to the 0°-180° plane. The bending component approximately corresponds to the 90°-270° plane;

[0082] 2) Rotationally invariant amplitude and coarse phase:

[0083] To explicitly provide bending strength information to the network, rotationally invariant amplitude features are constructed. and a rough phase angle : in, With bending angle It shows a positive correlation and does not change with the rotation of the boom, making it a key indicator for judging "strong bend / weak bend" working conditions. A priori estimate of the yaw angle is provided, but it includes a fixed deviation due to assembly errors. ;

[0084] 3) Final feature vector assembly:

[0085] To enhance robustness to outliers and noise, the statistical feature "sum of absolute values" is introduced. ) and "the sum of the two largest terms" The periodic angles are triangulated to construct an 11-dimensional input feature vector. for: in, The sum of absolute values The sum of the two largest terms; this eigenvector After standardization (normalization), the signal is input into the neural network.

[0086] Step 2, Overall Architecture of Deep Neural Network

[0087] The neural network designed in this invention employs a multi-task learning architecture, containing two independent parallel branches, each used to predict non-periodic bending angles. and periodic yaw angle Each branch not only outputs the angle prediction, but also forces the output of the log-variance of the prediction through the Heteroscedastic Regression Head to quantify the uncertainty.

[0088] More specifically, the bending angle The predicted branch is specifically: bending angle. Non-periodic variables (range) However, in weak bending ( Due to poor observability under certain operating conditions, a hybrid architecture of "triangular coding + direct regression + gating fusion" was added.

[0089] 1. Hybrid output head design

[0090] This branch contains three parallel output headers:

[0091] Trig header: Outputs unit circle vector Geometric constraints are used to ensure the continuity of the predicted values;

[0092] Direct head: Directly outputs scalar angle Capture local linear relationships;

[0093] Uncertainty Header: Outputs the log-variance of the predictions ;

[0094] 2. Gating weight generation:

[0095] To adaptively select the optimal prediction strategy under different bending intensities, a gating network is designed based on the amplitude. Generate weights : in, Using the Sigmoid activation function, the network is trained to trust the Direct sensor more under strong bending conditions and the Trig sensor more under weak bending conditions.

[0096] 3-branch final output: First, backproject the Trig head into an angle. Then, a weighted fusion is performed to obtain the final predicted value. : Simultaneously, the variance is directly output for subsequent confidence calculation: in, It is determined by the gating weight Adjusted fusion coefficient;

[0097] The yaw angle φ prediction branch is specifically as follows:

[0098] Yaw angle It is a periodic variable ( ), directly reverting to the challenges The transition problem includes the following:

[0099] 1. First stage: Geometric baseline calculation: utilizing the aforementioned principal axis difference features By using a pre-fitted polynomial design matrix Perform ridge regression and calculate the geometric baseline vector. : in, For the regression coefficient matrix fitted offline, the geometric baseline vector model provides a stable but systematic preliminary estimate;

[0100] 2. Second Stage: Expert Residual Learning and Uncertainty Estimation

[0101] The neural network aims to learn the residual angle relative to the baseline. The network adopts a "hybrid expert model (MoE)" structure, which... Divided into 8 sectors, the gated network is based on features Calculate the activation probability of experts in each sector and output the comprehensive residual vector. And uncertainty: in, For the first A local residual vector predicted by an expert;

[0102] 3. Third Stage: Adaptive Fusion Output:

[0103] To address the issue of unobservable yaw angles under weak curvature, an amplitude-based approach is adopted. Adaptive weights The baseline vector is fused with the residual vector predicted by the neural network: Finally, the predicted yaw angle is obtained through back projection. and its variance: The yaw angle φ prediction branch ensures that the system automatically reverts to a stable geometric baseline during weak curves, while compensating for strong curves using the high-precision residuals of the neural network. It can accurately reflect the reliability of the current prediction (e.g., outputting maximum variance in weak curves).

[0104] More specifically, the training strategy and online self-calibration module of the present invention include the following steps:

[0105] The training strategy and online self-calibration module describe the system's learning framework. This invention employs a hybrid strategy of "offline course-based pre-training + online incremental adaptation." In the offline phase, high-precision VINS data is used as "teacher signals" to build a robust base model; in the online phase, high-confidence samples are selected for self-iteration, endowing the system with long-term survivability against sensor aging, temperature drift, and environmental changes.

[0106] Step 1, Offline Training Strategy: Uncertainty-Based Supervised Learning

[0107] Offline training aims to establish an initial mapping from strain features to pose and to give the model self-evaluation capabilities (i.e., output variance):

[0108] 1) Data Construction and "Strong Bending Priority" Course Learning: In order to solve the problem of flexible arms in a straight line state ( To address the training divergence problem caused by poor observability, this invention employs a course-based learning strategy.

[0109] First, define the strong bend dataset. : in, The bending amplitude threshold (e.g., taking the 20th percentile of the data).

[0110] In the early stages of training only The model is iterated on to ensure it quickly learns the geometric features of the principal bending plane; then the full dataset is gradually added. Fine-tuning is performed, and noise in weak curve data is automatically processed using a gating mechanism;

[0111] 2) Bending angle probability loss function

[0112] For the bending angle branch, the training objective is to simultaneously optimize prediction accuracy and uncertainty estimation. Heteroscedastic negative log-likelihood (NLL) loss is used as the core objective, supplemented by geometric constraints, and defined as follows: Total channel loss : Geometric consistency loss Combining triangular head cosine similarity with direct head Huber loss ensures the geometric continuity of predicted values ​​on the unit circle; uncertainty loss. Forced model learning predicts variance : The above formula forces the model to operate with relatively large errors ( Larger regions actively predict larger variances To reduce losses, thereby achieving "self-awareness of uncertainty";

[0113] 3) Yaw angle The residual probability loss function:

[0114] For the yaw angle branch, the network learns the residuals relative to the geometric baseline, and heteroscedastic negative log-likelihood (NLL) loss is used to supervise the network's output. The definition is... Channel loss : This loss function ensures that the expert network can not only correct the deviation of the geometric baseline, but also accurately assess the reliability of the current sector prediction;

[0115] 4) Self-supervision of gating networks:

[0116] To enable the gating network to automatically learn to distinguish between "strong bend" and "weak bend" conditions, an amplitude-based approach is introduced. Soft tag supervision: in for The quantiles guide the network to automatically increase the weights on the Direct Head during strong bends;

[0117] Step 2, Online self-calibration mechanism: Anti-drift continuous learning

[0118] During operation, high-confidence segments from VINS are used as "pseudo-true values" to continuously fine-tune the "strain-visual model":

[0119] 1) High-confidence sample filter

[0120] Not all VINS data is suitable for training. The system monitors VINS confidence levels in real time. The data at the current time is valid only if the following conditions are met. Only then will it be collected: Furthermore, the IMU / reprojection error is stable, which ensures the purity of the signal itself and prevents the positioning error of the VINS from being transmitted to the strain model.

[0121] 2) Experience Replay Buffer

[0122] To prevent "catastrophic forgetting" in online learning, a fixed-capacity first-in-first-out (FIFO) queue is constructed. : When sampling batches for training, an amplitude-based approach is used. The bucket sampling strategy ensures that the training data is evenly distributed across different degrees of curvature, thus avoiding overfitting of the model to a specific pose.

[0123] 3) Incremental parameter fine-tuning

[0124] A mixed-batch strategy is used for parameter updates, with each iteration using the training set... From the new data currently collected and historical data in the buffer pool Composition, online optimization target for: in, With a very small learning rate, the model can correct the zero-point shift caused by strain gauge creep or temperature changes in real time through this small-step, fast-paced approach.

[0125] 4) Safety rollback and EMA protection

[0126] To prevent online training divergence, the system maintains an exponential moving average (EMA) model for actual inference and monitors validation error. Once a continuous increase in verification error is detected, or If the system remains sluggish for an extended period, it will immediately discard the current update and roll back to the last stable parameter checkpoint to ensure the absolute safety of flight operations.

[0127] Furthermore, the vision-strain pose fusion and state estimation layer specifically comprises:

[0128] The vision-strain pose fusion and state estimation layer elaborates on the core decision-making mechanism of the system—the adaptive fusion mechanism of heterogeneous information sources. Based on the "self-evaluation uncertainty" output by the strain network in the strain-vision module and the environmental confidence of the VINS system, a Bayesian inference framework based on inverse variance weighting is constructed. Through adaptive extended Kalman filtering (Adaptive EKF) for observation noise, smooth and optimal estimation of the pose and angular velocity of the flexible arm end effector under varying operating conditions is achieved.

[0129] Step 1: Unified Quantification of Uncertainty of Heterogeneous Information Sources

[0130] To achieve effective integration of the visual system (external observation) and the strain system (proprioception), it is first necessary to map the confidence indices of the visual system and the strain system to statistical observation variances, using this as the sole criterion for measuring the "credibility weight" of their respective data.

[0131] 1) Variance mapping of strain-visual models:

[0132] Based on the heteroscedastic neural network architecture built with the strain-vision module, the strain-vision model can directly output the variance of the predicted values ​​during the inference phase according to the distribution of the input features.

[0133] For bending angle and yaw angle Define its observation variance as follows: The physical meaning of this formula is: the variance of the network output. It directly quantifies the model's "self-doubt" about the current prediction results; for example, when the flexible continuum robot arm 2 is in an extreme posture with insufficient training data coverage or when there are abnormal fluctuations in the strain signal, the network will automatically output a very large variance value, thereby actively reducing its own weight in subsequent fusion.

[0134] 2) Equivalent variance modeling of the VINS system:

[0135] Confidence level of VINS system output This mainly reflects the number of feature points, tracking stability, and reprojection error. To align with the variance of the strain model, an inverse proportional mapping model from confidence level to variance needs to be established, defining the equivalent observation variance of VINS. as follows: in, To adjust the scaling factor between the dimensions of the vision and strain systems; To prevent tiny constants with a denominator of zero; As the penalty factor in the near-extended state, it can be seen that when the visual environment deteriorates ( When the variance decreases, its equivalent variance will increase exponentially until it approaches infinity.

[0136] Step 2, Static Fusion: Weighted Inverse Variance Based on Maximum Likelihood

[0137] Without considering time correlation, assuming that the errors of two independent sources follow a Gaussian distribution, the optimal fusion of the two should follow the "inverse variance weighting" rule to maximize the posterior probability.

[0138] 1) Bending angle Scalar fusion: For non-periodic bending angles, the reciprocals of the two differences are used as weights for linear combination to calculate the fused bending angle. : The above formula ensures that the fusion result always favors the side with smaller variance; for example, when visual loss ( When ), the limit of the formula is . This achieves a "soft switching" effect where the strain model seamlessly takes over.

[0139] 2) Yaw angle Vector domain fusion:

[0140] For the periodicity of yaw angle ( To avoid the cutting error caused by arithmetic averaging (the jump problem), we map it to the unit circle space and perform vector weighting, defining a weighted fusion vector. : The yaw angle after fusion was then extracted using inverse trigonometric functions. : Guarantee in The nearby fusion trajectory maintains geometric continuity, eliminating the impact of angle jumps on the control system;

[0141] Step 3, Dynamic Fusion: Observation Noise Adaptive EKF

[0142] To obtain the angular velocity information of the flexible arm's end effector and filter out high-frequency noise using kinematic constraints, an observation noise-adaptive extended Kalman filter (Adaptive EKF) was designed based on static fusion.

[0143] 1) State prediction and kinematic modeling

[0144] Select a state vector that includes the angle and its first derivative. A constant angular velocity model is used for time updates;

[0145] The state prior prediction equation is as follows: By using the angular velocity of the previous moment to predict the current pose, a smoothing constraint in the time dimension is introduced, which can effectively suppress instantaneous glitch noise from the sensor.

[0146] 2. Real-time construction of the observation noise matrix

[0147] Static fusion: The fused pose is calculated based on the inverse variance weighted by maximum likelihood. As observed values, their corresponding synthetic variances are used to dynamically fill the observation noise covariance matrix. Construct the real-time observation noise matrix : The magnitude of this matrix value reflects the reliability of the current fusion result in real time. When both visual and strain data are unreliable (e.g., due to severe vibration and weak texture), The surge in Kalman gain leads to a decrease in the filter's accuracy, causing it to automatically refuse updates and instead rely on prior predictions. This prevents the output from diverging.

[0148] 3) Posterior state update

[0149] Using the calculated Kalman gain The prior state is corrected to obtain the final optimal estimate. The state update equation is as follows: Final output It not only contains the fused high-precision pose information, but also the smoothed angular velocity estimate, which is directly used as the feedback signal for the subsequent closed-loop control system.

[0150] The above description is only used to illustrate the technical solution of the present invention and is not intended to limit it. Any other modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention, as long as they do not depart from the spirit and scope of the technical solution of the present invention, should be covered within the scope of the claims of the present invention.

Claims

1. A vision-strain fusion-based aerial flexible arm pose estimation and online self-correction system, characterized in that: It includes a physical system framework and an information processing framework: The physical system framework includes a quadcopter drone as a mobile base, a flexible continuum manipulator mounted on the quadcopter drone as an execution carrier, a resistance strain gauge mounted on the base of the flexible continuum manipulator, a depth camera module mounted at the end of the flexible continuum manipulator, and a drive unit for driving the flexible continuum manipulator. An embedded computing platform for real-time data processing is installed between the quadcopter drone and the flexible continuum manipulator. The information processing framework is a layered, modular software framework, including a perception layer, a vision-strain pose fusion and state estimation layer, and a decision and control layer, specifically: The perception layer includes a VINS module, a strain-visual model module, and a training strategy and online self-calibration module. VINS module: Receives image and IMU data streams from the depth camera module, runs the VINS-Fusion algorithm, and outputs the 6-DOF pose of the flexible arm's end effector in real time. The arm flexion angle and yaw angle were extracted by pose. And a confidence level that quantifies the quality of the current estimate. , among which angle All are defined in the base coordinate system Below, the angle of arm bending and the direction angle are indicated; Strain-Vision Module: Receives four strain signal vectors from resistance strain gauges. It is then fed into a pre-trained deep neural network model, which outputs the predicted pose of the arm's end effector in real time. And the uncertainty of the model in this prediction ; Training strategy and online self-calibration module: Based on a robust model foundation of offline course-based pre-training, a closed-loop mechanism of sample selection, experience replay, and safety rollback is adopted: the system continuously monitors the confidence level of the VINS module. ,when When the value exceeds the preset threshold, it indicates that the visual positioning is extremely reliable. At this time, the VINS pose and strain reading data are considered as high-quality samples. A mixed training set is constructed by combining the empirical replay buffer. The strain-pose model is fine-tuned online and incrementally using heteroscedasticity probability loss and geometric consistency constraints. At the same time, exponential moving average and validation monitoring are introduced. Once a decline in model performance is detected, rollback is automatically triggered to ensure that the system maintains absolute stability while continuously resisting sensor aging and temperature drift. The visual-strain pose fusion and state estimation layer is responsible for intelligently processing and fusing information from the perception layer to obtain the optimal estimate of the system's current state. This layer includes a confidence-weighted fusion module and a dynamic filtering module, specifically: Confidence-weighted fusion module: This module will be based on From the adaptive weight allocation, the visual confidence and the prediction variance of the strain network are first quantified into the observation variance, and the intermediate fused pose is calculated by using an inverse variance weighting strategy. The dynamic filtering module, based on a unified uncertainty quantification mechanism for heterogeneous information sources, maps the confidence score of the vision system to an equivalent observation variance with the same dimensions as the prediction variance of the strain network, and constructs an adaptive observation noise matrix accordingly. Subsequently, it uses a kinematic model to perform Kalman filtering smoothing and differentiation on the pose, ultimately outputting the system's optimal state vector containing angles and angular velocities. ; Decision and control layer: Responsible for generating and executing control commands based on optimal state estimation and task objectives.

2. The aerial flexible arm pose estimation and online self-correction system based on vision-strain fusion according to claim 1, characterized in that: The physical system framework is specifically as follows: Quadcopter UAVs: Serving as a mobile base, they provide a wide range of movement and hovering capabilities within three-dimensional space. The UAV's coordinate system is denoted as... Its origin is located at the center of mass of the spacecraft; Flexible continuum robotic arms: Composed of multiple serially connected, tendon-driven continuum structures, internally containing a semi-rigid central skeleton and equidistant disks, covered by an elastic cladding. Flexible continuum robotic arms are lightweight, highly compliant, and inherently safe. The base coordinate system of the flexible continuum robotic arm is denoted as... ,and Rigid connection, with only fixed vertical displacement. The attitude of the flexible continuum robotic arm is provided by the inertial measurement unit of the quadcopter drone; Drive unit: Four high-precision digital servos are mounted on the top of the flexible continuum robotic arm, and the length of four low-elasticity tendons is controlled by a reel. The tendon passes through and is fixed to the distal disc at 0°, 90°, 180°, and 270° along the circumference of the arm, changing the position of the tendon. This allows the arm to bend and extend in any direction, with the driving space denoted as... ; Depth camera module: also known as external perception-visual inertial unit, is installed on the end effector of a flexible continuum robotic arm. Its visual and inertial information jointly drive the VINS-Fusion algorithm to output high-precision 6-DOF pose. and real-time confidence level The depth camera has an RGB resolution of 1920×1080@30fps and a depth resolution of 1280×720@90fps, with a built-in IMU sampling at 400Hz. Resistance strain gauges, also known as internal sensing-body strain units, are installed on the base of the flexible continuous manipulator. Four resistance strain gauges are symmetrically attached along the circumference of the flexible continuous manipulator in orthogonal directions of 0°, 90°, 180°, and 270°. The resistance strain gauges can accurately measure the minute strain generated on the surface of the flexible continuous manipulator root due to overall bending. The orthogonal layout can decouple the deformation information of the arm body on the two main bending planes to the greatest extent, providing a structured and information-rich raw input for subsequent machine learning models. A full-bridge circuit configuration is used, and sampling is performed through a 24-bit high-precision analog-to-digital converter at a frequency of 1000Hz.

3. The aerial flexible arm pose estimation and online self-correction system based on vision-strain fusion according to claim 1, characterized in that: The decision-making and control layer includes an inverse kinematics module and a closed-loop controller module, specifically: Inverse kinematics module: Receives the desired pose It is analyzed as the expected change in length of the four tendons in the driving space. ; Closed-loop controller module: It receives the output of the inverse kinematics module as a feedforward signal, and at the same time receives the optimal pose and angular velocity estimate from the fusion layer as a full-state feedback signal. Subsequently, it executes a two-layer closed-loop control law, that is, it uses angular velocity for feedforward compensation and D-term damping control to calculate the precise command to drive the four servo motors, thereby completing the precise tracking of the pose of the flexible arm end effector.

4. The aerial flexible arm pose estimation and online self-correction system based on vision-strain fusion according to claim 1, characterized in that: The VINS module performs the following steps: Step 1, Pose parameterization and angle extraction: VINS outputs end pose To serve the control of the flexible continuous body robotic arm, it is necessary to... Extraction from the system and ; Coordinate transformation: ; Bending angle : ; Yaw angle : ; Singular value handling: When At that time, it was believed Unstable, take the previous moment Or reduce its gain to avoid jumps in the near-extended state; Step 2, Dynamic Filtering and Timing Smoothing To suppress high-frequency jitter in visual localization and eliminate outliers, the angle sequence output by VINS is smoothed and preprocessed. A Kalman filter with a constant velocity model is constructed to suppress visual observation noise before the data is fed into the fusion layer, ensuring the continuity of the input to the subsequent fusion module. During the update, the angle difference is normalized around: use A smooth filter can be obtained by completing the filter update. and its angular velocity.

5. The aerial flexible arm pose estimation and online self-correction system based on vision-strain fusion according to claim 1, characterized in that: The strain-vision module includes the following steps: Step 1: Feature Vector Construction and Physical Preprocessing 1) Baseline correction and orthogonal decoupling: First, baseline correction is performed on the original voltage signal, and the corrected voltage signal is defined. as follows: in, This is the static bridge baseline voltage. To collect voltage in real time; Subsequently, by utilizing the orthogonal arrangement of resistance strain gauges on the circumference of the flexible continuum robotic arm, a principal axis differential feature reflecting the deformation of the two principal bending planes was constructed. and : in, The bending component approximately corresponds to the 0°-180° plane. The bending component approximately corresponds to the 90°-270° plane; 2) Rotationally invariant amplitude and coarse phase: Construct rotation-invariant amplitude characteristics and a rough phase angle : in, With bending angle It shows a positive correlation and does not change with the rotation of the boom, making it a key indicator for judging strong / weak bending conditions. A priori estimate of the yaw angle is provided, but it includes a fixed deviation due to assembly errors. ; 3) Final feature vector assembly: By introducing the sum of the absolute values ​​of statistical features and the sum of the two largest terms, and by performing triangular encoding on the periodic angles, an 11-dimensional input feature vector is finally constructed. for: in, The sum of absolute values The sum of the two largest terms; this eigenvector After standardization, the data is input into the neural network. Step 2, Overall Architecture of Deep Neural Network It contains two independent parallel branches, one for predicting the other for predicting the non-periodic bending angle. and periodic yaw angle Each branch not only outputs the predicted angle, but also forces the output of the logarithmic variance of the prediction through the heteroscedastic regression head to quantify the uncertainty.

6. The aerial flexible arm pose estimation and online self-correction system based on vision-strain fusion according to claim 5, characterized in that: The bending angle The prediction branch is as follows: Bending angle It is a non-periodic variable, but in weak bending Due to poor observability under operating conditions, a hybrid architecture combining triangular coding, direct regression, and gating fusion was added. 1) Hybrid output head design It contains three parallel output heads: Trig header: Outputs unit circle vector Geometric constraints are used to ensure the continuity of the predicted values; Direct head: Directly outputs scalar angle Capture local linear relationships; Uncertainty Header: Outputs the log-variance of the predictions ; 2) Gating weight generation: To adaptively select the optimal prediction strategy under different bending intensities, a gating network is designed based on the amplitude. Generate weights : in, Using the Sigmoid activation function, the network is trained to trust the Direct sensor more under strong bending conditions and the Trig sensor more under weak bending conditions. 3) Final output of the branch: First, back-project the Trig head as an angle. Then, a weighted fusion is performed to obtain the final predicted value. : Simultaneously, the variance is directly output for subsequent confidence calculation: in, It is determined by the gating weight Adjusted fusion coefficient; The yaw angle φ prediction branch specifically includes the following: 1) First stage: Geometric baseline calculation: Utilizing the aforementioned principal axis difference features By using a pre-fitted polynomial design matrix Perform ridge regression and calculate the geometric baseline vector. : in, For the regression coefficient matrix fitted offline, the geometric baseline vector model provides a stable but systematic preliminary estimate; 2) Second stage: Expert residual learning and uncertainty estimation The neural network aims to learn the residual angle relative to the baseline. The network adopts a hybrid expert model structure, which will Divided into 8 sectors, the gated network is based on features Calculate the activation probability of experts in each sector and output the comprehensive residual vector. And uncertainty: in, For the first A local residual vector predicted by an expert; 3) Third stage: Adaptive fusion output: Using amplitude-based Adaptive weights The baseline vector is fused with the residual vector predicted by the neural network: Finally, the predicted yaw angle is obtained through back projection. and its variance: The yaw angle φ prediction branch ensures that the system automatically reverts to a stable geometric baseline during weak curves, while compensating for strong curves using the high-precision residuals of the neural network. It can accurately reflect the credibility of current predictions.

7. The aerial flexible arm pose estimation and online self-correction system based on vision-strain fusion according to claim 1, characterized in that: The training strategy and online self-calibration module include the following steps: Step 1, Offline Training Strategy: Uncertainty-Based Supervised Learning Offline training aims to establish an initial mapping from strain features to pose and to give the model self-evaluation capabilities: 1) Data Construction and Strong Bending Priority Course Learning: First, define the strong bend dataset. : in, This is the threshold for bending amplitude; In the early stages of training only The model is iterated on to ensure it quickly learns the geometric features of the principal bending plane; then the full dataset is gradually added. Fine-tuning is performed, and noise in weak curve data is automatically processed using a gating mechanism; 2) Bending angle probability loss function For the bending angle branch, the training objective is to simultaneously optimize prediction accuracy and uncertainty estimation. Heteroscedasticity negative log-likelihood loss is used as the core objective, supplemented by geometric constraints, and defined as follows: Total channel loss : ; Geometric consistency loss By combining triangular head cosine similarity with direct head Huber loss, the geometric continuity of the predicted values ​​on the unit circle is ensured. Uncertainty loss Forced model learning predicts variance : The above formula forces the model to actively predict larger variances in regions with larger errors. To reduce losses; 3) Yaw angle The residual probability loss function: For the yaw angle branch, the network learns the residuals relative to the geometric baseline, and heteroscedastic negative log-likelihood loss is used to supervise the network's output. The definition is... Channel loss : This loss function ensures that the expert network can not only correct the deviation of the geometric baseline, but also accurately assess the reliability of the current sector prediction; 4) Self-supervision of gating networks: Introducing amplitude-based Soft tag supervision: in for The quantiles guide the network to automatically increase the weight of the direct regression head during strong bends; Step 2, Online self-calibration mechanism: Anti-drift continuous learning During operation, high-confidence segments from VINS are used as pseudo-true values ​​to continuously fine-tune the strain-visual model: 1) High-confidence sample filter The system monitors VINS confidence level in real time. The data at the current time is valid only if the following conditions are met. Only then will it be collected: And the IMU / reprojection error is stable. This ensures the purity of the signal itself and prevents the positioning errors of the VINS from being transmitted to the strain model; 2) Experience replay buffer pool To prevent catastrophic forgetting in online learning, a fixed-capacity first-in-first-out queue is constructed. : When sampling batches for training, an amplitude-based approach is used. The bucket sampling strategy ensures that the training data is evenly distributed across different degrees of curvature, thus avoiding overfitting of the model to a specific pose. 3) Incremental parameter fine-tuning A hybrid batch strategy is used for parameter updates, with each iteration using the training set... From the new data currently collected and historical data in the buffer pool Composition, online optimization target for: in, With a very small learning rate, the model can correct zero-point drift caused by strain gauge creep or temperature changes in real time through this small-step, fast-paced approach. 4) Safety rollback and EMA protection To prevent divergence during online training, the system maintains an exponential moving average model for actual inference and monitors validation error. Once a continuous increase in verification error is detected, or If the system remains sluggish for an extended period, it will immediately discard the current update and roll back to the last stable parameter checkpoint to ensure the absolute safety of flight operations.

8. The aerial flexible arm pose estimation and online self-correction system based on vision-strain fusion according to claim 1, characterized in that: The vision-strain pose fusion and state estimation layer is specifically as follows: Step 1: Unified Quantification of Uncertainty of Heterogeneous Information Sources To achieve effective integration of the vision system and the strain system, it is first necessary to map the confidence indices of the vision system and the strain system to statistical observation variance, using this as the sole criterion for measuring the reliability weight of their respective data. 1) Variance mapping of strain-visual models: During the inference phase, the strain-visual model can directly output the variance of the predicted values ​​based on the distribution of the input features. For bending angle and yaw angle Define its observation variance as follows: The physical meaning of this formula is: the variance of the network output. It directly quantifies the model's degree of self-doubt regarding the current prediction results; 2) Equivalent variance modeling of the VINS system: Confidence level of VINS system output This mainly reflects the number of feature points, tracking stability, and reprojection error. To align with the variance of the strain model, an inverse proportional mapping model from confidence level to variance needs to be established, defining the equivalent observation variance of VINS. as follows: in, To adjust the scaling factor between the dimensions of the vision and strain systems; To prevent tiny constants with a denominator of zero; As the penalty factor in the near-extended state, it can be seen that when the visual environment deteriorates, its equivalent variance will increase exponentially until it approaches infinity. Step 2, Static Fusion: Weighted Inverse Variance Based on Maximum Likelihood 1) Bending angle Scalar fusion: For non-periodic bending angles, the reciprocals of the two differences are used as weights for linear combination to calculate the fused bending angle. : The above formula ensures that the fusion result always favors the side with smaller variance; 2) Yaw angle Vector domain fusion: To address the periodicity of the yaw angle and avoid clipping errors caused by arithmetic averaging, it is mapped to a unit circle space and weighted by vector. Define a weighted fusion vector : The yaw angle after fusion was then extracted using inverse trigonometric functions. : Guarantee in The nearby fusion trajectory maintains geometric continuity, eliminating the impact of angle jumps on the control system; Step 3, Dynamic Fusion: Observation Noise Adaptive EKF 1) State prediction and kinematic modeling Select a state vector that includes the angle and its first derivative. A constant angular velocity model is used for time updates; The state prior prediction equation is as follows: By using the angular velocity of the previous moment to predict the current pose, a smoothing constraint in the time dimension is introduced, which can effectively suppress instantaneous glitch noise from the sensor. 2) Real-time construction of the observation noise matrix Static fusion: The fused pose is calculated based on the inverse variance weighted by maximum likelihood. As observed values, their corresponding synthetic variances are used to dynamically fill the observation noise covariance matrix. Construct the real-time observation noise matrix : The magnitude of this matrix value reflects the reliability of the current fusion result in real time. When both vision and strain are unreliable, The surge in Kalman gain leads to a decrease in the filter's accuracy, causing it to automatically refuse updates and instead rely on prior predictions. This prevents the output from diverging. 3) Posterior state update Using the calculated Kalman gain The prior state is corrected to obtain the final optimal estimate. The state update equation is as follows: Final output It not only contains the fused high-precision pose information, but also the smoothed angular velocity estimate, which is directly used as the feedback signal for the subsequent closed-loop control system.

Citation Information

Patent Citations

  • Multi-rotor and airborne mechanical arm combined position and posture control method based on visual servo control

    CN106363646A

  • Flexible body large-deformation space pose sensor and flexible body robot

    CN211205342U