Visual servo tracking method for marine target

Through multimodal sensor fusion and adaptive visual tracking algorithm, combined with servo control and visual collaboration mechanism, the tracking failure problem of marine target visual servo tracking in complex sea conditions is solved, high-precision and real-time target tracking is achieved, and it adapts to lighting changes and platform shaking, ensuring the stability and continuity of tracking.

CN120669761APending Publication Date: 2025-09-19HAINAN UNIV
View PDF 0 Cites 18 Cited by

Patent Information

Application Number
CN202510715257.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing marine target visual servo tracking technology has difficulty ensuring real-time and robustness under complex sea conditions, and cannot adapt to tracking failure problems in scenarios such as lighting changes, platform shaking, and target occlusion.

Method used

By fusing data from visual sensors, inertial measurement units, and global navigation satellite systems through a multimodal sensor fusion module, and using the extended Kalman filter algorithm for joint estimation, combined with an adaptive visual tracking algorithm and servo control and visual collaboration mechanism, and adopting an improved target detection model and occlusion processing module, accurate tracking and re-identification of targets can be achieved.

Benefits of technology

It effectively compensates for the effects of ship shaking and illumination changes caused by waves, improves the accuracy of target detection and tracking stability, ensures real-time performance and reliability in complex marine environments, and reduces tracking delays and mistracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669761A_ABST
    Figure CN120669761A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of marine monitoring, in particular to a marine target visual servo tracking method, which comprises the steps of multi-modal sensor fusion, a self-adaptive visual tracking algorithm, a servo control and visual collaboration mechanism and a shielding processing and target re-identification strategy. The problem that tracking is unstable under the conditions of illumination change, ship body shaking, target shielding and the like in a traditional method is solved. The IMU, the GNSS and the visual data are fused through extended Kalman filtering, ship body shaking is compensated, and the target state estimation precision is improved; the improved D-Fi ne target detection model is combined with an online feature updating mechanism to dynamically adapt to the appearance change of the target; the prediction and correction control strategy and the double-closed-loop PI D controller cooperate to adjust the camera holder, and the tracking delay is reduced; the multi-clue shielding detection and space-time joint feature matching technology ensures accurate re-identification of the target after shielding is removed. The real-time performance and robustness of the tracking system on an embedded platform are improved, and an efficient and stable target tracking solution is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ocean monitoring, and in particular to a method for visual servo tracking of ocean targets. Background Art

[0002] Maritime safety needs to be ensured through ocean monitoring, and unmanned boats are increasingly used in maritime monitoring tasks. In anti-smuggling and anti-smuggling scenarios, unmanned boats need to use the visual sensors they carry to track suspicious targets in real time and provide data support for subsequent law enforcement. However, the tracking accuracy and stability are still insufficient.

[0003] First, the complex lighting conditions in the ocean environment. Factors such as the alternation of day and night, cloud cover, and water reflections can cause dramatic changes in image quality, making target feature extraction more difficult. In strong light, targets may appear overexposed, while in low light, features may be lost due to noise interference, thus affecting the accuracy of the tracking algorithm.

[0004] Secondly, the motion of the waves causes the unmanned boat platform to continuously shake, resulting in dynamic distortion in the images captured by the visual sensors. Traditional tracking methods based on fixed cameras struggle to compensate for image offsets caused by platform shaking, easily leading to target loss. Furthermore, the target's rapid movement (such as a speedboat escaping) or changes in posture (such as turning or accelerating) can also increase tracking difficulty, making existing algorithms often unable to adjust tracking strategies in a timely manner.

[0005] Furthermore, targets at sea may temporarily disappear from view due to obstructions such as other ships and buoys. Traditional tracking algorithms struggle to quickly re-identify targets when they re-enter the field of view, leading to tracking interruptions. Furthermore, background interference in complex sea conditions (such as wave patterns and floating objects) can easily cause algorithms to mistake background features for targets, further reducing tracking reliability.

[0006] At present, some studies have attempted to improve tracking performance by improving tracking algorithms (such as introducing deep learning models) or combining other sensors (such as radar), but the following problems still exist: deep learning models have low running efficiency on edge computing devices (such as embedded systems carried by unmanned boats), and real-time performance is difficult to guarantee; multi-sensor fusion strategies often rely on pre-calibration and cannot adapt to the rapid changes in the dynamic environment at sea; the servo control system and the visual algorithm lack coordination, resulting in a lag in camera posture adjustment, affecting tracking continuity.

[0007] Therefore, there is an urgent need for a marine target visual servo tracking technology that can adapt to complex sea conditions and take into account both real-time and robustness, so as to solve the tracking failure problem of existing methods in scenarios such as lighting changes, platform shaking, and target occlusion. Summary of the Invention

[0008] In view of this, the purpose of the present invention is to propose a marine target visual servo tracking method to solve the technical problems of low operating efficiency, difficulty in ensuring real-time performance, and inability to adapt to rapid changes in the marine dynamic environment in the existing technology.

[0009] Based on the above objectives, the present invention provides a method for visual servo tracking of a marine target, comprising:

[0010] Step 1: Data from the visual sensor, inertial measurement unit (IMU), and global navigation satellite system (GNSS) are fused through a multimodal sensor fusion module. The extended Kalman filter (EKF) algorithm utilizes this multi-source fusion data to jointly estimate the UAV's attitude and target motion state. Images captured by the visual sensor are processed using an improved D-Fine target detection model to accurately extract the target's position within the image. The IMU acquires the platform's acceleration and angular velocity in real time to compensate for wave-induced hull motion; the GNSS provides the UAV's position and heading information to assist in correcting the visual sensor's viewing angle deviation. By fusing multi-source data, the accuracy of target state estimation is improved.

[0011] Step 2: An adaptive visual tracking algorithm dynamically extracts the target's color, texture, and contour features, constructing an online updated feature template. This model employs a feature pyramid network-based object detection model, combined with an online feature update mechanism. This model improves upon D-Fine to enhance detection of small and low-contrast targets. When lighting or target pose changes, the feature template is updated online using a particle filter algorithm to ensure robust tracking.

[0012] Step 3: The camera's gimbal angle is adjusted through servo control and visual coordination to compensate for the effects of the UAV's sway and target motion. Dual closed-loop control is introduced: the outer loop is a position loop, which calculates the gimbal adjustment based on the target's positional deviation in the image; the inner loop is a velocity loop, which optimizes the dynamic response of the gimbal motion using a PID controller. This mechanism monitors the gimbal's motion in real time, enabling the camera to quickly and smoothly follow the target, reducing tracking delays.

[0013] Step 4: The occlusion handling and target re-identification module maintains trajectory prediction when the target is occluded and resumes tracking through spatiotemporal joint feature matching after the occlusion is removed. When the target is occluded, the target's position is predicted using historical trajectory data and optical flow is used to track the motion trends of background features to distinguish between the occluder and the target. Once the target reappears, feature matching based on the improved DeepSORT algorithm is performed, combining spatiotemporal context information to confirm the target's identity and avoid mistracking.

[0014] The entire process transmits tracking data in real time and improves system performance through online optimization.

[0015] In step 1, the multimodal sensor fusion module primarily consists of state and observation model construction, extended Kalman filtering, a dynamic weight allocation mechanism, tight coupling optimization between vision and inertia, and real-time optimization. The improved D-Fine target detection model is used to detect images captured by the visual sensor and accurately obtain the position information of the target in the image.

[0016] The construction of the state model and observation model mainly includes:

[0017] 1. State vector definition

[0018] The system state vector is defined as:

[0019] Where, φ, θ, ψ are the roll angle, pitch angle, and bow angle of the USV (unit: rad), respectively; v x ,v y is the velocity component of the unmanned boat in the geographic coordinate system (unit: m / s); x ,p y is the position coordinate of the unmanned boat (unit: m); The target's motion velocity component in the image coordinate system (unit: pixel / s).

[0020] 2. State transfer equation

[0021] Considering the wave interference and the dynamic characteristics of the unmanned boat, a nonlinear state transition model is established:

[0022] x k+1 =f(x k ,u k )+w k ,

[0023] Where, f(·) is the nonlinear state transfer function, including the UAV attitude kinematic equation, velocity integral and target motion prediction; u k The acceleration and angular velocity input measured by the IMU; w k is the process noise, modeled as Gaussian white noise, and the covariance matrix is ​​Q k .

[0024] Aiming at the hull shaking caused by waves, a random wave model (ITTC wave spectrum) is introduced to convert it into a disturbance term of the attitude angle, thereby enhancing the adaptability of the state model to complex sea conditions.

[0025] 3. Observation equation

[0026] Fuse multi-source observation data and construct observation vectors:

[0027] z k =[z IMU ,z GNSS,z Vision ] T ,

[0028] in, It is the attitude angle directly measured by IMU; Position and heading information provided by GNSS; Vision =[x t ,y t ] T The center coordinates of the target detected by the vision sensor in the image coordinate system.

[0029] The observation equation is expressed as:

[0030] z k =h(x k )+v k ,

[0031] Among them, h(·) is the nonlinear observation function, v k is the observation noise, and the covariance matrix is ​​R k For visual observation, the world coordinates of the target are converted into image coordinates through the camera projection model, and a noise model is introduced to describe the influence of factors such as lighting and occlusion.

[0032] The extended Kalman filter (EKF) mainly includes:

[0033] 1. Prediction Step

[0034] 1.1. State Prediction

[0035] Estimated based on current status and IMU input u k , predict the next state through the nonlinear state transfer function f(·):

[0036]

[0037] 1.2. Covariance prediction

[0038] Calculate the state transition matrix And update the forecast covariance matrix:

[0039]

[0040] 2. Update steps

[0041] 2.1. Observational prediction

[0042] According to the predicted status The predicted observation value is calculated by the observation function h(·):

[0043]

[0044] 2.2. Kalman gain calculation

[0045] Calculate the observation Jacobian matrix And solve for the Kalman gain:

[0046]

[0047] 2.3. Status Update

[0048] Combined with the measured observation value z k+1 , update the state estimate:

[0049]

[0050] 2.4. Covariance Update

[0051] P k+1 =(IK k+1 H k+1 )P k+1|k

[0052] The dynamic weight allocation mechanism described above addresses the issue of sensor reliability in maritime environments, where sea conditions vary. By calculating the residual covariance between each sensor's measured and predicted values, the observation noise covariance matrix is ​​adaptively adjusted. When severe sea waves increase IMU noise, the weight of the IMU observation is reduced, increasing the confidence level of the GNSS and visual data, thereby improving the robustness of the fusion results.

[0053] This tightly coupled optimization of vision and inertia compares to traditional fusion methods that often process visual and inertial data separately. This new method combines the optimization of target motion and the attitude of the unmanned vehicle, directly incorporating the visually detected target position deviation into the state equation. The target's velocity is incorporated into the state vector, and optical flow is used to estimate the target's motion in the image. This, combined with IMU-derived hull motion compensation, enables more accurate target trajectory prediction.

[0054] The real-time optimization is implemented using a lightweight EKF to meet the computing resource constraints of the unmanned boat embedded system. By simplifying the nonlinear terms in the state transfer function and accelerating matrix operations using a parallel computing framework, the running time of the fusion algorithm is shortened at a high frame rate visual sampling frequency.

[0055] In the step 2, the adaptive visual tracking algorithm mainly includes an improved target detection model, multi-dimensional feature extraction and fusion, an online feature update mechanism and a multi-scale tracking strategy.

[0056] The improved target detection model has been improved to address the small size and low contrast of marine targets (such as small speedboats and buoys) at long distances or in low light conditions. The following improvements have been made to D-Fine:

[0057] 1. Feature Pyramid Enhancement: Introducing the attention mechanism into the Feature Pyramid Network (FPN), the response to small target features is enhanced through an efficient coordinate self-attention module.

[0058] 2. Loss function optimization: Design multi-task loss function, including target classification loss L class , bounding box regression loss L box and scale-aware loss L scale :

[0059] L total =αL class +βL box +γL scale ,

[0060] Among them, α, β, γ are weight coefficients, L scale Use logarithmic scale error to improve sensitivity to changes in target size.

[0061] 3. Lightweight design: Through pruning and quantization technology, the number of model parameters is reduced to meet the real-time computing requirements of the unmanned boat embedded system.

[0062] The multi-dimensional feature extraction and fusion mainly includes:

[0063] 1. Feature Representation and Fusion

[0064] The characteristic vector of the target is defined as:

[0065] f=[f color ,f texture ,f shape ] T

[0066] Among them, f color is the histogram feature of the HSV color space, which describes the target appearance by calculating the color distribution of pixels; f texture is the Histogram of Oriented Gradients (HOG) feature, which captures the edge and texture information of the target; f shape It is a contour moment feature (such as Hu moment) that characterizes the geometric shape of the target.

[0067] 2. Use weighted fusion strategy

[0068] f fusion =w1f color +w2f texture +w3f shape ,

[0069] Among them, the weights w1, w2, and w3 are dynamically adjusted through online learning to increase the weight of color features when the lighting changes drastically.

[0070] The online feature update mechanism introduces a particle filter algorithm to update the feature template online, which mainly includes:

[0071] 1. State transition model: Define the target state as position x, y, scale s and rotation angle θ, and the state transition equation is:

[0072] x k+1 =x k +v k +w k ,

[0073] Among them, v k is the target motion speed, w k is the process noise.

[0074] 2. Observation model: Calculate the Bhattacharyya coefficient between the candidate region and the template feature as a similarity measure:

[0075]

[0076] The higher the coefficient, the more similar the candidate region is to the target template.

[0077] 3. Particle weight update: Update particle weights based on observation likelihood:

[0078]

[0079] Among them, λ is the temperature parameter, which controls the concentration of weight distribution.

[0080] 4. Template update: When the tracking confidence is lower than the threshold, the target template is updated using the weighted average method:

[0081]

[0082] Among them, η is the learning rate, which is adaptively adjusted according to the target motion speed.

[0083] The multi-scale tracking strategy mentioned above mainly includes multi-scale tracking and occlusion processing, which are detailed as follows:

[0084] 1. Multi-scale pyramid tracking

[0085] To cope with the change in target scale, the size of the UAV increases as it approaches the target. An image pyramid is constructed to track the target at different scale levels:

[0086] Step 1: Generate a three-layer pyramid for the current frame image (original scale, 0.8 times, 1.2 times).

[0087] Step 2: Run the tracking algorithm independently on each layer of the pyramid to obtain multi-scale target state estimation.

[0088] Step 3: Combine the UAV motion information provided by the multimodal sensor fusion module and select the tracking result of the optimal scale layer to reduce computational redundancy.

[0089] 2. Occlusion Detection and Restoration

[0090] An occlusion detection method based on spatiotemporal context is designed. It calculates the feature similarity between the target region in the current frame and previous frames. If the similarity for three consecutive frames falls below a threshold, occlusion is detected. During occlusion, the target trajectory is maintained using the predicted state of a particle filter, and background motion is tracked using optical flow to eliminate interference from obstructions. After occlusion is removed, an improved DeepSORT algorithm is used for feature matching, combining historical trajectory and appearance features to reconfirm the target's identity.

[0091] The dynamic feature weight allocation is a feature weight dynamic adjustment strategy based on the entropy weight method. By calculating the information entropy E of each feature dimension j , reflecting the uncertainty of the feature. Define feature weight:

[0092]

[0093] When the lighting is stable, the entropy of the color feature is low and the weight increases; when the target maneuvers quickly, the weight of the texture and shape features increases.

[0094] This lightweight deep feature extraction method addresses the computing power limitations of embedded systems by designing a lightweight deep feature extraction network. Using MobileNetV3 as the backbone network, it reduces computational complexity through grouped convolutions and a linear bottleneck structure. An attention pooling layer is added at the end of the network to enhance feature focus on the target area.

[0095] In step 3, the servo control and vision collaboration mechanism achieves efficient collaboration between the camera pan / tilt and vision algorithm through prediction and correction control strategy, dual closed-loop PID control and motion compensation algorithm. The architecture of the servo control and vision collaboration mechanism mainly includes

[0096] 1. Visual processing module: outputs the position deviation (Δx, Δy) of the target in the image in real time.

[0097] 2. Servo controller: Calculates the gimbal angle adjustment (Δα, Δβ) based on the deviation.

[0098] 3. PTZ actuator: The motor drives the camera to move horizontally (azimuth angle α) and tilted (pitch angle β).

[0099] The control goal is to make the target always located in the center of the image, that is, to minimize the position deviation:

[0100]

[0101] The prediction and correction control strategy mentioned above uses the target motion prediction results obtained by multimodal sensor fusion to compensate for the delay of the visual system and the influence of hull sway, and calculates the angle that the gimbal needs to adjust in advance. It mainly includes:

[0102] 1. Target motion prediction: Utilize the target velocity provided by the multimodal sensor fusion module Predict the position of the target in the image at the next moment:

[0103]

[0104] Where T is the control period.

[0105] 2. Hull sway compensation: Calculate the offset of the camera view based on the roll angle φ and pitch angle θ measured by the IMU:

[0106] Δα comp =-k φ ·φ,Δβ comp =-k θ ·θ

[0107] Among them, k φ ,k θ is the compensation coefficient, which is adjusted through experiments.

[0108] 3. Calculation of control amount: Comprehensively predict the position deviation and shake compensation to obtain the total adjustment amount of the gimbal:

[0109] Δα=Δα pred +Δα comp ,Δβ=Δβ pred +Δβ comp

[0110] The dual closed-loop PID control algorithm is mainly divided into an outer loop (position loop) and an inner loop (speed loop) to improve the dynamic performance of the pan-tilt motion. The outer loop calculates the expected angular velocity based on the target position deviation. The inner loop uses a PID controller to adjust the motor drive voltage so that the actual angular velocity tracks the expected value. It mainly includes:

[0111] 1. Position ring design

[0112] Define the position deviation as:

[0113] e α =Δx·K p ,e β =Δy·K p

[0114] Among them, K pis the scaling factor that converts pixel deviation to angle deviation (unit: rad).

[0115] The desired angular velocity is calculated using the PD control algorithm:

[0116]

[0117] Among them, K p1 ,K d1 are the proportional and differential coefficients of the position loop.

[0118] 2. Speed ​​loop design

[0119] Define the speed deviation as:

[0120]

[0121] The motor drive voltage is calculated using the PID control algorithm:

[0122]

[0123] Among them, K p2 ,K i2 ,K d2 are the proportional, integral and differential coefficients of the speed loop.

[0124] The adaptive parameter tuning described above is designed based on fuzzy logic to meet the control requirements under different sea conditions:

[0125] 1. Define the input variables as position deviation e and deviation change rate The output variable is the PID parameter K p ,K i ,K d .

[0126] 2. Establish a fuzzy rule base. When e is large and hours, increase K p To respond quickly; when e is small and When large, increase K p To suppress overshoot.

[0127] 3. Real-time parameter adjustment through online fuzzy reasoning can improve the robustness of the system under wave interference.

[0128] The feedforward-feedback composite control introduces feedforward control to further compensate for the effects of hull sway and target acceleration:

[0129] 1.According to the acceleration measurement of IMU (a x ,a y ), calculate the pan-tilt disturbance caused by the hull shaking:

[0130]

[0131] Among them, k a1 ~k a4 is the disturbance compensation coefficient.

[0132] 2. Add the feedforward compensation to the input of the speed loop:

[0133]

[0134] This design reduces tracking errors caused by shaking.

[0135] In step 4, occlusion processing and target re-identification are achieved through multi-cue occlusion detection, trajectory prediction maintenance and spatiotemporal joint feature matching, thereby achieving robust tracking when occlusion occurs and accurate re-identification after occlusion is released.

[0136] The occlusion detection and tracking maintenance mainly include:

[0137] 1. Multi-cue occlusion detection mechanism

[0138] In order to accurately judge occlusion events, a multi-cue fusion detection method based on feature similarity and optical flow consistency is designed. It mainly includes the following contents:

[0139] 1.1. Feature Similarity Detection

[0140] Calculate the feature vector of the target area of ​​the current frame and the previous frame (color histogram f color , HOG feature f texture )’s cosine similarity:

[0141]

[0142] If S of N consecutive frames (e.g. N=3) feat Below the threshold τ feat , it is initially determined to be occlusion.

[0143] 1.2. Optical flow consistency detection

[0144] Apply the optical flow method to the target area and calculate the motion field v(x,y) of the pixel points. If the average amplitude of the optical flow field is Below the threshold τ flow , and the entropy value of the motion field H flow Above the threshold τ entropy (indicating an increase in motion disorder), it is confirmed that occlusion has occurred.

[0145] 1.3. Multi-cue fusion decision-making

[0146] Based on the feature similarity and optical flow detection results, occlusion is determined by logical threshold:

[0147]

[0148] 2. Trajectory Prediction and State Maintenance

[0149] During occlusion, Kalman filtering is used to predict the target position and maintain the tracking state.

[0150] 2.1. State Model

[0151] Define the target state as position (x, y), speed and acceleration The state transfer equation is:

[0152] x k+1 =Fx k +Bu k +w k

[0153] Among them, F is the state transfer matrix, B is the control input matrix (set to 0 here), w k is the process noise.

[0154] Observation model

[0155] During occlusion, visual observation fails and only relies on predicted state updates.

[0156] z k+1 =Hx k+1|k +v k

[0157] Among them, H is the observation matrix, v k is the observation noise (here the covariance matrix is ​​set to a large value to reduce the prediction weight).

[0158] 2.3. Status Update

[0159] The target state estimate is iteratively updated through Kalman filtering to ensure the continuity of the trajectory during occlusion.

[0160] The target re-identification and identity confirmation are mainly as follows:

[0161] 1. Joint spatiotemporal feature matching

[0162] After the occlusion is removed, the improved DeepSORT algorithm is used for target re-identification, combining appearance features with spatiotemporal context:

[0163] 1.1. Deep Appearance Feature Extraction

[0164] A lightweight convolutional neural network (MobileNetV3) is designed to extract a 256-dimensional feature vector of the target. An attention mechanism is introduced into the network to enhance the focus on the target's contour and texture.

[0165] 1.2. Spatial-temporal similarity calculation

[0166] Define comprehensive similarity S total is the appearance similarity S app and spatiotemporal similarity S st The weighted sum of:

[0167] S total =αS app +(1-α)S st

[0168] Among them, the appearance similarity Based on cosine distance calculation;

[0169] Spatiotemporal similarity Binding position distance d pos and speed difference d vel .

[0170] 1.3. Cascade Matching Strategy

[0171] Prioritize matching recent trajectories. If a match fails, expand the search range. Use the Hungarian algorithm to solve multi-target matching problems and ensure identity uniqueness.

[0172] 2. Dynamic template update mechanism

[0173] In order to adapt to the change of target posture, a template update strategy based on confidence is designed:

[0174] When the re-identification confidence is higher than the threshold τ conf When , the exponential weighted method is used to update the target template:

[0175]

[0176] Among them, γ is the learning rate, which is adaptively adjusted according to the target motion speed (at high speed, γ is increased).

[0177] If the re-identification fails for M consecutive times, the target detection is triggered to reinitialize the tracking.

[0178] The sea-condition-based occlusion prediction system combines multimodal sensor data (ship attitude from the IMU and drone heading from the GNSS) to establish an occlusion prediction model. A machine learning algorithm (random forest) is used to train the relationship between occlusion probability and sea condition parameters (wave height and visibility). When the predicted occlusion probability is high, the Kalman filter process noise covariance is increased in advance to improve the robustness of trajectory prediction.

[0179] The lightweight re-identification network proposed model compression and quantization technology to address the computing power limitations of the unmanned boat embedded system:

[0180] 1. Prune the deep feature extraction network and remove redundant connections.

[0181] 2. Using knowledge distillation technology, the knowledge of the teacher network (such as ResNet50) is transferred to the student network (such as MobileNetV3), while maintaining high recognition accuracy and improving the inference speed.

[0182] The multi-target occlusion distinction is designed in a multi-target scene based on trajectory intersection analysis:

[0183] 1. When two target trajectories intersect, the direction of motion of the occluded area is determined by the optical flow method.

[0184] 2. If the motion direction of the occluded area is consistent with the historical motion of one of the targets, then the target is determined to be occluded and the other target is an occluder.

[0185] The beneficial effects of this invention are as follows: 1. It effectively compensates for the effects of wave-induced hull sway and illumination changes. Real-time IMU correction of the unmanned boat's attitude and GNSS-assisted correction of visual perspective deviations, combined with visual target detection results, achieves high-precision estimation of the target's motion state.

[0186] 2. Enhanced anti-interference capabilities are achieved through an adaptive visual tracking algorithm. The target detection algorithm, combined with an online feature update mechanism, dynamically extracts the target's color, texture, and contour features, and optimizes the feature template in real time through particle filtering. The algorithm adaptively adjusts feature weights to ensure the effectiveness of the tracking model when sudden changes in illumination or rapid changes in the target's posture occur. A multi-scale tracking strategy further addresses the issue of varying target scale, improving tracking success rates even when the target's size varies. The lightweight design ensures real-time algorithm execution on embedded systems, meeting the real-time requirements of maritime missions.

[0187] 3. Servo control and vision are used in synergy to reduce tracking latency. A predictive and corrective control strategy, combined with a dual closed-loop PID controller, preemptively compensates for the effects of ship sway and target motion, enabling the gimbal to quickly and smoothly follow the target. Adaptive parameter tuning and feedforward and feedback control further enhance the system's dynamic response, shortening adjustment time during target acceleration and ensuring stable camera lock on the target.

[0188] 4. Tracking continuity is ensured through occlusion handling and target re-identification. A multi-cue occlusion detection mechanism (a fusion of feature similarity and optical flow consistency) accurately identifies occlusion events. Combined with trajectory prediction using a Kalman filter, it effectively maintains the target's state during occlusion. After occlusion is removed, a spatiotemporal joint feature matching algorithm achieves high-precision target re-identification by dually verifying deep appearance features and trajectory information.

[0189] 5. Through model pruning, quantization, and lightweight network design, the algorithm's operational efficiency on the unmanned vehicle embedded system was optimized, inference speed was increased, and real-time requirements were met. Mechanisms such as dynamic weight allocation and adaptive parameter tuning enable the system to automatically adjust its strategy based on changing sea conditions, increasing the weight of color features when illumination changes dramatically and enhancing the fusion weight of IMU data when waves interfere, thereby comprehensively improving the system's adaptability and reliability in complex marine environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0190] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0191] Figure 1 It is a schematic diagram of the overall flow of the marine target visual servo tracking technology and method of the present invention.

[0192] Figure 2 4 is a flow chart of a multimodal sensor fusion module in an embodiment of the present invention.

[0193] Figure 3 4 is a flow chart of an adaptive visual tracking algorithm in an embodiment of the present invention.

[0194] Figure 4 It is a flow chart of the servo control and visual collaboration mechanism in an embodiment of the present invention.

[0195] Figure 5 2 is a flow chart of occlusion processing and target re-identification in an embodiment of the present invention. DETAILED DESCRIPTION

[0196] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to specific embodiments.

[0197] It should be noted that, unless otherwise defined, the technical or scientific terms used in the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0198] The present invention Figure 1-Figure 5 As shown in the figure, a multimodal sensor fusion module fuses data from a visual sensor, an inertial measurement unit (IMU), and a global navigation satellite system (GNSS), and uses the extended Kalman filter (EKF) algorithm to jointly estimate the UAV's attitude and target motion state. The IMU acquires the platform's acceleration and angular velocity in real time to compensate for wave-induced hull motion. The GNSS provides the UAV's position and heading information, assisting in correcting the visual sensor's viewing angle deviation. By fusing multi-source data, the accuracy of target state estimation is improved.

[0199] In specific implementation, the system initialization in step 1 mainly includes:

[0200] 1. Hardware preparation

[0201] Select a suitable unmanned aerial vehicle platform equipped with industrial-grade visual sensors (such as high-definition cameras), a high-precision inertial measurement unit (IMU), a global navigation satellite system (GNSS), and an embedded processing unit with sufficient computing power. Install a servo gimbal that can achieve horizontal and pitch motion, and secure the visual sensor to the gimbal, ensuring it is securely mounted and has an unobstructed field of view.

[0202] 2. Software system deployment

[0203] Install the operating system and necessary drivers in the embedded processing unit to ensure proper communication between hardware devices. Deploy algorithms such as multimodal sensor fusion, adaptive visual tracking, servo control and visual collaboration, as well as occlusion handling and object re-identification.

[0204] 3. Sensor calibration

[0205] The visual sensor is calibrated using a checkerboard calibration plate. By shooting checkerboard images at different angles and positions, the camera's intrinsic parameters (such as focal length and distortion coefficient) and extrinsic parameters (such as rotation and translation matrices) are calculated to establish a mapping relationship between the image coordinate system and the world coordinate system.

[0206] Perform zero bias calibration and scale factor calibration on the IMU to improve the accuracy of attitude measurement.

[0207] Initialize the GNSS settings to ensure that it can accurately obtain the position and heading information of the unmanned boat.

[0208] Synchronize the timestamps of each sensor to ensure the time consistency of multi-source data.

[0209] In specific implementation, the multimodal sensor fusion in step 1, Figure 2 FIG. 1 is a flow chart of a multimodal sensor fusion module in an embodiment of the present invention. Figure 2 As shown, it mainly includes:

[0210] 1. Data Collection

[0211] The visual sensor collects sea surface image data at a certain frame rate (such as 30 frames per second).

[0212] The IMU collects the acceleration and angular velocity data of the unmanned boat in real time.

[0213] GNSS periodically obtains the position and heading data of the unmanned boat.

[0214] 2. Data Preprocessing

[0215] Denoise the visual image, such as using Gaussian filtering to remove noise; perform contrast enhancement, such as using histogram equalization to improve image clarity. Use the improved D-Fine target detection model to detect images captured by the visual sensor and accurately obtain the location information of the target in the image.

[0216] Filter the IMU data, such as using Kalman filtering or complementary filtering to remove high-frequency noise and drift errors.

[0217] Solve GNSS data, such as coordinate conversion and differential processing, to improve the accuracy of position measurement.

[0218] 3. Data Fusion

[0219] Construct the system state vector, which includes the attitude (roll, pitch, and bow), speed, position, and motion parameters of the unmanned boat.

[0220] The Extended Kalman Filter (EKF) algorithm is used to fuse multi-source data. The attitude and motion state of the unmanned vehicle are predicted based on the IMU data, the position deviation is corrected using GNSS data, and the target position data from the visual sensor is combined to update the target motion state estimate.

[0221] S2 uses an adaptive visual tracking algorithm to dynamically extract the target's color, texture, and contour features, constructing an online updated feature template. This model employs a feature pyramid network-based target detection model, combined with an online feature update mechanism. This model improves upon D-Fine to enhance detection of small and low-contrast targets. When lighting or target pose changes, the feature template is updated online using a particle filter algorithm to ensure robust tracking.

[0222] In specific implementation, the adaptive visual tracking in step 2, Figure 3 FIG. 1 is a flow chart of an adaptive visual tracking algorithm in an embodiment of the present invention. Figure 3 As shown, it mainly includes:

[0223] 1. Object Detection

[0224] The preprocessed visual image is fed into the improved D-Fine object detection model. This model, based on the original D-Fine, combines the Feature Pyramid Network (FPN) and the attention mechanism to enhance the detection capabilities of small and low-contrast objects.

[0225] The model outputs the location, category, and confidence information of the target.

[0226] 2. Feature Extraction

[0227] For the detected target area, the color histogram features are extracted to reflect the color distribution information of the target.

[0228] Extract HOG (Histogram of Oriented Gradients) texture features to describe the texture information of the target.

[0229] Contour moment features are extracted to represent the shape information of the target.

[0230] 3. Online feature template update

[0231] A particle filter algorithm is used to dynamically adjust the weights of particles based on the similarity between the current frame's target features and historical feature templates. When tracking confidence decreases, the target's feature template is updated using a weighted average method to adapt to changes in lighting and target posture.

[0232] 4. Multi-scale tracking strategy

[0233] An image pyramid is constructed for the current frame, including images at different scales, including the original scale, 0.8x scale, and 1.2x scale. The tracking algorithm is run independently on each scale layer. The target motion prediction results obtained through multimodal sensor fusion are combined, and the results from the scale layer with the best tracking performance are selected as the final tracking result.

[0234] S3 uses servo control and vision coordination to adjust the camera's gimbal angle, compensating for the effects of the UAV's sway and target motion. Dual closed-loop control is introduced: the outer loop is a position loop, which calculates gimbal adjustment based on the target's positional deviation in the image; the inner loop is a velocity loop, which uses a PID controller to optimize the dynamic response of the gimbal motion. This mechanism enables the camera to quickly and smoothly follow the target, reducing tracking delay.

[0235] In specific implementation, the servo control and vision in step 3 are coordinated. Figure 4 : is a flow chart of the servo control and visual collaboration mechanism in an embodiment of the present invention. Figure 4 As shown, it mainly includes:

[0236] 1. Position deviation calculation

[0237] Based on the position of the target in the image, calculate its pixel deviation from the center of the image.

[0238] 2. Predictive and corrective control

[0239] The target motion prediction results obtained by multimodal sensor fusion are used to calculate the angle that the gimbal needs to adjust in advance.

[0240] Combined with the hull attitude data measured by IMU, the camera viewing angle offset caused by waves is compensated.

[0241] 3. Double closed-loop PID control

[0242] Outer loop (position loop): converts pixel deviation into angle deviation and uses PD (proportional-differential) control algorithm to calculate the desired angular velocity of the gimbal.

[0243] Inner loop (velocity loop): Based on the deviation between the desired angular velocity and the actual angular velocity, a PID (proportional-integral-differential) control algorithm is used to generate the motor drive voltage, which drives the gimbal motor to rotate, allowing the camera to quickly and accurately align with the target.

[0244] Monitor the movement status of the gimbal in real time and dynamically adjust the PID control parameters according to actual conditions to ensure the stability and accuracy of the gimbal movement.

[0245] S4, an occlusion handling and target re-identification module, maintains trajectory prediction when the target is occluded and resumes tracking through spatiotemporal joint feature matching after the occlusion is removed. When the target is occluded, it uses historical trajectory prediction to predict the target's position and uses optical flow to track the motion trends of background features to distinguish between the occluder and the target. Once the target reappears, feature matching based on an improved DeepSORT algorithm is performed, combining spatiotemporal context to confirm the target's identity and avoid mistracking.

[0246] In the specific implementation, the occlusion processing and target re-identification in step 4 are Figure 5 FIG. 1 is a flow chart of occlusion processing and target re-identification in an embodiment of the present invention. Figure 5 As shown, it mainly includes:

[0247] 1. Occlusion detection

[0248] Calculate the feature similarity between the target area of ​​the current frame and the previous frame, such as using cosine similarity to measure the similarity of color and texture features.

[0249] Analyze the optical flow consistency of the target area and calculate the average amplitude and entropy of the optical flow field.

[0250] The feature similarity and optical flow consistency are combined to determine whether the target is occluded through multi-cue fusion method.

[0251] 2. Occlusion processing

[0252] When it is detected that the target is occluded, the Kalman filter algorithm is activated to predict its position during the occlusion period based on the target's historical motion state and maintain the target's tracking state.

[0253] 3. Occlusion removal judgment

[0254] Continuously monitor whether the target reappears in the field of view, and when certain re-identification conditions are met (such as the confidence of the target feature reaches a threshold), the occlusion is determined to be lifted.

[0255] 4. Target Re-Identification

[0256] After the occlusion is removed, target detection is performed again to obtain candidate targets.

[0257] Extract the deep appearance features of candidate targets and use lightweight convolutional neural networks for feature extraction.

[0258] The appearance similarity and spatiotemporal similarity between the candidate target and the historical target are calculated, and the identity of the target is confirmed by the spatiotemporal joint feature matching method.

[0259] If the match is successful, the target's feature template is updated and normal tracking is resumed; if the match fails, the tracking process is reinitialized.

[0260] S5. Transmit tracking data in real time and improve system performance through online optimization.

[0261] During specific implementation, system optimization and result output are as follows:

[0262] 1. Online Optimization

[0263] Regularly analyze performance indicators such as tracking error and gimbal response time, and use fuzzy logic or machine learning algorithms to dynamically adjust control parameters and algorithm strategies to adapt to different sea conditions and target motion states.

[0264] 2. Result output

[0265] The target's position, speed, tracking confidence and other information are transmitted to the control center through the wireless communication module to provide a basis for subsequent decision-making.

[0266] Store key data, such as image data, sensor data, control instructions, etc., for subsequent algorithm optimization and task review.

[0267] Through the above detailed implementation steps, this patented technology can enable unmanned boats to stably and accurately track moving targets in complex sea conditions, providing strong technical support for tasks such as maritime law enforcement.

[0268] In summary, the present invention solves the problems of traditional maritime target tracking, such as susceptibility to environmental interference, lack of real-time performance, and weak occlusion recovery capability, through the integration of multiple technologies and innovative design. It provides efficient and stable technical support for maritime law enforcement tasks such as anti-smuggling and stowaway, and has significant engineering application value.

[0269] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present invention (including the claims) is limited to these examples. Within the scope of the present invention, the above embodiments or technical features in different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the present invention as described above, which are not provided in detail for the sake of simplicity.

[0270] The present invention is intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for visual servo tracking of marine targets, characterized in that: The steps include: Step 1: First, the multimodal sensor fusion module fuses data from the visual sensor, inertial measurement unit, and global navigation satellite system. The improved D-Fine target detection model is used to obtain target position information in the image. The extended Kalman filter algorithm uses the multi-source fusion data to jointly estimate the UAV's attitude and target motion state, improving the accuracy of target state estimation. Step 2: Use the adaptive visual tracking algorithm to dynamically extract the target's color, texture, and contour features, build an online updated feature template, and use a target detection model based on a feature pyramid network, combined with an online feature update mechanism, to output accurate target location information. Step 3: Combining the target features obtained in steps 1 and 2, the camera gimbal angle is adjusted through servo control and visual coordination mechanisms, and dual closed-loop control is introduced to compensate for the effects of the UAV's shaking and the target's motion. The dual closed-loop control includes an outer position loop and an inner speed loop. The outer position loop calculates the gimbal adjustment amount based on the position deviation of the target in the image, and the inner speed loop optimizes the dynamic response of the gimbal motion through a PID controller. Step 4: Use the occlusion processing and target re-identification module to track the motion trend of background features through optical flow when the target is occluded, distinguish between the occluder and the target, maintain trajectory prediction, and resume tracking through spatiotemporal joint feature matching after the occlusion is released.

2. A method for visual servo tracking of marine targets according to claim 1, characterized in that: In the step 1, the multimodal sensor fusion module is constructed with the state model and observation model as the core. The state model first defines the state vector and considers the wave interference and the dynamic characteristics of the unmanned vehicle to establish a nonlinear state transfer model. The observation model fuses multi-source observation data to construct the observation vector and observation equation. The extended Kalman filter is used to fuse the acceleration and angular velocity data of the IMU, the position and heading data of the GNSS, and the target position data of the vision sensor to achieve joint estimation of the unmanned boat's attitude and target motion state.

3. A method for visual servo tracking of marine targets according to claim 1 or 2, characterized in that: In step 1, performing joint estimation using the extended Kalman filter algorithm includes the following steps: S1: prediction step, specifically including state prediction and covariance prediction; S2: Update step, which specifically includes state observation prediction, Kalman gain calculation, state update and covariance update.

4. A method for visual servo tracking of marine targets according to claim 1, characterized in that: In step 2, the adaptive visual tracking algorithm includes: an improved D-Fine target detection model, multi-dimensional feature extraction and fusion, an online feature update mechanism, and a multi-scale tracking strategy; Specific improvements to the improved D-Fine target detection model include feature pyramid enhancement, loss function optimization, and lightweight design. The multidimensional feature extraction and fusion include feature representation and fusion and adopt a weighted fusion strategy. The online feature update mechanism introduces a particle filter algorithm to update the feature template online. The multi-scale tracking strategy includes multi-scale pyramid tracking and occlusion detection and recovery processing.

5. A method for visual servo tracking of marine targets according to claim 4, characterized in that: The multi-scale pyramid tracking method is to cope with the change of target scale. The size of the unmanned vehicle increases when it approaches the target. An image pyramid is constructed to track the target at different scale layers. Specifically, the following steps are included: S1: Generate three layers of pyramid for the current frame image, which are original scale, 0.8 times and 1.2 times respectively; S2: Run the tracking algorithm independently on each layer of the pyramid to obtain multi-scale target state estimation; S3: Combined with the UAV motion information provided by the multimodal sensor fusion module, the tracking results of the optimal scale layer are selected to reduce computational redundancy.

6. A method for visual servo tracking of marine targets according to claim 4, characterized in that: The occlusion detection and restoration process is based on the occlusion detection method of the spatiotemporal context, which calculates the feature similarity between the target area of ​​the current frame and the previous frames. If the similarity of three consecutive frames is lower than the threshold, it is determined to be occlusion; During occlusion, the predicted state of the particle filter is used to maintain the target trajectory, and the background motion is tracked by the optical flow method to eliminate the interference of the occlusion. After the occlusion is released, the improved DeepSORT algorithm is used for feature matching, combining the historical trajectory and appearance features to reconfirm the target identity.

7. A method for visual servo tracking of marine targets according to claim 1, characterized in that: In step three, the architecture of the servo control and vision collaboration mechanism includes a vision processing module, a servo controller and a pan-tilt actuator. The vision processing module outputs the position deviation of the target in the image in real time, the servo control calculates the angle adjustment amount of the pan-tilt according to the deviation, and the pan-tilt actuator realizes the horizontal and pitch movement of the camera through motor drive.

8. A method for visual servo tracking of marine targets according to claim 1 or 7, characterized in that: The servo control and visual coordination mechanism also needs to cooperate with the prediction and correction control strategy, which is to compensate for the delay of the visual system and the influence of hull shaking, and mainly includes target motion prediction, hull shaking compensation and control amount calculation.

9. A method for visual servo tracking of marine targets according to claim 1, characterized in that: In step 4, occlusion processing and target re-identification are achieved through multi-cue occlusion detection, trajectory prediction maintenance and spatiotemporal joint feature matching, so as to achieve robust tracking when occlusion occurs and accurate re-identification after occlusion is released. In the multi-target tracking scenario, the occlusion relationship is judged through trajectory intersection analysis and optical flow motion direction to distinguish between the target and the occluder.

10. The method for visual servo tracking of a marine target according to claim 1, wherein: It also includes calibrating the visual sensor during the system initialization phase, establishing the mapping relationship between the image coordinate system and the world coordinate system, testing the motion range and response characteristics of the servo gimbal, optimizing control parameters, and transmitting tracking data in real time throughout the entire process, and improving system performance through online optimization.

Citation Information

Cited By

  • Offshore multi-target tracking detection method for unmanned ship

    CN121305346A

  • Multi-camera shooting holder control system

    CN121325979A

  • Visual target tracking method based on multi-feature scoring

    CN121330014A

  • A visual target tracking method based on multi-feature scoring

    CN121330014B

  • Edge calculation enabling Internet of Things visual tracking collaboration method

    CN121504976A