A Multi-Mode Fusion Tracking and Localization Method for Video Targets to Suppress Correlation Filter Distortion

By using a video target multi-mode fusion tracking and localization method to suppress correlation filter distortion, the problem of joint tracking and localization of underwater robots in dynamic flow fields using binocular cameras and imaging sonar was solved, achieving stable tracking, localization, and autonomous control of underwater targets.

CN115601400BActive Publication Date: 2025-11-14SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211361510.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-02
Publication Date
2025-11-14
Estimated Expiration
2042-11-02

AI Technical Summary

Technical Problem

Existing underwater robots lack effective methods for joint tracking and localization using binocular cameras and imaging sonar in dynamic flow fields, which makes autonomous underwater control difficult. In particular, problems such as random disturbances in dynamic flow fields, apparent nonlinear changes in underwater targets, and multipath scattering from imaging sonar have not been effectively solved.

Method used

A video target multi-mode fusion tracking and localization method with correlation filter distortion suppression is adopted. By constructing an underwater video target multi-mode detection system, a moving target detection method based on multi-mode fusion is established, a tracking and localization model with correlation filter distortion suppression is constructed, and the correlation filter learning minimization problem with correlation filter distortion suppression is solved. The tracking and localization is performed by combining an ultra-lightweight convolutional neural network and a sensitivity-free Kalman filtering algorithm.

Benefits of technology

It achieves stable tracking and positioning of underwater targets in dynamic flow fields, provides effective three-dimensional position information for underwater unmanned systems, and supports autonomous operation and control of underwater robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115601400B_ABST
    Figure CN115601400B_ABST
Patent Text Reader

Abstract

This invention discloses a multi-mode fusion tracking and localization method for video targets that suppresses correlation filter distortion, relating to the fields of underwater robots and video target tracking. The method includes the following steps: constructing an underwater video target multi-mode detection and tracking system, comprising multiple sensors such as imaging sonar, binocular camera, depth gauge, and inertial measurement unit (IMU); establishing an underwater moving target detection method based on multi-mode fusion; constructing a tracking and localization model based on correlation filter distortion suppression; using the Alternating Direction Multiplier Method (ADMM) algorithm to solve the correlation filter learning minimization problem based on correlation filter distortion suppression; updating the target appearance model based on linear weighting; and completing the performance evaluation and analysis of multi-mode tracking and localization of video targets. This invention can improve the feasibility and effectiveness of target detection, tracking, and localization in underwater unmanned systems, providing theoretical and technical support for the development of underwater robots, deep-sea exploration, and space robot operation and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of underwater robots and video target tracking, and in particular to a video target multi-mode fusion tracking and localization method that suppresses correlation filter distortion. Background Technology

[0002] Video target tracking has broad application prospects in fields such as intelligent visual navigation and intelligent transportation. Underwater unmanned operation systems, such as underwater robots and UUVs, are one of the core key technologies.

[0003] Underwater robots, also known as remotely operated vehicles (ROVs), are robots designed for extreme underwater operations. Given the harsh and dangerous underwater environment and the limited diving depth of humans, underwater robots have become an important tool for ocean exploration.

[0004] There are two main types of unmanned remotely operated vehicles (ROVs): tethered ROVs and untethered ROVs. Tethered ROVs are further divided into three types: self-propelled, towed, and crawling on seabed structures.

[0005] However, most current underwater robots do not adequately consider the joint tracking and localization problem of binocular cameras and imaging sonar in random currents. This problem poses a significant technical challenge to subsequent tasks such as dynamic docking, manipulator operation, and control. In other words, the lack of joint tracking and localization information from binocular cameras and imaging sonar makes underwater autonomous control extremely challenging.

[0006] A literature review of existing technologies reveals that the unmanned, intelligent, and modular characteristics of current underwater robots make underwater target tracking and localization increasingly important. Currently, there is no research on the joint tracking and localization problem using binocular cameras and imaging sonar in dynamic flow fields. Specifically, existing tracking and localization methods only address tracking and localization using a single sensor and do not consider challenges such as random disturbances in dynamic flow fields, apparent nonlinear changes in underwater targets, multipath scattering from imaging sonar, and rapid movement.

[0007] Therefore, those skilled in the art are dedicated to developing a video target multi-mode fusion tracking and localization method that suppresses correlation filter distortion, providing effective and stable three-dimensional position information for underwater unmanned systems. Summary of the Invention

[0008] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by the present invention is an underwater target tracking and positioning system and method in a dynamic flow field.

[0009] To achieve the above objectives, this invention provides a video target multi-mode fusion tracking and localization method that suppresses correlation filter distortion, comprising the following steps:

[0010] Step 1: Design and build an underwater video target multi-mode detection and tracking system;

[0011] Step 2: Establish an underwater moving target detection method based on multi-mode fusion;

[0012] Step 3: Construct a tracking and localization model based on correlation filter distortion suppression;

[0013] Step 4: Solve the correlation filter learning minimization problem based on correlation filter distortion suppression;

[0014] Step 5: Target appearance model update method based on linear weighting;

[0015] Step 6: Complete the multi-mode tracking and positioning performance evaluation and analysis of video targets.

[0016] Further, step 1 includes the following steps:

[0017] Step 1.1: Construct an underwater binocular camera sensor, including the preparation of light-transmitting materials and the design of a watertight structure. Then, complete the calibration of the underwater binocular camera and set the camera exposure time, resolution, and IP address parameters.

[0018] Step 1.2: Construct a depth sensor, set the depth measurement range and signal gain, and achieve the fusion of depth measurement, inertial measurement unit and imaging sonar measurement through structural design to complete the composite positioning of underwater targets;

[0019] Step 1.3: Construct an imaging sonar sensor, set the detection working distance, signal gain, and signal gamma value, and align the underwater binocular camera, inertial measurement unit, depth gauge, and imaging sonar through parameter calibration.

[0020] Furthermore, step 2 includes an underwater multimodal image target detection and moving target fusion detection method using an ultra-lightweight convolutional neural network.

[0021] Furthermore, the ultra-lightweight convolutional neural network for underwater multi-modal image target detection establishes an imaging sonar dataset based on static underwater target and moving target imaging sonar video data, which can be used to train and validate the ultra-lightweight convolutional neural network for underwater multi-modal image target detection. Based on the Fully Convolutional One-Stage Object Detection framework, improvements are introduced: label matching strategy, multi-scale feature fusion strategy, composite training method, and Generalized Focal Loss, to realize an ultra-lightweight convolutional neural network model for underwater multi-modal image target detection.

[0022] Furthermore, the moving target fusion detection method, based on imaging sonar and underwater binocular camera datasets, integrates the initial trajectory, target attributes, and target velocity characteristics during the detection process. It addresses external interferences such as complex lighting, random ray echo interference, multipath scattering interference, and random turbulence by combining frame difference method and adaptive background subtraction method. Through decision fusion, it achieves fusion detection of moving targets.

[0023] Furthermore, in step 3, given a vectorized target sample z with N channels, that is... The ideal vectorized response is denoted as Define B∈ W×H This is a selection matrix used to filter the middle W elements of each channel of the input target sample, while... The correlation filter component is learned from the nth channel; W << H; To address the correlation filter distortion problem during the learning process, response graphs P1 and P2 are defined to suppress filter distortion:

[0024]

[0025] Where u and v represent the local deviations of the two two-dimensional response plot peaks; the symbol {φ u,v} represents a two-dimensional translation operation; by minimizing the above expression, the distortion of the relevant filter can be identified or located, which is to suppress the filter distortion regularization term;

[0026] The objective function is:

[0027]

[0028] Where the subscripts k and k-1 represent the k-th and k-1-th frames, respectively; the third term in the formula is used to suppress the distortion of the correlation filter; the parameter γ represents the distortion penalty parameter; and the selection matrix B is used to ensure a sufficient search area.

[0029] Transform the objective function of formula (2) into matrix form.

[0030]

[0031] Among them, Z k The input target sample z k The matrix form; I N It is an N×N identity matrix; the symbols ⊙ and T represent the Kronecker product and conjugate transpose operations, respectively; P k-1 This represents the response map of the previous frame, and its value is equal to Z. k-1 (I N ⊙B T )q k ;

[0032] Transform equation (3) to the frequency domain, then

[0033]

[0034] in, The superscript wavy line indicates that the variable has been transformed to the frequency domain, which is to say, using the Discrete Fourier Transform. Similar definitions.

[0035] Furthermore, in step 4, based on the enhanced Lagrangian optimization framework, the objective function is obtained.

[0036]

[0037] Where μ is the penalty factor, and the Fourier form of the Lagrangian vector ω is:

[0038] Furthermore, in step 5,

[0039]

[0040] Where ε is the learning rate of the apparent model.

[0041] Furthermore, in step 6, the correlation filter learning minimization problem based on correlation filter distortion suppression is solved. Based on the joint calibration of the visible light camera, imaging sonar and depth gauge, the insensitive Kalman filtering algorithm is used to filter the data after tracking and positioning to complete the tracking and positioning.

[0042] Furthermore, in step 6, the mean and variance statistical analysis measures are used to evaluate and analyze the filtering performance of the results after multi-mode tracking and positioning.

[0043] In a preferred embodiment of the present invention, the present invention provides a video target multi-mode fusion tracking and localization system and method for suppressing correlation filter distortion:

[0044] Step 1: Design and build an underwater video target multi-mode detection and tracking system;

[0045] Step 2: Establish an underwater moving target detection method based on multi-mode fusion;

[0046] Step 3: Construct a tracking and localization model based on correlation filter distortion suppression;

[0047] Step 4: Solve the correlation filter learning minimization problem based on correlation filter distortion suppression;

[0048] Step 5, update method based on linearly weighted target appearance model;

[0049] Step 6: Complete the multi-mode tracking and positioning performance evaluation and analysis of the video target.

[0050] A video target multi-mode fusion tracking and localization system and method for suppressing correlation filter distortion, as described in the embodiments of the present invention, is as follows: Figure 1 As shown, step 1 specifically includes:

[0051] Step 1.1: Construct an underwater binocular camera, including the preparation of light-transmitting materials and the design of a watertight structure. Then, complete the calibration of the underwater binocular camera, set parameters such as camera exposure time, resolution, and IP address, and complete the configuration and online operation of the visual positioning software.

[0052] Step 1.2: Construct a depth sensor, set the depth measurement range, signal gain, etc., and achieve the fusion of depth measurement, inertial measurement unit and imaging sonar measurement through structural design to complete the composite positioning of underwater targets.

[0053] Step 1.3: Construct an imaging sonar sensor, set its detection working distance, signal gain, signal gamma value, etc., and align the underwater binocular camera, inertial measurement unit, depth gauge and imaging sonar through parameter calibration.

[0054] A video target multi-mode fusion tracking and localization system and method for suppressing correlation filter distortion, as described in the embodiments of the present invention, is as follows: Figure 2 As shown, step 2 specifically includes: underwater multimodal image target detection and moving target fusion detection using an ultra-lightweight convolutional neural network. The specific details are as follows:

[0055] Step 2.1: Construct a stereo camera and imaging sonar image dataset. These images have resolutions ranging from 640 to 1920x320 to 1080. The stereo camera and imaging sonar image dataset contains more than 1000 images.

[0056] Step 2.2, Underwater Multimodal Image Target Detection Using an Ultra-Lightweight Convolutional Neural Network. The key feature is the establishment of an imaging sonar dataset for training and validating underwater multimodal image target detection using an ultra-lightweight convolutional neural network, based on existing static and moving underwater target imaging sonar video data. The underwater multimodal image target detection algorithm of the ultra-lightweight convolutional neural network is characterized by three main modules: the detection network head, the backbone network, and the neck. Specifically, based on the Fully Convolutional One-Stage Object Detection framework, four improvements are introduced: label matching strategy, multi-scale feature fusion strategy, composite training method, and Generalized Focal Loss, to realize an ultra-lightweight convolutional neural network model for underwater multimodal image target detection. Specifically, this invention uses multimodal target detection algorithms based on the Anchor-Free framework, including DenseBox, YOLO, CornerNet, ExtremeNet, FSAF, FCOS, FoveaBox, etc. Specifically, this invention selects detection algorithms such as nanodet.

[0057] Step 2.3, Moving Target Fusion Detection. The key to its implementation lies in leveraging existing imaging sonar and underwater binocular camera datasets, comprehensively considering characteristics such as initial trajectory, target attributes, and target velocity during the detection process, and addressing external interferences such as complex lighting, random ray echo interference, multipath scattering interference, and random turbulence, by combining frame difference methods and adaptive background subtraction. Specifically, this invention achieves moving target fusion detection through decision fusion.

[0058] A video target multi-mode fusion tracking and localization system and method for suppressing correlation filter distortion, as described in the embodiments of the present invention, is as follows: Figure 3 As shown, step 3 specifically constructs a tracking and localization model based on correlation filter distortion suppression.

[0059] Given a vectorized target sample z with N channels, that is... The ideal vectorized response is denoted as We define B∈ W×H The selection matrix is ​​used to filter the middle W elements of each channel of the input target sample, while... This is the correlation filter component learned from the nth channel. Generally, W << H. To address the correlation filter distortion problem during the learning process, we define response plots P1 and P2 to suppress filter distortion:

[0060]

[0061] Here, u and v represent the local deviations of the peaks in the two-dimensional response plots. The symbol {φ} u,v} represents a two-dimensional translation operation. This invention achieves the identification or localization of correlation filter distortion by minimizing the above equation.

[0062] Based on the above considerations, the present invention has the following objective function:

[0063]

[0064] Here, the subscripts k and k-1 represent the k-th and (k-1)-th frames, respectively. The third term in the formula is used to suppress distortion from the correlation filter. The parameter γ represents the distortion penalty parameter. Matrix B is chosen to ensure a sufficient search area.

[0065] For ease of representation, the objective function of formula (2) is transformed into matrix form.

[0066]

[0067] Among them, Z k The input target sample z k The matrix form. I N It is an N×N identity matrix. The symbols ⊙ and T represent the Kronecker product and conjugate transpose operations, respectively. k-1 This represents the response map of the previous frame, and its value is equal to Z. k-1 (I N ⊙B T )q k .

[0068] To speed up the calculation, this invention transforms formula (3) into the frequency domain, so we have the following formula.

[0069]

[0070] in, The superscript wavy line indicates that the variable has been transformed to the frequency domain, which is to say, using the Discrete Fourier Transform. They have similar definitions.

[0071] An underwater target tracking and localization system and method based on non-parametric classification and composite confidence, as described in the embodiments of the present invention, is as follows: Figure 1 As shown, step 4 specifically includes solving the correlation filter learning minimization problem based on correlation filter distortion suppression, characterized by obtaining the following objective function based on the enhanced Lagrangian optimization framework.

[0072]

[0073] Where μ is the penalty factor, and the Fourier form of the Lagrangian vector ω is: like Figure 3 As shown.

[0074] According to the video target multi-mode fusion tracking and localization system and method for suppressing correlation filter distortion as described in the embodiments of the present invention, step 5 specifically involves a target appearance model update method based on linear weighting, characterized in that...

[0075]

[0076] Where ε is the learning rate of the target appearance model.

[0077] The underwater video target tracking and positioning system based on imaging sonar and binocular cameras features an interference-free configuration, consisting of four standard working units and two locking mechanisms.

[0078] Compared with the prior art, the present invention has the following obvious substantive features and significant advantages:

[0079] This invention presents a video target multi-mode fusion tracking and positioning method, addressing the urgent need for autonomous underwater operations. Based on correlation filtering theory and methods, it proposes a video target multi-mode fusion tracking and positioning system and method to suppress correlation filter distortion. A tracking and positioning model based on correlation filter distortion suppression is constructed, including a selection matrix and a distortion suppression filter regularization term. The problem of minimizing the correlation filter learning based on correlation filter distortion suppression is solved, and the underwater target tracking and positioning task of an unmanned underwater platform in a dynamic flow field is completed, providing key technical support for the operation and control of underwater robots.

[0080] The following will further explain the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention. Attached Figure Description

[0081] Figure 1 This is a preferred embodiment of the video target multi-mode fusion tracking and positioning system and method for suppressing correlation filter distortion according to the present invention;

[0082] Figure 2 This is a flowchart of an underwater moving target fusion detection based on an ultra-lightweight convolutional neural network, according to a preferred embodiment of the present invention.

[0083] Figure 3 This is a preferred embodiment of the present invention: a correlation filter based on composite confidence for underwater targets. Detailed Implementation

[0084] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.

[0085] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of some components has been appropriately exaggerated in the drawings.

[0086] This invention discloses a video target multi-mode fusion tracking and localization system and method for suppressing correlation filter distortion, comprising the following steps:

[0087] Step 1: Design and build an underwater video target multi-mode detection and tracking system;

[0088] Step 2: Establish an underwater moving target detection method based on multi-mode fusion;

[0089] Step 3: Construct a tracking and localization model based on correlation filter distortion suppression;

[0090] Step 4: Solve the correlation filter learning minimization problem based on correlation filter distortion suppression;

[0091] Step 5, update method based on linearly weighted target appearance model;

[0092] Step 6: Complete the multi-mode tracking and positioning performance evaluation and analysis of the video target.

[0093] Figure 1 This invention provides a preferred embodiment of a video target multi-mode fusion tracking and localization system and method for suppressing correlation filter distortion. Specifically, (a) is an underwater video target tracking and localization system based on imaging sonar and a binocular camera; (b) is an underwater moving target detection method based on multi-mode fusion; (c) is a tracking and localization model based on correlation filter distortion suppression; and (d) is a joint tracking and localization result based on a binocular camera and imaging sonar.

[0094] Step 1 specifically includes:

[0095] Step 1.1: Construct an underwater binocular camera sensor, including the preparation of light-transmitting materials and the design of a watertight structure. Then, complete the underwater binocular camera calibration by setting parameters such as camera exposure time, resolution, and IP address.

[0096] Step 1.2: Construct a depth sensor, set the depth measurement range, signal gain, etc., and achieve the fusion of depth measurement, inertial measurement unit and imaging sonar measurement through structural design to complete the composite positioning of underwater targets.

[0097] Step 1.3: Construct an imaging sonar sensor, set its detection working distance, signal gain, signal gamma value, etc., and align the underwater binocular camera, inertial measurement unit, depth gauge and imaging sonar through parameter calibration.

[0098] Figure 2 This is a preferred embodiment of the present invention. The flowchart of the underwater moving target fusion detection based on ultra-lightweight convolutional neural network proposed in the present invention is as follows: Step 2 establishes an underwater moving target detection method based on multi-mode fusion, which includes underwater multi-mode image target detection based on ultra-lightweight convolutional neural network and moving target fusion detection method.

[0099] Step 2.1, Underwater Multimodal Image Target Detection Using an Ultra-Lightweight Convolutional Neural Network. Its key feature is the establishment of an imaging sonar dataset for training and validating underwater multimodal image target detection using an ultra-lightweight convolutional neural network, based on existing static and moving underwater target imaging sonar video data. The underwater multimodal image target detection algorithm of the ultra-lightweight convolutional neural network is characterized by three main modules: the detection network head, the backbone network, and the neck. Specifically, based on the Fully Convolutional One-Stage Object Detection framework, four improvements are introduced: label matching strategy, multi-scale feature fusion strategy, composite training method, and Generalized Focal Loss, to realize an ultra-lightweight convolutional neural network model for underwater multimodal image target detection.

[0100] Step 2.2, Moving Target Fusion Detection Method. Its key feature is that it is based on existing imaging sonar and underwater binocular camera datasets, comprehensively considering characteristics such as initial trajectory, target attributes, and target velocity during the detection process. It addresses external interferences such as complex lighting, random ray echo interference, multipath scattering interference, and random turbulence by combining frame difference methods and adaptive background subtraction. Specifically, this invention achieves moving target fusion detection through decision fusion.

[0101] Figure 3 This is a preferred embodiment of the correlation filter based on composite confidence for underwater targets, wherein the method given in step 3 includes:

[0102] Given a vectorized target sample z with N channels, that is... The ideal vectorized response is denoted as Define B∈ W×H The selection matrix is ​​used to filter the middle W elements of each channel of the input target sample, while... This is the correlation filter component learned from the nth channel. Generally, W << H. To address the correlation filter distortion problem during the learning process, response plots P1 and P2 are defined to suppress filter distortion:

[0103]

[0104] Here, u and v represent the local deviations of the peaks in the two-dimensional response plots. The symbol {φ} u,v} represents a two-dimensional translation operation. This invention achieves the identification or location of correlation filter distortion by minimizing the above equation, that is, by suppressing the filter distortion regularization term.

[0105] Based on the above considerations, the present invention has the following objective function:

[0106]

[0107]

[0108] Here, the subscripts k and k-1 represent the (k-1)th and (k-1)th frames, respectively. The third term in the formula is used to suppress distortion from the correlation filter. The parameter γ represents the distortion penalty parameter. Matrix B is chosen to ensure a sufficient search area.

[0109] For ease of representation, the objective function of formula (2) is transformed into matrix form.

[0110]

[0111] Among them, Z k The input target sample z k The matrix form. I N It is an N×N identity matrix. The symbols ⊙ and T represent the Kronecker product and conjugate transpose operations, respectively. k-1 This represents the response map of the previous frame, and its value is equal to Z. k-1 (I N ⊙B T )q k .

[0112] To speed up the calculation, this invention transforms formula (3) into the frequency domain, resulting in the following formula:

[0113]

[0114] in, The superscript wavy line indicates that the variable has been transformed to the frequency domain, which is to say, using the Discrete Fourier Transform. They have similar definitions.

[0115] Solving the correlation filter learning minimization problem based on correlation filter distortion suppression, and using the enhanced Lagrangian optimization framework, the following objective function is obtained.

[0116]

[0117] Where μ is the penalty factor, and the Fourier form of the Lagrangian vector ω is:

[0118] The target appearance model update method based on linear weighting is characterized by:

[0119]

[0120] Where ε is the learning rate of the apparent model.

[0121] This invention evaluates and analyzes the performance of multi-mode tracking and localization of video targets. It solves a correlation filter learning minimization problem based on correlation filter distortion suppression. Using a joint calibration of a visible light camera, imaging sonar, and depth gauge, a sensitivity-free Kalman filter algorithm is employed to filter the tracking and localization data, thus completing the tracking and localization process. Furthermore, this invention uses statistical analysis quantities such as mean and variance to evaluate and analyze the filtering performance of the multi-mode tracking and localization results.

[0122] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A video target multi-mode fusion tracking and localization method for suppressing correlation filter distortion, characterized in that, Includes the following steps: Step 1: Design and build an underwater video target multi-mode detection and tracking system; Step 2: Establish an underwater moving target detection method based on multi-mode fusion; Step 3: Construct a tracking and localization model based on correlation filter distortion suppression; Step 4: Solve the correlation filter learning minimization problem based on correlation filter distortion suppression; Step 5: Target appearance model update method based on linear weighting; Step 6: Complete the multi-mode tracking and positioning performance evaluation and analysis of the video target; In step 3, a vectorized target sample is given. And there are One channel, that is ; The ideal vectorized response is denoted as ;definition This is a selection matrix used to filter the middle values ​​of each channel of the input target sample. One element, at the same time From the first The correlated filter components learned by each channel; To address the issue of filter distortion during the learning process, a response graph is defined to suppress filter distortion. and : (1) in, Represents the local deviation of the two peaks in the two-dimensional response plots; symbol This represents a two-dimensional translation operation; by minimizing the above equation, the identification or location of the distortion of the relevant filter can be achieved, which is to suppress the filter distortion regularization term. The objective function is: (2) Among them, subscript and They represent the first and the Frame; the third term in the formula is used to suppress distortion from the correlation filter; parameters Represents the distortion penalty parameter; selection matrix To ensure a sufficient search area; Transform the objective function of formula (2) into matrix form. (3) in, Input target sample Matrix form; It is The identity matrix; symbol as well as These represent the Kronecker product and the conjugate transpose operation, respectively. This represents the response graph of the previous frame, and its value is equal to... ; Transform formula (3) to the frequency domain, then (4) in, The superscript wavy line indicates that the variable has been transformed to the frequency domain, which is to say, the Discrete Fourier Transform is used. Similar definitions; Step 4, based on the enhanced Lagrangian optimization framework, yields the objective function. (5) in, It is a penalty factor, a Lagrangian vector. The Fourier form is .

2. The video target multi-mode fusion tracking and localization method for suppressing correlation filter distortion as described in claim 1, characterized in that, Step 1 includes the following steps: Step 1.1: Construct an underwater binocular camera sensor, including the preparation of light-transmitting materials and the design of a watertight structure. Then, complete the calibration of the underwater binocular camera and set the camera exposure time, resolution, and IP address parameters. Step 1.2: Construct a depth sensor, set the depth measurement range and signal gain, and achieve the fusion of depth measurement, inertial measurement unit and imaging sonar measurement through structural design to complete the composite positioning of underwater targets; Step 1.3: Construct an imaging sonar sensor, set the detection working distance, signal gain, and signal gamma value, and align the underwater binocular camera, inertial measurement unit, depth gauge, and imaging sonar through parameter calibration.

3. The video target multi-mode fusion tracking and localization method for suppressing correlation filter distortion as described in claim 1, characterized in that, Step 2 includes an underwater multimodal image target detection and moving target fusion detection method using an ultra-lightweight convolutional neural network.

4. The video target multi-mode fusion tracking and localization method for suppressing correlation filter distortion as described in claim 3, characterized in that, The underwater multimodal image target detection of the ultra-lightweight convolutional neural network is based on imaging sonar video data of static underwater targets and moving targets, and an imaging sonar dataset that can be used to train and validate the underwater multimodal image target detection of the ultra-lightweight convolutional neural network is established. Based on the fully convolutional single-stage object detection framework, improvements are introduced: label matching strategy, multi-scale feature fusion strategy, composite training method and Generalized Focal Loss, to realize an ultra-lightweight convolutional neural network model for underwater multi-modal image object detection.

5. The video target multi-mode fusion tracking and localization method for suppressing correlation filter distortion as described in claim 3, characterized in that, The proposed moving target fusion detection method, based on imaging sonar and underwater binocular camera datasets, integrates the initial trajectory, target attributes, and target velocity characteristics during the detection process. It addresses external interferences such as complex lighting, random ray echo interference, multipath scattering interference, and random turbulence by combining frame difference and adaptive background subtraction methods. Through decision fusion, it achieves fusion detection of moving targets.

6. The video target multi-mode fusion tracking and localization method for suppressing correlation filter distortion as described in claim 1, characterized in that, Step 5, (6) in It is the learning rate of the apparent model.

7. The video target multi-mode fusion tracking and localization method for suppressing correlation filter distortion as described in claim 1, characterized in that, Step 6 involves solving the correlation filter learning minimization problem based on correlation filter distortion suppression. On the basis of joint calibration of visible light camera, imaging sonar and depth gauge, the sensitivity-free Kalman filtering algorithm is used to filter the data after tracking and positioning to complete the tracking and positioning.

8. The video target multi-mode fusion tracking and localization method for suppressing correlation filter distortion as described in claim 1, characterized in that, In step 6, mean and variance statistical analysis measures are used to evaluate and analyze the filtering performance of the results after multi-mode tracking and positioning.

Citation Information

Patent Citations

  • Visual tracking method and device based on adaptive correlation filtering feature fusion learning

    CN113538509A

  • Underwater target multi-model tracking method based on underwater acoustic sensor network

    CN114415157A