Scale-space discriminative tracking method for underwater video targets based on multi-model fusion

Through the multi-model fusion underwater video target scale space discriminant tracking system, combined with imaging sonar and binocular camera, the accuracy and stability problems of imaging sonar detection and tracking in deep sea environment are solved, and efficient and reliable tracking of targets is achieved.

CN114898202BActive Publication Date: 2025-09-05SHANGHAI JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210343863.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-09-05
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

Existing imaging sonars are susceptible to multipath effects, reverberation interference, ocean noise, etc. in complex deep-sea environments, resulting in insufficient accuracy and stability in target detection and tracking.

Method used

An underwater video target scale-space discriminative tracking system based on multi-model fusion is adopted. Combining imaging sonar, binocular camera and multi-sensor module, the least squares estimation method, YOLO detection model, Gaussian mixture detection model and unscented Kalman filter are used to achieve accurate tracking of target position and scale.

Benefits of technology

The accuracy and stability of target detection are improved, the robustness is enhanced, and it can effectively track non-cooperative targets in deep-sea environments with stable and reliable performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114898202B_ABST
    Figure CN114898202B_ABST
Patent Text Reader

Abstract

This invention discloses a scale-space discriminant tracking method for underwater video targets based on multi-model fusion, which relates to the fields of underwater robots and video target tracking. The method comprises the following steps: 1. multi-sensor registration; 2. target detection using a multi-model fusion detection algorithm based on the YOLO detection model and the Gaussian mixture detection model; 3. input into a scale-space discriminant tracker for tracking, obtaining the relative position and scale information of the target; 4. calculation of the position and scale information of the target in the underwater multi-sensor detection and tracking system; and 5. filtering the target position information using an unscented Kalman filter to obtain the final target state and motion trajectory. This invention can improve the accuracy, stability, and effectiveness of video target detection and tracking in complex deep-sea environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of underwater robots and visual target tracking, and in particular to an underwater video target scale space discriminant tracking method based on multi-model fusion. Background Art

[0002] my country is currently in a critical period of accelerating its development into a maritime power, with significant development needs in the operation and control of underwater robots. Estimating the target's motion state for underwater robot operations is a key technical bottleneck. Imaging sonar, a crucial instrument for ocean exploration, is an electronic device that utilizes the underwater propagation characteristics of sound waves through electroacoustic conversion and information processing to perform underwater exploration and communication tasks. While imaging sonar boasts strong wavelength detection, target recognition capabilities, and concealment, it also has numerous drawbacks, including susceptibility to multipath effects, reverberation interference, ocean noise, self-noise, target reflection characteristics, or radiated noise intensity, which can lead to positioning errors.

[0003] Therefore, technicians in this field are committed to proposing an underwater video target scale space discriminant tracking system and method based on multi-model fusion. Summary of the Invention

[0004] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is how to overcome the shortcomings of the original imaging sonar image detection and tracking technology and improve the accuracy, stability and effectiveness of target detection and tracking in complex deep-sea environments.

[0005] To achieve the above objectives, the present invention provides an underwater video target scale space discriminant tracking system based on multi-model fusion, characterized in that the system includes an imaging sonar, a binocular camera, a multi-sensor module, and an underwater video target tracking processing board, wherein the imaging sonar and the binocular camera are integrated into a multi-sensor module, and the underwater video target tracking processing board is connected to the multi-sensor module.

[0006] A scale-space discriminative tracking method for underwater video targets based on multi-model fusion, characterized in that the method comprises the following steps:

[0007] Step 1: Use the least squares estimation method to complete the multi-sensor registration of imaging sonar and binocular camera;

[0008] Step 2: A multi-model fusion detection algorithm based on the YOLO detection model and the Gaussian mixture detection model is used to detect underwater video targets input from sonar video and camera video to obtain the relative position information and relative scale information of the target;

[0009] Step 3: Use the target relative position information and relative scale information from step 2 as initial values ​​to initialize the scale space discriminant tracker, use the position filter to predict the target relative position information of each frame in the video, and use the scale filter to predict the scale of the target in the neighborhood of the predicted position of each frame to obtain the relative position information and relative scale information of the target in each frame;

[0010] Step 4: Based on the multi-sensor imaging characteristics, the position and scale information in the sonar video is converted to the binocular camera coordinate system and fused with the position and scale information in the camera video to obtain the position and scale information of the target in each frame in the camera coordinate system;

[0011] Step 5: Based on the target position information of each frame obtained in step 4, the target position is filtered using an unscented Kalman filter to obtain the final target state and motion trajectory.

[0012] Furthermore, step 2 specifically includes:

[0013] Step 2.1: Use the YOLO detection algorithm and Gaussian mixture detection algorithm to detect the input sonar and camera video images respectively to obtain the relative position and scale information of the target;

[0014] Step 2.2: Fuse the detection results of the YOLO detection algorithm with the detection results of the Gaussian mixture detection algorithm.

[0015]

[0016] Among them, w is the weight of each detection algorithm, y represents the relative position information or scale information of the target, and Y is the fused relative position and scale information of the target.

[0017] Furthermore, step 3 specifically includes:

[0018] Step 3.1, the target relative position and target relative scale in step 2 are used as the target initial value of the first frame of the tracking sequence;

[0019] Step 3.2: Extract the features of the target location candidate window and transform it into the Fourier domain;

[0020] Step 3.3: Generate the target position regression matrix and transform it into the Fourier domain;

[0021] Step 3.4: Generate n-scale candidate boxes around the target initial box, extract features of the corresponding area for each candidate box, and transform the generated n features into the Fourier domain;

[0022] Step 3.5: Generate the scale regression matrix of the target and transform it into the Fourier domain;

[0023] Step 3.6: Train to obtain the position tracking template and scale tracking template;

[0024] Step 3.7: For the new frame, use the position tracking template to calculate the response on the candidate window. The position of the maximum response is the target position of the new frame. The formula for calculating the response is:

[0025]

[0026] Where R is the response, A and B are the numerator and denominator of the correlation filter respectively, Z is the feature map extracted from the target prediction position of the new frame, and λ is the regularization coefficient;

[0027] Step 3.8: Use the scale tracking template to calculate the responses of different scale multipliers at the desired position, and obtain the scale multiplier with the maximum response as the target scale for the new frame;

[0028] Step 3.9: Use the new scale and position to continue tracking the position of the next frame until all frames are predicted.

[0029] Furthermore, step 4 specifically includes:

[0030] Step 4.1, set the value of n to 30-40; set the range of imaging sonar to 10-60 meters; set the dual

[0031] The resolution and frame rate of the eye camera, as well as the peak power parameters of the visual signal processing board;

[0032] Step 4.2: Based on the imaging characteristics and installation positions of the imaging sonar and binocular camera, calculate the measurement value of the underwater video target in the coordinate system of the underwater multi-sensor detection and tracking system.

[0033] Furthermore, step 5 specifically includes:

[0034] Step 5.1, calculate and obtain the initial state estimate and estimated variance of the filter;

[0035] Step 5.2: Update the time to obtain the mean and covariance of the target state at time k+1 predicted at time k;

[0036] Step 5.3: Update the measurement to obtain the state estimate and estimated variance at time k+1.

[0037] Furthermore, the YOLO detection model algorithm in step 2.1 specifically includes:

[0038] The YOLO detection algorithm uses a separate convolutional neural network model to achieve end-to-end target detection. First, the input image is resampled, then sent to the convolutional neural network, and finally the network prediction result is processed to obtain the target detection result.

[0039] Furthermore, the target detection algorithm based on the Gaussian mixture model in step 2.1 includes:

[0040] The target detection algorithm based on Gaussian mixture model uses multiple single Gaussian models as the model of a pixel position, using the formula

[0041] |I(x, y, t), μ i (x, y, t-1)|<λ×σ i (x,y,t-1),i=1,2,…,K,

[0042] Judge the new pixel, where I is the pixel value of the new pixel, μ is the mean of the existing Gaussian model, σ is the standard deviation of the existing Gaussian model, and K represents the number of existing Gaussian models;

[0043] If the new pixel matches the single model, the pixel is judged to be the background, and the weight of the single model matching the new pixel is corrected; if there is no model matching the new pixel, the pixel is judged to be the foreground, and the single model with the least importance in the multi-model set is removed, and a new single model is added.

[0044] Furthermore, the detection algorithm based on the Gaussian mixture model in step 2.1 specifically includes:

[0045] Step 2.1.1 Define pixel model

[0046] Each pixel is described by multiple single models:

[0047] P(p)={[w i (x, y, t), u i (x, y, t), σ i (x, y, t) 2 ]}, i = 1, 2, ..., K, where K represents the number of single models in the Gaussian mixture model, w i (x, y, t) represents the weight of each model, μ is the mean of the existing Gaussian model, and σ is the standard deviation of the existing Gaussian model, satisfying:

[0048]

[0049] Step 2.1.2 Update parameters and perform foreground detection

[0050] Step 2.1.2.1,

[0051] If the pixel value of the corresponding point (x, y) of the newly input image satisfies:

[0052] |I(x, y, t)-μ i (x, y, t-1)|<λ×σ i(x,y,t-1),i=1,2,…,K,

[0053] The new pixel is matched with the single model, the pixel is judged to be background, and step 2.1.2.2 is performed; if there is no model matching the new pixel, the pixel is judged to be foreground, and step 2.1.2.3 is performed;

[0054] Step 2.1.2.2,

[0055] Correct the weights of the single model that matches the new pixel. The new weights are:

[0056] w i (x, y, t) = w i (x, y, t-1) + dw = w i (x, y, t-1)+α(1-w i (x, y, t-1)),

[0057] Where dw is the weight increment, the formula is

[0058] dw=α(1-w i (x, y, t-1)),

[0059] Similarly, correct the mean and variance of the single model matching the new pixel, and then proceed to step 2.1.2.4;

[0060] Step 2.1.2.3: If the new pixel does not match any existing single model, it is divided into the following two cases:

[0061] a. If the number of single models has reached the maximum allowed number, remove the single model with the least importance in the current multi-model set; the importance calculation formula is:

[0062]

[0063] b. If the number of current single models does not reach the maximum allowed number, add a new single model with a smaller weight, a mean of the new pixel value, and a given larger variance;

[0064] Step 2.1.2.4. Weight Normalization

[0065]

[0066] Where w represents the weight of each model, and W represents the weight of each model after weight normalization.

[0067] Furthermore, in step 3:

[0068] The steps for location tracking are:

[0069] Referring to the position of the template in the previous frame, a sample Z is extracted in the current frame at a size twice the target scale of the previous frame;

[0070] Using the sample Z and the numerator A and denominator B of the position model, according to the formula

[0071]

[0072] Calculate the response of the new position, and the place with the largest response is the new position of the target;

[0073] The scale tracking steps are:

[0074] Use the scale model to calculate the responses of different scale multipliers at the desired position, and obtain the scale multiplier with the maximum response as the target scale of the new frame;

[0075] The steps for training model update are:

[0076] Using the newly obtained relative position and scale of the target, new filters are trained to update the position model and scale model.

[0077] This paper proposes a scale-space discriminative tracking system and method for underwater video targets based on multi-model fusion. Taking into account the environmental characteristics of deep-sea environments and imaging sonar imaging, this system overcomes the shortcomings of existing imaging sonar image detection and tracking technologies. The system features stable performance, high accuracy, and strong robustness, promoting the development of underwater unmanned operation and control technology and possessing broad application prospects in deep-sea operations, space robotics, and other fields. First, the system uses a multi-model fusion detection algorithm based on the YOLO detection algorithm and the Gaussian mixture detection algorithm to detect targets in input video data. This multi-model fusion approach overcomes the problem of insufficient detection capabilities of a single model and improves target detection and recognition rates. Secondly, upon target detection, a scale-space discriminative tracker is immediately activated to track the target. Simultaneously, the detection algorithm is still used to detect the target. This allows the tracker to be reinitialized in the event of target loss during tracking, improving the robustness, stability, and effectiveness of the detection and tracking algorithm. Finally, an unscented Kalman filter is used to filter the tracked position, resulting in more reliable tracking results. In actual underwater target detection and tracking tests, the present invention has stable performance, strong reliability, and high accuracy, and has obtained good field experimental results. After testing on non-cooperative targets in the ocean, it is proved that the algorithm in the invention can effectively detect and track non-cooperative targets. The invention can be further applied in the deep sea field.

[0078] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 1 is a schematic diagram of a system and method for underwater video target scale space discriminant tracking based on multi-model fusion according to a preferred embodiment of the present invention;

[0080] Figure 2 1 is a neural network structure diagram of the YOLO detection algorithm adopted in a preferred embodiment of the present invention;

[0081] Figure 3 is a flow chart of a detection algorithm based on a Gaussian mixture model adopted in a preferred embodiment of the present invention;

[0082] Figure 4 is a flow chart of a scale-space discriminant tracking method used in a preferred embodiment of the present invention;

[0083] Figure 5 is data collected by a multi-frame imaging sonar used in a preferred embodiment of the present invention (frames 8, 40, 160, and 220);

[0084] Figure 6 Detection and tracking results (frames 201, 225, 270, and 299) obtained by testing data according to a preferred embodiment of the present invention;

[0085] Figure 7 This is a result graph obtained by smoothing the result using unscented Kalman filtering in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0086] The following describes several preferred embodiments of the present invention with reference to the accompanying drawings to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.

[0087] In the drawings, components with identical structures are denoted by the same reference numerals, and components with similar structures or functions are denoted by similar reference numerals. The size and thickness of each component shown in the drawings are arbitrary and are not limited by the present invention. For clarity, the thickness of components in some places in the drawings is appropriately exaggerated.

[0088] like Figure 1 As shown, a scale-space discriminative tracking system and method for underwater video targets based on multi-model fusion includes the following steps:

[0089] (1) Use a multi-model fusion detection algorithm based on the YOLO detection algorithm and the Gaussian mixture detection algorithm to detect targets in the input video data; use a multi-model fusion method to improve the detection rate and recognition rate of the target;

[0090] like Figure 2 As shown in the figure, the network of the YOLO detection algorithm used in the present invention includes: 24 convolutional layers for extracting image features and two fully connected layers for classification and positioning. The YOLO algorithm divides the input image into S×S

[0091] The network, each cell is responsible for detecting the target that falls into the cell, and each cell predicts B bounding boxes and their confidence. The accuracy of the bounding box can be expressed as the intersection of the predicted box and the true box.

[0092]

[0093] Finally, the predicted value of each bounding box is obtained, which contains 5 elements: (x, y, w, h, c), where the first 4 represent the size and position of the bounding box, and the last one represents the confidence.

[0094] like Figure 3 As shown in the figure, the flow chart of the detection algorithm based on Gaussian mixture model includes: background generation, background modeling, model update, foreground detection and foreground mask. The specific steps are as follows:

[0095] 1) Define pixel model

[0096] Each pixel is described by multiple single models:

[0097] P(p)={[w i (x, y, t), u i (x, y, t), σ i (x, y, t) 2 ]}, i = 1, 2, ..., K, where K represents the number of single models in the Gaussian mixture model, w i (x, y, t) represents the weight of each model, satisfying:

[0098]

[0099] 2) Update parameters and perform foreground detection

[0100] Step 1:

[0101] If the pixel value of the corresponding point (x, y) of the newly input image satisfies:

[0102] |I(x, y, t)-μ i (x, y, t-1)|<λ×σ i (x, y, t-1), i=1, 2,…,K, (13)

[0103] If the new pixel matches the single model, the point is judged to be background and step 2 is performed; if there is no model matching the new pixel, the point is judged to be foreground and step 3 is performed.

[0104] Step 2:

[0105] Correct the weights of the single model that matches the new pixel. The new weights are:

[0106] w i (x, y, t) = w i (x, y, t-1) + dw = w i (x, y, t-1)+α(1-w i (x, y, t-1)), (14)

[0107] Where dw is the weight increment, the formula is

[0108] dw=α(1-w i (x, y, t-1)), (15)

[0109] Similarly, correct the mean and variance of the single model that matches the new pixel, and then proceed to step 4.

[0110] Step 3:

[0111] If the new pixel does not match any single model, then:

[0112] If the number of current single models has reached the maximum allowed number, the single model with the least importance in the current multi-model set will be removed; the importance calculation formula is:

[0113]

[0114] Add a new single model with a smaller weight, a mean of the new pixel value, and a given larger variance.

[0115] Step 4:

[0116] Weight normalization

[0117]

[0118] After obtaining the YOLO algorithm detection results and the Gaussian mixture model detection results, the results are fused.

[0119]

[0120] Get the final test results.

[0121] (2) The detection results are input into the scale space discriminant tracker for tracking, the position is predicted using the position filter, and the target scale is predicted in the neighborhood of the predicted position using the scale filter to obtain the target's position information and scale information;

[0122] The algorithm process is as follows Figure 4 As shown, specifically including:

[0123] a. Use the detection results of the model fusion as the target initial value of the first frame of the sequence, including position and scale;

[0124] b. Extract the features of the candidate window of the target position and transform it into the Fourier domain;

[0125] c. Generate the target position regression matrix and transform it into the Fourier domain;

[0126] d. Generate n candidate boxes of different scales around the target initial box, extract features of the corresponding area for each candidate box, and transform the generated n features into the Fourier domain;

[0127] e. Generate the scale regression matrix of the target and transform it into the Fourier domain;

[0128] f. Train to obtain position tracking template and scale tracking template;

[0129] g. For the new frame, use the position tracking template to calculate the response on the candidate window and find the position of the maximum response.

[0130]

[0131] h. Use the scale tracking template to calculate the responses of different scale multipliers at the desired location, and obtain the scale multiplier with the maximum response as the new target scale;

[0132] i. The new scale and position are used to track the position of the next frame until all frames are predicted.

[0133] The steps of position tracking are:

[0134] Referring to the position of the template in the previous frame, a sample Z is extracted in the current frame at a size twice the target scale of the previous frame;

[0135] Using the sample Z and the numerator A and denominator B of the position model, according to the formula

[0136]

[0137] Calculate the response of the new position, and the place with the largest response is the new position of the target P t ;

[0138] The scale tracking steps are:

[0139] Use the scale model to calculate the responses of different scale multipliers at the desired position, and obtain the scale multiplier with the maximum response as the target scale of the new frame;

[0140] The steps for training model update are:

[0141] Using the newly obtained relative position and scale of the target, new filters are trained to update the position model and scale model.

[0142] (3) Based on the position information obtained in step 4, the target position is filtered using an unscented Kalman filter to obtain the final target state and motion trajectory. The specific process of the unscented Kalman filter is as follows:

[0143] a. Calculate the initial state estimate and estimated variance of the filter:

[0144]

[0145] b. Time update: Assume the state estimate at time k and the estimated variance P k|k , through the proportional correction symmetric sampling strategy, we get 2n+1 Sigma sampling points χ i ′ and the corresponding weights and Then the sampling point is

[0146] The nonlinear state function is transferred as follows:

[0147]

[0148] The mean and variance of the one-step state prediction are:

[0149]

[0150] c. Measurement update: Based on the mean and variance obtained from the time update and the sampling strategy formula, 2n+1 Sigma sampling points ζ′ can be obtained. i and the corresponding weights and After the nonlinear measurement function is transferred, we get:

[0151]

[0152] The measurement variables are further predicted with mean, variance and covariance:

[0153]

[0154] According to the measurement value z at time k+1 k+1 , the filter gain K can be calculated k+1, the state estimation and estimated variance at time k+1:

[0155]

[0156] In a preferred embodiment of the present invention, the above method uses sonar images for testing.

[0157] The following combination Figures 5 to 7 The scale space discriminant tracking method based on multi-model fusion detection of the present invention is analyzed from the aspects of video dynamic detection and tracking test results, basic performance and dynamic detection and tracking performance.

[0158] use Figure 5 The video data shown in the video dynamic detection and tracking test results show that Figure 5 The data in has the characteristics of large noise, unclear target features, and complex environment. Figure 6 The experimental results show that the algorithm of the present invention can detect the target in the 201st, 255th, 270th and 299th frames and the tracking effect is good. Figure 7 It can be seen that the target motion trajectory after unscented Kalman filtering is smoother and more reliable than the original result.

[0159] From the above overall video effects, statistical results of objective indicators and dynamic performance of video detection and tracking, it can be seen that the scale space discriminant tracking method based on multi-model fusion detection of the present invention has good visual effects and dynamic detection and tracking performance, and provides a very effective technical means for the field of dynamic image detection and tracking.

[0160] The preferred embodiments of the present invention have been described in detail above. It should be understood that numerous modifications and variations based on the concepts of the present invention are possible without inventive effort by those skilled in the art. Therefore, any technical solution that can be derived by one skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. A scale-space discriminative tracking method for underwater video targets based on multi-model fusion, characterized by: The method comprises the following steps: Step 1: Use the least squares estimation method to complete the multi-sensor registration of imaging sonar and binocular camera; Step 2: A multi-model fusion detection algorithm based on the YOLO detection model and the Gaussian mixture detection model is used to detect underwater video targets input from sonar video and camera video to obtain the relative position information and relative scale information of the target; Step 3: Initialize the scale space discriminant tracker using the target relative position information and relative scale information of step 2 as initial values, use the position filter to predict the target relative position information of each frame in the video, and use the scale filter to predict the scale of the target in the neighborhood of the predicted position of each frame to obtain the relative position information and relative scale information of the target in each frame; Step 4: Based on the multi-sensor imaging characteristics, the position and scale information in the sonar video is converted to the binocular camera coordinate system and fused with the position and scale information in the camera video to obtain the position and scale information of the target in each frame in the camera coordinate system; Step 5: Based on the target position information of each frame obtained in step 4, the target position is filtered using an unscented Kalman filter to obtain the final target state and motion trajectory.

2. The underwater video target scale space discriminant tracking method based on multi-model fusion according to claim 1, characterized in that: The step 2 specifically includes: Step 2.1: Use the YOLO detection algorithm and Gaussian mixture detection algorithm to detect the input sonar and camera video images respectively to obtain the relative position and scale information of the target; Step 2.2: Fuse the detection results of the YOLO detection algorithm with the detection results of the Gaussian mixture detection algorithm. Among them, w is the weight of each detection algorithm, y represents the relative position information or scale information of the target, and Y is the fused relative position and scale information of the target.

3. The underwater video target scale space discriminant tracking method based on multi-model fusion according to claim 1, characterized in that: The step 3 specifically includes: Step 3.1, taking the target relative position and target relative scale of step 2 as the target initial value of the first frame of the tracking sequence; Step 3.2: Extract the features of the target location candidate window and transform it into the Fourier domain; Step 3.3: Generate the target position regression matrix and transform it into the Fourier domain; Step 3.4: Generate n-scale candidate boxes around the target initial box, extract features of the corresponding area for each candidate box, and transform the generated n features into the Fourier domain; Step 3.5: Generate the scale regression matrix of the target and transform it into the Fourier domain; Step 3.6: Train to obtain the position tracking template and scale tracking template; Step 3.7: For the new frame, use the position tracking template to calculate the response on the candidate window. The position of the maximum response is the target position of the new frame. The formula for calculating the response is: Where R is the response, A and B are the numerator and denominator of the correlation filter respectively, Z is the feature map extracted from the target prediction position of the new frame, and λ is the regularization coefficient; Step 3.8: Use the scale tracking template to calculate the response of different scale multipliers at the desired location. Obtain the scale multiplier of the maximum response as the target scale of the new frame; Step 3.9: Use the new scale and position to continue tracking the position of the next frame until all frames are predicted.

4. The underwater video target scale space discriminant tracking method based on multi-model fusion according to claim 3, characterized in that: The step 4 specifically includes: In step 4.1, set the value of n to 30-40; set the imaging sonar range to 10-60 meters; set the resolution and frame rate of the binocular camera, and the peak power parameters of the visual signal processing board; Step 4.2: Based on the imaging characteristics and installation positions of the imaging sonar and binocular camera, calculate the measurement value of the underwater video target in the coordinate system of the underwater multi-sensor detection and tracking system.

5. The underwater video target scale space discriminant tracking method based on multi-model fusion according to claim 1, characterized in that: The step 5 specifically includes: Step 5.1, calculate and obtain the initial state estimate and estimated variance of the filter; Step 5.2: Update the time to obtain the mean and covariance of the target state at time k+1 predicted at time k; Step 5.3: Update the measurement to obtain the state estimate and estimated variance at time k+1.

6. The underwater video target scale space discriminant tracking method based on multi-model fusion according to claim 2, characterized in that: The YOLO detection model algorithm in step 2.1 specifically includes: The YOLO detection algorithm uses a separate convolutional neural network model to achieve end-to-end target detection. First, the input image is resampled, then fed into the convolutional neural network, and finally the network prediction result is processed to obtain the target detection result.

7. The underwater video target scale space discriminant tracking method based on multi-model fusion according to claim 2, characterized in that: The target detection algorithm based on the Gaussian mixture model in step 2.1 includes: The Gaussian mixture model-based target detection algorithm uses multiple single Gaussian models as the model of a pixel position, using the formula |I(x,y,t)-μ i (x,y,t-1)|<λ×σ i (x,y,t-1),i=1,2,…,K, Judge the new pixel, where I is the pixel value of the new pixel, μ is the mean of the existing Gaussian model, σ is the standard deviation of the existing Gaussian model, K represents the number of existing Gaussian models, and λ is the regularization coefficient; If the new pixel matches the single model, the pixel is judged to be the background, and the weight of the single model matching the new pixel is corrected; if there is no model matching the new pixel, the pixel is judged to be the foreground, and the single model with the least importance in the multi-model set is removed, and a new single model is added.

8. The underwater video target scale space discriminant tracking method based on multi-model fusion according to claim 7, characterized in that: The detection algorithm based on the Gaussian mixture model in step 2.1 specifically includes: Step 2.1.1 Define pixel model Each pixel is described by multiple single models: P(p)={[w i (x,y,t),u i (x,y,t),σ i (x,y,t) 2 ]}, i = 1, 2, ..., K, where K represents the number of single models in the Gaussian mixture model, w i (x, y, t) represents the weight of each model, μ is the mean of the existing Gaussian model, and σ is the standard deviation of the existing Gaussian model, satisfying: Step 2.1.2 Update parameters and perform foreground detection Step 2.1.2.1 If the pixel value of the corresponding point (x, y) of the newly input image satisfies: |I(x,y,t)-μ i (x,y,t-1)|<λ×σ i (x,y,t-1),i=1,2,…,K, The new pixel is matched with the single model, the pixel is judged to be background, and step 2.1.2.2 is performed; if there is no model matching the new pixel, the pixel is judged to be foreground, and step 2.1.2.3 is performed; Step 2.1.2.2, Correct the weights of the single model that matches the new pixel. The new weights are: w i (x,y,t)=w i (x,y,t-1)+dw=w i (x,y,t-1)+α(1-w i (x,y,t-1)), Where dw is the weight increment, the formula is dw=a(1-w i (x,y,t-1)), Similarly, correct the mean and variance of the single model matching the new pixel, and then proceed to step 2.1.2.4; Step 2.1.2.3: If the new pixel does not match any existing single model, it is divided into the following two cases: a. If the number of single models has reached the maximum allowed number, remove the single model with the least importance in the current multi-model set; the importance calculation formula is: b. If the number of current single models does not reach the maximum allowed number, add a new single model with the mean of the new model being the new pixel value; Step 2.1.2.

4. Weight Normalization Where w represents the weight of each model, and W represents the weight of each model after weight normalization.

9. The underwater video target scale space discriminant tracking method based on multi-model fusion according to claim 3, characterized in that: In step 3: The steps for location tracking are: Referring to the position of the template in the previous frame, a sample Z is extracted in the current frame at a size twice the target scale of the previous frame; Using the sample Z and the numerator A and denominator B of the position model, according to the formula Calculate the response of the new position, and the place with the largest response is the new position of the target; The scale tracking steps are: Use the scale model to calculate the responses of different scale multipliers at the desired position, and obtain the scale multiplier with the maximum response as the target scale of the new frame; The steps for training model update are: Using the newly obtained relative position and scale of the target, new filters are trained to update the position model and scale model.

Citation Information

Patent Citations

  • Trinocular underwater detection method based on acousto-optic imaging

    CN109143247A