Underwater target tracking classification method and system

By employing underwater acoustic image denoising, FCM and K-means dual clustering, and an improved SORT algorithm, combined with static-dynamic feature fusion, we have achieved efficient detection, tracking, and classification of underwater targets. This solves the problem of underwater target tracking and classification in existing technologies and improves detection accuracy and real-time performance.

CN121837751APending Publication Date: 2026-04-10SHANGHAI OCEAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI OCEAN UNIV
Filing Date
2025-12-29
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing underwater target tracking and classification methods and systems cannot efficiently detect, track, classify, and analyze underwater targets, and face problems such as poor image quality, difficulty in target recognition, high real-time requirements, and low degree of automation.

Method used

By employing underwater acoustic image denoising, FCM and K-means double clustering high overlap ROI extraction, FCM-PCNN composite segmentation, improved SORT multi-target tracking algorithm, and intelligent classification technology that combines static-dynamic feature fusion, along with background estimation, spatial filtering, bright spot suppression, and contrast enhancement, we can achieve efficient target detection, tracking, and classification.

Benefits of technology

It achieves accurate detection, stable tracking, and automatic classification of underwater biological targets, improves image quality, reduces human intervention, enhances real-time performance and automation, and solves the problem of efficient tracking and classification of underwater targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121837751A_ABST
    Figure CN121837751A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater target tracking classification method and system. The tracking classification method comprises the following steps: S3, carrying out ROI region identification, and identifying a potential target region by using an FCM and K-means biclustering algorithm; s4, generating a high-overlapping-degree ROI mask, and calculating the intersection-to-union ratio of an FCM result and a K-means result; s5, target segmentation is carried out, and target segmentation is realized in combination with FCM pre-segmentation and a PCNN neural network; s6, performing target detection, analyzing and detecting the segmented target based on a connected region, and extracting static features of the target; s7, multi-target tracking is carried out, and target tracking is realized based on an improved SORT algorithm; and S8, target intelligent classification is carried out, static features and dynamic features of the target are fused, and classification and identification of the underwater target are realized. According to the tracking classification method and system, accurate detection, stable tracking and automatic classification of underwater biological targets can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a kind of underwater target tracking classification method and system. BACKGROUND

[0002] In the field of marine fishery resources management, underwater ecological environment monitoring and intelligent breeding management, underwater acoustic imaging technology has become a key technical means to obtain underwater environment and biological information. Multi-beam imaging sonar equipment can generate high-resolution underwater images, providing important data sources for underwater target detection and analysis. However, due to the complexity and particularity of the underwater environment, underwater acoustic image processing and analysis still face many challenges:

[0003] 1. Image quality problems: underwater acoustic images often have strong reverberation background, various noise interference, uneven brightness and low contrast, which seriously affect the accuracy of subsequent analysis.

[0004] 2. Target recognition difficulty: underwater target boundary is fuzzy, shape is variable, and is easy to deform and overlap under the influence of water flow and its own movement.

[0005] 3. Real-time requirement: fishery resources monitoring and underwater target tracking require rapid processing of a large amount of continuous image data, which puts high requirements on the real-time performance of the algorithm.

[0006] 4. Low degree of automation: traditional underwater acoustic image processing and analysis relies on manual intervention, which is not only inefficient, but also the results are easily affected by subjective factors.

[0007] 5. Target classification complexity: the acoustic characteristics of different underwater organisms (such as fish, shrimp, jellyfish, etc.) are similar but have subtle differences, making accurate classification difficult.

[0008] In summary, the main problem at present is:

[0009] At present, due to various factors, the existing tracking classification method and system cannot efficiently detect, track, classify and analyze underwater targets. SUMMARY

[0010] The purpose of the present application is to provide a kind of underwater target tracking classification method and system, which can efficiently detect, track, classify and analyze underwater targets.

[0011] In order to achieve the above technical purpose, the present application adopts the following technical solutions:

[0012] A kind of underwater target tracking classification method, the tracking classification method includes:

[0013] S1, establish processing data set;

[0014] S2, underwater acoustic image noise reduction is performed;

[0015] S3, ROI region recognition is performed, and FCM and K-means double clustering algorithms are used to identify potential target regions;

[0016] S4, high overlap ROI mask generation is performed, and the intersection and union ratio of FCM and K-means results is calculated, the high overlap region is reserved and morphological optimization is performed;

[0017] S5, target segmentation is performed, and FCM pre-segmentation and PCNN neural network are combined to realize target segmentation;

[0018] S6, target detection is performed, and the segmented target is detected based on connected region analysis, and the static features of the target are extracted;

[0019] S7, multi-target tracking is performed, and the improved SORT algorithm is used to realize target tracking;

[0020] S8, target intelligent classification is performed, and the static and dynamic features of the target are fused to realize classification and identification of underwater targets;

[0021] S9, confidence evaluation is performed.

[0022] Further, the method for realizing the underwater acoustic image noise reduction comprises:

[0023] Background estimation: background estimation is performed using optimized median filtering, and the original image is subtracted, thereby effectively removing static background noise;

[0024] Spatial filtering: fast non-local mean filtering is used for noise reduction processing, which improves the processing speed while maintaining good noise reduction effect;

[0025] Bright spot suppression: adaptive threshold detection is used to suppress bright spot noise, so as to avoid the interference of strong reflection points on subsequent analysis;

[0026] Contrast enhancement: contrast adjustment method is used to improve the discrimination of target and background.

[0027] Further, the method for realizing the ROI region recognition comprises:

[0028] An optimized FCM clustering algorithm is realized,

[0029] K-means clustering is used to jointly determine the ROI region,

[0030] Potential target regions are identified based on pixel brightness and spatial distribution characteristics.

[0031] Furthermore, the specific method for generating a high-overlap ROI mask includes:

[0032] Calculate the intersection-union ratio (IUU) of the FCM and K-means results;

[0033] Only retain ROI regions with high overlap;

[0034] Use morphological operations to optimize mask quality.

[0035] Furthermore, the specific methods for performing target segmentation include:

[0036] Implement an improved FCM clustering algorithm for pre-segmentation;

[0037] Develop accurate segmentation based on PCNN;

[0038] By combining the advantages of both algorithms, segmentation accuracy can be improved.

[0039] Furthermore, the specific methods for performing target detection include:

[0040] Detection of segmented targets based on connected component analysis;

[0041] Calculate the static characteristics of the target, including its area, shape, and density;

[0042] Implement a nonmaximum suppression algorithm to filter overlapping targets;

[0043] Establish a target quality assessment mechanism to filter out low-quality test results.

[0044] Furthermore, the specific methods for performing multi-target tracking include:

[0045] Achieving efficient target tracking based on the improved SORT algorithm;

[0046] Use a linear prediction model to predict the target's position in the next frame;

[0047] Target matching is achieved based on distance and feature similarity;

[0048] ID management procedures for target appearance, disappearance, occlusion, and reappearance;

[0049] Maintain a complete historical record of the target trajectory.

[0050] Furthermore, the specific methods for performing target intelligent classification include:

[0051] Static feature extraction: Calculate the target's area-to-perimeter ratio, elongation, duty cycle, and contour roughness.

[0052] Dynamic feature extraction: Analyze the target's moving speed, acceleration, direction of motion, and brightness changes;

[0053] Multi-feature fusion classification: a rule-based classification algorithm that comprehensively evaluates various features;

[0054] Supports multi-category recognition: including fish, shrimp, jellyfish, and unknown categories;

[0055] Classification confidence score calculation: Quantifies the reliability of classification results.

[0056] Furthermore,

[0057] Step S3 also includes:

[0058] Large images are downsampled to speed up clustering. The original image is downsampled by average to obtain a smaller image. All subsequent clustering is performed on the new image.

[0059] The FCM algorithm is optimized by using vectorization operations to output the confidence membership matrix of the target region, thereby reducing computational complexity.

[0060] An optimized joint algorithm of FCM clustering and K-means clustering is implemented to identify target regions based on pixel brightness and spatial distribution features, and K-means clustering is used for joint judgment to ensure the reliability of detection results.

[0061] Step S4 also includes:

[0062] Contour extraction and minimum bounding rectangle construction: Calculate the minimum bounding rectangle for the binary mask output in step S3, and also calculate the minimum bounding rectangle for the maximum connected component obtained by K-means;

[0063] IoU Fast Calculation and 80% Threshold Filtering: Use the formula to calculate the high overlap ROI mask. If the Euclidean distance d between the centers of two rectangles is greater than half the length and width of any minimum bounding rectangle, skip it; otherwise, calculate the intersection area of ​​the rectangles.

[0064] Sub-region mask cropping and morphological optimization: The area where the grayscale of the original image intersects with the sub-region mask is cropped, and the target ROI area is optimized using erosion operation;

[0065] Step S5 also includes:

[0066] The improved FCM clustering algorithm is used for pre-segmentation to initially identify the outline of the target region within the ROI region;

[0067] We developed a precise segmentation method based on PCNN, which utilizes its pulse coupling characteristics to accurately extract target boundaries and gradually obtain a refined target region after multiple iterations.

[0068] Optimize PCNN parameters; implement a skip iteration mechanism to reduce the number of iterations; adaptively adjust the threshold to avoid oversegmentation or undersegmentation;

[0069] Step S6 also includes:

[0070] For each region with an area greater than or equal to 50, calculate 6-dimensional features: area-perimeter ratio, elongation, duty cycle, profile roughness, convex hull degree, and brightness contrast.

[0071] Implement a non-maximum suppression algorithm to find local maxima and filter out the remaining values ​​in the neighborhood to prevent a large amount of overlap in the subsequently identified target boxes;

[0072] Establish a target quality assessment mechanism to filter out low-quality test results: set a threshold for filtering, and if the area and elongation indicators are less than a fixed value, they will not enter the trajectory pool;

[0073] Step S7 also includes:

[0074] Using a linear prediction model to predict the target's position in the next frame improves the accuracy of target matching;

[0075] Target matching is achieved based on distance and feature similarity, taking into account factors such as location distance, predicted trajectory, and geometric features.

[0076] Implement a mechanism to handle the disappearance, occlusion, and reappearance of targets, and maintain a complete historical record of target trajectories;

[0077] Establish an intelligent ID management mechanism to assign temporary IDs to newly detected targets; assign permanent IDs to targets that appear continuously and stably and reach a threshold; set a target disappearance tolerance to support ID retention after short-term occlusion.

[0078] Step S8 also includes:

[0079] Static feature extraction: Concatenating multidimensional features to obtain a static vector;

[0080] Dynamic feature extraction: The dynamic vector is given by the trajectory cache in step S7. Data cleaning is performed, and the sample data is truncated and mean is calculated.

[0081] Multi-feature fusion classification: A rule-based classification algorithm that comprehensively evaluates various features and supports the identification of fish, shrimp, jellyfish and unknown categories;

[0082] Classification confidence calculation: Based on the weighted score of the matching degree of each feature, the confidence of the classification result is quantified;

[0083] Step S9 also includes:

[0084] Trajectory confidence: Based on the continuity, smoothness and consistency of the trajectory, the trajectory smoothness and prediction error are calculated using the 7 consecutive trajectory segments obtained in step S8.

[0085] Geometric confidence: Based on the stability evaluation of the target geometric features obtained from the multi-frame static features in step S8, including the rate of change of size, shape consistency and contour stability;

[0086] ROI Confidence: Based on the image quality assessment of the target region, the edge sharpness is quantified by statistically analyzing the proportion of edge pixels within the mask. The consistency of the region is evaluated by calculating the standard deviation of the internal gray-level distribution and introducing an exponential decay function. The gray-level contrast between the region and the external domain is calculated to measure the target-background differentiation capability. The ROI confidence is generated by fusion of the three image quality indicators.

[0087] An underwater target tracking and classification system, the tracking and classification system comprising:

[0088] Video frame extraction module: responsible for converting the input video into an image sequence;

[0089] Image preprocessing module: Implements dedicated noise reduction and enhancement algorithms for underwater acoustic images; the underwater acoustic image noise reduction algorithm in the image preprocessing module includes background estimation, fast nonlocal mean filtering, bright spot suppression, and contrast enhancement;

[0090] ROI extraction module: Identifies highly overlapping ROI regions based on FCM and K-means dual clustering algorithms;

[0091] Target segmentation module: Combines FCM pre-segmentation and PCNN neural network to achieve accurate target segmentation;

[0092] Target tracking module: Based on the improved SORT algorithm, it achieves efficient and stable multi-target tracking; the target tracking module can perform complete target lifecycle management;

[0093] Target classification module: This module integrates static and dynamic features to achieve intelligent classification of underwater targets. It supports the extraction of static and dynamic features of targets and uses rule-based classification algorithms to achieve multi-class recognition.

[0094] Results visualization module: Visually presents tracking results and classification information on the original image.

[0095] The main advantages of the underwater target tracking and classification method and system of the present invention compared with the prior art are as follows:

[0096] By optimizing underwater acoustic image denoising, extracting high-overlap ROIs through FCM and K-means double clustering, FCM-PCNN composite segmentation, improving the SORT multi-target tracking algorithm, and intelligent classification technology that integrates static and dynamic features, accurate detection, stable tracking, and automatic classification of underwater biological targets can be achieved. Attached Figure Description

[0097] Figure 1 This is an overall flowchart of the underwater target tracking and classification method and system of the present invention;

[0098] Figure 2 The image is the original image obtained via step S1 in the tracking and classification method of the present invention.

[0099] Figure 3 The image is the denoised image processed by step S2 in the tracking and classification method of the present invention.

[0100] Figure 4 This refers to the ROI region extracted and morphologically processed in step S3 in the tracking and classification method of the present invention.

[0101] Figure 5 This is the result of applying the ROI-mask extracted by the joint decision of FCM and k-means to a denoised image in step S4 of the tracking classification method of the present invention.

[0102] Figure 6 This is a diagram showing the result of joint segmentation by FCM and PCNN in step S5 of the tracking and classification method of the present invention.

[0103] Figure 7 This is a frame result image of the target detection and multi-target tracking performed in steps S6, S7, and S9 of the tracking and classification method of the present invention.

[0104] Figure 8 This is an image processed by the tracking and classification method of the present invention to perform multi-species classification by fusing the static and dynamic features of the target in step S8.

[0105] Figure 9 This refers to the final processing result of step S10 in the tracking classification method of the present invention, which outputs ID, tracking box, and trajectory history. Detailed Implementation

[0106] The following provides further details on specific embodiments of the present invention:

[0107] This embodiment provides an underwater target tracking and classification method that can efficiently detect, track, classify, and analyze underwater targets.

[0108] See Figure 1The tracking and classification method of this embodiment includes the following steps S1 to S10.

[0109] S1, perform video frame extraction, extract image frames from the original video, and build a processing dataset.

[0110] Specifically, see Figure 2 ,

[0111] This feature allows for batch processing of multiple video files by extracting data based on time range and frame interval. VideoCapture is used for video reading to ensure the quality and timing accuracy of the extracted image frames.

[0112] It should be noted that the data obtained by extracting frames from the acquired video is a noisy binary image obtained by post-processing a multibeam sonar image.

[0113] S2 performs underwater acoustic image denoising, thereby developing a dedicated denoising processing workflow suitable for underwater acoustic images.

[0114] Specifically, see Figure 3 The underwater acoustic image denoising specifically includes:

[0115] 1) Background estimation: Background estimation is performed using optimized median filtering and then subtracted from the original image, thereby effectively removing static background noise;

[0116] 2) Spatial filtering: Fast nonlocal mean filtering is used for noise reduction to replace the traditional BM3D algorithm, which improves processing speed while maintaining good noise reduction effect;

[0117] 3) Bright spot suppression: Adaptive threshold detection and suppression of bright spot noise are used to avoid interference from strong reflection points in subsequent analysis;

[0118] 4) Contrast Enhancement: Use simple and efficient contrast adjustment methods to improve the distinction between the target and the background.

[0119] It should be noted that,

[0120] Static background is estimated using optimized median filtering:

[0121]

[0122] in, pixel position The background grayscale value at that location, For filtering window, The gray levels are the input binarized sonar image.

[0123] Further noise reduction is performed using fast nonlocal mean filtering.

[0124]

[0125] in, Let x be the gray value of the filtered image at position x. This represents the similarity weight.

[0126] Adaptive threshold bright spot detection is performed using mean and variance, brightness threshold suppression is applied, and finally, linear stretching is used for contrast enhancement.

[0127] S3. Perform ROI region identification, using FCM and K-means bi-clustering algorithms to identify potential target regions.

[0128] Specifically, see Figure 4 The ROI region identification specifically includes:

[0129] 1) Implement an optimized FCM (Fuzzy C-means) clustering algorithm;

[0130] 2) Simultaneously use K-means clustering to jointly determine the ROI region;

[0131] 3) Identify potential target areas based on pixel brightness and spatial distribution characteristics.

[0132] More specifically,

[0133] ROI region identification: identifying potential target regions based on FCM and K-means dual clustering algorithms.

[0134] 1) Downsampling is performed on large images to speed up clustering. The original image is downsampled by average to obtain a smaller image. All subsequent clustering is performed on the new image.

[0135] 2) Optimize the FCM algorithm using vectorization operations to output the confidence membership matrix of the target region, thereby reducing computational complexity;

[0136] 3) Implement an optimized joint algorithm of FCM (fuzzy C-means) clustering and K-means clustering to identify target regions based on pixel brightness and spatial distribution features, and use K-means clustering for joint judgment to ensure the reliability of detection results.

[0137] It should be noted that,

[0138] We compared FCM clustering with K-means, setting the number of clusters to 2 (target / background), and the input features were a combination of pixel brightness and spatial coordinates:

[0139]

[0140] in, For pixel brightness, These are spatial coordinates.

[0141] S4 generates a high-overlap ROI mask, calculates the intersection-union ratio of FCM and K-means results, retains high-overlap regions, and performs morphological optimization.

[0142] Specifically, see Figure 5 The high overlap ROI mask generation specifically includes:

[0143] 1) Calculate the intersection-over-union ratio (IoU) of the FCM and K-means results;

[0144] 2) Only retain ROI regions with high overlap (>80%);

[0145] 3) Use morphological operations (erosion, dilation) to optimize mask quality.

[0146] More specifically,

[0147] 1) Contour extraction and minimum bounding rectangle construction: Calculate the minimum bounding rectangle for the binary mask output in step S3. Similarly, calculate the minimum bounding rectangle for the maximum connected component obtained by K-means.

[0148] 2) Fast IoU Calculation and 80% Threshold Filtering: The high overlap ROI mask is calculated using a formula. If the Euclidean distance d between the centers of two rectangles is greater than half the length and width of any minimum bounding rectangle, it is skipped directly; otherwise, the intersection area of ​​the rectangles is calculated.

[0149] 3) Sub-region mask cropping and morphological optimization: Cropping out the area where the grayscale of the original image intersects with the sub-region mask, and using erosion operation to optimize the target ROI area.

[0150] It should be noted that,

[0151] Extract bounding boxes from the FCM and K-means results respectively, and calculate the intersection-union ratio:

[0152]

[0153] in, The binary mask target regions obtained by the two algorithms retain only the regions with IoU>0.8 to eliminate the error of over-segmentation or over-merging by a single algorithm.

[0154] Morphological optimization: small regions are removed using opening and closing operations.

[0155] 1) Extract the contour of the largest connected component using cv2.findContours, and take the one with the largest area as the initial target;

[0156] 2) Record the initial area S0, initialize the corrosion count n = 0 for subsequent area decay judgment;

[0157] 3) Use cv2.erode to perform erosion once, with the structuring element being p×p and the size of p being 10;

[0158] 4) Extract the contour of the largest connected component again. After erosion, it may break, and only keep the largest block;

[0159] 5) Record the current area, and calculate the remaining target area by S;

[0160] 6) n = n + 1 for corrosion count;

[0161] 7) Make a loop judgment. If S < S0 / count and count = 5, then terminate the erosion loop.

[0162] For S5, perform target segmentation, and combine FCM pre-segmentation and PCNN neural network to achieve accurate target segmentation.

[0163] Specifically, see Figure 6 , the target segmentation specifically includes:

[0164] 1) Implement an improved FCM clustering algorithm for pre-segmentation;

[0165] 2) Develop an accurate segmentation based on PCNN (Pulse Coupled Neural Network);

[0166] 3) Combine the advantages of the two algorithms to improve the segmentation accuracy.

[0167] More specifically,

[0168] For target segmentation, combine FCM pre-segmentation and PCNN neural network to achieve accurate target segmentation:

[0169] 1) Use an improved FCM clustering algorithm for pre-segmentation to initially identify the contour of the target area within the ROI region;

[0170] 2) Develop an accurate segmentation based on PCNN (Pulse Coupled Neural Network), and use its pulse coupling characteristics to accurately extract the target boundary, and gradually obtain a fine target area after multiple loops;

[0171] 3) Optimize the PCNN parameters, including connection strength, threshold decay coefficient, etc.; implement a skip iteration mechanism to reduce the number of iterations; adaptively adjust the threshold to avoid over-segmentation or under-segmentation.

[0172] It should be noted that,

[0173] Run FCM again in the ROI region to obtain the bright area mask, and the formula for the output and threshold update is:

[0174]

[0175]

[0176] in For the nth frame pixels The binary decision result, It is the membership degree (target class) output by FCM. It is a pixel-level temporal threshold. It is a forgetting factor. It is an update gain.

[0177] S6. Perform target detection, detect segmented targets based on connected component analysis, and extract static features of the targets.

[0178] Specifically, see Figure 7 The target detection specifically includes:

[0179] 1) Detect segmented targets based on connected component analysis;

[0180] 2) Calculate the static characteristics of the target, such as its area, shape, and density;

[0181] 3) Implement the non-maximum suppression (NMS) algorithm to filter overlapping targets;

[0182] 4) Establish a target quality assessment mechanism to filter out low-quality test results.

[0183] More specifically,

[0184] Object detection, based on connected component analysis, detects segmented targets and extracts their static features:

[0185] 1) Calculate 6-dimensional features for each region with an area greater than or equal to 50: area-to-perimeter ratio, elongation, duty cycle, profile roughness, convex hull degree, and brightness contrast.

[0186] 2) Implement the non-maximum suppression (NMS) algorithm to find local maxima and filter out the remaining values ​​in the neighborhood to prevent a large amount of overlap in the subsequently identified target boxes;

[0187] 3) Establish a target quality assessment mechanism to filter out low-quality detection results, such as targets with too small an area or irregular shape: set a threshold for filtering, and if the area, elongation, or other indicators are less than a fixed value, they will not enter the trajectory pool.

[0188] It should be noted that,

[0189] By combining the input stable contour with the original ROI image, specific data on static and dynamic features are extracted. The following are the specific calculation formulas for key factors:

[0190]

[0191]

[0192]

[0193] in, For the target pixel area, For the elongation formula, For the duty cycle formula, It is the convex hull area, used to describe the density of the target.

[0194] S7 performs multi-target tracking, based on the improved SORT algorithm, including position prediction, target matching, and ID management.

[0195] Specifically, see Figure 7 The multi-target tracking specifically includes:

[0196] 1) Achieve efficient target tracking based on the improved SORT algorithm;

[0197] 2) Use a linear prediction model to predict the target's position in the next frame;

[0198] 3) Target matching is achieved based on distance and feature similarity;

[0199] 4) ID management procedures when the target appears, disappears, is obscured, or reappears;

[0200] 5) Maintain a complete historical record of the target trajectory.

[0201] It should be noted that,

[0202] The ID management specifically includes:

[0203] 1) Assign a temporary ID to the newly detected target;

[0204] 2) Assign a permanent ID to targets that appear continuously and stably for a period of time up to a threshold (default 5 frames);

[0205] 3) Achieve target lifecycle management: registration, update, and deregistration;

[0206] 4) Set the target disappearance tolerance (default 3 frames) to support ID retention after short-term occlusion.

[0207] More specifically,

[0208] Multi-target tracking, achieving efficient target tracking based on an improved SORT algorithm:

[0209] 1) Use a linear prediction model to predict the target's position in the next frame to improve the accuracy of target matching.

[0210] 2) Target matching is achieved based on distance and feature similarity, taking into account multiple factors such as location distance, predicted trajectory, and geometric features.

[0211] 3) Implement a mechanism for handling the disappearance, occlusion, and reappearance of targets, and maintain a complete historical record of target trajectories.

[0212] 4) Establish an intelligent ID management mechanism: assign temporary IDs to newly detected targets; assign permanent IDs to targets that appear continuously and stably for a period of time up to a threshold (default 5 frames); set a target disappearance tolerance (default 3 frames) to support ID retention after short-term occlusion;

[0213] It should be noted that,

[0214] Contour similarity is used in the calculation of the total matching cost.

[0215]

[0216]

[0217] in, For contour similarity, The distance is the Euclidean distance from the centroid.

[0218] S8 performs intelligent target classification, integrating the static and dynamic features of the target to achieve underwater target classification and recognition.

[0219] Specifically, see Figure 8 The target intelligent classification specifically includes:

[0220] 1) Static feature extraction: Calculate the target's area-to-perimeter ratio, elongation (length-to-width ratio), duty cycle, and contour roughness, etc.

[0221] 2) Dynamic feature extraction: Analyze the target's moving speed, acceleration, direction of motion, brightness changes, etc.;

[0222] 3) Multi-feature fusion classification: A rule-based classification algorithm that comprehensively evaluates various features;

[0223] 4) Supports multi-category recognition: including fish, shrimp, jellyfish, and unknown categories;

[0224] 5) Classification confidence calculation: Quantify the credibility of the classification results.

[0225] More specifically,

[0226] Intelligent target classification integrates static and dynamic features of targets to achieve underwater target classification and identification:

[0227] 1) Static feature extraction: Concatenate the multidimensional features to obtain a static vector, where the multidimensional features are given in step S6;

[0228] 2) Dynamic feature extraction: Similarly, the dynamic vector is given by the trajectory cache in step S7. In addition, data cleaning is required, and the sample data is truncated by 5% and the mean is calculated.

[0229] 3) Multi-feature fusion classification: A rule-based classification algorithm that comprehensively evaluates various features and supports the identification of fish, shrimp, jellyfish and unknown categories;

[0230] 4) Classification confidence calculation: Based on the weighted score of the matching degree of each feature, the confidence of the classification result is quantified.

[0231] It should be noted that,

[0232] Static features for target intelligent classification: area-to-perimeter ratio, elongation, duty cycle, and contour roughness; 4-dimensional dynamic features: velocity, acceleration, rate of change of orientation, and rate of change of brightness. The classification confidence formula is as follows:

[0233]

[0234] in, It is calculated from category confidence, comprehensive weight, and similarity score.

[0235] S9. Conduct confidence assessment and establish a multi-dimensional confidence assessment system, including trajectory confidence, geometric confidence, and ROI confidence.

[0236] Specifically, see Figure 7 The confidence assessment specifically includes:

[0237] 1) Trajectory confidence: assessed based on trajectory smoothness and continuity;

[0238] 2) Geometric confidence: Stability assessment based on the geometric features of the target;

[0239] 3) ROI confidence: Image quality assessment based on the target region;

[0240] 4) Overall confidence level: The weighted fusion of confidence levels from multiple dimensions provides a comprehensive measure of reliability for tracking and classification results.

[0241] More specifically,

[0242] Confidence assessment: Establish a multi-dimensional confidence assessment system, including:

[0243] 1) Trajectory confidence: Based on the continuity, smoothness and consistency of the trajectory, the trajectory smoothness and prediction error are calculated using the 7 consecutive trajectory segments obtained in step S8.

[0244] 2) Geometric confidence: Based on the stability evaluation of the target geometric features obtained from the multi-frame static features in step S8, including the rate of change of size, shape consistency and contour stability;

[0245] 3) ROI Confidence: Based on the image quality assessment of the target region, the edge sharpness is quantified by statistically analyzing the proportion of edge pixels within the mask. The consistency of the region is evaluated by calculating the standard deviation of the internal gray-level distribution and introducing an exponential decay function. The gray-level contrast between the region and the external domain is calculated to measure the target-background differentiation capability. The ROI confidence is generated by fusion of the three image quality indicators.

[0246] It should be noted that,

[0247] From the following formula,

[0248]

[0249] Based on trajectory confidence Geometric confidence level ROI confidence level Calculate the overall confidence level The stability index of target detection was obtained.

[0250] S10 visualizes the results, presenting the tracking results and classification information on the original image, and supports data export.

[0251] Specifically, see Figure 9 The system visualizes the results, applying tracking information and classification data back to the original image to draw target bounding boxes, ID labels, classification information, and confidence scores. It also enables historical trajectory plotting, supporting different colors to distinguish target categories. Statistical information is generated and displayed on the image, and the results can be exported to CSV format for subsequent analysis.

[0252] The most important innovation of the tracking and classification method in this embodiment is:

[0253] 1) In step S8, a "static-dynamic" dual-channel feature fusion rule is constructed: the static side uses four morphological quantities—area-perimeter ratio, elongation, duty cycle, and contour roughness—while the dynamic side introduces four kinematic quantities—velocity, acceleration, rate of change of direction, and rate of change of brightness. Joint judgments are made using indicators such as duty cycle and velocity-acceleration, and an interpretable confidence score is given. This "morphological + kinematic joint classification" method specifically addresses point 5 mentioned in the background technology: "similar acoustic characteristics of different underwater organisms, making classification difficult."

[0254] 2) In step S7, a multi-target tracking framework of "improved SORT + hierarchical ID management" is proposed:

[0255] ① Based on the linear motion model of standard SORT, the 6-dimensional static contour features (elongation, duty cycle, etc.) output in step S6 are added together with the center position to form a distance metric, forming a joint association cost matrix of "position + shape", which significantly reduces the ID change problem under dense fish groups;

[0256] ② Introduce a two-level lifecycle of "temporary ID → permanent ID": new targets are first assigned a temporary ID, and only after 5 consecutive frames of successful matching are they upgraded to a permanent ID; the disappearance tolerance is set to 3 frames, and the original ID is retained after a brief occlusion.

[0257] ③ For low-confidence trajectories (C_total < 0.5), cancel and reclaim the ID number in advance to avoid ID drift and ID switching.

[0258] This mechanism directly addresses the pain points mentioned in the background technology, namely, "high real-time requirements" and "low degree of automation and high degree of manual intervention". It reduces the ID switching frequency from 89 times per 100 frames in the traditional SORT to 32 times, while the single-core CPU consumes less time, thus realizing long-term, stable and unattended tracking in underwater dense biological swarm scenarios.

[0259] 3) In step S4, FCM fuzzy clustering and K-means hard clustering are used for joint "pixel-rectangle" dual-space decision-making: first, both clustering methods are run simultaneously on the downsampled image; then, "highly overlapping" regions are screened with an Intersection over Union (IoU) ratio ≥ 80% as a threshold, and isolated bright spots are removed using a morphological erosion-dilation cycle. This "dual-cluster mutual verification" mechanism directly overcomes the defects mentioned in point 2 of the background technology, namely "fuzzy target boundaries, over-segmentation or over-merging of a single algorithm," significantly improving the false detection rate of the initial ROI compared to traditional K-means clustering, and providing a highly reliable starting point for subsequent tracking.

[0260] 4) To address the persistent problem of "strong reverberation and low contrast" in underwater acoustic images, a four-level dedicated noise reduction chain is constructed in step S2: "background estimation - non-local mean - bright spot suppression - contrast stretching". First, static reverberation is removed by optimized median filtering, then random speckle noise is removed by fast non-local mean filtering, and strong reflective bright spots are suppressed by adaptive thresholding. Finally, linear stretching is performed to expand the gray-scale difference between the target and the background. This directly solves the problem of subsequent segmentation and tracking error amplification caused by the first point mentioned in the background technology, "poor image quality", and provides a high signal-to-noise ratio input guarantee for the entire process.

[0261] In summary, the above four innovations correspond to the core requirements of "effectiveness of signal-to-noise ratio improvement, reliability of ROI extraction, accuracy of type differentiation, and stability of target detection," ultimately achieving a breakthrough in "accurate detection, differentiation, and correct identification" of underwater targets. This solves the overall technical problem in the background technology that "existing tracking and classification methods cannot efficiently detect, track, classify, and analyze underwater targets."

[0262] The main advantage of the tracking and classification method in this embodiment is that:

[0263] 1) In the tracking and classification method of this embodiment, the underwater biological targets are accurately detected, stably tracked and automatically classified by means of optimized underwater acoustic image denoising, FCM and K-means double clustering high overlap ROI extraction, FCM-PCNN composite segmentation, improved SORT multi-target tracking algorithm and intelligent classification technology of static-dynamic feature fusion.

[0264] In addition, the tracking and classification method of this embodiment has other advantages, as follows:

[0265] 2) In the tracking classification method of this embodiment, the static contour features are embedded into the data association process and a two-level ID lifecycle of "temporary → permanent" is set by the "improved SORT + hierarchical ID management" strategy. This significantly improves the tracking continuity and automation level in dense occlusion scenarios, thereby eliminating the high maintenance cost of manual intervention to correct ID drift required by traditional methods.

[0266] 3) In the tracking and classification method of this embodiment, the "static-dynamic dual-channel feature rule tree" is used to complete the online classification of multiple species such as fish, shrimp, and jellyfish in one go, and outputs an interpretable confidence score. This solves the problems of high misclassification rate and lack of real-time classification methods caused by similar acoustic features in the prior art, and provides timely and reliable species-level data for fishery resource statistics and ecological monitoring.

[0267] This embodiment also provides an underwater target tracking and classification system that can implement the tracking and classification method described above.

[0268] Specifically, the tracking classification system includes:

[0269] 1) Video frame extraction module: responsible for converting the input video into an image sequence;

[0270] 2) Image preprocessing module: Implements dedicated noise reduction and enhancement algorithms for underwater acoustic images;

[0271] The underwater acoustic image denoising algorithm in the image preprocessing module includes background estimation, fast nonlocal mean filtering, bright spot suppression, and contrast enhancement.

[0272] 3) ROI extraction module: Identifies highly overlapping ROI regions based on FCM and K-means bi-clustering algorithms;

[0273] 4) Target segmentation module: Combines FCM pre-segmentation and PCNN neural network to achieve accurate target segmentation;

[0274] 5) Target tracking module: Achieves efficient and stable multi-target tracking based on the improved SORT algorithm;

[0275] This target tracking module implements complete target lifecycle management, including registration, update, deregistration, and hierarchical management of temporary and permanent IDs;

[0276] 6) Target classification module: Integrates static and dynamic features to achieve intelligent classification of underwater targets;

[0277] This target classification module supports the extraction of static features (area-perimeter ratio, elongation, duty cycle, contour roughness) and dynamic features (velocity, acceleration, direction of motion) of targets, and achieves multi-class recognition based on rule-based classification algorithms;

[0278] 7) Results visualization module: Visualizes the tracking results and classification information on the original image.

[0279] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An underwater target tracking and classification method, characterized in that: The tracking classification method includes: S1, Create and process the dataset; S2, perform underwater acoustic image noise reduction; S3, perform ROI region identification, and use FCM and K-means bi-clustering algorithms to identify potential target regions; S4, generate high-overlapping ROI masks, calculate the intersection-union ratio of FCM and K-means results, retain high-overlapping regions and perform morphological optimization; S5, Perform target segmentation, combining FCM pre-segmentation and PCNN neural network to achieve target segmentation; S6, Perform target detection, detect segmented targets based on connected component analysis, and extract static features of the targets; S7 performs multi-target tracking, using an improved SORT algorithm to achieve target tracking; S8 performs intelligent target classification, integrating the static and dynamic features of the target to achieve underwater target classification and recognition; S9, conduct a confidence assessment.

2. The underwater target tracking and classification method according to claim 1, characterized in that: The specific methods for underwater acoustic image denoising include: Background estimation: Background estimation is performed using optimized median filtering and then subtracted from the original image, thereby effectively removing static background noise; Spatial filtering: Fast nonlocal mean filtering is used for noise reduction, which improves processing speed while maintaining good noise reduction effect; Bright spot suppression: Adaptive thresholding is used to detect and suppress bright spot noise to avoid interference from strong reflection points in subsequent analysis; Contrast Enhancement: Use contrast adjustment methods to improve the distinction between the target and the background.

3. The underwater target tracking and classification method according to claim 2, characterized in that: The specific methods for performing ROI region identification include: Implement an optimized FCM clustering algorithm. Simultaneously, K-means clustering is used to jointly determine the ROI region. Identify potential target regions based on pixel brightness and spatial distribution characteristics.

4. The underwater target tracking and classification method according to claim 3, characterized in that: The specific methods for generating high-overlap ROI masks include: Calculate the intersection-union ratio (IUU) of the FCM and K-means results; Only retain ROI regions with high overlap; Use morphological operations to optimize mask quality.

5. The underwater target tracking and classification method according to claim 4, characterized in that: The specific methods for performing target segmentation include: Implement an improved FCM clustering algorithm for pre-segmentation; Develop accurate segmentation based on PCNN; By combining the advantages of both algorithms, segmentation accuracy can be improved.

6. The underwater target tracking and classification method according to claim 5, characterized in that: The specific methods for performing target detection include: Detection of segmented targets based on connected component analysis; Calculate the static characteristics of the target, including its area, shape, and density; Implement a nonmaximum suppression algorithm to filter overlapping targets; Establish a target quality assessment mechanism to filter out low-quality test results.

7. The underwater target tracking and classification method according to claim 6, characterized in that: The specific methods for performing multi-target tracking include: Achieving efficient target tracking based on the improved SORT algorithm; Use a linear prediction model to predict the target's position in the next frame; Target matching is achieved based on distance and feature similarity; ID management procedures for target appearance, disappearance, occlusion, and reappearance; Maintain a complete historical record of the target trajectory.

8. The underwater target tracking and classification method according to claim 7, characterized in that: The specific methods for performing target intelligent classification include: Static feature extraction: Calculate the target's area-to-perimeter ratio, elongation, duty cycle, and contour roughness. Dynamic feature extraction: Analyze the target's moving speed, acceleration, direction of motion, and brightness changes; Multi-feature fusion classification: a rule-based classification algorithm that comprehensively evaluates various features; Supports multi-category recognition: including fish, shrimp, jellyfish, and unknown categories; Classification confidence score calculation: Quantifies the reliability of classification results.

9. The underwater target tracking and classification method according to claim 8, characterized in that: Step S3 also includes: Large images are downsampled to speed up clustering. The original image is downsampled by average to obtain a smaller image. All subsequent clustering is performed on the new image. The FCM algorithm is optimized by using vectorization operations to output the confidence membership matrix of the target region, thereby reducing computational complexity. An optimized joint algorithm combining FCM clustering and K-means clustering is implemented to identify target regions based on pixel brightness and spatial distribution features. K-means clustering is used for joint judgment to ensure the reliability of detection results. Step S4 also includes: Contour extraction and minimum bounding rectangle construction: Calculate the minimum bounding rectangle for the binary mask output in step S3, and also calculate the minimum bounding rectangle for the maximum connected component obtained by K-means; IoU Fast Calculation and 80% Threshold Filtering: Use the formula to calculate the high overlap ROI mask. If the Euclidean distance d between the centers of two rectangles is greater than half the length and width of any minimum bounding rectangle, skip it; otherwise, calculate the intersection area of ​​the rectangles. Sub-region mask cropping and morphological optimization: The area where the grayscale of the original image intersects with the sub-region mask is cropped, and the target ROI area is optimized using erosion operation; Step S5 also includes: The improved FCM clustering algorithm is used for pre-segmentation to initially identify the outline of the target region within the ROI region; We developed a precise segmentation method based on PCNN, which utilizes its pulse coupling characteristics to accurately extract target boundaries and gradually obtain a refined target region after multiple iterations. Optimize PCNN parameters; implement a skip iteration mechanism to reduce the number of iterations; adaptively adjust the threshold to avoid oversegmentation or undersegmentation; Step S6 also includes: For each region with an area greater than or equal to 50, calculate 6-dimensional features: area-perimeter ratio, elongation, duty cycle, profile roughness, convex hull degree, and brightness contrast. Implement a non-maximum suppression algorithm to find local maxima and filter out the remaining values ​​in the neighborhood to prevent a large amount of overlap in the subsequently identified target boxes; Establish a target quality assessment mechanism to filter out low-quality test results: set a threshold for filtering, and if the area and elongation indicators are less than a fixed value, they will not enter the trajectory pool; Step S7 also includes: Using a linear prediction model to predict the target's position in the next frame improves the accuracy of target matching; Target matching is achieved based on distance and feature similarity, taking into account factors such as location distance, predicted trajectory, and geometric features. Implement a mechanism to handle the disappearance, occlusion, and reappearance of targets, and maintain a complete historical record of target trajectories; Establish an intelligent ID management mechanism to assign temporary IDs to newly detected targets; assign permanent IDs to targets that appear continuously and stably and reach a threshold; set a target disappearance tolerance to support ID retention after short-term occlusion. Step S8 also includes: Static feature extraction: Concatenating multidimensional features to obtain a static vector; Dynamic feature extraction: The dynamic vector is given by the trajectory cache in step S7. Data cleaning is performed, and the sample data is truncated and mean is calculated. Multi-feature fusion classification: A rule-based classification algorithm that comprehensively evaluates various features and supports the identification of fish, shrimp, jellyfish and unknown categories; Classification confidence calculation: Based on the weighted score of the matching degree of each feature, the confidence of the classification result is quantified; Step S9 also includes: Trajectory confidence: Based on the continuity, smoothness and consistency of the trajectory, the trajectory smoothness and prediction error are calculated using the 7 consecutive trajectory segments obtained in step S8. Geometric confidence: Based on the stability evaluation of the target geometric features obtained from the multi-frame static features in step S8, including the rate of change of size, shape consistency and contour stability; ROI Confidence: Based on the image quality assessment of the target region, the edge sharpness is quantified by statistically analyzing the proportion of edge pixels within the mask. The consistency of the region is evaluated by calculating the standard deviation of the internal gray-level distribution and introducing an exponential decay function. The gray-level contrast between the region and the external domain is calculated to measure the target-background differentiation capability. The ROI confidence is generated by fusion of the three image quality indicators.

10. An underwater target tracking and classification system, characterized in that: The tracking and classification system includes: Video frame extraction module: responsible for converting the input video into an image sequence; Image preprocessing module: Implements dedicated noise reduction and enhancement algorithms for underwater acoustic images; the underwater acoustic image noise reduction algorithm in the image preprocessing module includes background estimation, fast nonlocal mean filtering, bright spot suppression, and contrast enhancement; ROI extraction module: Identifies highly overlapping ROI regions based on FCM and K-means dual clustering algorithms; Target segmentation module: Combines FCM pre-segmentation and PCNN neural network to achieve accurate target segmentation; Target tracking module: Based on the improved SORT algorithm, it achieves efficient and stable multi-target tracking; the target tracking module can perform complete target lifecycle management; Target classification module: This module integrates static and dynamic features to achieve intelligent classification of underwater targets. It supports the extraction of static and dynamic features of targets and uses rule-based classification algorithms to achieve multi-class recognition. Results visualization module: Visually presents tracking results and classification information on the original image.