Underwater target detection method based on state space model
By using the state space model to construct a parallel path between global degradation modeling and local physical enhancement in underwater target detection, multi-scale features are extracted and deep-level features are integrated, the problem of underwater target detection technology degradation in complex environments is solved, and higher detection accuracy and accuracy are achieved.
Patent Information
- Application Number
- CN202510297865.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-05-30
AI Technical Summary
Underwater target detection technology has the problem of degradation in complex underwater environments, mainly due to media degradation, which destroys the statistical and structural information of the image and affects the accuracy of target detection.
The underwater object detection method based on the state space model is adopted. By constructing a parallel path between global degradation modeling and local physical enhancement, multi-scale feature maps are extracted and deep feature fusion is performed to obtain target classification and positioning information, and then the object detection model is trained to improve detection accuracy.
It effectively overcomes the impact of media degradation on target detection in underwater environments, improves the accuracy and accuracy of underwater target detection, and performs excellently in complex characteristic scenarios.
Smart Images

Figure CN120071114A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater target detection. Specifically, it particularly relates to an underwater target detection method based on a state space model. Background Art
[0002] Underwater target detection, as an important research field, is widely applied in multiple fields such as marine ecological monitoring, fishery resource investigation, seabed exploration, environmental protection, etc. With the continuous development of deep-sea exploration and underwater operations, precise underwater target detection technology has become increasingly important. Different from the onshore environment, underwater imaging is simultaneously affected by multiple degradation factors. They are like invisible hands, distorting the essential characteristics of the image, destroying the statistical and structural information of the image, obscuring the characteristic information of the real data, and affecting the visual perception of the human eye and the pattern recognition efficiency of the computer. In order to overcome the negative impact of the underwater environment on the accuracy of underwater target detection, researchers are committed to developing various technologies and methods aimed at enabling the underwater target detection system to more accurately identify and locate targets in complex underwater environments. Summary of the Invention
[0003] To solve the technical problems mentioned in the above background art, an underwater target detection method based on a state space model is provided. The present invention provides an underwater target detection method based on a state space model. The present invention mainly adopts the idea of a state space model, constructs a parallel path of global degradation modeling and local physical enhancement for underwater images, and accurately extracts several degradation factors existing in underwater images. Aiming at the problem that existing target detection is limited by the local receptive field and it is difficult to model long-range degradation dependencies for various types of image degradations, the present invention poses questions, and finally conducts training and inference based on an underwater dataset in a real scene to obtain accurate underwater target positioning information.
[0004] The technical means adopted by the present invention are as follows:
[0005] An underwater target detection method based on a state space model, comprising the following steps:
[0006] S01: Obtain an underwater low-quality target detection dataset, and randomly divide the underwater low-quality target detection dataset into a training set, a validation set, and a test set according to a certain ratio;
[0007] S02: Construct a backbone network according to the state space model, take the underwater low-quality images in the training set as inputs, and extract multi-scale feature maps of the underwater low-quality images; the multi-scale feature maps include: high-resolution feature maps, medium-resolution feature maps, and low-resolution feature maps;
[0008] S03: Input the multi-scale feature maps extracted in S02 into the PAFPN network to perform deep feature fusion and obtain the output of feature information. The PAFPN network includes: a bottom-up feature aggregation path, a top-down feature parsing, and a feature information output.
[0009] S04: Based on the output of the feature information extracted in S03, obtain the mapping classification and localization information according to the detection head network. The classification and localization information includes: target classification information and target localization box coordinate information.
[0010] S05: Calculate the loss according to the classification and localization information. Wait for the loss to converge, then complete the training of the target detection model and obtain the trained target detection model.
[0011] S06: Use the trained target detection model in S05 to perform inference on the underwater image to be detected, and output and draw the detection results.
[0012] Further, the backbone network composed of the state space model includes: feature convolution downsampling and feature extraction of the state space model SSM module.
[0013] The calculation method of feature convolution downsampling is:
[0014] X 1 = Conv(F in , [k = 3, s = 2, p = 1]);
[0015] where, F in represents the input feature map; Conv represents the convolution operation; k represents the convolution kernel size, s represents the convolution stride, and p is equal to the padding size; X 1 represents the feature map after feature convolution downsampling;
[0016] The calculation method of feature extraction of the state space model SSM module is:
[0017] X 2 , X 3 = Split(Conv(X 1 , [k = 1, s = 1, p = 0]));
[0018] X 4 = SS2D(LEFM(X 2 ));
[0019] X 5 , X 6 = Split(Conv(X 3 , [k = 3, s = 1, p = 0]));
[0020] X 7 = Conv(X6 , [k = 3, s = 1, p = 0]);
[0021] X = Concat(X 4 , X 5 , X 7 );
[0022] F out = Conv(X, [k = 1, s = 1, p = 0]);
[0023] Where Split means splitting the feature map in the channel dimension; LEFM means performing local contrast enhancement on the feature map; SS2D means performing state space modeling on the feature map; Concat means reconnecting the feature map in the channel dimension; X 2 , X 3 , X 4 , X 5 , X 6 , X 7 , X all represent the intermediate states of the corresponding feature maps; F out represents the final output of feature extraction by the SSM module.
[0024] Furthermore, the calculation method for the LEFM to extract local features in the image is:
[0025] X out = Silu(BatchNorm(DWConv(X 2 , [k = 5, s = 2, p = 3])));
[0026] Where DWConv represents the depthwise separable convolution operation; BatchNorm represents the batch normalization operation; Silu represents the activation function; X out represents the output of the LEFM module.
[0027] Furthermore, the calculation method for the SS2D is:
[0028] X 1 = CrossScan(X 0 );
[0029] X 2 = Mamba(X 1 );
[0030] Out = CrossMerge(X 2 );
[0031] Among them, CrossScan represents the cross-scanning operation, which scans the input two-dimensional feature map into a one-dimensional sequence; Mamba represents the basic block of the state space model; CrossMerge represents cross-merging the sequence processed by the state space model into the shape of the original feature map; Out represents the final output.
[0032] Furthermore, the calculation method of the state space model constituting the backbone network feature extraction is as follows:
[0033] F in1 ,F in2 ,F in3 = Backbone(F);
[0034] Among them, F represents the underwater low-quality image, Backbone represents the backbone network composed of the state space model, F in1 represents the low-resolution feature map, F in2 represents the medium-resolution feature map, F in3 represents the high-resolution feature map.
[0035] Furthermore, the calculation method of the bottom-up feature aggregation of the PAFPN network is as follows:
[0036] T 0 = Upsample(F in1 );
[0037] T 1 = Concat(T 0 , F in2 );
[0038] T 2 = SSM(Concat(Upsample(SSM(T 1 )),F in3 ));
[0039] Among them, Upsample represents the upsampling operation on the feature map; SSM represents the feature extraction of the feature map through the state space module, F in1 、F in2 and F in3 respectively represent different-scale feature maps output by the backbone network;
[0040] The calculation method of the top-down feature parsing and feature output of the PAFPN network is as follows:
[0041] F out1 = T 2 ;
[0042] F out2 = SSM(Concat(Conv(F out1), T 1 ));
[0043] F out3 = SSM(Concat(Conv(F out2 ), T 0 ));
[0044] Among them, Conv represents resizing the feature map by downsampling, and T 0 , T 1 and T 2 all represent the output of multi-scale feature maps for bottom-up feature aggregation of the PAFPN network, and F out1 , F out2 and F out3 represent the feature maps of different sizes finally output by the PAFPN network.
[0045] Further, the target classification information is the probability value of the predicted object category in the prediction box, and the calculation method is:
[0046]
[0047] Among them, s i represents the original score output by the detection head network; C represents the number of categories; P(class i ) represents the probability that the category is i.
[0048] Further, the coordinate information of the target localization box in S05 is expressed as:
[0049] (X, Y, W, H, confidence);
[0050] Among them, X and Y represent the coordinate values of the center of the prediction box relative to the upper left corner of the feature map, W and H respectively represent the ratio of the width and height of the prediction box to the width and height of the feature map; confidence is the confidence that there is a target in the prediction box.
[0051] Further, the calculation method of the loss function is:
[0052]
[0053] Among them, represents the predicted value, and Y i the true value; the calculation method of the SmoothL1 loss is:
[0054]
[0055] Among them, x represents the gap between the prediction and the true value; ABS(x) represents the absolute value of x;
[0056] Further, the certain ratio is 8:1:1.
[0057] Compared with the prior art, the present invention has the following advantages:
[0058] In order to solve the problem of the accuracy decline of underwater target detection technology caused by medium degradation, the present invention proposes a state space model block for global degradation modeling, constructs a local physical enhancement path parallel to the convolutional neural network, enables the interaction of global-local hierarchical features in a suitable data capacity, improves the detection accuracy in complex underwater images, and is instantiated as a new underwater detection network applicable to complex feature scenarios.
[0059] The present invention provides an LFEM module derived from a state space model block, which makes up for the problems that the state space model is sensitive to noise and insufficient in local feature extraction in sequence modeling. By constructing a multi-level heterogeneous convolutional filtering network, it completes the suppression of wide-area scattering noise and the enhancement of the contrast of local structures, enabling the SSM sequence encoding modeling to focus on physically consistent features. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0061] Figure 1 It is the model flow chart of the present invention.
[0062] Figure 2 They are the detection results of different methods in the nearshore underwater ecological scene. Among them, (a) is the detection result of our method, and (b) is the detection result of using the YOLOv8 method.
[0063] Figure 3 They are the detection results of different methods in the severely greenish low-light dense scene. (c) is the detection result of our method, and (d) is the detection result of using the YOLOv10 method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0065] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0066] As Figure 1 shown, the present invention provides an underwater target detection method based on a state space model, including the following steps:
[0067] S01: Obtain an underwater low-quality target detection data set, and randomly divide the underwater low-quality target detection data set into a training set, a validation set and a test set according to a certain ratio (8:1:1);
[0068] S02: Construct a backbone network according to the state space model, use the underwater low-quality images in the training set as inputs, and extract multi-scale feature maps of the underwater low-quality images; the multi-scale feature maps include: high-resolution feature maps, medium-resolution feature maps, and low-resolution feature maps;
[0069] As a preferred implementation, in the present application, the construction of the backbone network by the state space model includes: feature convolutional downsampling and feature extraction by the state space model SSM module;
[0070] The calculation method of feature convolutional downsampling is:
[0071] X 1 = Conv(F in , [k = 3, s = 2, p = 1]);
[0072] Among them, F in represents the input feature map; Conv represents the convolution operation; k represents the convolution kernel size, s represents the convolution stride, and p is equal to the padding size; X 1 represents the feature map after feature convolutional downsampling; the calculation method of feature extraction by the state space model SSM module is:
[0073] X 2 , X 3 = Split(Conv(X 1 , [k = 1, s = 1, p = 0]));
[0074] X4 = SS2D(LEFM(X 2 ));
[0075] X 5 , X 6 = Split(Conv(X 3 , [k = 3, s = 1, p = 0]));
[0076] X 7 = Conv(X 6 , [k = 3, s = 1, p = 0]);
[0077] X = Concat(X 4 , X 5 , X 7 );
[0078] F out = Conv(X, [k = 1, s = 1, p = 0]);
[0079] Among them, Split means splitting the feature map in the channel dimension; LEFM means performing local contrast enhancement on the feature map; SS2D means performing state space modeling on the feature map; Concat means reconnecting the feature map in the channel dimension; X 2 , X 3 , X 4 , X 5 , X 6 , X 7 , X all represent the intermediate states of the corresponding feature maps; F out represents the final output of feature extraction by the SSM module.
[0080] The calculation method of the LEFM for extracting local features in the image is as follows:
[0081] X out = Silu(BatchNorm(DWConv(X 2 , [k = 5, s = 2, p = 3])));
[0082] Among them, DWConv represents the depthwise separable convolution operation; BatchNorm represents the batch normalization operation; Silu represents the activation function; X out represents the output of the LEFM module.
[0083] S03: Input the multi-scale feature maps extracted in S02 into the PAFPN network to perform deep feature fusion to obtain feature information output; The PAFPN network includes: a bottom-up feature aggregation path, a top-down feature parsing, and a feature information output.
[0084] Preferably, the SS2D calculation method is as follows:
[0085] X 1 = CrossScan(X 0 );
[0086] X 2 = Mamba(X 1 );
[0087] Out = CrossMerge(X 2 );
[0088] Among them, CrossScan represents the cross-scanning operation, which scans the input two-dimensional feature map into a one-dimensional sequence; Mamba represents the basic block of the state space model; CrossMerge represents cross-merging the sequence processed by the state space model into the shape of the original feature map; Out represents the final output.
[0089] S04: Output the feature information extracted through S03, and obtain the mapping classification and localization information according to the detection head network; the classification and localization information includes: target classification information and target localization box coordinate information;
[0090] S05: Calculate the loss according to the classification and localization information, wait for the loss to converge, then complete the training of the target detection model, and obtain the trained target detection model;
[0091] S06: Infer the underwater image to be detected through the trained target detection model in S05, and output and draw the detection result.
[0092] As a preferred implementation manner, the calculation method for the state space model to constitute the backbone network feature extraction is as follows:
[0093] F in1 ,F in2 ,F in3 = Backbone(F);
[0094] Among them, F represents the underwater low-quality image, Backbone represents the backbone network composed of the state space model, F in1 represents the low-resolution feature map, F in2 represents the medium-resolution feature map, F in3 represents the high-resolution feature map.
[0095] In this application, the calculation method for the bottom-up feature aggregation of the PAFPN network is as follows:
[0096] T 0 = Upsample(F in1 );
[0097] R 1 = Concat(T 0 , F in2 );
[0098] T 2 = SSM(Concat(Upsample(SSM(T 1 ))), F in3 ));
[0099] Among them, Upsample represents the upsampling operation on the feature map; SSM represents the feature extraction of the feature map through the state space module, F in1 , F in2 and F in3 respectively represent the feature maps of different scales output by the backbone network;
[0100] The calculation method of top-down feature parsing and feature output of the PAFPN network is as follows:
[0101] F out1 = T 2 ;
[0102] F out2 = SSM(Concat(Conv(F out1 ), T 1 ));
[0103] F out3 = SSM(Concat(Conv(F out2 ), T 0 ));
[0104] Among them, Conv represents the downsampling size adjustment of the feature map, T 0 , T 1 and T 2 all represent the multi-scale feature map outputs of bottom-up feature aggregation of the PAFPN network, F out1 , F out2 and F out3 represent the feature maps of different sizes finally output by the PAFPN network.
[0105] Preferably, the target classification information is the probability value of the predicted object category in the prediction box, and the calculation method is:
[0106]
[0107] Among them, s i represents the original score output by the detection head network; C represents the number of categories; P(class i ) represents the probability that the category is i.
[0108] Preferably, the target positioning box coordinate information in S05 is represented as:
[0109] (X, Y, W, H, confidence);
[0110] Among them, X and Y represent the coordinate values of the center of the prediction box relative to the upper left corner of the feature map, W and H respectively represent the ratio of the width and height of the prediction box to the width and height of the feature map; confidence is the confidence level indicating whether there is a target within the prediction box.
[0111] Furthermore, the calculation method of the loss function is:
[0112]
[0113] Among them, represents the predicted value, Y i the true value; the SmoothL1 loss calculation method is:
[0114]
[0115] Among them, x represents the gap between the prediction and the true value; ABS(x) represents the absolute value of x;
[0116] Example 1
[0117] As Figure 2 (a), (b) shown, the present invention provides the detection results of the nearshore underwater ecological scene compared with other algorithms. It can be seen from the experimental effect diagrams that both algorithms detected the targets in the underwater images to a certain extent. However, Yolov8 regarded the sea urchin in the middle blue box as the background, resulting in a missed detection of the target. Our method still obtained the correct detection results in the case of poor image quality and difficult-to-distinguish targets.
[0118] As Figure 3 (c), (d) shown, the present invention provides the detection results of different methods in the severely greenish low-light dense scene. It can be seen from the experimental effect diagrams that both algorithms detected the targets in the underwater images to a certain extent. However, the YOLOv10 method failed to detect the target in the middle dark area, and there was also a missed detection on the left side, and a wrong detection occurred. Our method correctly detected all the targets even in such severely greenish scenes.
[0119] In this embodiment, the experimental results of different algorithms are compared from two objective indicators: AP and AP50. In object detection, the AP (Average Precision) indicator is an indicator used to evaluate the performance of an object detection model. The goal of the object detection task is to identify and locate objects in an image or video, and provide a bounding box and the corresponding category for each detected object. For different thresholds, the model can produce different precision rates and recall rates. Plotting these points and forming a PR curve can be used to intuitively understand the performance of the model at different thresholds. AP is a measure of the average value of the area under the PR curve, representing the comprehensive performance of the model across all categories. The calculation method of AP depends on whether multiple detection results (boxes) are allowed to exist for the same object instance. AP50 (Average Precision at 50) is a performance evaluation indicator in object detection, especially used to evaluate the average accuracy of the model when the IoU (Intersection over Union) threshold is 50%. In the object detection task, IoU refers to the ratio of the intersection to the union between the detection box (predicted box) and the true annotation box. AP75 (Average Precision at 75) is used to evaluate the average accuracy of the model when the IoU threshold is 75%. Therefore, the present invention has a significant improvement in the AP, AP50, and AP75 indicators of the original image, and is superior to other underwater object detection algorithms.
[0120] Table 1 Comparison of indicators of the model of the present invention and the processing results of other advanced algorithms
[0121]
[0122] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments. In the above embodiments of the present invention, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways.
[0123] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for underwater target detection based on a state space model, characterized in that: The following steps are involved: S01: Obtain an underwater low-quality target detection dataset, and randomly divide the underwater low-quality target detection dataset into a training set, a validation set, and a test set according to a certain ratio; S02: constructing a backbone network according to the state space model, taking the underwater low-quality image in the training set as input, and extracting a multi-scale feature map of the underwater low-quality image; the multi-scale feature map includes: a high-resolution feature map, a medium-resolution feature map, and a low-resolution feature map; S03: input the multi-scale feature map extracted in S02 into the PAFPN network, perform deep feature fusion to obtain feature information output; the PAFPN network includes: bottom-up feature aggregation path, top-down feature analysis and feature information output; S04: outputting the feature information extracted in S03, and obtaining mapping classification and positioning information according to the detection head network; the classification and positioning information includes: target classification information and target positioning frame coordinate information; S05: Calculate the loss according to the classification and positioning information, wait for the loss to converge, and then complete the target detection model training to obtain a trained target detection model; S06: Inferring the underwater image to be detected by the target detection model trained in S05, and outputting and drawing the detection result.
2. The underwater target detection method based on the state space model according to claim 1, characterized in that: The backbone network of the state space model includes: feature convolution down sampling and state space model SSM module feature extraction; The calculation method of feature convolution downsampling is: X1=Conv(F in ,[k=3,s=2,p=1]); Among them, F in Represents the input feature map; Conv represents the convolution operation; k represents the convolution kernel size, s represents the convolution step size, and p is equal to the padding size; X1 represents the feature map after feature convolution downsampling; The calculation method for feature extraction of the state space model SSM module is: X2, X3=Split(Conv(X1, [k=1, s=1, p=0])); X4 = SS2D (LEFM (X2)); X5, X6=Split(Conv(X3, [k=3, s=1, p=0])); X7=Conv(X6, [k=3, s=1, p=0]); X = Concat(X4, X5, X7); F out =Conv(X,[k=1,s=1,p=0]); Among them, Split means splitting the feature map in the channel dimension; LEFM means local contrast enhancement of the feature map; SS2D means state space modeling of the feature map; Concat means reconnecting the feature map in the channel dimension; X2, X3, X4, X5, X6, X7, X all represent the intermediate states of the corresponding feature maps; F out Represents the final output of the SSM module feature extraction.
3. The underwater target detection method based on the state space model according to claim 2 is characterized in that: The calculation method of LEFM for extracting local features in an image is: X out =Silu(BatchNorm(DWConv(X2,[k=5,s=2,p=3]))); Among them, DWConv represents the depth-separable convolution operation; BatchNorm represents the batch normalization operation; Silu represents the activation function; X out Represents the LEFM module output.
4. The underwater target detection method based on the state space model according to claim 2 is characterized in that: The SS2D calculation method is: X1 = CrossScan(X0); X2=Mamba(X1); Out = CrossMerge(X2); Among them, CrossScan represents the cross-scanning operation, which scans the input two-dimensional feature map into a one-dimensional sequence; Mamba represents the basic block of the state space model; CrossMerge represents the cross-merging of the sequence processed by the state space model into the shape of the original feature map; Out represents the final output.
5. The underwater target detection method based on the state space model according to claim 1 or 2, characterized in that: The calculation method for extracting the backbone network features of the state space model is: F in1 ,F in2 ,F in3 =Backbone(F); Among them, F represents the underwater low-quality image, Backbone represents the backbone network based on the state space model, and F in1 represents the low-resolution feature map, F in2 represents the medium resolution feature map, F in3 Represents a high-resolution feature map.
6. The underwater target detection method based on the state space model according to claim 1, characterized in that: The calculation method of the bottom-up feature aggregation of the PAFPN network is: T0=Upsample(F in1 ); T1=Concat(T0,F in2 ); T2=SSM(Concat(Upsample(SSM(t1)),F in3 )); Among them, Upsample means upsampling the feature map; SSM means extracting features from the feature map through the state space module, and F in1 、F in2 and F in3 They represent feature maps of different scales output by the backbone network respectively; The calculation method of the top-down feature analysis and feature output of the PAFPN network is: F out1 =T2; F out2 =SSM(Concat(Conv(F out1 ),T1)); F out3 =SSM(Concat(Conv(F out2 ),T0)); Among them, Conv represents the downsampling size adjustment of the feature map, T0, T1 and T2 all represent the multi-scale feature map output of the bottom-up feature aggregation of the PAFPN network, and F out1 、F out2 and F out3 Represents feature maps of different sizes that are finally output by the PAFPN network.
7. The underwater target detection method based on the state space model according to claim 1, characterized in that: The target classification information is the probability value of the predicted object category in the prediction box, and the calculation method is: Among them, s i represents the raw score output by the detection head network; C represents the number of categories; P(class i ) represents the probability of category i.
8. The underwater target detection method based on the state space model according to claim 1, characterized in that: The target positioning frame coordinate information in S05 is expressed as: (X, Y, W, H, confidence); Among them, X and Y represent the coordinate values of the center of the prediction box relative to the upper left corner of the feature map, W and H represent the ratio of the width and height of the prediction box to the width and height of the feature map respectively; confidence is the confidence of whether there is an object in the prediction box.
9. The underwater target detection method based on the state space model according to claim 1, characterized in that: The calculation method of the loss function is: in, represents the predicted value, Y i True value; SmoothL1 loss calculation method is: Among them, x represents the difference between the prediction and the true value; ABS(x) represents the absolute value of x.
10. The underwater target detection method based on the state space model according to claim 1, characterized in that: The certain ratio is 8:1:1.