Knowledge distillation-based foresight sonar target detection method, apparatus and device, and storage medium

By constructing a lightweight student model based on knowledge distillation and applying an adaptive loss function, the problem of low accuracy of underwater target detection foreview sonar is solved, and high-precision underwater target detection is achieved.

CN120298657APending Publication Date: 2025-07-11SOUTHERN MARINE SCIENCE & ENGINEERING GUANGDONG LABORATORY (ZHANJIANG)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510262523.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the accuracy of forward-looking sonar in underwater target detection is low, which limits its widespread promotion and use in civilian underwater detection.

Method used

Using a knowledge distillation method, the forward-view sonar image is extracted through the teacher model, a lightweight student model is constructed, and the knowledge distillation framework is used for distillation learning, replacing the attention mechanism in the student model as an adaptive loss function, and combining the adaptive threshold focus loss function to improve the target detection accuracy.

Benefits of technology

On the basis of maintaining lightweight, the accuracy of underwater target detection is significantly improved, the detection difficulties in low signal-to-noise ratio environments are overcome, and the detection performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298657A_ABST
    Figure CN120298657A_ABST
Patent Text Reader

Abstract

The invention discloses a forward-looking sonar target detection method, device and equipment based on knowledge distillation and a storage medium, and the method comprises the steps: obtaining a forward-looking sonar image, carrying out the feature extraction of the forward-looking sonar image through a teacher model, and obtaining the Logs information; constructing a lightweight student model, performing distillation learning on the lightweight student model by using a knowledge distillation framework based on Logs information to obtain a target student model, and replacing an attention mechanism in the lightweight student model with an adaptive loss function; compared with the prior art, the method has the advantages that the effective relation between the teacher model and the student model is established through the knowledge distillation framework, and the student model can realize high-precision underwater target detection on the basis of keeping light weight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning and computer vision, and particularly to a forward-looking sonar target detection method, device, equipment and storage medium based on knowledge distillation. Background Art

[0002] Due to the influence of scattering effects and complex underwater scenes on optical imaging in the underwater environment, the effect of traditional optical imaging technology in underwater target detection is limited. Therefore, most underwater detection systems rely on acoustic instruments. Especially in autonomous cruising robots, side-scan sonar is often used for underwater target imaging to obtain high-resolution images. However, compared with forward-looking sonar, side-scan sonar has a higher cost, which restricts its popularization in civilian underwater detection and limits its wide promotion and use in practical applications. However, the imaging effect of forward-looking sonar is poor and the accuracy of underwater target detection is low.

[0003] Therefore, there is an urgent need for a forward-looking sonar target detection method based on knowledge distillation, which can effectively improve the accuracy of underwater target detection in forward-looking sonar images. Summary of the Invention

[0004] The main purpose of the present invention is to provide a forward-looking sonar target detection method, device, equipment and storage medium based on knowledge distillation, aiming to solve the technical problem of low accuracy of underwater target detection in forward-looking sonar images in the prior art.

[0005] To achieve the above purpose, the present invention provides a forward-looking sonar target detection method based on knowledge distillation, and the method includes the following steps:

[0006] Obtain a forward-looking sonar image, extract features from the forward-looking sonar image through a teacher model, and obtain Logits information;

[0007] Construct a lightweight student model, and perform distillation learning on the lightweight student model based on the Logits information using a knowledge distillation framework to obtain a target student model. In the lightweight student model, the attention mechanism is replaced with an adaptive loss function;

[0008] Use the target student model to perform target detection on the forward-looking sonar image to obtain a target detection result.

[0009] Optionally, the step of constructing a lightweight student model and performing distillation learning on the lightweight student model based on the Logits information using a knowledge distillation framework to obtain a target student model includes:

[0010] Replace the attention mechanism in the student model with an adaptive loss function to obtain a lightweight student model;

[0011] Construct a knowledge distillation loss function based on the output information of the teacher model and the lightweight student model. The knowledge distillation loss function includes a classification loss function, a bounding box regression loss function, and an object score loss function;

[0012] Use the knowledge distillation loss function in the knowledge distillation framework based on the Logits information to perform distillation learning on the lightweight student model to obtain a target student model.

[0013] Optionally, the classification loss function determines the soft label difference between the teacher model and the lightweight student model through KL divergence;

[0014] The bounding box regression loss function determines the difference in the predicted bounding box coordinates between the teacher model and the lightweight student model through mean squared error;

[0015] The object score loss function aligns the predicted confidence levels of the teacher model and the lightweight student model through binary cross-entropy.

[0016] Optionally, the step of using the knowledge distillation loss function in the knowledge distillation framework based on the Logits information to perform distillation learning on the lightweight student model to obtain a target student model includes:

[0017] Use the knowledge distillation loss function in the knowledge distillation framework based on the Logits information to perform knowledge transfer between the teacher model and the lightweight student model;

[0018] After realizing the knowledge transfer between the teacher model and the lightweight student model, in the knowledge distillation framework, use the adaptive loss function in the lightweight student model to train the lightweight student model to obtain a training result;

[0019] Optimize the parameters of the lightweight student model based on the training result to obtain a target student model.

[0020] Optionally, the adaptive loss function is a forward-looking sonar adaptive threshold focal loss function, and the forward-looking sonar adaptive threshold focal loss function is:

[0021]

[0022] In the formula, the loss value of the forward-looking sonar adaptive threshold focal loss function, p t represents the current average prediction probability value, represents the average prediction probability value of the next epoch, and λ is a hyperparameter.

[0023] Optionally, the step of obtaining the forward-looking sonar image and extracting features from the forward-looking sonar image through a teacher model to obtain Logits information includes:

[0024] Collect the original sonar data and convert the original sonar data into a single-frame forward-looking sonar image through the interpolation drawing method;

[0025] Extract features from the forward-looking sonar image through the teacher model to obtain Logits information.

[0026] Optionally, the teacher model includes a feature extraction backbone network, an attention mechanism module, a multi-scale feature pyramid network, and a decoupling head;

[0027] The lightweight student model is connected to the teacher model through a knowledge distillation framework. The lightweight student model includes a feature extraction backbone network, an adaptive loss function, a multi-scale feature pyramid network, and a decoupling head.

[0028] In addition, to achieve the above object, the present invention also provides a forward-looking sonar target detection device based on knowledge distillation, and the device includes:

[0029] An information extraction module, configured to obtain a forward-looking sonar image, and extract features from the forward-looking sonar image through a teacher model to obtain Logits information;

[0030] A model training module, configured to construct a lightweight student model, and perform distillation learning on the lightweight student model based on the Logits information by using a knowledge distillation framework to obtain a target student model. In the lightweight student model, the attention mechanism is replaced with an adaptive loss function;

[0031] A target detection module, configured to perform target detection on the forward-looking sonar image by using the target student model to obtain a target detection result.

[0032] In addition, to achieve the above object, the present invention also provides a forward-looking sonar target detection device based on knowledge distillation. The device includes: a memory, a processor, and a forward-looking sonar target detection program based on knowledge distillation stored on the memory and executable on the processor. The forward-looking sonar target detection program based on knowledge distillation is configured to implement the steps of the forward-looking sonar target detection method based on knowledge distillation as described above.

[0033] In addition, to achieve the above object, the present invention also provides a storage medium. A forward-looking sonar target detection program based on knowledge distillation is stored on the storage medium. When the forward-looking sonar target detection program based on knowledge distillation is executed by a processor, the steps of the forward-looking sonar target detection method based on knowledge distillation as described above are implemented.

[0034] The present invention discloses a method for obtaining forward-looking sonar images. Feature extraction is performed on the forward-looking sonar images through a teacher model to obtain Logits information. A lightweight student model is constructed, and based on the Logits information, distillation learning is performed on the lightweight student model using a knowledge distillation framework to obtain a target student model. In the lightweight student model, the attention mechanism is replaced with an adaptive loss function. The target student model is used to perform target detection on the forward-looking sonar images to obtain target detection results. Since the present invention performs distillation learning on the lightweight student model using the knowledge distillation framework based on the Logits information output by the teacher model according to the forward-looking sonar images to obtain a target student model, and then uses the target student model to perform target detection on the forward-looking sonar images, compared with the prior art, the present invention establishes an effective connection between the teacher model and the student model through the knowledge distillation framework, and enables the student model to achieve high-precision underwater target detection while maintaining lightweight. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a schematic flowchart of the first embodiment of the forward-looking sonar target detection method based on knowledge distillation of the present invention;

[0036] Figure 2 It is an overall framework diagram of the forward-looking sonar target detection method based on knowledge distillation of the present invention;

[0037] Figure 3 It is a schematic structural diagram of the teacher model of the present invention;

[0038] Figure 4 It is a schematic structural diagram of the feature extraction backbone network in the teacher model of the present invention;

[0039] Figure 5 It is a schematic structural diagram of the lightweight student model of the present invention;

[0040] Figure 6 It is a schematic structural diagram of the multi-scale feature pyramid network in the teacher model of the present invention;

[0041] Figure 7 It is a schematic flowchart of the second embodiment of the forward-looking sonar target detection method based on knowledge distillation of the present invention;

[0042] Figure 8 It is a structural block diagram of the first embodiment of the forward-looking sonar target detection device based on knowledge distillation of the present invention;

[0043] Figure 9 It is a schematic structural diagram of a knowledge distillation-based forward-looking sonar target detection device, which is the hardware operating environment involved in the embodiment solution of the present invention.

[0044] The realization, functional features and advantages of the objectives of the present invention will be further described in conjunction with the embodiments and with reference to the accompanying drawings. Detailed implementation manners

[0045] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0046] An embodiment of the present invention provides a forward-looking sonar target detection method based on knowledge distillation. Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the forward-looking sonar target detection method based on knowledge distillation of the present invention.

[0047] In this embodiment, the forward-looking sonar target detection method based on knowledge distillation includes steps S10 to S30:

[0048] Step S10: Obtain a forward-looking sonar image, and perform feature extraction on the forward-looking sonar image through a teacher model to obtain Logits information.

[0049] It should be noted that the execution subject of this embodiment can be a computer server device with data processing, network communication and program running functions applied to the underwater target detection scenario, such as a server, a tablet computer, a personal computer, etc., or an electronic device, a forward-looking sonar target detection device based on knowledge distillation that can implement the above functions. Hereinafter, the forward-looking sonar target detection device based on knowledge distillation will be taken as an example to illustrate this embodiment and the following embodiments.

[0050] It should be explained that the above-mentioned forward-looking sonar target detection device based on knowledge distillation can at least include a forward-looking sonar device, an upper computer, a network switch, a memory and a processor, etc., and can also include other devices. This embodiment does not limit this, and in addition, the above-mentioned forward-looking sonar target detection device based on knowledge distillation is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present application.

[0051] It should be understood that the original sonar data captured by the forward-looking sonar underwater can be collected first, and then the original sonar data can be converted into a single-frame forward-looking sonar image through the interpolation drawing method, and then the teacher model can perform feature extraction on the forward-looking sonar image to obtain Logits information.

[0052] It should be noted that the interpolation drawing method is a process of estimating or calculating new data points based on known data points through specific algorithms or methods and drawing these points into a graph.

[0053] It should be noted that the teacher model in this embodiment includes a feature extraction backbone network, an attention mechanism module, a multi-scale feature pyramid network, and a decoupling head. The teacher model is a complex and high-precision model used to guide the lightweight student model in the knowledge distillation framework. In specific implementation, the teacher model learns the target features (such as tiny targets, noise suppression, etc.) in the forward-looking sonar image through training, outputs high-precision detection results (Logits information) as the learning target of the lightweight student model, and uses the attention mechanism and the multi-scale feature pyramid network to extract key information from the forward-looking sonar image with low signal-to-noise ratio, improving the detection accuracy of underwater targets.

[0054] It should be understood that Logits information in deep learning refers to the original prediction output value of the model before the final activation function (such as Softmax), which is the unnormalized score in the classification task. In the knowledge distillation framework, Logits information is the core carrier for the teacher model to transfer knowledge to the lightweight student model.

[0055] Step S20: Construct a lightweight student model, and use the knowledge distillation framework to perform distillation learning on the lightweight student model based on the Logits information to obtain a target student model. In the lightweight student model, the attention mechanism is replaced with an adaptive loss function.

[0056] It should be noted that in this embodiment, the teacher model uses a relatively large spatial complexity coefficient, saves the output of the teacher model as Logits information, and uses the original attention mechanism module in the teacher model to obtain information from the sonar image features because it does not need to be deployed in the teacher model. In the student model, for the lightweight deployment problem, this embodiment proposes a lightweight student model. The lightweight student model is based on the student model and replaces the attention mechanism in the student model with an adaptive loss function, thus using the adaptive loss function to replace the attention mechanism for extracting small targets. Therefore, for the noise in the underwater environment of the forward-looking sonar image in this embodiment, the lightweight student model used is a lightweight model specifically for the forward-looking sonar.

[0057] Reference Figure 2 , Figure 2 is the overall framework diagram of the forward-looking sonar target detection method based on knowledge distillation of the present invention. Figure 2It consists of three main parts, namely the teacher model, a lightweight student model that can be deployed on a portable robot, and a knowledge distillation training process. Input the forward-looking sonar image frame: This is the starting point of the process, indicating that the input data is a single-frame forward-looking sonar image, which is usually used in scenarios such as underwater or obstacle detection. Teacher model: This is a pre-trained model used to generate Logits information (i.e., the raw scores output by the model, not processed by the activation function). The teacher model plays a role in guiding the training of the lightweight student model here. Save as Logits results: The output of the teacher model is saved as Logits results, which will be used to train the lightweight student model. Teacher model storage pool: This is a place to store the Logits results (i.e., Logits information) generated by the teacher model, which may be used for subsequent training of the lightweight student model. Local model training backpropagation (connecting the teacher model and the lightweight student model): This means that the lightweight student model learns the Logits information of the teacher model through the backpropagation algorithm to optimize its own parameters. Lightweight student model: This is a model under training that improves its performance by learning the Logits information of the teacher model. Output the results of the lightweight student model: After the lightweight student model is trained, it outputs its prediction results. Local model training backpropagation (connecting the lightweight student model): This means that the lightweight student model also uses the backpropagation algorithm to optimize its parameters during training.

[0058] For example, refer to Figure 3 , Figure 3 which is a schematic structural diagram of the teacher model of the present invention. The teacher model includes a feature extraction backbone network, an attention mechanism module, a multi-scale feature pyramid network, and a decoupled head. In the teacher model, the tiny targets in the forward-looking sonar image are tracked through the feature extraction backbone network and the attention mechanism module respectively, and finally the type and regression linear position of the target are output through the decoupled head. The output Logits information is stored in a new folder as the output result of the teacher model. Then, the forward-looking sonar image is put into the lightweight student model for training, and distillation learning is carried out in combination with the Logits information output by the teacher model.

[0059] Furthermore, the details of the forward-looking sonar image passing through the teacher model are as follows: First, it passes through the feature extraction backbone network, then through an original version of the attention mechanism module, then through the multi-scale feature pyramid network, and finally through three decoupled heads for output, respectively performing regression and classification tasks. Finally, the output result Logits information of the teacher model is formed to facilitate the lightweight student model to extract posterior knowledge using the knowledge distillation framework.

[0060] Even further, the feature extraction backbone network in the teacher model is as Figure 4 shown.Figure 4 The schematic diagram of the structure of the feature extraction backbone network in the teacher model of the present invention is composed of convolution, pooling and residual modules. The specific process is as follows Figure 4 As shown below: Assuming the size of the input image is [3, 640, 360], the input image is first subjected to two layers of two-dimensional convolution to extract features and obtain features of [32, 640, 360]. Then, the features are compressed to [32, 256, 256] through one two-dimensional convolution. Then, a layer of two-dimensional convolution is used to obtain low-dimensional features of [64, 128, 128]. Finally, considering that an overly complex model structure is not required, high-dimensional features are extracted through four repeated two-dimensional residual blocks, and finally the internal features of a single image of [1024, 20, 20] are formed. Furthermore, the two-dimensional residual block is as follows: Figure 3 As shown in the black box Add, it consists of a layer of two-dimensional convolution and a summation operation. That is, after the input features are compressed by the convolution layer, the compressed features and the input features are superimposed by a summation operation.

[0061] It is understandable that in order to improve the target detection performance, this embodiment first applies the attention mechanism before each image feature output by the feature extraction backbone network is input into the multi-scale feature pyramid network. This mechanism processes the image features by weighting, so that the model focuses on the key areas, thereby enhancing the expression of target information and improving the target detection accuracy.

[0062] After outputting the internal features of the single-frame image to the attention mechanism module, it is input into the multi-scale feature pyramid network. Its main structure is as follows Figure 6 As shown in the figure: The scale feature pyramid network first adjusts the feature maps of three different sizes ([1024, 40, 40], [512, 20, 20] and [256, 80, 80]) to feature maps of the same size through upsampling, and then performs feature concatenation. Then, the merged feature map undergoes three convolution operations to finally obtain a feature map of size [256, 80, 80]. Subsequently, the size of the feature map is adjusted to [512, 40, 40] through upsampling, and then undergoes three convolution layers to gradually reduce the size of the feature map, and finally outputs a high-dimensional feature map of [1024, 20, 20].

[0063] And as Figure 5As shown, the original attention mechanism is abandoned in the lightweight student model to ensure an extremely lightweight structure. The lightweight student model is connected to the teacher model through a knowledge distillation framework. The lightweight student model includes a feature extraction backbone network, an adaptive loss function, a multi-scale feature pyramid network, and a decoupling head. In this embodiment, since entity deployment does not need to be considered in the teacher model, it is different from other student model architectures that have added attention mechanisms. Other student model architectures that have added attention mechanisms usually need to balance spatial complexity and accuracy and modify the attention mechanism module. However, this embodiment mainly focuses on the deployment of the student model. Therefore, the teacher model can use the original attention mechanism to ensure output accuracy.

[0064] It should be understood that underwater forward-looking sonar images in the marine environment are usually disturbed by marine noise, resulting in a low signal-to-noise ratio (SNR), which makes it very difficult to express target information. Compared with side-scan sonar, forward-looking sonar is more susceptible to these noises. This low SNR characteristic limits the manifestation of details in the image, especially the part containing key target features. To address this problem, the current common practice is to introduce an attention mechanism into the model to help the model focus on the target area and reduce the impact of noise. However, although autonomous underwater vehicles (AUVs) integrate multiple algorithms to improve detection performance, since these algorithms usually occupy a large amount of storage space, there are certain spatial and performance limitations in practical applications.

[0065] Moreover, the main part of underwater forward-looking sonar images is the background, and the target usually occupies only a small part of the area. During the training process, it is easier to learn background features than target features. Therefore, the background is regarded as an easy sample, while the target is a difficult sample. Although the background information is familiar and occupies most of the image, it dominates the gradient update direction, resulting in the target information being submerged by the background.

[0066] To overcome the above problems, this embodiment proposes an adaptive loss function - the forward-looking sonar adaptive threshold focal loss (Forward-sonar-ATFL) function. This adaptive loss function abandons the traditional attention mechanism in the student model, avoiding additional computational and storage overhead. At the same time, by adaptively adjusting the loss function, the lightweight student model can effectively improve the performance of target detection when facing underwater forward-looking sonar images with low SNR.

[0067] Among them, the traditional cross-entropy loss function can be expressed as:

[0068] L BCE = -(ylog(p)+(1 - y)log(1 - p));

[0069] In the formula, p is the predicted probability, and y is the true label. This cross-entropy loss function can be simplified and expressed as:

[0070] L BCE = -log(p t );

[0071] where p t can be understood as the probability of easy and hard samples:

[0072]

[0073] Since the traditional cross-entropy loss function cannot solve the problem of sample imbalance, focal loss introduces a modulation factor (1 - p t ) α , and reduces the loss contribution of easy-to-classify samples by adjusting the focusing hyperparameter α. Therefore, focal loss can be expressed as:

[0074] FL(p t )=-(1 - p t ) α log(p t );

[0075] Focal loss can adjust the value of the focusing hyperparameter α to reduce the loss weight of easy samples. However, while the modulation factor reduces the loss of easy samples, it also reduces the loss value of hard samples, which is not conducive to the learning of hard samples.

[0076] To solve the above problems, a Forward-sonar-TFL (Forard-sonar-Threshold Focal Loss) function can be proposed based on focal loss. This function effectively reduces its impact by reducing the loss weight of simple samples, while increasing the loss weight assigned to hard samples.

[0077]

[0078] where η, α, and λ (>1) are hyperparameters.

[0079] For easy samples, in this embodiment, it is expected that the loss value decreases as p t increases, further reducing the loss generated by easy samples. At the beginning of training, even for simple samples, their predicted probabilities are relatively low and gradually increase as the training process progresses, and α should gradually approach 0. The predicted probability value of the true target can be used for mathematical modeling in the model training process and can be predicted by the moving average method.

[0080]

[0081] where represents the predicted value of the next epoch, and p tRepresents the current average predicted probability value, Represents the average predicted probability value for each training epoch. According to Shannon's information theory, the greater the probability value of an event, the smaller the amount of information it brings; conversely, the greater the amount of information. Therefore, the adaptive modulation factor α can be expressed as:

[0082]

[0083] However, in the later stage of model training, if the expected probability value is too large, the proportion of difficult samples will be reduced. In this embodiment, η can be expressed as:

[0084] η = -ln(p t );

[0085] To sum up, the adaptive loss function of this embodiment is the forward-looking sonar adaptive threshold focal loss function, and the forward-looking sonar adaptive threshold focal loss function is:

[0086]

[0087] In the formula, the loss value of the forward-looking sonar adaptive threshold focal loss function, p t Represents the current average predicted probability value, Represents the average predicted probability value of the next epoch, and λ is a hyperparameter.

[0088] It should be noted that this adaptive loss function is used for the separation of targets and backgrounds in the lightweight student model. The loss value is adaptively adjusted according to the predicted probability value to improve the detection performance of non-significant targets in the forward-looking sonar.

[0089] Step S30: Use the target student model to perform target detection on the forward-looking sonar image to obtain a target detection result.

[0090] This embodiment discloses obtaining a forward-looking sonar image, extracting features from the forward-looking sonar image through a teacher model to obtain Logits information; constructing a lightweight student model, and performing distillation learning on the lightweight student model based on the Logits information using a knowledge distillation framework to obtain a target student model, where the attention mechanism in the lightweight student model is replaced with an adaptive loss function; using the target student model to perform target detection on the forward-looking sonar image to obtain a target detection result. Since this embodiment performs distillation learning on the lightweight student model based on the Logits information output by the teacher model according to the forward-looking sonar image using a knowledge distillation framework to obtain a target student model, and then uses the target student model to perform target detection on the forward-looking sonar image, compared with the prior art, this embodiment establishes an effective connection between the teacher model and the student model through the knowledge distillation framework, and enables the student model to achieve high-precision underwater target detection while maintaining lightweight.

[0091] Reference Figure 7 , Figure 7 is a schematic flowchart of the second embodiment of the forward-looking sonar target detection method based on knowledge distillation of the present invention.

[0092] Based on the above first embodiment, in this embodiment, the step S20 includes steps S201 to S203:

[0093] Step S201: Replace the attention mechanism in the student model with an adaptive loss function to obtain a lightweight student model.

[0094] Step S202: Based on the output information of the teacher model and the lightweight student model, construct a knowledge distillation loss function, where the knowledge distillation loss function includes a classification loss function, a bounding box regression loss function, and an object score loss function.

[0095] Step S203: Based on the Logits information, perform distillation learning on the lightweight student model using the knowledge distillation loss function in the knowledge distillation framework to obtain a target student model.

[0096] It should be noted that the classification loss function determines the soft label difference between the teacher model and the lightweight student model through KL divergence; the bounding box regression loss function determines the difference in the predicted bounding box coordinates between the teacher model and the lightweight student model through mean square error; the object score loss function aligns the predicted confidence levels of the teacher model and the lightweight student model through binary cross entropy.

[0097] In a specific implementation, the attention mechanism in the student model can be replaced with an adaptive loss function to obtain a lightweight student model; based on the output information of the teacher model and the lightweight student model, a knowledge distillation loss function is constructed, and the knowledge distillation loss function includes a classification loss function, a bounding box regression loss function, and an object score loss function; based on the Logits information, the lightweight student model is distilled using the knowledge distillation loss function in the knowledge distillation framework to obtain a target student model.

[0098] It should be noted that the step of distilling the lightweight student model using the knowledge distillation loss function in the knowledge distillation framework based on the Logits information to obtain a target student model includes: performing knowledge transfer between the teacher model and the lightweight student model using the knowledge distillation loss function in the knowledge distillation framework based on the Logits information; after realizing the knowledge transfer between the teacher model and the lightweight student model, in the knowledge distillation framework, the lightweight student model is trained using the adaptive loss function in the lightweight student model to obtain a training result; based on the training result, the parameters of the lightweight student model are optimized to obtain a target student model.

[0099] It should be noted that the classification loss function is:

[0100]

[0101] In the formula, k represents the number of outputs of each sample category in the total training samples, represents the class soft label of the j-th input passing through the teacher model, represents the class soft label of the j-th input passing through the student model output.

[0102] Furthermore, similar to L KD_cls Applied to the class output tasks of the teacher model and the lightweight student model, in order to effectively transfer the knowledge of the teacher model to the lightweight student model, two different knowledge distillation loss functions are designed in the bounding box regression and target prediction tasks, namely the bounding box regression loss function and the object score loss function. First, for the bounding box regression loss function, the mean squared error (MSE) loss is used to minimize the difference between the bounding box r0 predicted by the lightweight student model and the output t0 of the teacher model. The bounding box regression loss function is:

[0103]

[0104] In the formula, N represents the number of the coordinates (x1, y1, x2, y2) of the label boxes of the total training samples, represents the coordinates of the bounding box predicted by the i-th lightweight student model, Represent the bounding box coordinates predicted by the i-th teacher model.

[0105] Furthermore, in order to transfer the judgment confidence knowledge of the teacher model to the student model in the object prediction task, an object score loss function designed based on the binary cross-entropy (BCE) loss function is adopted to align the output s2 of the lightweight student model with the output t2 of the teacher model. The object score loss function is:

[0106]

[0107] In the formula, represents the probability predicted by the lightweight student model, represents the probability output by the teacher model. The BCE loss ensures that the prediction of the lightweight student model is highly consistent with the output of the teacher model, and optimizes the output of the lightweight student model by penalizing the deviation between positive and negative predictions.

[0108] For the knowledge distillation (KD) framework process, each decoupled head (Neck) of the teacher model and the lightweight learning model calculates different loss functions, including tasks such as classification, bounding box regression, and object score. In this knowledge distillation (KD) framework, the KD process involves calculating the loss functions output by each FPN (multi-scale feature pyramid network), specifically including the classification loss function, the bounding box regression loss function, and the object score loss function, and comparing the lightweight student model with the teacher model.

[0109] The above classification loss function, bounding box regression loss function, and object score loss function are output through different output channels and FPN (feature pyramid network), ensuring that the loss of each type of task is accurately calculated, and thus realizing the knowledge transfer between the lightweight student model and the teacher model.

[0110] The difference between the knowledge distillation framework in this embodiment and the rest of the knowledge distillation architectures is that: in the general knowledge distillation architecture object detection framework, only the final output result is saved for a single-frame image input, while in this embodiment, through the above classification loss function, bounding box regression loss function, and object score loss function output through different output channels and FPN (multi-scale feature pyramid network), it is ensured that the loss of each type of task is accurately calculated, and thus the knowledge transfer between the lightweight student model and the teacher model is realized. As follows:

[0111] In specific implementation, the KD loss function is divided into the following parts:

[0112] Classification loss function:

[0113] L KD_cls = L KD_cls_FPN0 + L KD_cls_FPN1+L KD_cls_FPN2 / batch_size * 3。

[0114] Bounding box regression loss function:

[0115] L KD_bbox = L KD_bbox_FPN0 + L KD_bbox_FPN1 + L KD_bbox_FPN2 / batch_size * 3。

[0116] Object score loss function:

[0117] L KD_obj = L KD_obj_FPN0 + L KD_obj_FPN1 + L KD_obj_FPN2 / batch_size * 3。

[0118] Among them, L KD_cls_FPN0 , L KD_bbox_FPN0 , L KD_obj_FPN0 , L KD_cls_FPN1 , L KD_bbox_FPN1 , L KD_obj_FPN1 , L KD_cls_FPN2 , L KD_bbox_FPN2 and L KD_obj_FPN2 These are the losses at multiple stages or levels in the knowledge distillation process (FPN stands for Feature Pyramid Network, a multi-scale feature pyramid network). Different FPN layers represent the learning processes of the lightweight student model and the teacher model at different levels. They may respectively correspond to the outputs of the teacher model and the lightweight student model at different FPN layers, and the specific losses are calculated through the predictions of these layers. batch_size refers to the amount of data in each batch during training, usually the number of samples input into the model in each round of training.

[0119] Furthermore, the total student loss function is defined as:

[0120] L student = L hard + L soft 。

[0121] Among them, L hard hard label loss, usually refers to the standard training loss using the true labels (cross-entropy loss in this embodiment). It describes how the model is optimized according to the original labels. L softThis is the soft label loss, which usually refers to using the predicted probabilities of the teacher model as labels for loss calculation. Soft labels provide richer classification information than hard labels and can help the lightweight student model better learn the knowledge of the teacher model. In fact, the soft label loss is the knowledge distillation loss, that is, L soft = L KD .

[0122] This embodiment discloses replacing the attention mechanism in the student model with an adaptive loss function to obtain a lightweight student model; constructing a knowledge distillation loss function based on the output information of the teacher model and the lightweight student model, where the knowledge distillation loss function includes a classification loss function, a bounding box regression loss function, and an object score loss function; using the knowledge distillation loss function in the knowledge distillation framework based on the Logits information to perform distillation learning on the lightweight student model to obtain a target student model. Compared with the prior art, this embodiment constructs a knowledge distillation loss function based on the output information of the teacher model and the lightweight student model. The knowledge distillation loss function includes a classification loss function, a bounding box regression loss function, and an object score loss function. Using the knowledge distillation loss function in the knowledge distillation framework based on the Logits information to perform distillation learning on the lightweight student model to obtain a target student model. Compared with the prior art, this embodiment ensures that the loss of each type of task is accurately calculated, thereby realizing the knowledge transfer between the student model and the teacher model and further improving the target detection accuracy.

[0123] In addition, an embodiment of the present invention further proposes a storage medium, on which a forward-looking sonar target detection program based on knowledge distillation is stored. When the forward-looking sonar target detection program based on knowledge distillation is executed by a processor, the steps of the forward-looking sonar target detection method based on knowledge distillation as described above are implemented.

[0124] Referring to Figure 8 , Figure 8 is a structural block diagram of the first embodiment of the forward-looking sonar target detection device based on knowledge distillation of the present invention.

[0125] As Figure 8 shown, the forward-looking sonar target detection device proposed by the embodiment of the present invention includes: an information extraction module 801, a model training module 802, and a target detection module 803.

[0126] The information extraction module 801 is used to obtain a forward-looking sonar image and perform feature extraction on the forward-looking sonar image through the teacher model to obtain Logits information.

[0127] The model training module 802 is used to construct a lightweight student model, and perform distillation learning on the lightweight student model based on the Logits information by using a knowledge distillation framework to obtain a target student model. In the lightweight student model, the attention mechanism is replaced with an adaptive loss function.

[0128] The target detection module 803 is used to perform target detection on the forward-looking sonar image by using the target student model to obtain a target detection result.

[0129] This embodiment of the device discloses obtaining a forward-looking sonar image, extracting features from the forward-looking sonar image through a teacher model to obtain Logits information; constructing a lightweight student model, and performing distillation learning on the lightweight student model based on the Logits information by using a knowledge distillation framework to obtain a target student model. In the lightweight student model, the attention mechanism is replaced with an adaptive loss function; performing target detection on the forward-looking sonar image by using the target student model to obtain a target detection result. Since this embodiment of the device performs distillation learning on the lightweight student model based on the Logits information output by the teacher model according to the forward-looking sonar image by using a knowledge distillation framework to obtain a target student model, and then uses the target student model to perform target detection on the forward-looking sonar image. Compared with the prior art, this embodiment of the device establishes an effective connection between the teacher model and the student model through the knowledge distillation framework, and enables the student model to achieve high-precision underwater target detection while maintaining lightweight.

[0130] Based on the first embodiment of the forward-looking sonar target detection device based on knowledge distillation of the present invention, a second embodiment of the forward-looking sonar target detection device based on knowledge distillation of the present invention is proposed.

[0131] In this embodiment, the model training module 802 is further used to replace the attention mechanism in the student model with an adaptive loss function to obtain a lightweight student model; construct a knowledge distillation loss function based on the output information of the teacher model and the lightweight student model. The knowledge distillation loss function includes a classification loss function, a bounding box regression loss function, and an object score loss function; perform distillation learning on the lightweight student model based on the Logits information by using the knowledge distillation loss function in the knowledge distillation framework to obtain a target student model.

[0132] Other embodiments or specific implementation manners of the forward-looking sonar target detection device based on knowledge distillation of the present invention may refer to the above method embodiments, and will not be elaborated here.

[0133] The present application provides a forward-looking sonar target detection device based on knowledge distillation. The forward-looking sonar target detection device based on knowledge distillation includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the forward-looking sonar target detection method based on knowledge distillation in Embodiment 1 above.

[0134] In addition, the forward-looking sonar target detection device based on knowledge distillation in the present application may further include: at least one forward-looking sonar device, at least one host computer, and at least one network switch.

[0135] Next, refer to Figure 9 , which shows a schematic structural diagram of a forward-looking sonar target detection device suitable for implementing the embodiments of the present application. The forward-looking sonar target detection device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The forward-looking sonar target detection device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0136] As Figure 9As shown, the forward-looking sonar target detection device based on knowledge distillation may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1002 or the program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the forward-looking sonar target detection device based on knowledge distillation are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the forward-looking sonar target detection device based on knowledge distillation to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a forward-looking sonar target detection device with various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems can be implemented or had.

[0137] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiments disclosed in the present application are executed.

[0138] The forward-looking sonar target detection device provided by the present application adopts the forward-looking sonar target detection method in the above embodiments, and can solve the technical problem of low accuracy in detecting underwater targets in forward-looking sonar images in the prior art. Compared with the prior art, the beneficial effects of the forward-looking sonar target detection device provided by the present application are the same as those of the forward-looking sonar target detection method provided by the above embodiments, and other technical features in the forward-looking sonar target detection device are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.

[0139] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0140] As mentioned above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0141] It should be noted that in this text, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or system including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or system including that element.

[0142] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0143] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a read-only memory / random access memory, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0144] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A forward-looking sonar target detection method based on knowledge distillation, characterized in that, The method includes: Obtain a forward-looking sonar image, extract features from the forward-looking sonar image through a teacher model, and obtain Logits information; Construct a lightweight student model, and use the knowledge distillation framework to perform distillation learning on the lightweight student model based on the Logits information to obtain a target student model. In the lightweight student model, the attention mechanism is replaced with an adaptive loss function; Use the target student model to perform object detection on the forward-looking sonar image to obtain an object detection result.

2. The method for forward-looking sonar target detection based on knowledge distillation according to claim 1, wherein, The steps of constructing a lightweight student model and using the knowledge distillation framework to perform distillation learning on the lightweight student model based on the Logits information to obtain a target student model include: Replace the attention mechanism in the student model with an adaptive loss function to obtain a lightweight student model; Construct a knowledge distillation loss function based on the output information of the teacher model and the lightweight student model. The knowledge distillation loss function includes a classification loss function, a bounding box regression loss function, and an object score loss function; Use the knowledge distillation loss function in the knowledge distillation framework to perform distillation learning on the lightweight student model based on the Logits information to obtain a target student model.

3. The forward-looking sonar target detection method based on knowledge distillation according to claim 2, wherein, The classification loss function determines the soft label difference between the teacher model and the lightweight student model through KL divergence; The bounding box regression loss function determines the difference in the predicted bounding box coordinates between the teacher model and the lightweight student model through mean square error; The object score loss function makes the teacher model and the lightweight student model align the predicted confidence through binary cross-entropy.

4. The forward-looking sonar target detection method based on knowledge distillation according to claim 2, wherein The steps of using the knowledge distillation loss function in the knowledge distillation framework to perform distillation learning on the lightweight student model based on the Logits information to obtain a target student model include: Use the knowledge distillation loss function in the knowledge distillation framework based on the Logits information to perform knowledge transfer between the teacher model and the lightweight student model; After realizing the knowledge transfer between the teacher model and the lightweight student model, in the knowledge distillation framework, use the adaptive loss function in the lightweight student model to train the lightweight student model to obtain a training result; Optimize the parameters of the lightweight student model based on the training result to obtain a target student model.

5. The forward-looking sonar target detection method based on knowledge distillation according to claim 1, characterized in that, The adaptive loss function is a forward-looking sonar adaptive threshold focal loss function, and the forward-looking sonar adaptive threshold focal loss function is: In the formula, the loss value of the forward-looking sonar adaptive threshold focal loss function, p t represents the current average prediction probability value, represents the average prediction probability value of the next epoch, and λ is a hyperparameter.

6. The forward-looking sonar target detection method based on knowledge distillation according to claim 1, characterized in that The steps of obtaining a forward-looking sonar image, extracting features from the forward-looking sonar image through a teacher model, and obtaining Logits information include: Collect original sonar data, and convert the original sonar data into a single-frame forward-looking sonar image through interpolation drawing; Extract features from the forward-looking sonar image through a teacher model to obtain Logits information.

7. The forward-looking sonar target detection method based on knowledge distillation according to any one of claims 1-6, characterized in that The teacher model includes a feature extraction backbone network, an attention mechanism module, a multi-scale feature pyramid network, and a decoupled head; The lightweight student model is connected to the teacher model through a knowledge distillation framework. The lightweight student model includes a feature extraction backbone network, an adaptive loss function, a multi-scale feature pyramid network, and a decoupled head.

8. A forward-looking sonar target detection device based on knowledge distillation, characterized in that, The device includes: An information extraction module, configured to obtain a forward-looking sonar image, and perform feature extraction on the forward-looking sonar image through the teacher model to obtain Logits information; A model training module, configured to construct a lightweight student model, and perform distillation learning on the lightweight student model based on the Logits information by using the knowledge distillation framework to obtain a target student model, where the attention mechanism in the lightweight student model is replaced by an adaptive loss function; A target detection module, configured to perform target detection on the forward-looking sonar image by using the target student model to obtain a target detection result.

9. A forward-looking sonar target detection device based on knowledge distillation, characterized in that, The device includes: a memory, a processor, and a forward-looking sonar target detection program based on knowledge distillation stored on the memory and executable on the processor. The forward-looking sonar target detection program based on knowledge distillation is configured to implement the steps of the forward-looking sonar target detection method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, A forward-looking sonar target detection program based on knowledge distillation is stored on the storage medium. When the forward-looking sonar target detection program based on knowledge distillation is executed by a processor, the steps of the forward-looking sonar target detection method according to any one of claims 1 to 7 are implemented.