An Underwater Target Recognition Method Based on Quantization Distillation
Through quantitative distillation technology, the neural network is compressed and deployed, and the problems of large size and resource consumption in underwater target recognition are solved, achieving efficient and real-time underwater target recognition.
Patent Information
- Application Number
- CN202310426376.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-20
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-04-20
AI Technical Summary
The existing underwater target recognition methods based on neural networks have numerous neural network weights and large network size, making it difficult to deploy on hardware devices with limited resources.
Using a quantitative distillation-based method, the teacher network is compressed through a multi-teacher knowledge distillation method, the student network is obtained, and quantified to deploy on the FPGA.
It realizes the lightweight neural network, maintains high recognition accuracy, and speeds up inference, facilitates hardware equipment deployment, and improves the real-time nature of underwater target recognition.
Smart Images

Figure CN116524341B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater target recognition, and in particular to an underwater target recognition method based on quantization distillation. Background Art
[0002] With the development of the economy, the exploitation of marine resources has attracted wide attention. As the main tool for marine exploration and underwater operations, underwater robots can work in dangerous underwater environments and can adapt to changing underwater environments. Currently, underwater robots are mainly applied in the fields of marine data monitoring, maintenance of underwater equipment, acquisition and recognition of marine targets, etc., which are closely related to the recognition and positioning of targets by underwater robots. Underwater target recognition technology, as an auxiliary means for various underwater tasks, can help underwater operators and underwater vehicles to better carry out underwater tasks. With the increasing development of deep learning, deep learning has also been more and more applied in the fields of machine vision, face recognition, video analysis, etc. Deep neural networks have high recognition and clear semantic segmentation capabilities in details, and can better improve the poor semantic segmentation of traditional recognition methods and improve the image recognition effect. Compared with traditional methods, deep neural networks can more accurately identify and locate underwater targets through accurate extraction of features by deep networks, with higher recognition rates. However, since electromagnetic waves attenuate very quickly underwater and have relatively weak ability to transmit data through seawater, the deep neural network must be locally deployed in the device.
[0003] Currently, there are many underwater target recognition methods based on neural networks, such as "Liu Yuhao. Research on underwater small target detection and recognition technology based on Faster R-CNN [D]. Dalian University of Technology, 2021. DOI: 10.26991 / d.cnki.gdllu.2021.003381.", "Liu Dali, Shen Wenhao, Feng Hongwei, Tian Yu, Su Yuchen, Tian Wenfeng. An underwater target recognition method with a small number of samples based on deep learning [P]. Tianjin: CN115508838A, 2022-12-23.", "Liu Zhangqiang. Research on target detection and recognition method based on underwater image enhancement [D]. Shandong University, 2021. DOI: 10.27272 / d.cnki.gshdu.2021.002570.". However, these methods have problems such as a large number of neural network weights, a huge network volume, and difficulty in deploying them on hardware devices with limited resources. Therefore, the present invention proposes an underwater target recognition technology based on quantization distillation, aiming to make the accuracy of the neural network not decrease through an effective network compression method, and at the same time make the compressed classification network have a lightweight volume, speed up its inference speed and facilitate deployment on hardware devices. Summary of the Invention
[0004] The present invention aims to address the drawbacks and deficiencies of the existing technology and provide an underwater target recognition method based on quantization distillation. This method can collect underwater images and use a lightweight network trained by quantization distillation to determine and label the types of underwater targets.
[0005] The technical solution of the present invention is as follows: An underwater target recognition method based on quantization distillation, comprising the following steps:
[0006] Step 1: Obtain underwater biological image data samples and construct a sample data set; the sample data set is randomly formed into a training data set and a test data set in a certain proportion;
[0007] Step 2: Implement a data augmentation method based on sample replication on the training data set to obtain an augmented training data set;
[0008] Step 3: Input the augmented training data set into the Resnet50 network and pre-train multiple Resnet50 networks as teacher networks; the teacher network structure is divided into 7 parts. The first part includes a convolutional layer, an activation layer, and a max pooling layer, which perform convolution, activation, and max pooling calculations on the input augmented training data set. The second to fifth parts have the same structure and are intermediate layers, including residual blocks, convolutional blocks, and classifiers. The classifier includes a fully connected layer and a softmax activation function, which are used to represent the classification information of the intermediate layer. The sixth part is an average pooling layer, which converts the input content into a feature vector. The seventh part is a fully connected layer, which calculates the feature vector and outputs the class probability;
[0009] Step 4: Compress the teacher network based on the multi-teacher knowledge distillation method to obtain a student network; the structure of the student network is the same as that of the teacher network; the student network uses the convolutional neural network Resnet18, and the convolutional neural network Resnet18 includes 17 convolutional layers and a fully connected layer;
[0010] Step 4.1: represents the augmented training data set, where N is the number of samples, and x i is the i-th input variable; y i is the true label value corresponding to the i-th input variable;
[0011] Calculate the prediction result σ(Z c ) of the teacher network according to the Softmax function, and soften the output probability distribution through the hyperparameter distillation temperature T:
[0012]
[0013] where Z cis the output of the c-th node in the fully connected layer of the teacher network, T is the distillation temperature parameter, and C represents the total number of outputs of the fully connected layer of the teacher network;
[0014] Step 4.2: Calculate the loss function between the teacher network and the student network;
[0015] Step 4.2.1: To effectively aggregate the prediction distributions of multiple teachers, calculate the cross-entropy loss between the prediction results of multiple teacher networks and the true label values, and assign different weights to reflect the sample confidence of different teacher networks in the student network;
[0016]
[0017]
[0018] Among them, k represents the k-th teacher network; K represents the total number of teacher networks; c represents the output of the c-th node in the fully connected layer of a single teacher network; C represents the total number of outputs of the fully connected layer of a single teacher network; represents the output of the fully connected layer of the k-th teacher network; y c represents the true label value of the teacher network; represents the cross-entropy loss between the prediction of the teacher network in the fully connected layer and the basic true label; represents the sample confidence of the output probability values of the fully connected layers of each teacher network;
[0019] Multiply the predictions of multiple teacher networks by their corresponding sample confidences, and the loss function of the output features of the fully connected layer of the teacher network is:
[0020]
[0021] Among them, represents the output of the fully connected layer of the student network; represents the prediction result of the student network;
[0022] Step 4.2.2: Extend the distillation method to the intermediate layer of the teacher network, and the calculation of its feature matching is as follows:
[0023]
[0024]
[0025]
[0026] Among them, is the classifier of the k-th teacher network; h S is the feature vector of the last layer of the student network; It represents the feature classification result of the k-th teacher network on the feature vector of the last layer of the student network; is the cross-entropy loss between the prediction of the teacher network in the intermediate layer and the ground truth label; K represents the number of teacher networks; represents the sample confidence of the output probability value of the intermediate layer of the teacher network;
[0027] To stabilize the knowledge transfer process, the student network focuses more on imitating the teacher network with a similar feature space, and the loss function of the intermediate layer features of the student network and the teacher network is:
[0028]
[0029] where r(·) is a feature dimension alignment function used to increase the dimension of the student network to match the dimension of the teacher network; F S is the feature output of the feature map of the student network; is the feature output of the feature map of the k-th teacher network; represents the norm, which is a distance metric for the intermediate features of the teacher network and the student network;
[0030] Step 4.2.3: Calculate the regularized cross-entropy loss function with the ground truth label as:
[0031]
[0032] where y c represents the ground truth label value of the teacher network, is the output of the fully connected layer of the student network;
[0033] Step 4.2.4: Finally, the overall loss function of multi-teacher network knowledge distillation is:
[0034]
[0035] where α and β are hyperparameters that balance the loss of the output features of the fully connected layer of the teacher network and the loss of the intermediate layer features of the student network and the teacher network;
[0036] Step 4.3: Initialize the student network model, use multi-teacher knowledge distillation for the student network; calculate the output of the teacher network and soften the probability distribution as shown in Step 4.1, add the temperature coefficient T to soften the output probability distribution to provide more supervision information; the loss functions of the teacher network and the student network are as shown in Step 4.2, and use the overall loss function of multi-teacher distillation for the student network to learn the knowledge of the teacher network; finally, adjust the hyperparameters distillation temperature T and the overall distillation loss function The student network is further converged to obtain the distilled and compressed student network;
[0037] Step 5: Quantize the distilled and compressed student network and deploy it on the FPGA:
[0038] Step 5.1: Let the maximum and minimum values of the parameters to be quantized be set as x max , x min , and let the maximum and minimum values of the quantized parameters be set as q max , q min , and adopt the asymmetric quantization method with the quantization interval set as [0, 255];
[0039] Step 5.2: Calculate the quantization scale factor B and zero point O, and the formulas are respectively:
[0040]
[0041]
[0042] where round(·) represents the rounded value of the floating-point number within the parentheses;
[0043] Step 5.3: Quantize the floating-point parameters, and the quantization formula is:
[0044] r = B(q - O)
[0045]
[0046] where r represents the floating-point real number, q represents the quantized fixed-point integer; the meaning of round(·) is the same as above;
[0047] Step 5.4: During the process of calculating the student network, the operations of the convolutional layer and fully connected layer of the student network are reduced to matrix multiplication operations; the steps to convert a floating-point matrix into a fixed-point matrix are as follows;
[0048] Step 5.4.1: The sets m 1 and m 2 are two floating-point weight matrices in the intermediate layer of the student network, both with a size of N×N; the matrix m 3 is the result of multiplying m 1 m 2 ; the elements 3 in m are calculated by the formula:
[0049]
[0050] where i, r, k are respectively the row number and column number of the corresponding matrix;
[0051] Step 5.4.2: Obtained by organizing the formulas in Steps 5.3 and 5.4.1:
[0052]
[0053] where B 1 and O 1 are the scale factor and zero point corresponding to the m 1 matrix; B 2 and O 2 are the scale factor and zero point corresponding to the m 2 matrix; B 3 and O 3 are the scale factor and zero point corresponding to the m 3 matrix;
[0054] Step 5.4.3: Quantize each element 3 in matrix q as:
[0055]
[0056] Step 5.4.4: Convert the in Step 5.4.3 into fixed-point operations as:
[0057]
[0058] where M 0 = 2 n M, M 0 is a real number, and the value of n, M 0 is selected as the optimal n, M 0 according to the initial values of M and P and the optimization algorithm, -n such that MP = 2 0 M 3
[0059] 3 in matrix q after optimized calculation as:
[0060]
[0061] Step 5.5: In Step 4, the distilled and compressed student network is represented as: H = (x, y); x is the input matrix of the student network, and q is the output matrix of the student network; calculate the matrix multiplication in the convolutional layer and fully connected layer as:
[0062]
[0063]
[0064] where, a1 is the feature map matrix of the convolutional layer of the student network, w 1 is the weight parameter matrix of the convolutional layer of the student network, and x is the input matrix of the convolutional layer of the student network; a 2 is the input matrix of the fully connected layer of the student network, w 2 is the weight parameter matrix of the fully connected layer of the student network, and q is the feature map matrix of the fully connected layer of the student network; b 1 and b 2 are the bias parameter matrices;
[0065] Step 5.6: According to the formulas in Step 5.4 and Step 5.5, the quantization calculations of the convolutional layer and the fully connected layer in the student network are respectively:
[0066]
[0067]
[0068] where a 1 , x, and w 1 have the same meanings as above; q a1 and q x and are the quantization matrices of a 1 , x, and w 1 ; O x and are the zero-point matrices of the matrices x, w 1 , and a 1 ;
[0069] q, a 2 , and w 2 have the same meanings as above; q q and q a2 and are the quantization matrices of q, a 2 , and w 2 ; O q is the zero-point matrix of the matrices a 2 , w 2 , and q;
[0070] b 1 and b 2 are the bias parameter matrices, and are the quantization matrix values of b 1 and b 2 ; f represents the output of the student network; M a1 , M a2 and the method of obtaining M 0 in Step 5.4.4 are the same;
[0071] Step 5.7: In the student network activation layer, the activation function is:
[0072]
[0073] where a 2 is the feature map matrix of the student network activation layer, a 1 is the input matrix of the student network activation layer, q a1 and q a2 are the quantization matrices of a 1 and a 2 ; is the zero-point matrix of matrix a 1 ; The max(·) function represents comparing the values of the input quantization matrix and the zero-point matrix of the activation layer and returns the larger value at the corresponding positions of the two matrices;
[0074] Step 5.8: Through the calculations in Step 5.6 and Step 5.7, a full-precision student network is obtained. The entire student network is quantized, and finally, the student network is deployed on the FPGA to achieve hardware-accelerated computing;
[0075] Step 6: Update the distilled and quantized student network to a new object detection network, and input images in real time into the updated object detection network to achieve real-time underwater object recognition.
[0076] The specific content of Step 2 is as follows:
[0077] Step 2.1: Select pictures from the training dataset where the volume of the seabed creature category samples accounts for less than 5% of the entire picture, with the number being 1 - 2 and the edges being clear. These pictures in the target set are used for sample replication;
[0078] Step 2.2: Draw the true bounding boxes and class labels for supervised training on the pictures in the target set to determine the specific positions and categories of the seabed creatures in the pictures of the target set;
[0079] Step 2.3: Perform contour annotation on the pictures in the target set according to the determined specific positions and categories of the seabed creatures to obtain the physical contour annotation text;
[0080] Step 2.4: Read in the pictures obtained in Step 2.3 in sequence, traverse the annotation names obtained from the physical contour annotation text. When there is an annotation name that matches the name of the currently read picture, extract the contour annotation of this picture from the physical contour annotation text; Copy all the pixel values within the contour of the current picture to a random area with the same shape to complete sample replication;
[0081] Step 2.5: Use OpenCV to calculate the area U of the minimum circumscribed quadrilateral of the physical contour annotation in the target set of pictures obtained in Step 2.3, and calculate the intersection over union of U and the areas of all the true bounding boxes for supervised training in the currently completed sample replication picture. The bounding box with the largest intersection over union with U is the bounding box of the replicated target.
[0082] Step 2.6: Assign the bounding box filtered out in Step 2.5 to the samples randomly replicated in Step 2.4, and at the same time add class labels for supervised training. After the training is completed, an augmented training data set will be obtained.
[0083] The underwater biological image data samples are used to capture target images in different angles and directions from different backgrounds. There is at least one shape similar to the object to be detected in the target image, or an image of the side of the object to be detected. The size of the captured image is close to the size of the usage scenario; the sample data set includes more than 2000 pictures; the sample data set includes the target images to be detected and they are labeled. Randomly select 70% of the image data in the sample data set to form the training data set, and 30% to form the test data set.
[0084] Furthermore, the Resnet50 network contains 49 convolutional layers and one fully connected layer. In the Resnet50 network structure, each residual block has three convolutional layers. The input of the neural network is 224×224×3. After the convolutional calculations in the first 5 parts, the output is 7×7×2048. The average pooling layer will convert it into a feature vector, and finally the fully connected layer will calculate this feature vector and output the class probability.
[0085] Advantages of the present invention: The present invention proposes an underwater target recognition method based on quantization distillation, enabling the neural network to achieve good recognition results in the field of underwater target recognition. Based on the proposed multi-teacher knowledge distillation structure, by constructing a lightweight small model student network and using the supervision information of multiple better-performing large model teacher networks to train this small model, better performance and accuracy can be achieved. The lightweight small model student network has similar accuracy and performance to the large model teacher network. The accuracy of the student network in underwater recognition is improved, but the size of the student network model is much smaller than that of the large model teacher network, which is convenient for our hardware deployment; based on the proposed quantization deployment structure, the weight parameters in the student network are quantized from floating-point numbers to integers, which is convenient for accelerating operations on the hardware FPGA. After quantization and hardware acceleration, the student network obtained after distillation has improved speed in underwater recognition and enhanced real-time performance. Therefore, the underwater target recognition method based on quantization distillation proposed by the present invention can improve the accuracy and real-time performance of the neural network in underwater recognition, greatly enhancing the underwater recognition ability. Description of the Drawings
[0086] Figure 1 This is a flowchart of an underwater target recognition method based on quantization distillation according to the present invention;
[0087] Fig. 2(a) is a schematic structural diagram of the teacher network Resnet50 selected according to the present invention;
[0088] Fig. 2(b) is a schematic structural diagram of the student network Resnet18 selected according to the present invention;
[0089] Figure 3 This is a schematic structural diagram of the multi-teacher knowledge distillation network according to the present invention. Detailed implementation manners
[0090] The present invention will be further described below in conjunction with embodiments; it should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0091] The following further elaborates on the above and other technical features and advantages of the present invention with reference to the accompanying drawings.
[0092] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0093] Knowledge distillation is a commonly used method for network compression. Different from pruning and quantization in network compression, knowledge distillation trains a lightweight small network by using the supervision information of a large network with better performance, in order to achieve better performance and accuracy. It was first proposed by Hinton in 2015 and applied to classification tasks for the first time. This large network is called the teacher network, and the small network is called the student network. The supervision information output from the teacher network is called knowledge, and the process by which the student learns and transfers the supervision information from the teacher is called distillation.
[0094] Compared with traditional underwater recognition, the underwater target recognition method based on quantization distillation can greatly improve the recognition ability of the neural network and enhance the recognition accuracy. At the same time, it overcomes the problems of slow recognition speed and difficult deployment caused by excessive data volume.
[0095] Based on the above algorithm, the present invention discloses an underwater target recognition method based on quantization distillation, as Figure 1 shown, including the following steps:
[0096] Step 1: Obtain underwater biological image data samples and construct a sample data set. We can select some existing underwater images or images captured in real time underwater by an underwater camera.
[0097] The selection requirements for the dataset are that the biological image data samples should capture target images from different angles and backgrounds. The object data detected in the images should be concentrated, and there should be at least one shape similar to the detected object or an image of the side of the detected object in the image; the size of the collected images should be close to the size of the usage scenario; the sample dataset includes more than 2000 pictures; the sample dataset should include the target images to be detected and label them. Randomly select 70% of the image data in the sample dataset to form the training dataset and 30% to form the test dataset;
[0098] Step 2: Implement a data augmentation method based on sample replication on the training dataset obtained in Step 1. Data augmentation is to make the teacher network more powerful and improve the accuracy of underwater recognition.
[0099] Step 2.1: Select pictures in the training dataset with small class sample volume, small quantity, and clear edges as the target set for sample replication.
[0100] Step 2.2: Use an open-source visualization tool to draw the real bounding boxes and class labels for supervised training on the pictures in the target set, and determine the specific positions and classes of objects in the blurred underwater environment;
[0101] Step 2.3: Use a web tool to perform contour annotation on the class samples that need to be amplified in quantity in the pictures of the target set according to the determined specific positions and classes, and obtain the physical contour annotation text.
[0102] Step 2.4: Read in the pictures of the target set obtained in Step 2.3 in sequence, traverse the annotation names obtained from the contour annotation text. When there is an annotation name that matches the current read-in picture name, then take out the contour annotation of this picture from the physical contour annotation text; copy all the pixel values within the contour of the current picture to a random area of the same shape to complete the sample replication.
[0103] Step 2.5: Use OpenCV to calculate the area U of the minimum circumscribed quadrilateral of the physical contour annotation in the pictures of the target set obtained in Step 2.3, and then calculate the intersection over union of this U and the areas of all the bounding boxes for supervised training within the current picture. The bounding box with the largest intersection over union with U is regarded as the bounding box of the replicated target.
[0104] Step 2.6: Assign the bounding boxes screened in Step 2.5 to the randomly replicated samples in Step 2.4, and add class labels for supervised training.
[0105] Step 3: Input the image dataset after data augmentation in Step 2 into the target detection network, and pre-train multiple complex networks as teacher networks, such as Figure 3Teacher Network 1, Teacher Network 2, and Teacher Network 3 are all networks trained on the input image dataset.
[0106] The complex network uses the convolutional neural network Resnet50 as the teacher network. The Resnet50 network contains 49 convolutional layers and one fully connected layer. The neural network structure can be divided into 7 parts. The first part does not contain residual blocks and mainly performs calculations of convolution, regularization, activation function, and max pooling on the input. The second to fifth parts of the structure all contain residual blocks. In the Resnet50 network structure, each residual block has three convolutional layers. The neural network has a total of 49 convolutional layers, and adding the final fully connected layer makes it a total of 50 layers. The teacher network's final classifier is added to the middle layer part. The classifier consists of a fully connected layer and a softmax activation function, aiming to represent the classification information of the middle layer. The input of the neural network is 224×224×3. After the convolutional calculations of the first 5 parts, the output is 7×7×2048. The pooling layer will convert it into a feature vector, and finally the fully connected layer will calculate this feature vector and output the class probability.
[0107] Step 4: Distill and compress the multiple teacher networks that have been trained in Step 3 using the multi-teacher network knowledge distillation method to obtain the student network, and add hyperparameters such as the distillation temperature and the distillation loss function, as Figure 3 shown in the multi-teacher knowledge distillation network.
[0108] Step 4.1: represents the training dataset after data augmentation in Step 2, where N is the number of samples, and x i is the i-th input variable, representing the input feature; y i is the i-th output variable, representing the corresponding true label value.
[0109] The prediction result σ(Z c ) of the teacher network is calculated by the Softmax function:
[0110]
[0111] where Z c is the output of the c-th node of the fully connected layer, T is the distillation temperature of the loss function, and C represents the total number of outputs of the teacher network's fully connected layer.
[0112] Step 4.2: Calculate the loss function between the teacher network and the student network;
[0113] Step 4.2.1: To effectively aggregate the prediction distributions of multiple teachers, by calculating the cross-entropy loss between the predicted values of the teacher network output and the ground truth labels, different weights are assigned to reflect their sample confidence levels.
[0114]
[0115]
[0116] Among them, k represents the k-th teacher network; K represents the number of teacher networks; c represents the c-th output of the fully connected layer of the teacher network; C represents the number of outputs of the fully connected layer of a single teacher network; represents the output of the fully connected layer of the k-th teacher network; y c represents the true label value of the teacher network; represents the cross-entropy loss between the predicted value of the teacher network in the output layer and the basic true label; represents the sample confidence of the output probability value of the fully connected layer of each teacher network.
[0117] Multiply the overall prediction of the teacher network by its corresponding sample confidence to obtain the loss function of the output features of the fully connected layer as:
[0118]
[0119] Among them, for the teacher network with a prediction closer to the true label, the smaller, the corresponding the larger, indicating that the sample confidence of the teacher network with a prediction close to the true label is greater.
[0120] Step 4.2.2: Extend the distillation method to the intermediate layer of the teacher network to mine more information. The calculation of the intermediate feature matching is as follows:
[0121]
[0122]
[0123]
[0124] Among them, is the final classifier of the k-th teacher network; h S is the feature vector of the last layer of the student network; represents the result of the k-th teacher network classifying the feature vector features of the last layer of the student network; is the cross-entropy loss between the prediction of the teacher network in the intermediate layer and the basic true label; K represents the number of teacher networks; represents the sample confidence of the output probability value of the intermediate layer of the teacher network.
[0125] To stabilize the knowledge transfer process, the student network focuses more on imitating the teacher network with a similar feature space, and the loss function of the intermediate layer features is obtained as:
[0126]
[0127] Among them, r(·) is a function for aligning the feature dimensions, aiming to increase the dimension of the student network to match the dimension of the teacher network; F S is the feature output of the feature map of the student network; is the feature output of the k-th feature map of the teacher network; denotes the norm, which is used as a distance metric for intermediate features; The meaning is the same as above.
[0128] Step 4.2.3: Calculate the regularized cross-entropy loss function with true labels as:
[0129]
[0130] Among them, y c denotes the result predicted by the teacher network, is the output of the fully connected layer of the student network.
[0131] Step 4.2.4: Finally, the overall loss function of multi-teacher knowledge distillation is:
[0132]
[0133] Among them, α and β are hyperparameters for balancing the loss of the output features of the fully connected layer of the teacher network and the loss of the intermediate layer features between the student network and the teacher network.
[0134] According to the above steps, the teacher network is distilled and trained, and finally the distilled and compressed student network is obtained. The distilled student network has similar recognition ability to the teacher network, but the network structure is smaller than that of the teacher network.
[0135] Step 5: Quantize the distilled and compressed student network and deploy it on the FPGA:
[0136] Step 5.1: Set the maximum and minimum values of the parameters to be quantized as x max and x min respectively, and set the maximum and minimum values of the quantized parameters as q max and q min respectively, and adopt the asymmetric quantization method with the quantization interval set as [0, 255];
[0137] Step 5.2: Calculate the quantization scale factor B and zero point O, and the formulas are respectively:
[0138]
[0139]
[0140] Among them, round(·) represents the rounded value of the floating-point number within the parentheses;
[0141] Step 5.3: Quantize the floating-point parameters, and the quantization formula is:
[0142] r = B(q - O)
[0143]
[0144] Among them, r represents the floating-point real number, q represents the quantized fixed-point integer; the meaning of round(·) is the same as above;
[0145] Step 5.4: During the process of calculating the student network, the operations of the convolutional layer and the fully connected layer of the student network are reduced to matrix multiplication operations; the steps to convert a floating-point matrix into a fixed-point matrix are as follows;
[0146] Step 5.4.1: Set m 1 and m 2 are two floating-point weight matrices in the middle layer of the student network, both with a size of N×N; the matrix m 3 is the result of multiplying m 1 m 2 ; the elements in m 3 are calculated by the formula:
[0147]
[0148] Among them, i, r, and k are the row number and column number of the corresponding matrix respectively;
[0149] Step 5.4.2: Obtained by organizing the formulas in Steps 5.3 and 5.4.1:
[0150]
[0151] Among them, B 1 , O 1 are the scale factor and zero point corresponding to the matrix m 1 ; B 2 , O 2 are the scale factor and zero point corresponding to the matrix m 2 ; B 3 , O 3 are the scale factor and zero point corresponding to the matrix m 3 ;
[0152] Step 5.4.3: Quantize each element 3 in the matrix q as:
[0153]
[0154] Step 5.4.4: Convert the in Step 5.4.3 into fixed-point operation as:
[0155]
[0156] where, M 0 is a real number, and the values of n and M 0 are selected as the optimal n and M according to the initial values of M and P and the optimization algorithm 0 such that MP = 2 -n M 0 P holds;
[0157] Step 5.4.5: Quantize each element 3 in matrix q through optimized calculation as:
[0158]
[0159] Step 5.5: In Step 4, the distilled and compressed student network is represented as: H = (x, y); where, x is the input matrix of the student network, and q is the output matrix of the student network; calculate the matrix multiplications in the convolutional layer and the fully connected layer as:
[0160]
[0161]
[0162] where, a 1 is the feature map matrix of the convolutional layer of the student network, w 1 is the weight parameter matrix of the convolutional layer of the student network, and x is the input matrix of the convolutional layer of the student network; a 2 is the input matrix of the fully connected layer of the student network, w 2 is the weight parameter matrix of the fully connected layer of the student network, and q is the feature map matrix of the fully connected layer of the student network; b 1 and b 2 are bias parameter matrices;
[0163] Step 5.6: According to the formulas in Step 5.4 and Step 5.5, the quantization calculations of the convolutional layer and the fully connected layer in the student network are respectively:
[0164]
[0165]
[0166] where, the meanings of a 1 , x, and w 1 are the same as above; q a1, q x , is a 1 , x, w 1 's quantization matrix; O x , is the matrix x, w 1 , a 1 's zero matrix;
[0167] q, a 2 , w 2 has the same meaning as above; q q , q a2 , is the quantization matrix of q, a 2 , w 2 ; O q is the zero matrix of the matrix a 2 , w 2 , q;
[0168] b 1 , b 2 are bias parameter matrices, and are the quantization matrix values of b 1 and b 2 ; f represents the output of the student network; M a1 , M a2 and the M obtained in step 5.4.4 0 is the same method;
[0169] Step 5.7: In the activation layer of the student network, the activation function is:
[0170]
[0171] where a 2 is the feature map matrix of the activation layer of the student network, a 1 is the input matrix of the activation layer of the student network, q a1 , q a2 is a 1 , a 2 's quantization matrix, is the zero matrix of the matrix a 1 ; the max(·) function represents comparing the values of the input quantization matrix and the zero matrix of the activation layer, and returns the larger value at the corresponding positions of the two matrices;
[0172] Step 5.8: Through the calculations in steps 5.6 and 5.7, a full-precision student network is obtained, and the entire student network is quantized. Finally, the student network is deployed on the FPGA to achieve hardware-accelerated computing;
[0173] Step 6: Implement the automatic recognition of the target object in the image by inputting the images collected by the underwater laser camera into the distilled and quantized student network accelerated by the hardware platform.
[0174] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. An underwater target recognition method based on quantization distillation, characterized in that, it includes the following steps: Step 1: Obtain seabed biological image data samples and construct a sample data set; The sample data set is randomly formed into a training data set and a test data set in proportion; Step 2: Implement a data augmentation method based on sample replication on the training data set to obtain an augmented training data set; Step 3: Input the augmented training data set into the Resnet50 network and pre-train multiple Resnet50 networks as teacher networks; Step 4: Compress the teacher network based on the multi-teacher knowledge distillation method to obtain a student network; the structure of the student network is the same as that of the teacher network; the student network uses the convolutional neural network Resnet18, and the convolutional neural network Resnet18 includes 17 convolutional layers and a fully connected layer; Step 5: Quantize the distilled and compressed student network and deploy it on the FPGA: Step 5.1: Set the maximum and minimum values of the parameter to be quantized as x max and x min respectively. Set the maximum and minimum values of the quantized parameter as q max and q min respectively, and adopt the asymmetric quantization method with the quantization interval set as [0, 255]. Step 5.2: Calculate the quantization scale factor B and zero point O, and the formulas are respectively: where round(·) represents the rounded value of the floating-point number in the parentheses; Step 5.3: Quantize the floating-point parameters, and the quantization formula is: r = B(q - O) where r represents the floating-point real number, q represents the quantized fixed-point integer; the meaning of round(·) is the same as above; Step 5.4: During the calculation of the student network, the operations of the convolutional layer and the fully connected layer of the student network are reduced to matrix multiplication operations; the steps to convert a floating-point matrix into a fixed-point matrix are as follows; Step 5.4.1: Aggregate m 1 and m 2 are two floating-point weight matrices in the middle layer of the student network, both of size N×N; matrix m 3 is m 1 m 2 the result of multiplying m 3 The formula for each element in m is as follows: where i, r, k are the number of rows and columns of the corresponding matrix respectively; Step 5.4.2: Obtained by sorting out the formulas in Steps 5.3 and 5.4.1: Among them, B 1 , O 1 are the scale factor and zero point corresponding to the m 1 matrix. B 2 , O 2 are the scale factor and zero point corresponding to the m 2 matrix; B 3 , O 3 are the scale factor and zero point corresponding to the m 3 matrix; Step 5.4.3: Quantization matrix q 3 each element in is as follows: Step 5.4.4: Convert the in Step 5.4.3 into fixed-point operation as follows: Among them, M 0 is a real number, and the values of n and M 0 are selected optimally according to the initial values of M and P and the optimization algorithm 0 such that MP = 2 -n M 0 P holds; Step 5.4.5: Quantization matrix q 3 for each element is optimized and calculated as: Step 5.5: In Step 4, the distilled and compressed student network is expressed as: H = (x, y); x is the input matrix of the student network, and q is the output matrix of the student network; calculate the matrix multiplication in the convolutional layer and the fully connected layer as: Among them, a 1 is the feature map matrix of the convolutional layer of the student network, w 1 is the weight parameter matrix of the convolutional layer of the student network, and x is the input matrix of the convolutional layer of the student network; a 2 is the input matrix of the fully connected layer of the student network, w 2 is the weight parameter matrix of the fully connected layer of the student network, and q is the feature map matrix of the fully connected layer of the student network; b 1 and b 2 are bias parameter matrices; Step 5.6: According to the formulas in Steps 5.4 and 5.5, the quantization calculations of the convolutional layer and the fully connected layer in the student network are respectively: Among them, a 1 , x, w 1 have the same meanings as above; q a1 , q x , is the quantization matrix of a 1 , x, w 1 ; O x , is the zero matrix of matrices x, w 1 , a 1 ; q, a 2 , w 2 has the same meaning as above; q q , q a2 , is the quantization matrix of q, a 2 , w 2 ; O q is the zero matrix of matrix a 2 , w 2 , q; b 1 and b 2 are bias parameter matrices, and are the quantization matrix values of b 1 and b 2 ; f represents the output of the student network; M a1 and M a2 are obtained in the same way as M 0 in step 5.4.4; Step 5.7: In the activation layer of the student network, the activation function is: Among them, a 2 is the feature map matrix of the student network activation layer, and a 1 is the input matrix of the student network activation layer. q a1 and q a2 are the quantization matrices of a 1 and a 2 . is the zero-point matrix of matrix a 1 . The max(·) function represents comparing the values of the input quantization matrix and the zero-point matrix of the activation layer and returns the larger value at the corresponding positions of the two matrices. Step 5.8: Through the calculations in Steps 5.6 and 5.7, a full-precision student network is obtained, and the entire student network is quantized. Finally, the student network is deployed on the FPGA to achieve hardware acceleration calculation; Step 6: Update the distilled and quantized student network to a new target detection network, and input the image into the updated target detection network in real time to achieve real-time underwater target recognition.
2. The underwater target recognition method based on quantization distillation according to claim 1, characterized in that, the seabed biological image data samples are used to capture target images in different angular directions from different backgrounds, and at least one of the target images has a shape similar to the object to be detected, or is an image of the side of the object to be detected; the sample data set includes the target images to be detected and annotates them.
3. The underwater target recognition method based on quantization distillation according to claim 1, characterized in that, the specific content of Step 2 is: Step 2.1: Select pictures from the training dataset where the volume of the seabed creature category samples accounts for less than 5% of the whole picture, with the number being 1 - 2 and the edges being clear. These pictures in the target set are used for sample replication; Step 2.2: Draw the true bounding boxes and class labels for supervised training on the pictures in the target set to determine the specific positions and categories of the seabed creatures in the target set pictures; Step 2.3: Perform contour annotation on the pictures in the target set according to the determined specific positions and categories of the seabed creatures to obtain the physical contour annotation text; Step 2.4: Read in the pictures in the target set obtained in Step 2.3 in sequence, traverse the annotation names obtained from the physical contour annotation text. When there is an annotation name that matches the name of the currently read-in picture, extract the contour annotation of this picture from the physical contour annotation text; copy all the pixel values within this contour of the current picture to a random area of the same shape to complete sample replication; Step 2.5: Use OpenCV to calculate the area U of the minimum bounding rectangle of the physical contour annotation in the pictures in the target set obtained in Step 2.3, calculate the intersection over union of U and the areas of all the true bounding boxes for supervised training within the currently completed sample replication pictures. The bounding box with the largest intersection over union with U is the bounding box of the replicated target; Step 2.6: Assign the bounding box selected in Step 2.5 to the randomly replicated samples in Step 2.4, and add class labels for supervised training. After the training is completed, an augmented training dataset will be obtained.
4. The underwater target recognition method based on quantization distillation according to claim 1, characterized in that, The teacher network structure in Step 3 is divided into 7 parts. The first part includes a convolutional layer, an activation layer, and a max pooling layer, which perform convolution, activation, and max pooling calculations on the input augmented training dataset; the structures of the second to fifth parts are the same and are intermediate layers, including residual blocks, convolutional blocks, and classifiers; the classifier includes a fully connected layer and a softmax activation function, which are used to represent the classification information of the intermediate layer; the sixth part is an average pooling layer, which converts the input content into a feature vector; the seventh part is a fully connected layer, which calculates the feature vector and outputs the class probability.
5. The underwater target recognition method based on quantization distillation according to claim 1, characterized in that, Step 4 is specifically as follows: Step 4.1: represents the training data set after data augmentation, where N is the number of samples, and x i is the i-th input variable; y i is the true label value corresponding to the i-th input variable; The prediction result σ(Z c ) of the teacher network is calculated according to the Softmax function, and the output probability distribution is softened by the hyperparameter distillation temperature T: Among them, Z c is the output of the c-th node in the fully connected layer of the teacher network, T is the distillation temperature, and C represents the total number of outputs of the fully connected layer of the teacher network; Step 4.2: Calculate the loss function between the teacher network and the student network; Step 4.2.1: Calculate the cross-entropy loss between the prediction results of multiple teacher networks and the true label values, and assign different weights to reflect the sample confidence of different teacher networks for the student network; Among them, k represents the k-th teacher network; K represents the total number of teacher networks; c represents the output of the c-th node in the fully connected layer of a single teacher network; C represents the total number of outputs of the fully connected layer of a single teacher network; represents the output of the fully connected layer of the k-th teacher network; y c represents the true label value of the teacher network; represents the cross-entropy loss between the prediction of the fully connected layer of the teacher network and the basic true label; represents the sample confidence of the output probability values of the fully connected layers of each teacher network; Multiply the predictions of multiple teacher networks by their corresponding sample confidences, and the loss function of the output features of the fully connected layer of the teacher network is: Among them, represents the output of the fully connected layer of the student network; represents the prediction result of the student network; Step 4.2.2: Extend the distillation method to the intermediate layer of the teacher network, and the calculation of its feature matching is as follows: Among them, is the classifier of the k-th teacher network; h S is the feature vector of the last layer of the student network; represents the feature classification result of the k-th teacher network on the feature vector of the last layer of the student network; is the cross-entropy loss between the prediction of the teacher network in the intermediate layer and the ground truth label; K represents the number of teacher networks; represents the sample confidence of the output probability value of the intermediate layer of the teacher network; The loss function of the intermediate layer features of the student network and the teacher network is: Among them, r(·) is a feature dimension alignment function used to increase the dimension of the student network to match the dimension of the teacher network; F S is the feature output of the feature map of the student network; is the feature output of the k-th feature map of the teacher network; denotes the norm, which is a distance metric for the intermediate features of the teacher network and the student network; Step 4.2.3: Calculate the regular cross-entropy loss function with true labels as: Among them, y c represents the true label value of the teacher network, which is the output of the fully connected layer of the student network; Step 4.2.4: Finally, the overall loss function of multi-teacher network knowledge distillation is as follows: where α and β are hyperparameters that balance the loss of the output features of the fully connected layer of the teacher network and the loss of the intermediate layer features of the student network and the teacher network; Step 4.3: Initialize the student network model and use multi-teacher knowledge distillation for the student network; calculate the output of the teacher network and soften the probability distribution as shown in Step 4.1, adding the distillation temperature T to soften the output probability distribution and provide more supervision information; the loss functions of the teacher network and the student network are as shown in Step 4.2, and the overall loss function using multi-teacher distillation for the student network to learn the knowledge of the teacher network; finally, by adjusting the hyperparameters of the distillation temperature T and the overall distillation loss function the student network is further converged to obtain the distilled and compressed student network.
Citation Information
Patent Citations
Small sample number underwater target identification method based on deep learning
CN115508838A
2-exponential power deep neural network quantification method based on knowledge distillation training
CN111985523A
Knowledge distillation method and system
CN112508169A