An underwater target detection method based on a differential pyramid neural network of high and low frequency features

By adopting high- and low-frequency feature differential pyramid neural network in underwater target detection technology, the problem of insufficient information recognition speed and accuracy in underwater target detection is solved, and high-precision and high-speed underwater target detection and more effective underwater scene monitoring are achieved.

CN115761467BActive Publication Date: 2025-05-27DALIAN MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211456179.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2025-05-27
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

The existing underwater target detection technology has bottlenecks in terms of information recognition speed and target detection accuracy. Especially when there is insufficient underwater light source, poor image quality leads to blurred target features, affecting the detection effect.

Method used

The underwater object detection method based on the differential pyramid neural network of high and low frequency feature is adopted. By acquiring and preprocessing the underwater image data, wavelet transformation is performed to separate high and low frequency features, a high and low frequency dual branch network is constructed, a multi-scale feature pyramid is extracted, and feature fusion is performed to achieve detection of targets of different scales of underwater images.

Benefits of technology

It improves the accuracy and speed of underwater target detection, realizes more effective underwater scene monitoring, saving time and capital costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761467B_ABST
    Figure CN115761467B_ABST
Patent Text Reader

Abstract

The present invention provides an underwater target detection method based on a high-low frequency feature difference pyramid neural network, including: acquiring underwater image data, and performing preprocessing and wavelet transform operations on the acquired underwater image data to obtain the input data of the network; constructing a high-low frequency double-branch network to extract features from the underwater image after preprocessing and wavelet transform operations; respectively extracting 4 intermediate feature layers from the features extracted by the high-low frequency double-branch network for splicing and differential operations of residual subtraction to construct a double-branch multi-scale feature pyramid; performing feature fusion on the double-branch multi-scale feature pyramid to construct a multi-scale output feature map, and realizing the detection of underwater targets of different scales through the output results after each feature fusion. The technical solution of the present invention can effectively achieve high-precision and high-speed detection of underwater targets, and at the same time realize more effective monitoring of underwater scenes, saving time and capital costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection technology, and in particular to an underwater target detection method based on a high- and low-frequency feature differential pyramid neural network. Background Art

[0002] At present, underwater target detection technology is a hot research direction in the field of underwater image processing. With the rapid development of artificial intelligence and computer vision technology, the bottlenecks of low information recognition speed and poor target detection accuracy in traditional underwater target detection technology have been significantly broken through. Different from traditional target detection algorithms, underwater target detection technology based on deep learning realizes category recognition and position box labeling of underwater targets through deep network learning mapping relationship, but the accuracy and real-time detection speed of current mainstream underwater target detection technology still have a lot of room for improvement.

[0003] The difficulty of underwater target detection lies in the fact that the underwater environment is more complex than the environment on land. Due to the lack of underwater light sources, artificial light sources are required when performing tasks such as target detection. The propagation of light sources underwater will be affected by water absorption, reflection, and refraction, resulting in serious attenuation of the effect. As a result, the images collected underwater appear blurred, low contrast, uneven lighting, and inconsistent colors. The emergence of these problems makes the features of the target in the image blurred and difficult to identify, and the image edges merge with the background water area, which greatly affects the device's judgment of the target in the underwater image. Therefore, the current target detection algorithm based on deep learning has many problems in the application of underwater target detection and classification. Summary of the invention

[0004] For target detection based on underwater images, in order to overcome the above-mentioned defects of the current algorithm, improve the accuracy and speed of underwater target detection, and at the same time, achieve more effective monitoring of underwater scenes and save time and financial costs, the present invention proposes an underwater target detection method based on high and low frequency feature difference pyramid neural network.

[0005] The technical means adopted by the present invention are as follows:

[0006] An underwater target detection method based on high- and low-frequency feature difference pyramid neural network, comprising:

[0007] Acquire underwater image data, and perform preprocessing and wavelet transform operations on the acquired underwater image data to obtain input data of the network;

[0008] A high- and low-frequency dual-branch network is constructed to extract features from underwater images after preprocessing and wavelet transform operations;

[0009] The four intermediate feature layers of the features extracted by the high- and low-frequency dual-branch network are respectively extracted for concatenation and differential operation of residual subtraction to construct a dual-branch multi-scale feature pyramid.

[0010] The dual-branch multi-scale feature pyramid is fused to construct a multi-scale output feature map. The output results after each feature fusion are used to detect targets of different scales in underwater images.

[0011] Furthermore, the underwater image data is obtained, and preprocessing and wavelet transforming are performed on the obtained underwater image data to obtain input data of the network, specifically including:

[0012] Preprocessing operation: taking the acquired underwater image data as the original image, cropping and sharpening the original image;

[0013] Wavelet transform operation: Through the wavelet transform decomposition algorithm, each preprocessed original image is divided into a high-frequency image and a low-frequency image as input data.

[0014] Furthermore, the high- and low-frequency dual-branch network is constructed to extract features from the underwater image after preprocessing and wavelet transform operation, specifically including:

[0015] Input the high-frequency image and the low-frequency image into two weight-shared ResNet50 networks to obtain a high- and low-frequency dual-branch network;

[0016] After a series of convolution, batch normalization, ReLU activation, and maximum pooling, the high-frequency image and the low-frequency image pass through one Conv block and two Identity blocks, and the intermediate features at this time are extracted as the first high-frequency branch intermediate feature layer and the first low-frequency branch intermediate feature layer;

[0017] After passing through one Conv block and three Identity blocks, the intermediate feature layer at this time is extracted as the second high-frequency branch intermediate feature layer and the second low-frequency branch intermediate feature layer;

[0018] After one more Conv block and five Identity blocks, the intermediate feature layer at this time is extracted as the third high-frequency branch intermediate feature layer and the third low-frequency branch intermediate feature layer;

[0019] After passing through one Conv block and two Identity blocks, the fourth high-frequency branch intermediate feature layer and the fourth low-frequency branch intermediate feature layer of the intermediate feature layer are extracted.

[0020] Furthermore, the four intermediate feature layers extracted from the features extracted by the high- and low-frequency dual-branch network are respectively concatenated and differentially operated by residual subtraction to construct a dual-branch multi-scale feature pyramid, specifically including:

[0021] The first high-frequency branch intermediate feature layer and the first low-frequency branch intermediate feature layer, the second high-frequency branch intermediate feature layer and the second low-frequency branch intermediate feature layer, the third high-frequency branch intermediate feature layer and the third low-frequency branch intermediate feature layer, and the fourth high-frequency branch intermediate feature layer and the fourth low-frequency branch intermediate feature layer are respectively concatenated and 1×1 convolved to form the first branch four-scale feature pyramid;

[0022] Residual subtraction operations are performed on the first high-frequency branch intermediate feature layer and the first low-frequency branch intermediate feature layer, the second high-frequency branch intermediate feature layer and the second low-frequency branch intermediate feature layer, the third high-frequency branch intermediate feature layer and the third low-frequency branch intermediate feature layer, and the fourth high-frequency branch intermediate feature layer and the fourth low-frequency branch intermediate feature layer as the second branch four-scale feature pyramid.

[0023] Furthermore, the dual-branch multi-scale feature pyramid is subjected to feature fusion to construct a multi-scale output feature map, and the detection of targets of different scales in underwater images is realized through the output results after each feature fusion, which specifically includes:

[0024] Based on the first branch four-scale feature pyramid and the second branch four-scale feature pyramid, respectively extract the feature maps of the largest scale, and perform a splicing operation on the feature maps of the largest scale;

[0025] Perform a 1×1 convolution operation on the feature map of the largest scale after the splicing operation;

[0026] Based on the first-branch four-scale feature pyramid and the second-branch four-scale feature pyramid, feature maps of other scales are respectively extracted, and the feature maps of other scales are spliced;

[0027] A 1×1 convolution operation is performed on the feature maps of other scales after the splicing operation to obtain four-scale output feature maps.

[0028] Furthermore, the underwater target detection method based on high- and low-frequency feature difference pyramid neural network is used to identify underwater target images, comprising the following steps:

[0029] Divide the underwater image dataset into a training set and a test set;

[0030] High and low frequency feature difference pyramid neural network model training:

[0031] Determine the high- and low-frequency feature difference pyramid neural network model, input the training set images into the high- and low-frequency feature difference pyramid neural network model, and the high- and low-frequency feature difference pyramid neural network model explores the mapping relationship between the original underwater image target prediction rectangular frame position and the underwater image target rectangular frame position with the mark according to the input data in the training set, and preliminarily learns the target detection capability;

[0032] Determine the loss function of the high- and low-frequency feature difference pyramid neural network model, and take the logarithm of the ratio of the intersection and union of the rectangular box of the detection result and the rectangular box of the sample annotation and then take the negative value (IOU loss) as the loss function;

[0033] The high- and low-frequency feature difference pyramid neural network model is trained using all training set data, and the interaction ratio values ​​of the positions of the rectangular frame output by the high- and low-frequency feature difference pyramid neural network model for actual underwater target detection and the target frame marked in the training set are compared. The loss function value is calculated and recorded, and the change of the loss function curve is observed until the high- and low-frequency feature difference pyramid neural network model has converged.

[0034] High and low frequency feature difference pyramid neural network model test:

[0035] After the high- and low-frequency feature difference pyramid neural network model is trained, the underwater image test set data is input into the high- and low-frequency feature difference pyramid neural network model; the interaction ratio loss function value of the target rectangular box position output by the high- and low-frequency feature difference pyramid neural network model and the target rectangular box position marked in the test set and the evaluation index of the test high- and low-frequency feature difference pyramid neural network model are analyzed to determine whether the high- and low-frequency feature difference pyramid neural network model has the ability to detect underwater targets. If so, enter the step of running the high- and low-frequency feature difference pyramid neural network model; if not, return to the high- and low-frequency feature difference pyramid neural network model training step for further training;

[0036] Run the high and low frequency feature difference pyramid neural network model:

[0037] The image data to be used for underwater target detection is input into the high- and low-frequency feature difference pyramid neural network model, and the high- and low-frequency feature difference pyramid neural network model is run to obtain and record the loss function value of the interaction ratio with the input underwater target candidate frame position calculated by the high- and low-frequency feature difference pyramid neural network model and the results of the network model evaluation index.

[0038] Compared with the prior art, the present invention has the following advantages:

[0039] 1. The underwater target detection method based on high- and low-frequency feature differential pyramid neural network provided by the present invention can effectively realize high-precision and high-speed detection of underwater targets.

[0040] 2. The underwater target detection method based on high- and low-frequency feature differential pyramid neural network provided by the present invention can achieve more effective monitoring of underwater scenes and save time and financial costs.

[0041] Based on the above reasons, the present invention can be widely promoted in the fields of target detection and the like. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0043] Figure 1 It is a flow chart of the underwater target detection method based on high- and low-frequency feature differential pyramid neural network of the present invention;

[0044] Figure 2 A detailed schematic diagram of extracting intermediate feature layers of different scales based on ResNet50 in the present invention;

[0045] Figure 3 An example of a high-low frequency dual-branch feature pyramid neural network provided by an embodiment of the present invention;

[0046] Figure 4 An illustration of a method for differentially constructing a dual-branch feature pyramid network and performing feature fusion provided by an embodiment of the present invention;

[0047] Figure 5 It is a flowchart of training the network model and its actual application in the present invention. DETAILED DESCRIPTION

[0048] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0049] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0051] Unless otherwise specifically stated, the relative arrangement of the parts and steps described in these embodiments, the numerical expressions and numerical values ​​do not limit the scope of the present invention. At the same time, it should be clear that, for ease of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship. The technology, methods and equipment known to ordinary technicians in the relevant field may not be discussed in detail, but in appropriate cases, the technology, methods and equipment should be regarded as part of the authorization specification. In all examples shown and discussed here, any specific value should be interpreted as merely exemplary, rather than as a limitation. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0052] In the description of the present invention, it is necessary to understand that the directions or positional relationships indicated by directional words such as "front, back, up, down, left, right", "lateral, vertical, perpendicular, horizontal" and "top, bottom" are usually based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description. Unless otherwise specified, these directional words do not indicate or imply that the device or element referred to must have a specific direction or be constructed and operated in a specific direction. Therefore, they cannot be understood as limiting the scope of protection of the present invention: the directional words "inside and outside" refer to the inside and outside relative to the contours of each component itself.

[0053] For ease of description, spatially relative terms such as "above", "above", "on the upper surface of", "above", etc. may be used here to describe the spatial positional relationship between a device or feature and other devices or features as shown in the figure. It should be understood that spatially relative terms are intended to include different orientations of the device in use or operation in addition to the orientation described in the figure. For example, if the device in the accompanying drawings is inverted, the device described as "above other devices or structures" or "above other devices or structures" will be positioned as "below other devices or structures" or "below their position devices or structures". Thus, the exemplary term "above" can include both "above" and "below". The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and the spatially relative descriptions used here are interpreted accordingly.

[0054] In addition, it should be noted that the use of terms such as "first" and "second" to limit components is only for the convenience of distinguishing the corresponding components. If not otherwise stated, the above terms have no special meaning and therefore cannot be understood as limiting the scope of protection of the present invention.

[0055] like Figure 1 As shown, the present invention provides an underwater target detection method based on a high- and low-frequency feature difference pyramid neural network, comprising:

[0056] S1, acquiring underwater image data, and performing preprocessing and wavelet transform operations on the acquired underwater image data to obtain input data of the network;

[0057] S2, constructing a high-low frequency dual-branch network to extract features from underwater images after preprocessing and wavelet transform operations;

[0058] S3, respectively extracting the four intermediate feature layers from the features extracted by the high- and low-frequency dual-branch network for concatenation and differential operation of residual subtraction to construct a dual-branch multi-scale feature pyramid;

[0059] S4. Perform feature fusion on the dual-branch multi-scale feature pyramid to construct a multi-scale output feature map, and detect targets of different scales in underwater images through the output results after each feature fusion.

[0060] In specific implementation, as a preferred embodiment of the present invention, in step S1, underwater image data is obtained, and preprocessing and wavelet transform operations are performed on the obtained underwater image data to obtain input data of the network, which specifically includes:

[0061] S11, preprocessing operation: taking the acquired underwater image data as the original image, and performing cropping and sharpening processing on the original image;

[0062] In this embodiment, underwater image data can be obtained through channels such as Internet open source databases, papers or technical reports. The underwater images obtained are first cropped, that is, the original images are reshaped (Resize) by direct stretching, and the image sizes are collectively integrated to 512×512. The purpose of this step is to prevent the data set from being too large and resulting in insufficient video memory during training. The cropped image is then sharpened, where the original image sharpening algorithm uses a relative total variation model to perform image sharpening based on the structural layer and texture layer of the underwater image to effectively balance the image's hue, saturation and clarity. The purpose of this step is to improve the visual effect of the image, highlight information that is useful for human and machine analysis, suppress useless information, and improve the image's use value.

[0063] S12, wavelet transform operation: through the wavelet transform decomposition algorithm, each pre-processed original image is divided into a high-frequency image and a low-frequency image as input data.

[0064] In this embodiment, a wavelet transform function is used to perform a wavelet transform operation on the "preprocessed" image. The specific implementation details are to convert each input image into a high-frequency image and a low-frequency image respectively; the high-frequency features of the image refer to the places where the image intensity changes dramatically, that is, the edges and contours of the image, and the low-frequency features of the image refer to the places where the image intensity changes slowly, that is, the large color blocks in the image; by separating the high- and low-frequency information of the image, the required image information can be differentially extracted in advance, further enhancing the deep learning ability of the image.

[0065] In step S1, the underwater image data set can be divided according to a certain quantity ratio. For example, according to common classifications in the field of artificial intelligence, 80% of all image pairs are randomly selected as the training set of the network model to improve the model's target detection ability; the other 20% is used as a test set for the target detection model to test the model's actual target detection effect.

[0066] In specific implementation, as a preferred embodiment of the present invention, in step S2, a high-low frequency dual-branch network is constructed to extract features from the underwater image after preprocessing and wavelet transform operation, specifically including:

[0067] S21, input the high-frequency image and the low-frequency image into two weight-shared ResNet50 networks to obtain a high- and low-frequency dual-branch network;

[0068] In this embodiment, the ResNet50 network is used as the feature extraction network, wherein the ResNet50 network is also called a residual network. The design of the deep residual network is to overcome the problem that the learning efficiency becomes low and the accuracy cannot be effectively improved due to the deepening of the network depth; the ResNet50 network has two basic blocks, namely the Conv block and the Identity block. The detailed network structure is as follows Figure 2 As shown in , the function of the Conv block is to change the dimension of the network and cannot be connected in series. The function of the Identity block is to deepen the network without changing the dimension and can be connected in series, such as Figure 2 As shown, the overall ResNet50 network structure diagram.

[0069] S22, high-frequency images and low-frequency images are processed as follows Figure 2 After a series of convolutions, batch normalization, ReLU activation, and maximum pooling, a Conv block and two Identity blocks are passed through, and the intermediate features at this time are extracted as the first high-frequency branch intermediate feature layer and the first low-frequency branch intermediate feature layer;

[0070] S23, after continuing through one Conv block and three Identity blocks, extracting the intermediate feature layer at this time as the second high-frequency branch intermediate feature layer and the second low-frequency branch intermediate feature layer;

[0071] S24, after continuing through one Conv block and five Identity blocks, extracting the intermediate feature layer at this time as the third high-frequency branch intermediate feature layer and the third low-frequency branch intermediate feature layer;

[0072] S25. After passing through one Conv block and two Identity blocks, extract the fourth high-frequency branch intermediate feature layer and the fourth low-frequency branch intermediate feature layer of the intermediate feature layer at this time.

[0073] In the specific implementation, as a preferred embodiment of the present invention, in step S3, four intermediate feature layers of the features extracted by the high- and low-frequency dual-branch network are respectively extracted for splicing and differential operation of residual subtraction to construct a dual-branch multi-scale feature pyramid, such as Figure 3 As shown, the purpose of this embodiment is to perform the first stage feature fusion on the four feature layers obtained in the high and low frequency feature pyramids in the previous step, perform concatenation and differential operations of residual subtraction connection according to different scales, divide the features in detail, and finally obtain a dual-branch output feature pyramid, which provides an innovative idea for constructing feature maps in the field of target detection. The detailed steps are as follows:

[0074] S31, respectively concatenate and convolve the first high-frequency branch intermediate feature layer and the first low-frequency branch intermediate feature layer, the second high-frequency branch intermediate feature layer and the second low-frequency branch intermediate feature layer, the third high-frequency branch intermediate feature layer and the third low-frequency branch intermediate feature layer, and the fourth high-frequency branch intermediate feature layer and the fourth low-frequency branch intermediate feature layer as the first branch four-scale feature pyramid;

[0075] In this embodiment, since step S21 uses the same ResNet50 network and the positions of the extracted intermediate feature layers are the same, the sizes and depths of the first high-frequency branch intermediate feature layer and the first low-frequency branch intermediate feature layer, the second high-frequency branch intermediate feature layer and the second low-frequency branch intermediate feature layer, the third high-frequency branch intermediate feature layer and the third low-frequency branch intermediate feature layer, and the fourth high-frequency branch intermediate feature layer and the fourth low-frequency branch intermediate feature layer are the same, and a concatenation operation is used to connect the two intermediate feature layers. This operation can fuse the two intermediate feature layers in a specified dimension.

[0076] S32, respectively perform residual subtraction operations on the first high-frequency branch intermediate feature layer and the first low-frequency branch intermediate feature layer, the second high-frequency branch intermediate feature layer and the second low-frequency branch intermediate feature layer, the third high-frequency branch intermediate feature layer and the third low-frequency branch intermediate feature layer, and the fourth high-frequency branch intermediate feature layer and the fourth low-frequency branch intermediate feature layer as the second branch four-scale feature pyramid. That is, in the same dimension, the corresponding data is subtracted. The purpose of this step is to reduce the complexity of the model to reduce overfitting and prevent gradient disappearance.

[0077] In specific implementation, as a preferred embodiment of the present invention, in step S4, Figure 4 As shown in the figure, the dual-branch multi-scale feature pyramid is fused to construct a multi-scale output feature map. The output results after each feature fusion are used to detect targets of different scales in underwater images, including:

[0078] S41, based on the first branch four-scale feature pyramid and the second branch four-scale feature pyramid, respectively extract the feature map of the maximum scale, and perform a splicing operation on the feature map of the maximum scale; this step can fuse different information under the same scale feature map.

[0079] S42. Perform a 1×1 convolution operation on the feature map of the largest scale after the concatenation operation, in order to reduce the dimension, reduce the computational complexity, and reduce the amount of computation.

[0080] S43, based on the first-branch four-scale feature pyramid and the second-branch four-scale feature pyramid, respectively extract feature maps of other scales, and perform a splicing operation on the feature maps of other scales;

[0081] S44, perform 1×1 convolution operation on the feature maps of other scales after the splicing operation to obtain a four-scale output feature map. The four-scale output feature map obtained in this embodiment can be further classified and regressed to perform multi-scale object detection.

[0082] In this embodiment, the purpose of step S4 is to perform the second-stage feature fusion on the dual-branch four-scale feature pyramid obtained in step S3 to obtain a new four-scale output feature pyramid. This operation can complementarily fuse the previous high- and low-frequency features to form a new feature pyramid map, which is convenient for continuously updating the weights under the constraints of the subsequent loss function. At the same time, the multi-scale feature map can increase the detection capability of targets of different sizes.

[0083] In specific implementation, as a preferred implementation mode of the present invention, in this embodiment, the following step S5 adopts the general idea for model training and testing in the field of artificial intelligence. Figure 5 The flowchart of training, testing and running the network model to obtain the segmentation result in this embodiment is shown. The target detection method based on the high-low frequency feature difference pyramid neural network provided in the above embodiment is used to identify the underwater target image, that is, step S5: training, testing and running the network model to obtain the target detection result, specifically including:

[0084] The training set determined in step S1 is input into the network model for training and testing, so as to help the network model initially acquire the ability to detect multi-scale underwater targets. The underwater image dataset is divided into a training set and a test set;

[0085] S51. Build a network model with high and low frequency feature difference pyramid neural network model

[0086] After the underwater images for training and testing are respectively established, the network model is determined, and the training set images are input into the network model. The network model will infer the mapping relationship between the input data in the training set and the true value data, so as to learn the ability to detect multi-scale underwater targets. In this embodiment, it is necessary to build the relevant network model described in steps S1 to S4, build it layer by layer according to the relevant details, add the required relevant processing functions, and then instantiate it as a training network for underwater images. After the training data is input into the above-mentioned network model, according to the general ideas in the field of artificial intelligence, the network model can explore the mapping relationship between the target position of the original underwater image and the position of the marked underwater target frame based on the input data and the true value data in the training set. The network model's self-exploration of the mapping relationship between the original underwater image and the image with the manually annotated target frame is specifically manifested in that all parameters in the network can be adjusted autonomously during the training process.

[0087] S52. Determine the loss function of the high and low frequency feature difference pyramid neural network model

[0088] Determine the loss function of the network model to compare the difference between the image target category and its target position candidate frame coordinate output predicted by the network model and the manually annotated target category and its candidate frame position coordinate. This difference is related to the effect of autonomous adjustment of network parameters. The smaller the difference, the more similar the target detection category and target position candidate frame coordinate predicted by the network model are to the real manually annotated target category and target position candidate frame coordinate. This can indirectly prove that the effect of autonomous adjustment of the network model is conducive to the ability to detect multi-scale underwater targets.

[0089] The classification loss selected in this embodiment adopts cross entropy loss (CE loss). For cross entropy loss, binary classification (binary CE loss) is an extreme case of it. Cross entropy can measure the degree of difference between two different probability distributions in the same random variable. In machine learning, it is expressed as the difference between the true probability distribution and the predicted probability distribution. The smaller the value of cross entropy, the better the model prediction effect. Cross entropy is often standard with softmax in classification problems. Softmax processes the output structure so that the sum of its multiple classification prediction values ​​is 1, and then calculates the loss through cross entropy. The cross entropy loss function is a prior art in the current field of artificial intelligence, so it will not be described in detail here.

[0090] The regression loss selected in this embodiment is the intersection-over-union (IOU) loss function, where IOU refers to the ratio of the intersection and union of the "predicted bounding box" and the "ground truth". The calculation formula of IOU is:

[0091]

[0092] That is, IOU is equivalent to the result obtained by dividing the intersection of two areas by the union of the two areas. When the system is training the model, it usually uses the IOU ratio between the "predicted bounding box" and the "ground truth" to determine whether the result is a good one. Generally, IOU>0.5 is considered a good result. Similarly, the IOU loss function is also an existing technology in the current field of artificial intelligence, so it will not be described in detail here.

[0093] In order to maximize the learning ability of the network model, this embodiment further sets the hyperparameters in the network training model before the network model training.

[0094] When inputting training data, the order in which the training set is input into the network model should be random. Then, according to the common setting method in the field of artificial intelligence, the three hyper parameters of the batch size, learning rate, and the number of training cycles (epoch) are set before the network model training, so that the network knows the specific parameter autonomous adjustment method. It should be noted that the above three hyper parameters do not have a unique setting value, but can select common empirical values. For example, in this embodiment, the batch size in the network model training process is set to 32 (it can also be set to 4, 8, 16, 64, etc.), the learning rate is 0.0001 (it can also be set to 0.01, 0.001, 0.00001, etc.), and 100 training cycles are executed (it can also be set to 50, 150, 200, etc.), so as to realize the training of the network model, so as to improve the network model's ability to judge the mapping relationship between the image prediction target frame position coordinates and the manually annotated target frame position coordinates, thereby realizing the network model's ability to detect underwater multi-scale targets.

[0095] S53, high and low frequency feature difference pyramid neural network model training:

[0096] S531, determining a high- and low-frequency feature difference pyramid neural network model, and after inputting the training set image into the high- and low-frequency feature difference pyramid neural network model, the high- and low-frequency feature difference pyramid neural network model explores the mapping relationship between the original underwater image target prediction rectangular frame position and the underwater image target rectangular frame position with a mark according to the input data in the training set, and preliminarily learns the target detection capability;

[0097] S532, determining the loss function of the high- and low-frequency feature difference pyramid neural network model, taking the logarithm of the ratio of the intersection and union of the rectangular box of the detection result and the rectangular box of the sample annotation and then taking the negative value (IOU loss) as the loss function;

[0098] S533, using all the training set data to train the high- and low-frequency feature difference pyramid neural network model, comparing the interaction ratio values ​​of the positions of the rectangular frame output by the high- and low-frequency feature difference pyramid neural network model for actual underwater target detection and the target frame marked in the training set, calculating and recording the intersection-and-union ratio loss function value, and observing the change of the loss function curve until the high- and low-frequency feature difference pyramid neural network model has converged;

[0099] In this step, the network model constructed above is trained using all the training set data, and then the intersection-over-union (IOU) loss function curve is observed. When the curve gradually converges to the point where no violent fluctuations occur, the training is stopped, and it is preliminarily determined that the training of the network model in the current state has been completed.

[0100] It should be noted that the above steps are general steps for judging the convergence of network models in the current field of artificial intelligence. Under normal circumstances, the curve of any form of loss function will experience very obvious oscillations in the early stage of model training. As the number of training cycles increases, the loss function curve gradually decreases and the oscillation effect is significantly weakened. In the end, the loss function curve hardly fluctuates (or even does not fluctuate at all). When this state remains unchanged with the increase of training cycles, it can be concluded that the network model has converged.

[0101] S54, high and low frequency feature difference pyramid neural network model test:

[0102] After the high- and low-frequency feature difference pyramid neural network model is trained, in order to further determine whether the network model has truly completed training and has the ability to detect multi-scale targets in underwater images, the underwater image test set data is input into the high- and low-frequency feature difference pyramid neural network model; the interaction ratio value of the target rectangular box position output by the high- and low-frequency feature difference pyramid neural network model and the target rectangular box position marked in the test set and the evaluation index of the test high- and low-frequency feature difference pyramid neural network model are analyzed to determine whether the high- and low-frequency feature difference pyramid neural network model has the ability to detect underwater targets. If so, enter the step of running the high- and low-frequency feature difference pyramid neural network model; if not, return to the high- and low-frequency feature difference pyramid neural network model training step for further training;

[0103] In this embodiment, the predicted target category and candidate box position coordinates of the output image are compared with the actual target category and position coordinates manually marked. This is done by setting an evaluation index. When the evaluation index is higher than or not lower than a preset value, it is considered that the network model has the ability to detect multi-scale underwater targets. The evaluation index selected in this embodiment is the mAP comprehensive evaluation index recognized in the field of target detection, where AP refers to the average precision (Average Precision). For the task of target detection, each class can calculate its accuracy (Precision) and recall (Recall). Through reasonable calculation, each class can obtain a PR curve (with Precision as the horizontal coordinate and Recall as the vertical coordinate). The area under the curve is the AP value, which is used to measure the quality of a class detection. The evaluation index here is a mature evaluation index in the field of target detection, so it is not described in detail here. For example, the evaluation index is set to 60. When the above evaluation indicators are all higher than 60, it can be determined that the trained network model has the ability to detect multi-scale targets in underwater images. If the evaluation index is not higher than 60, you can return to step S5 and retrain. Obviously, the closer the above indicators are to 100, the more similar the output of the network model is to the ideal result, which indirectly proves that the network model has a stronger ability to detect multi-scale targets in underwater images. When the mAP evaluation index of the test network model is higher than 60, it is considered that the network model has the ability to detect multi-scale targets in underwater images.

[0104] S55. Run the high and low frequency feature difference pyramid neural network model:

[0105] The image data to be used for underwater target detection is input into the high- and low-frequency feature difference pyramid neural network model, and the high- and low-frequency feature difference pyramid neural network model is run. The interaction ratio of the input underwater target candidate frame position and the network model evaluation index calculated by the high- and low-frequency feature difference pyramid neural network model are obtained and recorded, and the underwater target detection work is completed.

[0106] Due to the current limitations in the field of artificial intelligence, the underwater target detection capability of the network model cannot be judged by directly observing the internal parameters of the network. It can only be indirectly identified by observing the intersection over union (IOU) loss function curve in step S5 and the test network model evaluation index mAP. Therefore, when the intersection over union (IOU) loss function curve converges stably and the test network model evaluation index mAP meets the standard, it can be judged that the current network model has the ability to detect underwater targets. At this time, according to the actual situation, the original underwater image data that needs to be detected for underwater targets can be input into the network model. The network model can automatically infer the expected underwater target detection results, thereby realizing the current underwater target detection work.

[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An underwater target detection method based on a high-low frequency feature differential pyramid neural network, characterized in that, it includes: Obtain underwater image data, and perform preprocessing and wavelet transform operations on the obtained underwater image data to obtain the input data of the network; Construct a high-low frequency dual-branch network to extract features from the underwater image after preprocessing and wavelet transform operations; Extract 4 intermediate feature layers from the features extracted by the high-low frequency dual-branch network respectively for splicing and differential operations of residual subtraction, and construct a dual-branch multi-scale feature pyramid; Perform feature fusion on the dual-branch multi-scale feature pyramid to construct a multi-scale output feature map, and detect underwater image targets at different scales through the output results after each feature fusion.

2. The underwater target detection method based on a high-low frequency feature differential pyramid neural network according to claim 1, characterized in that, the step of obtaining underwater image data, and performing preprocessing and wavelet transform operations on the obtained underwater image data to obtain the input data of the network specifically includes: Preprocessing operation: Use the obtained underwater image data as the original image, and perform cropping and sharpening on the original image; Wavelet transform operation: Through the wavelet transform decomposition algorithm, divide each preprocessed original image into a high-frequency image and a low-frequency image as the input data.

3. The underwater target detection method based on a high-low frequency feature differential pyramid neural network according to claim 2, characterized in that, the step of constructing a high-low frequency dual-branch network to extract features from the underwater image after preprocessing and wavelet transform operations specifically includes: Input the high-frequency image and the low-frequency image into two ResNet50 networks with shared weights to obtain a high-low frequency dual-branch network; After the high-frequency image and the low-frequency image go through a series of convolutional, batch normalization, ReLU activation, and max pooling operations, and then go through 1 Conv block and 2 Identity blocks, extract the intermediate features at this time as the first high-frequency branch intermediate feature layer and the first low-frequency branch intermediate feature layer; Continue to go through 1 Conv block and 3 Identity blocks, and extract the intermediate feature layer at this time as the second high-frequency branch intermediate feature layer and the second low-frequency branch intermediate feature layer; Continue to go through 1 Conv block and 5 Identity blocks, and extract the intermediate feature layer at this time as the third high-frequency branch intermediate feature layer and the third low-frequency branch intermediate feature layer; Continue to go through 1 Conv block and 2 Identity blocks, and extract the intermediate feature layer at this time as the fourth high-frequency branch intermediate feature layer and the fourth low-frequency branch intermediate feature layer.

4. The underwater target detection method based on a high-low frequency feature differential pyramid neural network according to claim 3, characterized in that, the step of respectively extracting 4 intermediate feature layers from the features extracted by the high-low frequency dual-branch network for splicing and differential operations of residual subtraction to construct a dual-branch multi-scale feature pyramid specifically includes: Perform concatenation operations and 1×1 convolution operations on the intermediate feature layers of the first high-frequency branch and the first low-frequency branch, the intermediate feature layers of the second high-frequency branch and the second low-frequency branch, the intermediate feature layers of the third high-frequency branch and the third low-frequency branch, and the intermediate feature layers of the fourth high-frequency branch and the fourth low-frequency branch respectively, as the four-scale feature pyramid of the first branch; Perform residual subtraction operations on the intermediate feature layers of the first high-frequency branch and the first low-frequency branch, the intermediate feature layers of the second high-frequency branch and the second low-frequency branch, the intermediate feature layers of the third high-frequency branch and the third low-frequency branch, and the intermediate feature layers of the fourth high-frequency branch and the fourth low-frequency branch respectively, as the four-scale feature pyramid of the second branch.

5. The underwater target detection method based on the high-low frequency feature difference pyramid neural network according to claim 4, characterized in that, The feature fusion of the double-branch multi-scale feature pyramid is performed to construct a multi-scale output feature map, and the detection of underwater images with different scales of targets is realized through the output results after each feature fusion, specifically including: Based on the four-scale feature pyramid of the first branch and the four-scale feature pyramid of the second branch, extract the feature maps of the largest scale respectively, and perform a concatenation operation on the feature maps of the largest scale; Perform a 1×1 convolution operation on the feature map of the largest scale after the concatenation operation; Based on the four-scale feature pyramid of the first branch and the four-scale feature pyramid of the second branch, extract the feature maps of other scales respectively, and perform a concatenation operation on the feature maps of other scales; Perform a 1×1 convolution operation on the feature map of other scales after the concatenation operation to obtain a four-scale output feature map.

6. The underwater target detection method based on the high-low frequency feature difference pyramid neural network according to claim 1, characterized in that, The target detection method based on the high-low frequency feature difference pyramid neural network is used for target recognition of underwater target images, including the following steps: Divide the underwater image dataset into a training set and a test set; Training of the high-low frequency feature difference pyramid neural network model: Determine the high-low frequency feature difference pyramid neural network model. After inputting the training set images into the high-low frequency feature difference pyramid neural network model, the high-low frequency feature difference pyramid neural network model explores the mapping relationship between the position of the predicted rectangular box of the original underwater image target and the position of the rectangular box of the underwater image target with labels in the training set, and initially learns the target detection ability; Determine the loss function of the high-low frequency feature difference pyramid neural network model, and take the negative value of the logarithm of the ratio of the intersection and union of the rectangular box of the detection result and the rectangular box of the sample annotation as the loss function; Use all the training set data to train the high-low frequency feature difference pyramid neural network model, compare the intersection ratio value of the actual underwater target detection output rectangular box of the high-low frequency feature difference pyramid neural network model and the position of the target box marked in the training set, and observe the change of the loss function curve until the high-low frequency feature difference pyramid neural network model has converged; Testing of High-Frequency and Low-Frequency Feature Difference Pyramid Neural Network Model: After the training of the high-frequency and low-frequency feature difference pyramid neural network model is completed, the data of the underwater image test set is input into the high-frequency and low-frequency feature difference pyramid neural network model; analyze the intersection ratio value between the position of the target rectangle frame output by the high-frequency and low-frequency feature difference pyramid neural network model and the position of the target rectangle frame marked in the test set, as well as the evaluation index of the test high-frequency and low-frequency feature difference pyramid neural network model, and judge whether the high-frequency and low-frequency feature difference pyramid neural network model has the ability of underwater target detection. If it has, enter the step of running the high-frequency and low-frequency feature difference pyramid neural network model. If not, return to the training step of the high-frequency and low-frequency feature difference pyramid neural network model for further training; Running the High-Frequency and Low-Frequency Feature Difference Pyramid Neural Network Model: Input the image data to be detected for underwater targets into the high-frequency and low-frequency feature difference pyramid neural network model, run the high-frequency and low-frequency feature difference pyramid neural network model, and obtain and record the intersection ratio with the position of the underwater target candidate box calculated by the high-frequency and low-frequency feature difference pyramid neural network model itself and the results of the network model evaluation index.

Citation Information

Patent Citations

  • Fast optical flow field calculation method based on error-distributed multilayer grid

    CN103247058A

  • High-speed target detection method and system based on TridentNet structure and Cascade-RCNN structure

    CN112365497A