Underwater target detection and identification method and device, electronic equipment and storage medium
By using an annular dynamic optimization algorithm and a generative adversarial network based on Riemann manifold optimization in the underwater target detection and recognition method, the problem of low target detection accuracy in underwater dense scenarios is solved, and efficient and accurate underwater target recognition is achieved.
Patent Information
- Application Number
- CN202510395209.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The prior art is difficult to achieve efficient and accurate object detection and image processing in underwater dense scenarios, mainly due to scarce data, weak model generalization ability and low object detection accuracy.
The neural network model is optimized by a circular dynamic optimization algorithm, combined with a generative adversarial network based on Riemann manifold optimization for data expansion, and a regional convolutional neural network based on deep feature space mapping is used for object detection.
The model's adaptability and target detection accuracy in complex underwater environments are improved, the risk of overfitting is reduced, and efficient and accurate target detection is achieved.
Smart Images

Figure CN120236190A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to an underwater target detection and recognition method, device, electronic device, and storage medium. Background Art
[0002] With the progress of underwater detection technology, a large amount of high-quality underwater image data has been used in fields such as marine biology research, underwater archaeology, and environmental monitoring. However, due to the particularity of the underwater environment, such as unstable lighting conditions, many visual interferences, and complex scenes, it has become a great challenge to accurately identify and count targets from these images. The existing target detection and image processing methods face problems such as data scarcity, weak model generalization ability, and low target detection accuracy in underwater dense scenes. Summary of the Invention
[0003] The present invention aims to solve the problems of related technical limitations to at least a certain extent. For this purpose, the present invention provides an underwater target detection and recognition method, device, electronic device, and storage medium, which can efficiently and accurately detect and recognize underwater targets.
[0004] On the one hand, an embodiment of the present invention provides an underwater target detection and recognition method, including the following steps:
[0005] Obtain underwater image data; the underwater image data carries annotation information for targets in the underwater environment;
[0006] Based on the underwater image data, optimize the parameters of a preset neural network model through a circular dynamic optimization algorithm to obtain a feature extraction model; wherein, the feature extraction model is feature image data according to the processing result of the underwater image data;
[0007] Based on the feature image data and the annotation information, perform target detection training on a preset region convolutional neural network to obtain a target detection model;
[0008] Use the feature extraction model and the target detection model to detect and recognize targets in the underwater image to be recognized, and obtain a target detection result.
[0009] Optionally, the method further includes the following steps:
[0010] Perform data augmentation on the underwater image data based on a generative adversarial network optimized by Riemannian manifold.
[0011] Optionally, the generative adversarial network includes a generator and a discriminator; the method further includes the following steps:
[0012] Initialize the first parameter of the generator and the second parameter of the discriminator;
[0013] Determine environmental parameters according to the annotation information, perform encoding processing on the environmental parameters to obtain an encoded vector; the environmental parameters include the number of targets in the underwater environment, the underwater depth, and the average size.
[0014] Combine the encoded vector with the random noise vector of the underwater image data as the input data of the generator.
[0015] Based on the first parameter, process the input data through the generator to obtain a generated image; construct a generator loss based on the generated image.
[0016] Based on the second parameter, evaluate the underwater image data and its corresponding generated image through the discriminator, and construct a discriminator loss according to the evaluation results.
[0017] Construct a curvature regularization loss based on the generated image.
[0018] Based on the generator loss, the curvature regularization loss, and the discriminator loss, update the first parameter and the second parameter using backpropagation and the gradient descent method; the gradient applied by the gradient descent method is obtained through the chain rule.
[0019] Return to execute the step of determining the environmental parameters according to the annotation information for repeated iteration until the first stop iteration condition is satisfied.
[0020] Optionally, the process of parameter optimization includes a growth period, a synthesis period, an evaluation period, a prophase of division, a division period, and a dormancy period; optimize the parameters of a preset neural network model through a circular dynamic optimization algorithm to obtain a feature extraction model, including the following steps:
[0021] Initialize the neural network model using a convolutional neural network; among them, the neural network model includes multiple convolutional layers, pooling layers, fully connected layers, and deconvolutional layers.
[0022] Initialize the model parameters of the neural network model based on random values of a preset normal distribution; the model parameters include weights and biases.
[0023] In the growth period, update the model parameters using the accelerated gradient descent algorithm.
[0024] In the synthesis period, synthesize multiple weights through a synthesis function based on a preset importance coefficient.
[0025] In the evaluation period, obtain the evaluation results of the model parameters based on a preset performance evaluation function; when the evaluation results do not meet the preset performance threshold, return to execute the steps of the growth period, otherwise, enter the steps of the prophase of division.
[0026] In the prophase of division, use a preset perturbation method to obtain the adjustment increments of the weights and biases to fine-tune the model parameters.
[0027] During the division phase, weight pruning is used to constrain the weights;
[0028] During the dormancy phase, the learning rate is adjusted downward based on a preset decay coefficient;
[0029] Return to the steps of the growth phase for repeated iteration until the second iteration stop condition is met.
[0030] Optionally, the region convolutional neural network includes a feature secondary extractor, a region proposal network, and a classification and regression network; based on the feature image data combined with the annotation information, target detection training is performed on the preset region convolutional neural network to obtain a target detection model, including the following steps:
[0031] The feature image data is adjusted by the feature secondary extractor to obtain a feature map;
[0032] Based on the feature map and its true class at the corresponding sample of the feature image data, a feature secondary extractor loss is constructed; the true class is obtained according to the annotation information;
[0033] The region proposal network is used to extract candidate regions on the feature map through a sliding window, and the first predicted bounding box is obtained using the anchor box mechanism;
[0034] Based on the prediction result probability and true label of the candidate region, a first classification loss is constructed; based on the first predicted bounding box and the true bounding box, a first regression loss is constructed; the true label and true bounding box are obtained according to the annotation information;
[0035] Based on the first classification loss and the first regression loss, a region proposal network loss is constructed;
[0036] The adaptive hierarchical fusion strategy is used to fuse the feature maps of the candidate regions to obtain fused features;
[0037] The fused features are input into the classification and regression network to obtain a prediction output; the prediction output includes the class prediction probability and the second predicted bounding box of each candidate region;
[0038] Based on the class prediction probability and true class corresponding to each candidate region, a second classification loss is constructed; based on the second predicted bounding box and the true bounding box, a second regression loss is constructed;
[0039] Based on the second classification loss and the second regression loss, a classification and regression network loss is constructed;
[0040] Based on the feature secondary extractor loss, the region proposal network loss, and the classification and regression network loss, the network weights of the feature secondary extractor, the region proposal network, and the classification and regression network are optimized and adjusted accordingly;
[0041] Return and repeat the step of performing image feature adjustment on the feature image data through the feature secondary extractor until the third stop iteration condition is met.
[0042] Optionally, the region convolutional neural network includes a feature secondary extractor, a region proposal network, and a classification and regression network; using the feature extraction model and the object detection model, perform object detection and recognition on the underwater image to be recognized, and obtain the object detection result, including the following steps:
[0043] Perform feature extraction on the underwater image to be recognized through the feature extraction model to obtain a feature image;
[0044] Based on the feature image, generate a feature map using the feature secondary extractor;
[0045] Use the region proposal network to analyze and process the feature map to generate candidate target regions;
[0046] Perform adaptive hierarchical fusion on the candidate target regions, and then use the classification and regression network to perform object category recognition to obtain the object detection result.
[0047] Optionally, the object detection result includes the recognition categories of multiple detection objects in the underwater image to be recognized; the method further includes the following steps:
[0048] Based on the recognition category of each detection object, summarize to obtain the number of objects of each recognition category.
[0049] On the other hand, an embodiment of the present invention provides an underwater object detection and recognition device, including:
[0050] The first module is used to obtain underwater image data; the underwater image data carries annotation information for objects in the underwater environment;
[0051] The second module is used to optimize the parameters of the preset neural network model through the annular dynamic optimization algorithm based on the underwater image data to obtain a feature extraction model; wherein, the feature extraction model is the feature image data according to the processing result of the underwater image data;
[0052] The third module is used to perform object detection training on the preset region convolutional neural network based on the feature image data combined with the annotation information to obtain an object detection model;
[0053] The fourth module is used to perform object detection and recognition on the underwater image to be recognized using the feature extraction model and the object detection model to obtain the object detection result.
[0054] Optionally, the device further includes:
[0055] The fifth module is used to perform data augmentation on the underwater image data based on the Riemannian manifold optimization-based generative adversarial network.
[0056] Optionally, the generative adversarial network includes a generator and a discriminator; the apparatus further includes a sixth module, which is specifically configured to perform the following operations:
[0057] Initialize the first parameter of the generator and the second parameter of the discriminator;
[0058] Determine environmental parameters according to the annotation information, perform encoding processing on the environmental parameters to obtain an encoded vector; the environmental parameters include the number of targets in the underwater environment, the underwater depth, and the average size;
[0059] Combine the encoded vector with the random noise vector of the underwater image data as the input data of the generator;
[0060] Based on the first parameter, the generator processes the input data to obtain a generated image according to the input data; construct a generator loss based on the generated image;
[0061] Based on the second parameter, the discriminator evaluates the underwater image data and its corresponding generated image, and constructs a discriminator loss according to the evaluation result;
[0062] Construct a curvature regularization loss based on the generated image;
[0063] Based on the generator loss, the curvature regularization loss, and the discriminator loss, update the first parameter and the second parameter using backpropagation and the gradient descent method; the gradient applied by the gradient descent method is obtained through the chain rule;
[0064] Return to execute the step of determining the environmental parameters according to the annotation information for repeated iteration until the first stop iteration condition is satisfied.
[0065] Optionally, the object detection result includes the recognition categories of multiple detection objects in the underwater image to be recognized; the apparatus further includes:
[0066] A seventh module, configured to summarize the number of targets of each recognition category based on the recognition category of each detection object.
[0067] On the other hand, an embodiment of the present invention provides an electronic device, including: a processor and a memory; the memory is used to store a program; the processor executes the program to implement the above-mentioned underwater target detection and recognition method.
[0068] On the other hand, an embodiment of the present invention provides a computer storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to implement the above-mentioned underwater target detection and recognition method when executed by the processor.
[0069] In an embodiment of the present invention, underwater image data is acquired; the underwater image data carries annotation information for a target in an underwater environment; based on the underwater image data, a preset neural network model is optimized for parameters through a circular dynamic optimization algorithm to obtain a feature extraction model; wherein, the feature extraction model generates feature image data according to the processing result of the underwater image data; based on the feature image data and in combination with the annotation information, a preset region convolutional neural network is trained for target detection to obtain a target detection model; the feature extraction model and the target detection model are used to detect and identify a target in an underwater image to be recognized, and a target detection result is obtained. In an embodiment of the present invention, by introducing an advanced deep learning architecture and an optimization algorithm, the processes of feature extraction and target detection are precisely controlled, achieving efficient and accurate target detection. Especially in a complex and dense underwater environment, the performance of the target detection model is effectively improved, realizing more accurate and efficient underwater target recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The accompanying drawings are used to provide a further understanding of the technical solutions of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solutions of the present invention, and do not constitute a limitation to the technical solutions of the present invention.
[0071] Figure 1 FIG. is a schematic diagram of an implementation environment for underwater target detection and recognition provided by an embodiment of the present invention;
[0072] Figure 2 FIG. is a schematic flowchart of an underwater target detection and recognition method provided by an embodiment of the present invention;
[0073] Figure 3 FIG. is a schematic diagram of an extended process of the underwater target detection and recognition method provided by an embodiment of the present invention;
[0074] Figure 4 FIG. is a schematic diagram of an expanded process for obtaining a feature extraction model provided by an embodiment of the present invention;
[0075] Figure 5 FIG. is a schematic diagram of an expanded process of step S300 provided by an embodiment of the present invention;
[0076] Figure 6 FIG. is a schematic diagram of an expanded process of step S400 provided by an embodiment of the present invention;
[0077] Figure 7 FIG. is a schematic diagram of the overall process of the underwater target detection and recognition method provided by an embodiment of the present invention;
[0078] Figure 8 FIG. is a schematic flowchart of the training process of the neural network algorithm based on circular dynamic optimization provided by an embodiment of the present invention;
[0079] Figure 9 Structural schematic diagram of an underwater target detection and recognition device provided by an embodiment of the present invention;
[0080] Figure 10 Structural schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0081] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0082] It should be noted that although functional module division is performed in the system schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the system or the sequence in the flowchart. The terms "first / S100", "second / S200", etc. in the description and claims and the above drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence.
[0083] Referring to "embodiment" herein means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase does not necessarily refer to the same embodiment at various positions in the description, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0084] For the convenience of understanding the technical solutions of the present invention, the proprietary technical features that may be applied in the present invention will be explained first:
[0085] Machine learning is a subfield of artificial intelligence (AI) that gives computer systems the ability to learn from experience and improve their performance without being explicitly programmed. This means that machine learning models can automatically identify patterns and regularities in data and use this knowledge to make decisions or predictions. The following are some technical materials on machine learning related to the present invention:
[0086] Data-driven: Machine learning relies on large amounts of data to train models. The data can be labeled (supervised learning), unlabeled (unsupervised learning), or obtained through interaction with the environment (reinforcement learning).
[0087] Model: In machine learning, a model is a mathematical representation of a real-world problem and they can learn from data. Common models include decision trees, neural networks, support vector machines, etc.
[0088] Learning algorithm: The learning algorithm defines how the model is adjusted or "learned". The algorithm operates on data and is optimized based on the performance of the model. For example, by minimizing the difference between the prediction and the actual result.
[0089] Evaluation: The performance of the model needs to be measured through some form of evaluation. This usually involves dividing the data into a training set and a test set, where the training set is used for learning and the test set is used to evaluate the model's ability to generalize to unseen data.
[0090] Overfitting and underfitting: Overfitting means that the model performs well on the training data but cannot generalize to new data. Underfitting means that the model does not learn the features of the data sufficiently and cannot make effective predictions.
[0091] It can be understood that the underwater target detection and recognition method provided by the embodiments of the present invention can be applied to any computer device with data processing and computing capabilities, and this computer device can be various types of terminals or servers. When the computer device in the embodiment is a server, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal is a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto.
[0092] As Figure 1 shown, it is a schematic diagram of an implementation environment provided by the embodiments of the present invention. Referring to Figure 1 , this implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be network-connected wirelessly or wiredly to complete data transmission and exchange.
[0093] The server 101 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0094] In addition, the server 101 can also be a node server in a blockchain network. Among them, the blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms.
[0095] The terminal 102 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal 102 and the server 101 can be directly or indirectly connected through wired or wireless communication means, and the embodiments of the present invention do not make limitations here.
[0096] Exemplarily based on Figure 1 the described implementation environment, an embodiment of the present invention provides an underwater target detection and recognition method. Taking the application of this underwater target detection and recognition method in the server 101 as an example for illustration, it can be understood that this underwater target detection and recognition method can also be applied to the terminal 102.
[0097] Referring to Figure 2 , Figure 2 is a flowchart of the underwater target detection and recognition method applied to the server provided by the embodiment of the present invention. The execution subject of this underwater target detection and recognition method can be any of the foregoing computer devices (including the server or the terminal). Referring to Figure 2 , this method includes the following steps:
[0098] S100. Obtain underwater image data;
[0099] Among them, the underwater image data carries annotation information for the target in the underwater environment;
[0100] Exemplarily, in some specific implementation manners, the present invention collects high-quality underwater image data and precisely annotates it to train the subsequent deep learning model. The underwater image data comes from a specific underwater environment. In one embodiment, it may include coral reefs, ancient cultural relics, marine organisms, etc. The collected image data is stored in the jpg format. In this embodiment, the pixel is 1920x1080, and the number of channels is 3 RGB channels.
[0101] Further, the collected data is annotated. The annotation method of the data can be manual annotation, the annotation information is the individual of the target, and the total number of underwater targets is calculated through the number of annotations in a single image.
[0102] Among them, in some embodiments, the method may further include the following steps: performing data augmentation on the underwater image data based on the generative adversarial network optimized by the Riemannian manifold.
[0103] Exemplarily, in some specific embodiments, the acquisition, annotation, and preprocessing of training data are time-consuming and labor-intensive, and insufficient training samples can easily lead to poor generalization ability of the model and affect the accuracy of the model. The present invention uses a generative adversarial network optimized based on Riemannian manifolds for data augmentation. The generative adversarial network optimized based on Riemannian manifolds consists of two parts: a generator (G) and a discriminator (D). The generator is responsible for creating realistic image data, and the task of the discriminator is to distinguish between the generated images and the real images. To improve the quality and diversity of the generated images, the present invention uses a curvature regularization loss function to enable the generator to generate more adaptable image data according to specific environmental parameters.
[0104] Among them, in some embodiments, the generative adversarial network includes a generator and a discriminator; as Figure 3 shown, the method may further include the following steps: T100. Initialize the first parameters of the generator and the second parameters of the discriminator; T200. Determine the environmental parameters according to the annotation information, perform encoding processing on the environmental parameters to obtain an encoded vector; the environmental parameters include the number of targets in the underwater environment, the underwater depth, and the average size; T300. Combine the encoded vector with the random noise vector of the underwater image data as the input data of the generator; T400. Based on the first parameters, process the input data through the generator to obtain a generated image; construct a generator loss based on the generated image; T500. Based on the second parameters, evaluate the underwater image data and its corresponding generated image through the discriminator, and construct a discriminator loss according to the evaluation results; T600. Construct a curvature regularization loss based on the generated image; T700. Based on the generator loss, the curvature regularization loss, and the discriminator loss, update the first parameters and the second parameters using backpropagation and the gradient descent method; the gradient applied by the gradient descent method is obtained through the chain rule; T800. Return to execute the step of determining the environmental parameters according to the annotation information for iterative repetition until the first stop iteration condition is satisfied.
[0105] Exemplarily, in some specific embodiments, the training process of the generative adversarial network algorithm optimized based on Riemannian manifolds is as follows:
[0106] T1. Initialize the parameters of the generator and the discriminator. In one embodiment, the initialization method is expressed as:
[0107]
[0108] In the formula, and are the initial parameters of the generator and the discriminator respectively; represents a normal distribution with a mean of 0 and a standard deviation of 0.02 when initializing the parameters.
[0109] T2. Encode the environmental parameters into a vector, combine it with the random noise vector of the image data, and use it as the input of the generator, which is expressed as:
[0110] v env = E encode (env)
[0111] z = z random ⊕ v env
[0112] In the formula, v env is the encoded vector of the environmental parameter env; E encode () is the environmental parameter encoding function; z random is the random noise vector; ⊕ represents the concatenation operation of vectors; z is the input vector of the generator.
[0113] In one embodiment, the environmental parameter env includes the number of underwater environmental cultural relics contained in each image on average, the underwater depth, the average size of underwater cultural relic target, etc. In this embodiment, the E encode () function is an autoencoder neural network.
[0114] T3. The generator generates an image according to the input noise and conditional vector, and the generation method is expressed as:
[0115] x gen = G(z; Θ g )
[0116] Moreover, the loss calculation method of the generator is:
[0117] L G = -logD(x gen ; Θ d )
[0118] Furthermore, the discriminator evaluates the generated image and the real image, provides feedback to the generator, and the loss calculation method of the discriminator is:
[0119] L D = -[logD(x real ; Θ d ) + log(1 - D(x gen ; Θ d ))]
[0120] In the formula, x gen is the image generated by the generator G; G() is the generator function; D() is the discriminator function; x real is the real image; L D is the loss function of the discriminator; L G is the loss function of the generator; Θ g and Θ dThey are the parameters of the generator and the discriminator respectively.
[0121] T4. Use the curvature regularization loss function and the adversarial loss to jointly optimize the generator and the discriminator. The curvature regularization loss helps to maintain the geometric continuity and visual authenticity of the generated images. The calculation method is as follows:
[0122]
[0123] In the formula, L curve is the curvature regularization loss; λ curve is the weight of the curvature regularization; represents the second-order derivative of the image, that is, the Laplacian operator; ∑ i,j is the accumulation of the image pixels with indices i and j.
[0124] T5. Update the parameters of the generator and the discriminator using backpropagation and the gradient descent method. The update method is expressed as:
[0125]
[0126] In the formula, and are the updated parameters of the generator and the discriminator. η g and η d are the learning rates of the generator and the discriminator; and are the gradients of the loss functions of the generator and the discriminator with respect to their respective parameters. Preferably, η g and η d are set to 0.01 and 0.03 respectively.
[0127] Furthermore, the gradient is calculated by the chain rule and is expressed as:
[0128]
[0129] In the formula, and represent the partial derivatives of the generated image x gen with respect to the generation loss and the curvature regularization loss, and represents the partial derivative of the generated image with respect to the generator parameters.
[0130] T6. Repeat the above steps iteratively until the preset stop iteration condition is satisfied, which means the model training is completed. In one embodiment, the preset stop iteration condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations can be set to 1000 times.
[0131] S200. Optimize the parameters of a preset neural network model based on underwater image data through a circular dynamic optimization algorithm to obtain a feature extraction model;
[0132] Among them, the feature extraction model is the feature image data according to the processing result of the underwater image data;
[0133] It should be noted that the process of parameter optimization includes a growth period, a synthesis period, an evaluation period, a pre - division period, a division period, and a dormancy period; in some embodiments, as Figure 4 shown, optimizing the parameters of a preset neural network model through a circular dynamic optimization algorithm to obtain a feature extraction model may include the following steps: S201. Initialize the neural network model using a convolutional neural network; among them, the neural network model includes multiple convolutional layers, pooling layers, fully - connected layers, and de - convolutional layers; S202. Initialize the model parameters of the neural network model based on random values of a preset normal distribution; the model parameters include weights and biases; S203. In the growth period, use the accelerated gradient descent algorithm to update the model parameters; S204. In the synthesis period, synthesize multiple weights through a synthesis function based on a preset importance coefficient; S205. In the evaluation period, obtain the evaluation result of the model parameters based on a preset performance evaluation function; when the evaluation result does not meet the preset performance threshold, return to execute the steps in the growth period, otherwise, enter the steps in the pre - division period; S206. In the pre - division period, use a preset perturbation method to obtain the adjustment increments of the weights and biases to finely adjust the model parameters; S207. In the division period, use weight pruning to constrain the weights; S208. In the dormancy period, decrease the learning rate based on a preset decay coefficient; S209. Return to execute the steps in the growth period for iterative repetition until the second stop iteration condition is met.
[0134] Exemplarily, in some specific embodiments, the augmented data is input into the feature extraction model for training the feature extraction model. The traditional gradient descent method relies on global gradient information for weight update. Inspired by the periodic metabolic process of biological cells, different mechanisms are used in different stages to optimize the responses and adaptability of organisms. The present invention uses a circular dynamic optimization algorithm to optimize the parameters of the neural network. The circular dynamic optimization algorithm adjusts the optimization strategy in different training stages, simulating each stage in the cell cycle, and using different strategies in each stage to adjust the weights and biases of the neural network, enabling the algorithm to more flexibly adapt to the changes in training data, reducing the risk of overfitting, and effectively jumping out of local optimal solutions.
[0135] S300. Based on the feature image data combined with annotation information, perform object detection training on a preset region - based convolutional neural network to obtain an object detection model;
[0136] It should be noted that the region convolutional neural network includes a feature secondary extractor, a region proposal network, and a classification and regression network; in some embodiments, as Figure 5 shown, step S300 may include the following steps: S301. Adjust the image features of the feature image data through the feature secondary extractor to obtain a feature map; S302. Construct a feature secondary extractor loss based on the feature map and its true category at the corresponding sample of the feature image data; the true category is obtained according to the annotation information; S303. Use the region proposal network to extract candidate regions on the feature map through a sliding window, and obtain a first predicted bounding box using the anchor box mechanism; S304. Construct a first classification loss based on the prediction result probability and the true label of the candidate region; construct a first regression loss based on the first predicted bounding box and the true bounding box; the true label and the true bounding box are obtained according to the annotation information; S305. Construct a region proposal network loss based on the first classification loss and the first regression loss; S306. Use an adaptive hierarchical fusion strategy to perform feature layer fusion on the feature map of the candidate region to obtain a fused feature; S307. Input the fused feature into the classification and regression network to obtain a prediction output; the prediction output includes the class prediction probability and the second predicted bounding box of each candidate region; S308. Construct a second classification loss based on the class prediction probability and the true category corresponding to each candidate region; construct a second regression loss based on the second predicted bounding box and the true bounding box; S309. Construct a classification and regression network loss based on the second classification loss and the second regression loss; S310. Optimize and adjust the network weights of the feature secondary extractor, the region proposal network, and the classification and regression network based on the feature secondary extractor loss, the region proposal network loss, and the classification and regression network loss; S311. Return to execute the step of adjusting the image features of the feature image data through the feature secondary extractor for iterative repetition until the third iteration stop condition is met.
[0137] Exemplarily, in some specific embodiments, the data after feature extraction is input into a target detection model for target detection. The present invention uses a region convolutional neural network based on deep feature space mapping to accurately detect targets in an underwater dense scene. The region convolutional neural network based on deep feature space mapping includes three components: a feature secondary extractor, a region proposal network, and a classification and regression network.
[0138] Specifically, the training process of the region convolutional neural network algorithm based on deep feature space mapping can be implemented as follows:
[0139] S311. Use a pre-trained deep convolutional network as the feature secondary extractor and fine-tune the image after feature extraction. Specifically, optimize the network weight W qe to minimize the classification loss L qe , expressed as:
[0140]
[0141] In the formula, N is the number of training samples, y i is the true category of the i-th sample, x i is the corresponding input feature, P(y i |x i ; W qe ) represents the probability that the model predicts that x qe belongs to the category y i under the current weight W i .
[0142] S312. Train the region proposal network to generate high-quality candidate regions. The region proposal network extracts candidate regions on the feature map through a sliding window and uses the anchor box mechanism to predict the object boundary. Specifically, for the region proposal network, the goal is to learn a set of weights W qr such that the generated candidate regions are as close as possible to the true target regions. The loss function L qr of the region proposal network is calculated as follows:
[0143]
[0144] In the formula, q i is the probability that the candidate region i is predicted as the target, is the true label (1 for the target and 0 for the background), t i is the predicted bounding box parameter, is the true bounding box parameter, L cls and L reg are the classification loss and the regression loss respectively, and λ qr is the coefficient for balancing the two losses.
[0145] In one embodiment, the classification loss L cls is calculated using the cross-entropy loss and is expressed as:
[0146]
[0147] In the formula, q i is the probability of the predicted target existing, is the true label (1 represents the target and 0 represents the background).
[0148] Moreover, in this embodiment, the regression loss L reg is implemented using the smooth L1 loss and is expressed as:
[0149]
[0150] In the formula, t i and are the predicted and true bounding box parameters respectively.
[0151] S313. After feature extraction and candidate region generation, an adaptive hierarchical fusion strategy is adopted. The adaptive hierarchical fusion strategy dynamically selects appropriate feature layers for fusion by analyzing the size and shape complexity of the candidate regions, so as to optimize the detection effect for small targets and partially occluded targets. In one embodiment, the adaptive hierarchical fusion is achieved by calculating the weight α of feature maps at different levels. ql The weight is dynamically adjusted through the Softmax function to optimize the quality of the fused features. The calculation method is expressed as:
[0152]
[0153] In the formula, f l represents the feature map of the l-th layer, W ql is the corresponding weight matrix, and k is the total number of feature levels.
[0154] Moreover, in this embodiment, the weight matrix W ql is carried out through an optimization process. The goal is to minimize the loss between the fused feature map and the target detection task. Specifically, the weights are adjusted through the following optimization objective:
[0155]
[0156] In the formula, M is the total number of samples processed during the hierarchical fusion process, represents the feature calculated from the i-th sample through the feature map of the l-th layer, and y i is the corresponding ground truth label.
[0157] S314. Using the fused features, train a classification and regression network to accurately determine the category of each candidate region and adjust its position. This network not only identifies the target type but also refines the position and size of the target, improving the detection accuracy. The loss function L qc of the classification and regression network includes classification loss and bounding box regression loss, and the calculation method is:
[0158]
[0159] In the formula, c j is the predicted probability of the category of the j-th candidate region, is the ground truth category, b j is the predicted bounding box parameter, is the ground truth bounding box parameter, and λ qc is the balance coefficient.
[0160] S315. Repeat the above steps iteratively until the preset iteration stop condition is met, indicating that the model training is completed. In one embodiment, the preset iteration stop condition is reaching the preset maximum number of iterations. Preferably, the preset maximum number of iterations can be set to 1000 times.
[0161] S400. Use the feature extraction model and the target detection model to detect and identify the objects in the underwater image to be recognized, and obtain the target detection result.
[0162] It should be noted that the region convolutional neural network includes a feature secondary extractor, a region proposal network, and a classification and regression network. In some embodiments, as Figure 6 shown, step S400 may include the following steps: S401. Extract features from the underwater image to be recognized through the feature extraction model to obtain a feature image; S402. Based on the feature image, use the feature secondary extractor to generate a feature map; S403. Use the region proposal network to analyze and process the feature map to generate candidate target regions; S404. Perform adaptive hierarchical fusion on the candidate target regions, and then use the classification and regression network to perform target category recognition to obtain the target detection result.
[0163] Exemplarily, in some specific embodiments, the above trained model can be used to process new underwater image data. In one embodiment, for a newly acquired underwater image sample, first, the feature-extracted image data is obtained through feature extraction. Further, a feature map of the image is generated through feature secondary extraction. Further, the region proposal network is used to analyze the feature map to generate candidate target regions. Further, adaptive hierarchical fusion and the classification and regression network are used to perform target category recognition and position fine-tuning for each candidate region.
[0164] Among them, in some embodiments, the target detection result includes the recognition categories of multiple detection targets in the underwater image to be recognized. The method may further include the following steps: Based on the recognition categories of each detection target, summarize the number of targets for each recognition category.
[0165] Exemplarily, in some specific embodiments, the detected targets can be counted based on the recognition results and summarized.
[0166] To explain the principle of the technical solution of the present invention in detail, the overall process of the present invention will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.
[0167] First of all, it should be noted that a related technology proposes an underwater target recognition model and method based on laser image enhancement and restoration, belonging to the technical field of underwater target image enhancement and detection. The model structure includes an image enhancement module, a feature extraction and fusion module, and an underwater target recognition module connected in sequence; the underwater target recognition technology based on laser image enhancement and restoration uses the RGB intensity data obtained under single-frequency laser irradiation to recognize underwater targets and generate images, and then uses image enhancement and image restoration methods to enhance and restore the images to improve the visibility and recognizability of the targets, which makes the resolution and quality of the images higher, thus improving the visibility and recognizability of the targets. And the image processing algorithm used in the present invention can better remove the noise and interference in the images, thereby improving the recognition accuracy and recognition speed of the targets.
[0168] Another related technology proposes a method for recognizing forward-looking sonar images based on a convolutional neural network with data increment. This method uses five different underwater moving targets, namely single-column target, double-column target, triple-column target, quadruple-column target, and T-shaped target. The convolutional neural network forward-looking sonar image recognition technology is as follows: Step 1, extract the training set and test set from the original forward-looking sonar target images; Step 2, center-crop the target images and rotate them grayscale; Step 3, data increment; Step 4, input into the convolutional neural network for training; Step 5, obtain the trained convolutional neural network; Step 6, perform recognition and classification on the test set; The present invention does not require cumbersome feature engineering, can greatly reduce the labor cost and has stronger generalization ability and faster training speed.
[0169] There is also a related technology that proposes an underwater target classification method based on an improved grey wolf optimization algorithm, belonging to the field of underwater target recognition. The underwater target classification method first extracts features from the underwater target images through the principal component analysis method and realizes data dimensionality reduction; secondly, uses SVM classification to perform underwater target classification tasks; finally, adopts an improved grey wolf optimization algorithm to optimize the parameters of the support vector machine to achieve better classification effects, that is, optimizes and improves the intelligent optimization algorithm used by the SVM to improve the classification accuracy and efficiency. Using the underwater target classification method of the present invention can effectively improve the accuracy of target classification.
[0170] However, the disadvantages of the related technologies are as follows:
[0171] 1. In traditional methods, data scarcity and insufficient sample diversity often lead to insufficient model training. Especially in specific and complex environments such as underwater scenes, there is a lack of sufficiently diverse image data, which affects the generalization ability and accuracy of the model.
[0172] 2. Traditional feature extraction methods rely on static optimization strategies and fail to effectively address the dynamic changes in training data and overfitting problems, resulting in insufficient adaptability and accuracy of the model in practical applications.
[0173] 3. Existing object detection methods are inefficient in dealing with small objects and partially occluded objects, and cannot accurately identify multiple objects in dense scenes. Especially in complex underwater environments, detection performance is often poor due to inaccurate feature extraction and unreasonable object proposals.
[0174] Therefore, developing an effective underwater object detection and counting method that can improve the accuracy of recognition and the efficiency of processing is of great significance for promoting underwater scientific research and the development of related technologies. Specifically, the present invention aims to solve these technical problems and achieve more accurate and efficient underwater object recognition and counting by introducing advanced deep learning architectures and optimization algorithms. As Figure 7 shown, the method flow of the embodiments of the present invention can be implemented as follows:
[0175] S1. Data collection and annotation:
[0176] The present invention collects high-quality underwater image data and precisely annotates it to train subsequent deep learning models. The underwater image data is sourced from specific underwater environments. In one embodiment, it may include coral reefs, ancient cultural relics, marine organisms, etc. The collected image data is stored in jpg format. In this embodiment, the pixel size is 1920x1080, and the number of channels is RGB3 channels.
[0177] Furthermore, the collected data is annotated. The annotation method of the data is manual annotation, and the annotation information is the individuals of the objects. The total number of underwater objects is calculated based on the annotation quantity of a single image.
[0178] In one embodiment, there is a piece of underwater image data that contains 2 underwater cultural relic objects. In this embodiment, the coordinates of the 2 underwater cultural relic objects are annotated, and the image is annotated as:
[0179] Object 1: {x1, y1, x2, y2};
[0180] Object 2: {x3, y3, x4, y4};
[0181] S2. Data augmentation:
[0182] It is understandable that in the task of the present invention, the acquisition, annotation, and preprocessing of training data are time-consuming and laborious, and insufficient training samples easily lead to poor generalization ability of the model and affect the accuracy of the model. The present invention uses a generative adversarial network based on Riemannian manifold optimization for data augmentation. The generative adversarial network based on Riemannian manifold optimization consists of two parts: a generator (G) and a discriminator (D). The generator is responsible for creating realistic image data, and the task of the discriminator is to distinguish between the generated images and the real images. In order to improve the quality and diversity of the generated images, the present invention uses a curvature regularization loss function to enable the generator to generate more adaptable image data according to specific environmental parameters.
[0183] Specifically, the training process of the generative adversarial network algorithm based on Riemannian manifold optimization is as follows:
[0184] S21. Initialize the parameters of the generator and the discriminator. In one embodiment, the initialization method is expressed as:
[0185]
[0186] In the formula, and are the initial parameters of the generator and the discriminator respectively; represents a normal distribution with a mean of 0 and a standard deviation of 0.02 when initializing the parameters.
[0187] S22. Encode the environmental parameters into a vector and merge it with the random noise vector of the image data as the input of the generator, which is expressed as:
[0188] v env = E encode (env)
[0189] z = z random ⊕ v env
[0190] In the formula, v env is the encoded vector of the environmental parameter env; E encode () is the environmental parameter encoding function; z random is the random noise vector; ⊕ represents the concatenation operation of vectors; z is the input vector of the generator.
[0191] In one embodiment, the environmental parameter env includes the number of underwater cultural relics in each image on average, the underwater depth, the average size of underwater cultural relic target, etc. In this embodiment, the E encode () function is an autoencoder neural network.
[0192] S23. The generator generates images according to the input noise and conditional vector, and the generation method is expressed as:
[0193] x gen = G(z; Θ g )
[0194] And the loss calculation method of the generator is as follows:
[0195] L G = -logD(x gen ; Θ d )
[0196] Furthermore, the discriminator evaluates the generated image and the real image and provides feedback to the generator. The loss calculation method of the discriminator is as follows:
[0197] L D = -[logD(x real ; Θ d ) + log(1 - D(x gen ; Θ d ))]
[0198] In the formula, x gen is the image generated by the generator G; G() is the generator function; D() is the discriminator function; x real is the real image; L D is the loss function of the discriminator; L G is the loss function of the generator; Θ g and Θ d are the parameters of the generator and the discriminator respectively.
[0199] S24. Use the curvature regularization loss function and the adversarial loss to jointly optimize the generator and the discriminator. The curvature regularization loss helps to maintain the geometric continuity and visual authenticity of the generated image. The calculation method is as follows:
[0200]
[0201] In the formula, L curve is the curvature regularization loss; λ curve is the weight of the curvature regularization; represents the second-order derivative of the image, that is, the Laplace operator; ∑ i,j is the accumulation of the image pixels with indices i, j.
[0202] S25. Use backpropagation and the gradient descent method to update the parameters of the generator and the discriminator. The update method is expressed as:
[0203]
[0204] In the formula, and are the updated parameters of the generator and the discriminator, η g and ηd are the learning rates of the generator and the discriminator; and are the gradients of the loss functions of the generator and the discriminator with respect to their respective parameters. Preferably, η g and η d are set to 0.01 and 0.03 respectively.
[0205] Furthermore, the gradient is calculated by the chain rule and is expressed as:
[0206]
[0207] In the formula, and represent the partial derivatives of the generated image x gen with respect to the generation loss and the curvature regularization loss, represents the partial derivative of the generated image with respect to the generator parameters.
[0208] S26. Repeat the above steps iteratively until the preset iteration stop condition is met, which indicates that the model training is completed. In one embodiment, the preset iteration stop condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.
[0209] S3. Feature extraction model training:
[0210] Input the augmented data into the feature extraction model for training the feature extraction model. The traditional gradient descent method relies on global gradient information for weight update. Inspired by the periodic metabolic process of biological cells, different mechanisms are used at different stages to optimize the responses and adaptabilities of organisms. The present invention uses a circular dynamic optimization algorithm to optimize the parameters of the neural network. The circular dynamic optimization algorithm adjusts the optimization strategy at different training stages, simulating each stage in the cell cycle, and different strategies are used in each stage to adjust the weights and biases of the neural network, enabling the algorithm to more flexibly adapt to the changes in training data, reducing the risk of overfitting, and effectively jumping out of local optimal solutions.
[0211] Specifically, as Figure 8 shown, the training process of the neural network algorithm based on circular dynamic optimization is as follows:
[0212] S31. Initialize the neural network model. The present invention uses a convolutional neural network for feature extraction, specifically including 3 convolutional layers, 1 pooling layer, 1 fully connected layer, and 1 deconvolutional layer.
[0213] In one embodiment, the sample input into the feature extraction model is an underwater image data with a pixel of 1920x1080 and 3 RGB channels.
[0214] Moreover, through the deconvolution layer of the neural network model, the finally output data after feature extraction is single-channel image data with a pixel size of 1920x1080.
[0215] Furthermore, initialize the weights w r and biases b r , and set the initial learning rate In one embodiment, the setting method of the initial weights and biases is expressed as:
[0216]
[0217] In the formula, ~ means subject to a specific distribution, is a normal distribution with a mean of 0 and a variance of σ r 2 , and σ r 2 is the variance of the initial distribution. Preferably, σ r 2 is set to 0.01.
[0218] S32. During the growth period, the algorithm quickly adapts to the data characteristics by accelerating the fine-tuning of the weights. The present invention uses the accelerated gradient descent algorithm to update the neural network parameters, and the update method is expressed as:
[0219]
[0220] In the formula, γ r is the acceleration coefficient, L R is the loss function of the neural network training, is the gradient of the loss function with respect to the weight parameter, is the gradient of the loss function with respect to the bias parameter, is the weight parameter of the t-th iteration, is the bias parameter of the t-th iteration, is the weight parameter of the (t + 1)-th iteration, is the bias parameter of the (t + 1)-th iteration, is the partial derivative symbol, η r is the learning rate during the growth period. Preferably, the loss function L R uses cross-entropy loss, and η r is set to 0.01.
[0221] Furthermore, the acceleration coefficient γ r is set according to the descent speed and curvature of the loss function to avoid over-adjustment, and the calculation method is:
[0222]
[0223] In the formula, is the preset maximum acceleration coefficient, and min(,) is the minimum value function.
[0224] S33. During the synthesis period, the algorithm performs weight synthesis at this stage, that is, multiple weight parameters are synthesized into new parameters through specific operations to improve the generalization ability of the network. In one embodiment, the synthesis function S(w r ) is calculated as follows:
[0225]
[0226] In the formula, is the importance coefficient of the weight , and nr is the number of weight parameters.
[0227] Furthermore, the importance coefficient α r is dynamically calculated based on the gradient magnitude of w r in the previous stage. For the i-th weight parameter, the calculation method of its importance coefficient is expressed as:
[0228]
[0229] In the formula, σ wr is the decay speed parameter.
[0230] Furthermore, the calculation method of the decay speed parameter σ wr is as follows:
[0231]
[0232] S34. During the evaluation period, the algorithm evaluates the current network performance and decides whether to adjust the cycle optimization strategy. In one embodiment, a performance evaluation function P(w r , b r ) is set, and it outputs whether to enter the next cycle in the form of a decision function. The decision method is expressed as:
[0233]
[0234] In the formula, θ r is the performance threshold. Preferably, θ r is set to 0.5. If P(w r , b r ) is calculated as 1, then enter the next cycle, that is, enter the growth period; otherwise, enter the prophase of division.
[0235] S35. During the prophase of division, simulate the preparatory division stage of the cell and make fine adjustments to the weights. In one embodiment, the perturbation methods Δw r and Δb rAs the adjustment increments of the weights and biases, the calculation method is expressed as:
[0236]
[0237] In the formula, Δw r is the adjustment increment of the weight parameter, Δb r is the adjustment increment of the bias parameter, ξ r and β r are the amplitude and frequency parameters of fine-tuning. Preferably, ξ r is set to 5, and β r is set to 3.14.
[0238] Furthermore, the setting method of ξ r and β r is expressed as:
[0239]
[0240] In the formula, ρ r is the scaling factor of the perturbation amplitude. Preferably, ρ r is set to 0.95.
[0241] S36. During the splitting period, the algorithm performs a large structural adjustment. In one embodiment, the adjustment method is to use weight pruning to constrain the weights. The calculation method of the weight pruning function R(w r ) is expressed as:
[0242]
[0243] In the formula, λ r is the gradient threshold. Preferably, λ r is set to 0.01, that is, only the parameters with weights greater than 0.01 are retained.
[0244] S37. During the dormant period, the network gradually reaches a stable state based on the existing weights. In this stage, the learning rate η r is adjusted so that the learning rate decreases. The adjustment method is expressed as::
[0245]
[0246] In the formula, δ r is the attenuation coefficient, and k r is the number of completed cycles. Preferably, δ r is set to 0.95.
[0247] S38. Repeat the above steps iteratively until the preset stop iteration condition is met, which indicates that the model training is completed. In one embodiment, the preset stop iteration condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.
[0248] S4. Train the object detection model:
[0249] Input the data after feature extraction into the object detection model for object detection. The present invention uses a region-based convolutional neural network based on deep feature space mapping to accurately detect objects in an underwater dense scene. The region-based convolutional neural network based on deep feature space mapping includes three components: a feature secondary extractor, a region proposal network, and a classification and regression network.
[0250] Specifically, the training process of the region-based convolutional neural network algorithm based on deep feature space mapping is as follows:
[0251] S41. Use a pre-trained deep convolutional network as the feature secondary extractor and fine-tune the image after feature extraction. Specifically, optimize the network weight W qe to minimize the classification loss L qe , expressed as:
[0252]
[0253] In the formula, N is the number of training samples, y i is the true category of the i-th sample, x i is the corresponding input feature, P(y i |x i ; W qe ) represents the probability that the model predicts that x qe belongs to the category y i under the current weight W i .
[0254] S42. Train the region proposal network to generate high-quality candidate regions. The region proposal network extracts candidate regions on the feature map through a sliding window and uses the anchor box mechanism to predict the object boundary. Specifically, for the region proposal network, the goal is to learn a set of weights W qr such that the generated candidate regions are as close as possible to the true target regions. The loss function L qr of the region proposal network is calculated as:
[0255]
[0256] In the formula, q i is the probability that the candidate region i is predicted as the target, is the true label (the target is 1 and the background is 0), t iis the predicted bounding box parameter, is the ground truth bounding box parameter, L cls and L reg are the classification loss and regression loss respectively, and λ qr is the coefficient for balancing the two losses.
[0257] In one embodiment, the classification loss L cls is calculated using cross-entropy loss and is expressed as:
[0258]
[0259] In the formula, q i is the probability of the presence of the predicted target, is the ground truth label (1 represents the target and 0 represents the background).
[0260] Moreover, in this embodiment, the regression loss L reg is implemented using smooth L1 loss and is expressed as:
[0261]
[0262] In the formula, t i and are the predicted and ground truth bounding box parameters respectively.
[0263] S43. After feature extraction and candidate region generation, an adaptive hierarchical fusion strategy is adopted. The adaptive hierarchical fusion strategy dynamically selects appropriate feature layers for fusion by analyzing the size and shape complexity of the candidate regions to optimize the detection effect for small targets and partially occluded targets. In one embodiment, the adaptive hierarchical fusion is achieved by calculating the weights α ql of the feature maps at different levels. The weights are dynamically adjusted through the Softmax function to optimize the quality of the fused features, and the calculation method is expressed as:
[0264]
[0265] In the formula, f l represents the feature map of the l-th level, W ql is the corresponding weight matrix, and k is the total number of feature levels.
[0266] Moreover, in this embodiment, the weight matrix W ql is obtained through an optimization process. The goal is to minimize the loss between the fused feature map and the target detection task. Specifically, the weights are adjusted through the following optimization objective:
[0267]
[0268] In the formula, M is the total number of samples processed during the hierarchical fusion process, denote the features obtained by calculating the feature map of the i-th sample through the l-th layer, and y i is the corresponding true label.
[0269] S44. Using the fused features, train a classification and regression network to accurately determine the category of each candidate region and adjust its position. This network not only identifies the target type but also refines the position and size of the target to improve the detection accuracy. The loss function L of the classification and regression network qc includes classification loss and bounding box regression loss, and the calculation method is as follows:
[0270]
[0271] In the formula, c j is the predicted probability of the category of the j-th candidate region, is the true category, b j is the predicted bounding box parameter, is the true bounding box parameter, and λ qc is the balance coefficient.
[0272] S45. Repeat the above steps iteratively until the preset stop iteration condition is met, which indicates that the model training is completed. In one embodiment, the preset stop iteration condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.
[0273] S5. Underwater target counting in dense scenes:
[0274] The above-trained model is used to process new underwater image data to achieve real-time target counting. In one embodiment, for newly collected underwater image samples, first, the image data after feature extraction is obtained through feature extraction. Further, the feature map of the image is generated through secondary feature extraction. Further, the region proposal network is used to analyze the feature map to generate candidate target regions. Further, the adaptive hierarchical fusion and classification and regression network are used to identify the target category and fine-tune the position of each candidate region. Further, the detected targets are counted and summarized.
[0275] In summary, the present invention uses a generative adversarial network optimized based on Riemannian manifolds to perform data augmentation, and optimizes the generator through a curvature regularization loss function, solving the problems of data scarcity and insufficient sample diversity, and effectively improving the adaptability of the model to complex underwater environments. Moreover, the present invention uses a circular dynamic optimization algorithm to train the feature extraction model, imitating different stages of the biological cell cycle, optimizing the weight and bias update strategy of the neural network, and effectively coping with the changes in training data and overfitting problems. In addition, the present invention also uses a region convolutional neural network based on deep feature space mapping for object detection, and through feature secondary extraction and adaptive hierarchical fusion strategies, enhances the detection ability for small objects and partially occluded objects, and improves the efficiency and accuracy of the region proposal network and the classification and regression network.
[0276] Compared with the prior art, the present invention has at least the following beneficial effects:
[0277] 1. The optimization of the generative adversarial network significantly improves the visual authenticity and geometric continuity of the generated underwater images, thus providing higher-quality training data for the deep learning model.
[0278] 2. The periodic optimization strategy of the neural network parameters reduces the overfitting risk of the model, and enables the model to effectively adapt to the changes in data characteristics in various training stages, improving the generalization ability of the model.
[0279] 3. By precisely controlling the process of feature extraction and object detection, efficient and accurate object detection is achieved, especially in complex and dense underwater environments, effectively improving the performance of the object detection model.
[0280] On the other hand, as Figure 9 shown, an underwater target detection and recognition device 900 is provided in an embodiment of the present invention, which may include:
[0281] A first module 901, configured to obtain underwater image data; the underwater image data carries annotation information for targets in the underwater environment;
[0282] A second module 902, configured to optimize the parameters of a preset neural network model through a circular dynamic optimization algorithm based on the underwater image data to obtain a feature extraction model; wherein, the feature extraction model is feature image data according to the processing result of the underwater image data;
[0283] A third module 903, configured to perform object detection training on a preset region convolutional neural network based on the feature image data in combination with the annotation information to obtain an object detection model;
[0284] The fourth module 904 is configured to use a feature extraction model and an object detection model to detect and identify objects in the underwater image to be recognized, and obtain an object detection result.
[0285] In some embodiments, the apparatus may further include:
[0286] A fifth module for data augmentation of underwater image data based on a generative adversarial network optimized by Riemannian manifold.
[0287] In some embodiments, the generative adversarial network includes a generator and a discriminator; the apparatus may further include a sixth module, and the sixth module is specifically configured to perform the following operations:
[0288] Initialize the first parameter of the generator and the second parameter of the discriminator;
[0289] Determine the environmental parameters according to the annotation information, perform encoding processing on the environmental parameters, and obtain an encoded vector; the environmental parameters include the number of objects in the underwater environment, the underwater depth, and the average size;
[0290] Combine the encoded vector with the random noise vector of the underwater image data as the input data of the generator;
[0291] Based on the first parameter, the generator processes the input data to obtain a generated image according to the input data; construct a generator loss based on the generated image;
[0292] Based on the second parameter, the discriminator evaluates the underwater image data and its corresponding generated image, and constructs a discriminator loss according to the evaluation result;
[0293] Construct a curvature regularization loss based on the generated image;
[0294] Based on the generator loss, the curvature regularization loss, and the discriminator loss, update the first parameter and the second parameter by using backpropagation and the gradient descent method; the gradient applied by the gradient descent method is obtained through the chain rule;
[0295] Return to execute the step of determining the environmental parameters according to the annotation information for repeated iteration until the first stop iteration condition is satisfied.
[0296] In some embodiments, the apparatus may further include:
[0297] A seventh module for summarizing the number of objects of each recognized category based on the recognized category of each detected object.
[0298] The content of the method embodiments of the present invention is applicable to the apparatus embodiments of the present invention. The functions specifically implemented by the apparatus embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above methods.
[0299] On the other hand, an embodiment of the present invention further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned underwater target detection and recognition method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0300] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0301] As Figure 10 shown, Figure 10 FIG. 1000 schematically shows the hardware structure of an electronic device 1000 according to another embodiment. The electronic device 1000 includes:
[0302] A processor 1001, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;
[0303] A memory 1002, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of the present specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute the network node population optimization method of the embodiments of the present invention;
[0304] An input / output interface 1003, which is used to implement information input and output;
[0305] A communication interface 1004, which is used to implement communication interaction between the device and other devices, and can communicate through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);
[0306] A bus 1005, which transmits information between various components of the device (such as the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004);
[0307] Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 are communicatively connected to each other inside the device through the bus 1005.
[0308] The electronic device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0309] The content of the method embodiments of the present invention is applicable to the electronic device embodiments of the present invention. The functions specifically implemented by the electronic device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method.
[0310] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium. The storage medium stores a program, and the program is executed by a processor to implement the foregoing method.
[0311] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0312] The content of the method embodiments of the present invention is applicable to the embodiments of this computer-readable storage medium. The functions specifically implemented by the embodiments of this computer-readable storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method.
[0313] The embodiments of the present invention also disclose a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and these computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method described above.
[0314] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0315] It should be noted that although several modules of devices for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present invention, the features and functions of two or more of the above-described modules or units may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by a plurality of modules or units.
[0316] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which may be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present invention.
[0317] In some alternative embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.
[0318] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0319] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0320] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a predefined sequence list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution device, apparatus, or equipment (such as a computer-based device, a device including a processor, or other devices that can fetch and execute instructions from the instruction execution device, apparatus, or equipment), or in combination with these instruction execution devices, apparatus, or equipment. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution device, apparatus, or equipment.
[0321] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0322] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution device. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0323] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0324] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
[0325] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. A method for underwater target detection and identification, characterized in that: The following steps are involved: Acquire underwater image data; the underwater image data carries annotation information for targets in the underwater environment; Based on the underwater image data, a preset neural network model is optimized by a circular dynamic optimization algorithm to obtain a feature extraction model; wherein the feature extraction model is feature image data according to a processing result of the underwater image data; Based on the feature image data and the annotation information, a preset regional convolutional neural network is trained for target detection to obtain a target detection model; The feature extraction model and the target detection model are used to perform target detection and recognition on the underwater image to be recognized, and a target detection result is obtained.
2. The underwater target detection and identification method according to claim 1, characterized in that: The method further comprises the following steps: The underwater image data is expanded based on a generative adversarial network optimized on a Riemannian manifold.
3. The underwater target detection and identification method according to claim 2, characterized in that: The generative adversarial network includes a generator and a discriminator; the method further includes the following steps: Initialize a first parameter of the generator and a second parameter of the discriminator; Determine environmental parameters according to the annotation information, and encode the environmental parameters to obtain an encoding vector; the environmental parameters include the number, underwater depth and average size of the targets in the underwater environment; Combining the encoding vector with a random noise vector of the underwater image data as input data of the generator; Based on the first parameter, obtaining a generated image by the generator according to the input data; constructing a generator loss based on the generated image; Based on the second parameter, the underwater image data and the corresponding generated image are evaluated by the discriminator, and a discriminator loss is constructed according to the evaluation result; constructing a curvature regularization loss based on the generated image; Based on the generator loss, the curvature regularization loss and the discriminator loss, the first parameter and the second parameter are updated by back propagation and gradient descent method; the gradient used by the gradient descent method is obtained by the chain rule; Return to the step of determining the environmental parameters according to the annotation information and repeat the iteration until the first stop iteration condition is met.
4. The underwater target detection and identification method according to claim 1, characterized in that: The parameter optimization process includes a growth period, a synthesis period, an evaluation period, a pre-mitotic period, a mitotic period, and a dormant period; the parameter optimization of the preset neural network model by the circular dynamic optimization algorithm to obtain a feature extraction model includes the following steps: Initializing the neural network model using a convolutional neural network; wherein the neural network model includes multiple layers of convolutional layers, pooling layers, fully connected layers, and deconvolution layers; Initializing model parameters of the neural network model based on a preset normal distributed random value; the model parameters include weights and biases; During the growth period, the model parameters are updated using an accelerated gradient descent algorithm; During the synthesis period, a plurality of the weights are synthesized by a synthesis function based on a preset importance coefficient; In the evaluation period, obtaining the evaluation result of the model parameters based on a preset performance evaluation function; when the evaluation result does not meet the preset performance threshold, returning to the step of executing the growth period, otherwise, entering the step of the pre-mitosis period; In the early stage of the split, a preset perturbation method is used to obtain adjustment increments of the weight and the bias to fine-tune the model parameters; During the splitting period, the weights are constrained by weight pruning; During the dormant period, the learning rate is adjusted downward based on a preset decay coefficient; Return to execute the steps of the growth phase and repeat the iteration until the second iteration stop condition is met.
5. The underwater target detection and identification method according to claim 1, characterized in that: The regional convolutional neural network includes a feature secondary extractor, a regional proposal network and a classification and regression network; the target detection training is performed on the preset regional convolutional neural network based on the feature image data combined with the annotation information to obtain a target detection model, including the following steps: Performing image feature adjustment on the feature image data by the feature secondary extractor to obtain a feature map; Constructing a feature secondary extractor loss based on the feature map and its true category at the sample corresponding to the feature image data; the true category is obtained according to the annotation information; Using the region proposal network to extract candidate regions on the feature map through a sliding window, and using an anchor box mechanism to obtain a first predicted bounding box; Constructing a first classification loss based on the prediction result probability of the candidate region and the true label; constructing a first regression loss based on the first predicted border and the true border; the true label and the true border are obtained according to the annotation information; A region proposal network loss is constructed based on the first classification loss and the first regression loss; Using an adaptive hierarchical fusion strategy to perform feature layer fusion on the feature map of the candidate area to obtain a fused feature; Inputting the fusion feature into the classification and regression network to obtain a prediction output; the prediction output includes a category prediction probability and a second prediction bounding box for each candidate region; Constructing a second classification loss based on the category prediction probability corresponding to each candidate region and the true category; constructing a second regression loss based on the second predicted border and the true border; A classification and regression network loss is constructed based on the second classification loss and the second regression loss; Based on the loss of the feature secondary extractor, the loss of the region proposal network and the loss of the classification and regression network, the network weights of the feature secondary extractor, the region proposal network and the classification and regression network are correspondingly optimized and adjusted; Return to execute the step of adjusting the image features of the feature image data by the feature secondary extractor and repeat the iteration until a third stopping iteration condition is met.
6. The underwater target detection and identification method according to claim 1, characterized in that: The regional convolutional neural network includes a feature secondary extractor, a regional proposal network and a classification and regression network; the feature extraction model and the target detection model are used to perform target detection and recognition on the underwater image to be recognized to obtain the target detection result, including the following steps: Extracting features of the underwater image to be identified by using the feature extraction model to obtain a feature image; Based on the feature image, generating a feature map using the feature secondary extractor; Analyzing and processing the feature map using the region proposal network to generate a candidate target region; The candidate target regions are adaptively hierarchically fused, and the target category is then identified using the classification and regression network to obtain the target detection result.
7. The underwater target detection and identification method according to claim 1, characterized in that: The target detection result includes the identification categories of multiple detection targets in the underwater image to be identified; the method also includes the following steps: Based on the identification category of each of the detected targets, the number of targets of each identification category is obtained by summarizing.
8. An underwater target detection and identification device, characterized in that: include: The first module is used to obtain underwater image data; The underwater image data carries annotation information for targets in the underwater environment; The second module is used to optimize the parameters of the preset neural network model through a circular dynamic optimization algorithm based on the underwater image data to obtain a feature extraction model; wherein the feature extraction model is feature image data according to the processing result of the underwater image data; The third module is used to perform target detection training on a preset regional convolutional neural network based on the feature image data combined with the annotation information to obtain a target detection model; The fourth module is used to use the feature extraction model and the target detection model to perform target detection and recognition on the underwater image to be recognized, and obtain a target detection result.
9. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.
10. A computer storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to implement the method according to any one of claims 1 to 7 when executed by the processor.
Citation Information
Patent Citations
High-precision target detection model and method based on underwater image enhancement
CN117975251A
Crop disease and pest prediction method based on artificial intelligence
CN118378742A
Belt scratch and offset monitoring method based on machine vision
CN119206721A
Generating Anti-infective design spaces for selecting drug candidates
US20220165359A1
Methods and apparatuses for generating peptides by synthesizing a portion of a design space to identify peptides having non-canonical amino acids
US20220364166A1
Cited By
WiFi signal data processing method and device
CN120547603A