Aquatic organism image recognition counting method, device, equipment and medium

By introducing technologies such as multi-scale feature fusion, adaptive pooling and dense detection head in the object detection model, the problem of high feature information loss and missed detection rates in densely distributed small object detection is solved, and high-precision identification and counting of densely distributed aquatic organisms are achieved.

CN119964155AActive Publication Date: 2025-05-09TSINGHUA UNIVERSITY

Patent Information

Application Number
CN202510132086.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-09
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

When traditional object detection algorithms deal with densely distributed small targets, they face the problems of missing feature information, poor detection of densely distributed objects and high missed detection rates.

Method used

Using a deep learning-based object detection model, including feature extraction networks with the introduction of multi-scale feature fusion technology, convolutional neural networks with adaptive pooling and spatial pyramid pooling technology, and dense detection heads, the densely distributed aquatic organisms in images are identified and counted through improved non-maximum suppression strategies.

Benefits of technology

High-precision identification and counting of densely distributed aquatic organisms is achieved, missing and missed detection and error detection are reduced, and detection accuracy and efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964155A_ABST
    Figure CN119964155A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition and counting, in particular to an aquatic organism image recognition and counting method and device, equipment and a medium, and the method comprises the steps: obtaining an aquatic organism image inputted by a user, and carrying out the preprocessing of the aquatic organism image; and inputting the preprocessed aquatic organism image into a target detection model based on deep learning, and outputting a detection result of a target aquatic organism in the aquatic organism image by the target detection model, the target detection model comprises a feature extraction network introducing a multi-scale feature fusion technology, a convolutional neural network adopting adaptive pooling and spatial pyramid pooling technologies, and dense detection heads for densely distributed target detection; and post-processing the detection result of the target aquatic organisms, and identifying the number of the target aquatic organisms and the positions of the target aquatic organisms in the aquatic organism image according to the post-processed detection result. Therefore, the problems of feature information loss, poor detection of densely distributed objects, high omission ratio and the like in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image recognition and counting, and in particular to a method, device, equipment and medium for image recognition and counting of aquatic organisms. Background Art

[0002] In the natural environment, larvae are usually distributed in clusters, and the proportion of individual larvae in the microscope image is extremely small, which poses a huge challenge to their automatic identification and counting. Traditional target detection algorithms face the following major problems when dealing with these densely distributed small targets:

[0003] Feature information loss: In convolutional neural networks, the role of the pooling layer is to reduce data dimensions and computational complexity, but this also causes the feature information of small targets to be weakened layer by layer. In particular, for larvae, which account for a very small proportion of the image, their feature information may be severely lost after multiple layers of pooling operations, resulting in a decrease in detection accuracy.

[0004] Detection of densely distributed objects: Many current target detection algorithms perform poorly when dealing with densely distributed objects. These algorithms usually rely on the obvious features of a single target, but in the case of densely distributed larvae, the distance between targets is close and they occlude each other severely, making it difficult for the detection algorithm to distinguish and identify each individual target, resulting in a very high overall missed detection rate.

[0005] Missed detection problem: Due to the above two reasons, traditional target detection algorithms often have a high missed detection rate when dealing with densely distributed small targets. Missed detection not only affects the accuracy of recognition, but also brings great inconvenience to subsequent quantitative statistics and data analysis. Summary of the invention

[0006] The present application provides a method, device, equipment and medium for aquatic organism image recognition and counting to solve the problems of loss of relevant technical feature information, poor detection of densely distributed objects, high missed detection rate, etc.

[0007] The first aspect of the present application provides a method for identifying and counting aquatic organism images, comprising the following steps: obtaining an aquatic organism image input by a user and preprocessing the aquatic organism image; inputting the preprocessed aquatic organism image into a target detection model based on deep learning, and the target detection model outputs the detection result of the target aquatic organism in the aquatic organism image, wherein the target detection model includes a feature extraction network that introduces multi-scale feature fusion technology, a convolutional neural network that adopts adaptive pooling and spatial pyramid pooling technology, and a dense detection head for densely distributed target detection; post-processing the detection result of the target aquatic organism, and identifying the number of the target aquatic organism and the position of the target aquatic organism in the aquatic organism image according to the post-processed detection result.

[0008] Optionally, the feature extraction network extracts feature maps of different levels in the aquatic organism image, and in the feature extraction process, multi-scale feature fusion technology is used to fuse the feature maps of different levels; the convolutional neural network uses adaptive pooling and spatial pyramid pooling technology to perform pooling operations on the fused feature maps, and completes the prediction of category and position in a single forward propagation; the dense detection head recognizes multiple target aquatic organisms in the feature map after the pooling operation, and distinguishes adjacent target aquatic organisms among multiple target aquatic organisms through an improved non-maximum suppression strategy.

[0009] Optionally, before inputting the preprocessed aquatic organism image into the deep learning-based target detection model, it also includes: obtaining a training set, the training set including multiple aquatic organism images labeled with target aquatic organisms; expanding the training set through data enhancement technology, using the expanded training set to train the target detection model, and optimizing the loss function and adjusting hyperparameters during the training process.

[0010] Optionally, the data enhancement technology includes performing at least one operation of rotating, scaling and translating a plurality of aquatic organism images labeled with target aquatic organisms.

[0011] Optionally, obtaining the aquatic organism image input by the user includes: using a graphical user interface developed with a PyQt framework to receive the aquatic organism image input by the user.

[0012] Optionally, the graphical user interface is further used to display the number of target aquatic organisms and the positions of the target aquatic organisms in the aquatic organism image.

[0013] Optionally, the preprocessing method includes at least one of image size adjustment, color space conversion and noise removal.

[0014] The second aspect of the present application provides an aquatic organism image recognition and counting device, including: an acquisition module, used to acquire an aquatic organism image input by a user and pre-process the aquatic organism image; an input module, used to input the pre-processed aquatic organism image into a target detection model based on deep learning, and the target detection model outputs the detection result of the target aquatic organism in the aquatic organism image, wherein the target detection model includes a feature extraction network that introduces multi-scale feature fusion technology, a convolutional neural network that adopts adaptive pooling and spatial pyramid pooling technology, and a dense detection head for densely distributed target detection; a processing module, used to post-process the detection result of the target aquatic organism, and identify the number of the target aquatic organism and the position of the target aquatic organism in the aquatic organism image according to the post-processed detection result.

[0015] A third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the aquatic organism image recognition and counting method as described in the above embodiment.

[0016] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the aquatic organism image recognition and counting method as described in the above-mentioned embodiment.

[0017] Therefore, this application includes the following beneficial effects:

[0018] The embodiment of the present application obtains an aquatic organism image input by a user, pre-processes the aquatic organism image, inputs the pre-processed aquatic organism image into a target detection model based on deep learning, and the model outputs the detection result of the target aquatic organism in the aquatic organism image, and post-processes the detection result of the target aquatic organism. According to the post-processed detection result, the number of target aquatic organisms and the position of the target aquatic organisms in the aquatic organism image are identified, and the number and growth characteristics of the larvae of the marsh clam in the image can be quickly counted and accurately located, reducing missed detection and false detection. Thus, the problems of loss of relevant technical feature information, poor detection of densely distributed objects, and high missed detection rate are solved.

[0019] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0021] Figure 1 A flowchart of a method for identifying and counting aquatic organism images provided according to an embodiment of the present application;

[0022] Figure 2 A flowchart of a base algorithm YOLOv8 according to an embodiment of the present application;

[0023] Figure 3 A flowchart of an algorithm application provided according to an embodiment of the present application;

[0024] Figure 4 This is a graph showing the test results of the larvae of the example provided according to one embodiment of the present application;

[0025] Figure 5 A comparison diagram of the training curves of Mussel-ID and YOLOv8n provided according to an embodiment of the present application;

[0026] Figure 6 A graph showing the changing trend of loss functions and performance indicators during model training and verification according to an embodiment of the present application;

[0027] Figure 7 This is an example diagram of an aquatic organism image recognition and counting device provided according to an embodiment of the present application;

[0028] Figure 8 The present invention is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0030] The following describes the aquatic organism image recognition and counting method, device, equipment and medium of the embodiment of the present application with reference to the accompanying drawings. In view of the problems mentioned in the above background technology, such as the loss of relevant technical feature information, poor detection of densely distributed objects, high missed detection rate, etc., the present application provides an aquatic organism image recognition and counting method. In this method, by obtaining the aquatic organism image input by the user and preprocessing the aquatic organism image, the preprocessed aquatic organism image is input into a target detection model based on deep learning, the model outputs the detection result of the target aquatic organism in the aquatic organism image, and the detection result of the target aquatic organism is post-processed. According to the post-processed detection result, the number of target aquatic organisms and the position of the target aquatic organisms in the aquatic organism image are identified, which can realize the rapid statistics and precise positioning of the number and growth characteristics of the larvae of the swamp clam in the image, and reduce missed detection and false detection. Thus, the problems such as the loss of relevant technical feature information, poor detection of densely distributed objects, and high missed detection rate are solved.

[0031] Specifically, Figure 1 A schematic flow chart of a method for image recognition and counting of aquatic organisms provided in an embodiment of the present application.

[0032] like Figure 1 As shown, the aquatic organism image recognition and counting method comprises the following steps:

[0033] In step S101, an aquatic organism image input by a user is obtained, and the aquatic organism image is preprocessed.

[0034] The preprocessing method will be described in detail below and will not be repeated here.

[0035] In an embodiment of the present application, obtaining an aquatic organism image input by a user includes: using a graphical user interface developed with a PyQt framework to receive the aquatic organism image input by the user.

[0036] Among them, PyQt is a set of Python binding libraries for creating graphical user interfaces, which encapsulates the functions of the Qt library. Qt is a cross-platform C++ library that is widely used in developing desktop applications and touch user interfaces.

[0037] It can be understood that, in order to facilitate user operation, the embodiment of the present application provides an intuitive and friendly graphical user interface developed by the PyQt framework, which can receive aquatic organism images input by the user.

[0038] In the embodiment of the present application, the graphical user interface is also used to display the number of target aquatic organisms and the positions of the target aquatic organisms in the aquatic organism image.

[0039] It can be understood that the embodiment of the present application uses a graphical user interface developed based on the PyQt framework, so users can not only upload the aquatic organism images to be analyzed, but also directly view the processed results on the interface, which shows the number of target aquatic organisms and the position of the target aquatic organisms in the aquatic organism image.

[0040] In an embodiment of the present application, the preprocessing method includes at least one of image size adjustment, color space conversion and noise removal.

[0041] Among them, image resizing is to adjust the image to a size suitable for model input; color space conversion is to convert the image to a color space suitable for target detection, such as from RGB to grayscale or HSV space; noise removal is to remove noise in the image through filtering and other methods to improve image quality.

[0042] It can be understood that the embodiments of the present application preprocess the aquatic organism image, including methods such as image resizing, color space conversion and noise removal, so as to adjust the size of the aquatic organism image to a size suitable for the model input, change the color representation of the image, eliminate unnecessary random points or patches in the image, and obtain a preprocessed image that is convenient for analysis.

[0043] In step S102, the preprocessed aquatic organism image is input into a target detection model based on deep learning, and the target detection model outputs the detection result of the target aquatic organism in the aquatic organism image, wherein the target detection model includes a feature extraction network that introduces multi-scale feature fusion technology, a convolutional neural network that adopts adaptive pooling and spatial pyramid pooling technology, and a dense detection head for densely distributed target detection.

[0044] Among them, the multi-scale feature fusion technology combines low-level detail information and high-level semantic information to build a richer feature representation, thereby improving the perception of small targets; the feature extraction network is the first layer of the model, which uses multi-scale feature fusion technology and is responsible for extracting useful feature information from the input image; adaptive pooling and spatial pyramid pooling are two technologies used to process convolutional neural network feature maps, which help to improve the performance of the model; convolutional neural networks use convolution operations to automatically learn complex patterns in images.

[0045] It can be understood that the target detection model of the embodiment of the present application includes a feature extraction network that introduces multi-scale feature fusion technology, a convolutional neural network that adopts adaptive pooling and spatial pyramid pooling technology, and a dense detection head for densely distributed target detection. The preprocessed aquatic organism image is input into the target detection model based on deep learning, which can achieve accurate detection of target aquatic organisms in the aquatic organism image.

[0046] In an embodiment of the present application, a feature extraction network extracts feature maps of different levels in an image of aquatic organisms. During the feature extraction process, a multi-scale feature fusion technique is used to fuse feature maps of different levels. A convolutional neural network uses adaptive pooling and spatial pyramid pooling techniques to perform pooling operations on the fused feature maps, and completes the prediction of categories and positions in a single forward propagation. A dense detection head identifies multiple target aquatic organisms in the feature map after the pooling operation, and distinguishes adjacent target aquatic organisms among multiple target aquatic organisms through an improved non-maximum suppression strategy.

[0047] Among them, the improved non-maximum suppression strategy means that in the process of target detection, a large number of candidate boxes will be generated at the position of the same target, and these candidate boxes may overlap. At this time, non-maximum suppression is needed to find the best target bounding box and eliminate redundant bounding boxes. Specifically, the following steps are included:

[0048] a) Sort by confidence score;

[0049] b) Select the bounding box with the highest confidence and add it to the final output list, and delete it from the bounding box list;

[0050] c) Calculate the area of ​​all bounding boxes;

[0051] d) Calculate the IoU between the bounding box with the highest confidence and other candidate boxes;

[0052] e) Delete the bounding boxes whose IoU is greater than the threshold;

[0053] f) Repeat the above process until the bounding box list is empty.

[0054] Among them, IoU (intersection-over-union) is the intersection of two bounding boxes divided by their union.

[0055] It can be understood that when the embodiment of the present application inputs the preprocessed aquatic organism image into the target detection model based on deep learning, the feature extraction network first analyzes it, extracts feature maps of different levels, and fuses the feature maps of different levels using multi-scale feature fusion technology. The fused feature maps are input into the convolutional neural network, optimized by adaptive pooling and spatial pyramid pooling technology, and the category and position of each target aquatic organism are predicted in a single forward propagation. Finally, a dense detection head is used to identify multiple target aquatic organisms, and adjacent target aquatic organisms are distinguished by an improved non-maximum suppression strategy, providing high-precision and reliable detection results.

[0056] In an embodiment of the present application, before the preprocessed aquatic organism image is input into the target detection model based on deep learning, it also includes: obtaining a training set, the training set includes multiple aquatic organism images labeled with target aquatic organisms; processing and expanding the training set through data enhancement technology, using the expanded training set to train the target detection model, and optimizing the loss function and adjusting the hyperparameters during the training process. The embodiment of the present application is based on the mapping function relationship in Table 1, tests the optimization effect obtained by different functions, and finally selects the Guass function as the optimal loss function. The selected hyperparameters are shown in Table 2, where Table 1 is a corresponding table of function expressions selected for optimizing the loss function, and Table 2 is a training hyperparameter configuration table.

[0057] Table 1

[0058]

[0059] Table 2

[0060]

[0061] Among them, data enhancement technology refers to the use of data enhancement technology to expand the original training set in order to improve the generalization ability and robustness of the model. The specific method will be described in detail below and will not be repeated here.

[0062] It can be understood that before preparing to input the preprocessed aquatic organism images into the deep learning-based target detection model, the embodiment of the present application first needs to construct a high-quality training set, which is composed of multiple aquatic organism images labeled with target aquatic organisms, and in order to increase the diversity of training data and improve the generalization ability of the model, data enhancement technology is used to expand the original training set to generate more training samples. The expanded training set can be used to more effectively train the target detection model. During the training stage, the loss function is continuously optimized to ensure that the model output is as close to the actual annotation information as possible. At the same time, the hyperparameters are adjusted to improve the performance of the model, and ultimately the model can achieve efficient and accurate recognition of target aquatic organisms in aquatic organism images.

[0063] In an embodiment of the present application, the data enhancement technology includes at least one operation of rotating, scaling, and translating a plurality of aquatic organism images labeled with target aquatic organisms.

[0064] It can be understood that the data enhancement technology of the embodiments of the present application involves various transformations of the image, such as rotation, scaling, translation, etc., to generate more diverse samples, so that the model can maintain good performance when facing images under different conditions.

[0065] In step S103, the detection result of the target aquatic organisms is post-processed, and the number of the target aquatic organisms and the positions of the target aquatic organisms in the aquatic organism image are identified according to the post-processed detection result.

[0066] Among them, post-processing refers to the process of further improving these results after the model generates preliminary detection results, such as bounding boxes, category labels, and confidence scores, in order to remove redundancy, correct errors, and improve the quality of the final output.

[0067] It can be understood that after the target detection model of the embodiment of the present application outputs the detection results of the target aquatic organisms in the aquatic organism image, it is also necessary to perform post-processing operations on it to optimize the original detection results. After post-processing, the number of target aquatic organisms in the image can be more accurately identified and counted, and the precise location of each target in the aquatic organism image can be identified.

[0068] According to the aquatic organism image recognition and counting method proposed in the embodiment of the present application, an aquatic organism image input by a user is obtained and pre-processed, the pre-processed aquatic organism image is input into a target detection model based on deep learning, the model outputs the detection result of the target aquatic organism in the aquatic organism image, the detection result of the target aquatic organism is post-processed, and the number of target aquatic organisms and the position of the target aquatic organisms in the aquatic organism image are identified according to the post-processed detection result, so as to realize rapid statistics and precise positioning of the number and growth characteristics of the larvae of the marsh clam in the image, thereby reducing missed detections and false detections.

[0069] The aquatic organism image recognition and counting method is further described below through a specific embodiment.

[0070] The algorithm involved in this embodiment is mainly based on the deep learning Mussel-ID model, which is used for image recognition and counting of larvae. Mussel-ID is an end-to-end target detection algorithm suitable for processing complex image scenes. Figure 3 As shown, the basic situation of the algorithm is as follows:

[0071] 1. Multi-scale feature fusion

[0072] To address the problem that small targets account for a small proportion of the image and their features are easily ignored, this algorithm introduces a multi-scale feature fusion mechanism based on Mussel-ID. Multi-scale feature fusion ensures that even subtle feature information can be fully utilized by retaining and fusing feature maps of different levels in the model, thereby improving the detection capability of small targets.

[0073] 2. Improved pooling operation

[0074] In traditional convolutional neural networks, pooling operations usually lead to the loss of feature information of small targets. To solve this problem, this algorithm uses adaptive pooling and spatial pyramid pooling (SPPF) technology in the pooling layer. These technologies adaptively adjust the size of the pooling window or perform pooling operations at different scales to retain the key information of small targets as much as possible, so that these small targets can still be captured in the subsequent detection process.

[0075] 3. Detection strategy for densely distributed targets

[0076] Marsh clams are usually densely distributed in images, which poses a challenge to traditional target detection algorithms. To solve this problem, the algorithm introduces a dense detection head and adds a detection layer in the network to improve the model's ability to recognize dense targets. In addition, by improving the non-maximum suppression (NMS) strategy, the algorithm can effectively distinguish adjacent marsh clam larvae targets and reduce missed detections and false detections caused by dense distribution.

[0077] 4. Model training and optimization

[0078] In order to improve the accuracy and robustness of the algorithm, this model is trained using a large number of accurately labeled images of larvae of the marsh clam. Through data augmentation techniques (such as rotation, scaling, translation, etc.), the training set is expanded to enhance the generalization ability of the model. In addition, during the training process, the algorithm further improves the detection accuracy of the model by optimizing the loss function and adjusting the hyperparameters.

[0079] 5. System Integration

[0080] This embodiment integrates the improved Mussel-ID model with image processing technology to form a complete image recognition and counting system for clam larvae. The system can run on a variety of operating systems, supports a variety of image formats, and has good scalability. Users can upload images through a graphical user interface and view recognition and counting results in real time. The system automatically saves and exports data, greatly improving the efficiency and accuracy of clam monitoring.

[0081] Based on the above algorithm, this embodiment proposes a method for underwater larval image recognition and counting based on deep learning, aiming to solve the problem of feature information loss and high missed detection rate in traditional methods when identifying densely distributed small targets (such as larvae). By introducing the comprehensive YOLO v8 algorithm (see the algorithm flow) Figure 2 ), multi-scale feature fusion, optimized pooling operations, and the detection model Mussel-ID for densely distributed targets, such as Figure 5 The following is a comparison of the training curves of Mussel-ID and YOLOv8n. Figure 6 The figure shows the changing trend of the loss function and performance index during the model training and verification process. This embodiment significantly improves the detection accuracy and efficiency of small targets. The method includes the following steps:

[0082] 1. Image acquisition:

[0083] Use a microscope or other image acquisition tool to obtain images of the larvae. These images can include distribution scenes of the larvae at different angles and under different lighting conditions to ensure data diversity.

[0084] The acquired images are transmitted to the system for subsequent processing.

[0085] 2. Image preprocessing:

[0086] Preprocess the acquired images to improve the effect of subsequent target detection. The preprocessing steps include:

[0087] Image resizing: Resize images to a suitable size for model input.

[0088] Color space conversion: Convert the image to a color space suitable for target detection (such as from RGB to grayscale or HSV space).

[0089] Noise removal: Remove noise from the image through filtering and other methods to improve image quality.

[0090] 3. Feature extraction and multi-scale feature fusion:

[0091] The preprocessed image is fed into the Mussel-ID model, which uses convolutional layers to extract features from the image and identify areas that may contain larvae.

[0092] In the feature extraction process, multi-scale feature fusion technology is used to fuse feature maps at different levels to enhance the model's perception of larvae.

[0093] 4. Target detection and adaptive pooling:

[0094] The Mussel-ID model is used for object detection to identify the larvae in the image. The model predicts both the object category and the location in a single forward pass.

[0095] In order to retain the feature information of small targets, adaptive pooling and spatial pyramid pooling (SPPF) techniques are used to reduce the loss of small target feature information caused by pooling operations.

[0096] 5. Dense target detection strategy and non-maximum suppression (NMS):

[0097] In view of the dense distribution of Marsh clam larvae, the model uses dense detection heads and an improved NMS strategy to effectively distinguish and identify adjacent larval targets, reducing the occurrence of missed detections and false detections.

[0098] 6. Post-processing of test results:

[0099] Post-process the detected larvae to optimize the recognition results. Post-processing includes removing overlapping detection frames and merging adjacent targets to improve the final recognition accuracy.

[0100] 7. Result display and storage:

[0101] The identification and counting results are displayed on a graphical user interface, allowing users to intuitively view each detected larvae and their location and number in the image. The results can be exported to report files in multiple formats (such as text files, spreadsheets, and image files) for easy data analysis and management.

[0102] 8. System self-learning and optimization:

[0103] The system has a self-learning function, which can optimize the model and improve performance based on the feedback of recognition results, and continuously improve the accuracy and efficiency of recognition. As more data accumulates, the system can adaptively update the model parameters to meet the recognition needs in different scenarios.

[0104] 9. Extended application:

[0105] This method can be extended to the recognition and counting of other aquatic organisms or small targets and can be applied to different detection tasks by adjusting the model and training dataset.

[0106] The following takes the collected underwater swamp clam image as an example to illustrate the recognition and counting process of this embodiment, as follows:

[0107] First, we used underwater video equipment to obtain multiple images of the distribution area of ​​​​marsh clam larvae. In these images, the larvae of marsh clam are densely distributed and the proportion of a single marsh clam larva in the image is relatively small. In some images, the larvae of marsh clam may be difficult to identify due to factors such as environmental lighting.

[0108] 1. Preprocessing stage. The preprocessing step includes adjusting the image size to meet the input requirements of Mussel-ID, usually adjusted to a standard size of 640x640 pixels. In addition, the image needs to be converted to grayscale or HSV color space to enhance the characteristics of the larvae. At the same time, the noise in the image is removed by methods such as median filtering to further improve the image quality.

[0109] 2. Detection stage. After preprocessing, the image is input into the Mussel-ID model for feature extraction and target detection. In the feature extraction stage, multi-scale feature fusion technology is used to combine low-level detail features with high-level semantic features. The model uses adaptive pooling and spatial pyramid pooling techniques to ensure that the key features of the larvae are retained while reducing the data dimension. Through the improved non-maximum suppression (NMS) algorithm, the system can effectively distinguish and identify the densely distributed swamp clam larvae targets in the image, avoiding the problems of missed detection and false detection caused by the dense distribution of targets in traditional methods.

[0110] 3. Detection completion stage. The system intuitively displays the recognition results on the graphical user interface, such as Figure 4 As shown in the figure. Each identified clam is marked with a location box on the interface, and the system outputs the total number of clams in the image. Users can view the specific location of each clam through the interface and export the detection results as a report file.

[0111] Next, the aquatic organism image recognition and counting device proposed in accordance with the embodiment of the present application will be described with reference to the accompanying drawings.

[0112] Figure 7 It is a block diagram of the aquatic organism image recognition and counting device according to an embodiment of the present application.

[0113] like Figure 7 As shown, the aquatic organism image recognition and counting device 10 includes: an acquisition module 201 , an input module 202 and a processing module 203 .

[0114] Among them, the acquisition module 201 is used to obtain the aquatic organism image input by the user and pre-process the aquatic organism image; the input module 202 is used to input the pre-processed aquatic organism image into the target detection model based on deep learning, and the target detection model outputs the detection result of the target aquatic organism in the aquatic organism image, wherein the target detection model includes a feature extraction network that introduces multi-scale feature fusion technology, a convolutional neural network that adopts adaptive pooling and spatial pyramid pooling technology, and a dense detection head for densely distributed target detection; the processing module 203 is used to post-process the detection result of the target aquatic organism, and identify the number of target aquatic organisms and the position of the target aquatic organisms in the aquatic organism image according to the post-processed detection results.

[0115] In an embodiment of the present application, the input module 202 is further used for: the feature extraction network extracts feature maps of different levels in the aquatic organism image, and in the feature extraction process, the feature maps of different levels are fused using multi-scale feature fusion technology; the convolutional neural network uses adaptive pooling and spatial pyramid pooling technology to perform pooling operations on the fused feature maps, and completes the prediction of category and position in a single forward propagation; the dense detection head identifies multiple target aquatic organisms in the feature map after the pooling operation, and distinguishes adjacent target aquatic organisms among multiple target aquatic organisms through an improved non-maximum suppression strategy.

[0116] In an embodiment of the present application, a training module is also included, wherein the training module is further used to obtain a training set before inputting the preprocessed aquatic organism image into the target detection model based on deep learning, the training set including multiple aquatic organism images labeled with target aquatic organisms; the training set is expanded by processing through data enhancement technology, the target detection model is trained using the expanded training set, and the loss function is optimized and the hyperparameters are adjusted during the training process.

[0117] In an embodiment of the present application, the data enhancement technology includes at least one operation of rotating, scaling, and translating a plurality of aquatic organism images labeled with target aquatic organisms.

[0118] In the embodiment of the present application, the acquisition module 201 is further used to: acquire the aquatic organism image input by the user, and receive the aquatic organism image input by the user using a graphical user interface developed using the PyQt framework.

[0119] In the embodiment of the present application, the graphical user interface is also used to display the number of target aquatic organisms and the positions of the target aquatic organisms in the aquatic organism image.

[0120] In an embodiment of the present application, the preprocessing method includes at least one of image size adjustment, color space conversion and noise removal.

[0121] It should be noted that the above explanations of the embodiment of the aquatic organism image recognition and counting method are also applicable to the aquatic organism image recognition and counting device of this embodiment, and will not be repeated here.

[0122] According to the aquatic organism image recognition and counting device proposed in the embodiment of the present application, by obtaining an aquatic organism image input by a user and preprocessing the aquatic organism image, the preprocessed aquatic organism image is input into a target detection model based on deep learning, the model outputs the detection result of the target aquatic organism in the aquatic organism image, and the detection result of the target aquatic organism is post-processed. The number of target aquatic organisms and the position of the target aquatic organisms in the aquatic organism image are identified according to the post-processed detection result, which can realize the rapid statistics and precise positioning of the number and growth characteristics of the marsh clam larvae in the image, reducing missed detections and false detections.

[0123] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:

[0124] A memory 301 , a processor 302 , and a computer program stored in the memory 301 and executable on the processor 302 .

[0125] When the processor 302 executes the program, the aquatic organism image recognition and counting method provided in the above embodiment is implemented.

[0126] Furthermore, the electronic device further comprises:

[0127] The communication interface 303 is used for communication between the memory 301 and the processor 302 .

[0128] The memory 301 is used to store computer programs that can be run on the processor 302 .

[0129] The memory 301 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.

[0130] If the memory 301, the processor 302 and the communication interface 303 are implemented independently, the communication interface 303, the memory 301 and the processor 302 can be connected to each other through a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0131] Optionally, in a specific implementation, if the memory 301, the processor 302 and the communication interface 303 are integrated on a chip, the memory 301, the processor 302 and the communication interface 303 can communicate with each other through an internal interface.

[0132] The processor 302 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.

[0133] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned aquatic organism image recognition and counting method.

[0134] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0135] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0136] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0137] It should be understood that the various parts of the present application can be implemented in hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, the steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array, a field programmable gate array, etc.

[0138] A person of ordinary skill in the art may understand that all or part of the steps carried by the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the above-mentioned program may be stored in a computer-readable storage medium, which, when executed, includes one of the steps of the method embodiment or a combination thereof.

[0139] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A method for recognizing and counting aquatic organisms, characterized in that: The following steps are involved: Acquire an aquatic organism image input by a user, and pre-process the aquatic organism image; Inputting the preprocessed aquatic organism image into a target detection model based on deep learning, the target detection model outputs a detection result of the target aquatic organism in the aquatic organism image, wherein the target detection model includes a feature extraction network that introduces a multi-scale feature fusion technology, a convolutional neural network that adopts adaptive pooling and spatial pyramid pooling technology, and a dense detection head for densely distributed target detection; The detection result of the target aquatic organism is post-processed, and the number of the target aquatic organisms and the positions of the target aquatic organisms in the aquatic organism image are identified according to the post-processed detection result.

2. The aquatic organism image recognition and counting method according to claim 1, characterized in that: The feature extraction network extracts feature maps of different levels in the aquatic organism image. In the feature extraction process, the feature maps of different levels are fused using multi-scale feature fusion technology; the convolutional neural network uses adaptive pooling and spatial pyramid pooling technology to perform pooling operations on the fused feature maps, and completes the prediction of categories and positions in a single forward propagation; The dense detection head identifies multiple target aquatic organisms in the feature map after the pooling operation, and distinguishes adjacent target aquatic organisms among the multiple target aquatic organisms by improving the non-maximum suppression strategy.

3. The aquatic organism image recognition and counting method according to claim 1 or 2, characterized in that: Before the preprocessed aquatic life images are fed into the deep learning-based object detection model, the following steps are also included: Acquire a training set, wherein the training set includes a plurality of aquatic organism images labeled with target aquatic organisms; The training set is expanded by processing with data enhancement technology, the target detection model is trained using the expanded training set, and the loss function is optimized and the hyperparameters are adjusted during the training process.

4. The aquatic organism image recognition and counting method according to claim 3, characterized in that: The data enhancement technology includes at least one operation of rotating, scaling and translating a plurality of aquatic organism images labeled with target aquatic organisms.

5. The aquatic organism image recognition and counting method according to claim 1, characterized in that: The step of obtaining the aquatic organism image input by the user comprises: The graphical user interface developed using the PyQt framework receives user input of aquatic life images.

6. The aquatic organism image recognition and counting method according to claim 5, characterized in that: The graphical user interface is also used to display the number of target aquatic organisms and the positions of the target aquatic organisms in the aquatic organism image.

7. The aquatic organism image recognition and counting method according to claim 1, characterized in that: The preprocessing method includes at least one of image size adjustment, color space conversion and noise removal.

8. An aquatic organism image recognition and counting device, characterized in that: include: An acquisition module, used for acquiring an aquatic organism image input by a user, and preprocessing the aquatic organism image; An input module, used for inputting the preprocessed aquatic organism image into a target detection model based on deep learning, wherein the target detection model outputs a detection result of the target aquatic organism in the aquatic organism image, wherein the target detection model includes a feature extraction network that introduces a multi-scale feature fusion technology, a convolutional neural network that adopts adaptive pooling and spatial pyramid pooling technology, and a dense detection head for densely distributed target detection; The processing module is used to post-process the detection result of the target aquatic organism, and identify the number of the target aquatic organisms and the positions of the target aquatic organisms in the aquatic organism image according to the post-processed detection result.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the aquatic organism image recognition and counting method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed, the aquatic organism image recognition and counting method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Image front vehicle target identification method based on convolutional neural network

    CN115661767A

  • Shrimp seed counting method and system based on high resolution and target detection

    CN115937169A

  • Infrared small target detection method and device based on deep fusion of edge details and deep features

    CN116468980A

  • Early detection and early warning method for marine environmental pollution

    CN117611588A

  • Seed germination segmentation method and seed germination rate and germination rate identification method

    CN119339087A

Cited By

  • Aquatic organism detection system, method and device

    CN120385669A

  • Aquatic organism detection system, method and device

    CN120385669B