Bait counting method, device and equipment based on adaptive scale aggregation network and storage medium
By preprocessing and training near-infrared images through an adaptive scale aggregation network, the problem of accuracy in counting fish bait in low-light environments was solved, and high-quality bait recognition and counting were achieved.
Patent Information
- Application Number
- CN202510808839.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-12
AI Technical Summary
Existing bait counting methods based on computer vision technology have difficulty in accurately identifying fish bait particles in low-light environments and are easily affected by noise, resulting in reduced counting accuracy.
An adaptive scale aggregation network is used to preprocess near-infrared images through hybrid filtering and dynamic contrast limited adaptive histogram equalization. The models of feature map encoder and density map decoder are constructed, and they are trained using Euclidean loss and local pattern consistency loss to output bait density map and quantity.
The accuracy and robustness of fish bait counting in near-infrared images in low-light environments were improved, image contrast was enhanced, noise was removed, and model performance was optimized.
Smart Images

Figure CN120635670A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of fishery farming technology, and in particular to a bait counting method, device, equipment and storage medium based on an adaptive scale aggregation network. Background Art
[0002] With the rapid development of the global fisheries and aquaculture industries, traditional artificial aquaculture methods are gradually exposing limitations such as inefficiency and high costs. To meet growing production demands and improve aquaculture efficiency, intelligent aquaculture technologies have emerged and become a hot topic of research. In factory aquaculture, precise feeding is particularly important for improving efficiency and reducing costs. Effectively calculating the remaining feed in the aquaculture pond enables feeding the appropriate amount of feed at the right time, thus achieving intelligent aquaculture management.
[0003] Currently, research on intelligent fish farming focuses on using computer vision to count bait particles. Most existing studies analyze visible light images, using image processing and object detection algorithms to identify and count bait particles.
[0004] However, existing bait counting methods based on computer vision technology have many problems in low-light environments. First, the distribution density of bait particles in near-infrared images will change dynamically with feeding rate, water flow disturbances, etc., resulting in the formation of multi-scale characteristic patterns of the aggregation degree (spatial scale) and texture characteristics (local details) of bait in the water, which increases the difficulty of counting. Secondly, existing methods have difficulty in accurately identifying bait particles in low-light environments and are easily interfered by noise, further reducing the accuracy of counting. Therefore, how to improve the accuracy of fish bait counting in near-infrared images in low-light environments has become an urgent problem to be solved.
[0005] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0006] The purpose of this application is to provide a bait counting method, device, equipment and storage medium based on an adaptive scale aggregation network, aiming to solve the technical problem of how to improve the accuracy of fish bait counting in near-infrared images under low illumination environments.
[0007] To achieve the above objectives, the present application proposes a bait counting method based on an adaptive scale aggregation network, the method comprising:
[0008] The near-infrared image dataset is subjected to hybrid filtering and dynamic contrast limited adaptive histogram equalization to obtain an enhanced image dataset;
[0009] Constructing an adaptive scale aggregation network model including a feature map encoder and a density map decoder. The feature map encoder is composed of multiple adaptive scale aggregation modules to perform dynamic weight allocation of multi-scale convolution kernels.
[0010] Training the adaptive scale aggregation network model according to the enhanced image dataset, Euclidean loss, and local pattern consistency loss to obtain a target model;
[0011] The target bait image is input into the target model to obtain a bait density map and bait quantity.
[0012] In one embodiment, the step of constructing an adaptive scale aggregation network model including a feature map encoder and a density map decoder includes: constructing multiple cascaded adaptive scale aggregation modules including multiple convolution kernels; introducing an attention subnetwork into the adaptive scale aggregation module to obtain a feature map encoder, and the attention subnetwork is used to calculate the weight of the convolution kernel; constructing a density map decoder based on multiple convolution layers of different sizes and multiple transposed convolution layers; and connecting the feature map encoder with the density map decoder to form an adaptive scale aggregation network model.
[0013] In one embodiment, the target model includes a feature mapping encoder and a density mapping decoder; the step of inputting the target bait image into the target model to obtain a bait density map and the bait quantity includes: extracting features of the target bait image through the feature mapping encoder to obtain multi-scale features; reconstructing the resolution of the multi-scale features through the density mapping decoder to obtain a bait density map; performing region segmentation on the bait density map through a region growing algorithm to obtain connected regions; and determining the bait quantity based on the integral density value of the connected region.
[0014] In one embodiment, the step of determining the amount of bait based on the integral density value of the connected area includes: calculating the sum of the density values of all pixels in each of the connected areas to obtain the integral density value; obtaining the number of bait particles corresponding to each of the connected areas based on the ratio of the integral density value to a preset unit density threshold; and accumulating the number of bait particles in all the connected areas to obtain the amount of bait.
[0015] In one embodiment, the step of performing hybrid filtering and dynamic contrast limited adaptive histogram equalization on the near-infrared image dataset to obtain an enhanced image dataset includes: performing adaptive median filtering on the near-infrared image dataset to obtain a denoised image dataset; performing non-local mean filtering on the denoised image dataset to obtain a smoothed image dataset; and performing dynamic contrast limited adaptive histogram equalization and dark field correction on the smoothed image dataset to obtain an enhanced dataset.
[0016] In one embodiment, the step of performing dynamic contrast limited adaptive histogram equalization and dark field correction on the smoothed image dataset to obtain an enhanced dataset includes: dividing each frame image in the smoothed image dataset into sub-blocks of a preset size; calculating the local entropy value of the sub-block, and marking the sub-block as a high entropy area or a low entropy area based on the local entropy value and a preset entropy threshold; determining a contrast limiting threshold parameter based on the high entropy area and the low entropy area; performing histogram equalization on the sub-block according to the contrast limiting threshold parameter, and eliminating boundary artifacts between the sub-blocks through bilinear interpolation to obtain a reference dataset; generating a dark field template through a Gaussian blur kernel, and applying the dark field template to the reference dataset to obtain an enhanced dataset.
[0017] In one embodiment, the step of training the adaptive scale aggregation network model according to the enhanced image dataset, Euclidean loss and local pattern consistency loss to obtain a target model includes: iteratively training the adaptive scale aggregation network model according to the enhanced image dataset using an Adam optimizer with a preset learning rate and a preset batch size; in each iterative training, calculating the Euclidean loss and structural similarity index between the predicted density map output by the adaptive scale aggregation network model and the true density map in the enhanced image dataset; calculating the local pattern consistency loss according to the number of pixels in the true density map and the structural similarity index; weightedly fusing the Euclidean loss and the local pattern consistency loss according to the mean and standard deviation of the density distribution of the enhanced image dataset to obtain a joint loss; when the joint loss meets a preset condition, stopping the training to obtain the target model.
[0018] In addition, to achieve the above-mentioned purpose, the present application also proposes a bait counting device based on an adaptive scale aggregation network, the device comprising:
[0019] A data enhancement module is used to perform hybrid filtering and dynamic contrast limited adaptive histogram equalization on the near-infrared image dataset to obtain an enhanced image dataset;
[0020] A model building module for building an adaptive scale aggregation network model including a feature map encoder and a density map decoder. The feature map encoder is composed of multiple adaptive scale aggregation modules to perform dynamic weight allocation of multi-scale convolution kernels.
[0021] A model training module, configured to train the adaptive scale aggregation network model according to the enhanced image dataset, Euclidean loss, and local pattern consistency loss to obtain a target model;
[0022] The bait counting module is used to input the target bait image into the target model to obtain the bait density map and the bait quantity.
[0023] In addition, to achieve the above-mentioned purpose, the present application also proposes a bait counting device based on an adaptive scale aggregation network, the device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the bait counting method based on an adaptive scale aggregation network as described above.
[0024] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the bait counting method based on the adaptive scale aggregation network as described above are implemented.
[0025] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the bait counting method based on the adaptive scale aggregation network as described above.
[0026] One or more technical solutions proposed in this application have at least the following technical effects:
[0027] First, a near-infrared image dataset is subjected to hybrid filtering and dynamic contrast-limited adaptive histogram equalization to remove noise and enhance image contrast, resulting in an enhanced image dataset. This provides high-quality input data for subsequent model training. Next, an adaptive scale aggregation network model is constructed, consisting of a feature map encoder and a density map decoder. The feature map encoder dynamically allocates weights to multi-scale convolution kernels through multiple adaptive scale aggregation modules, improving the model's adaptability to features at different scales. The model is then trained on the enhanced image dataset using Euclidean loss and local pattern consistency loss to obtain a target model, which optimizes model performance. Finally, a target bait image is input into the target model to obtain a bait density map and bait count, thereby improving the accuracy of fish bait counting in near-infrared images under low-light conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0029] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0030] Figure 1 A flow chart of the first embodiment of the bait counting method based on the adaptive scale aggregation network of the present application is provided;
[0031] Figure 2 The original image and annotation example diagram provided in Example 1 of the bait counting method based on the adaptive scale aggregation network of this application;
[0032] Figure 3 A schematic diagram of the preprocessing effect provided in Example 1 of the bait counting method based on the adaptive scale aggregation network of this application;
[0033] Figure 4 A schematic diagram of the structure of an adaptive scale aggregation network provided in Example 1 of the bait counting method based on the adaptive scale aggregation network of the present application;
[0034] Figure 5 A schematic diagram of the structure of an adaptive scale aggregation module provided in Example 1 of the bait counting method based on an adaptive scale aggregation network of this application;
[0035] Figure 6 Schematic diagram of the attention sub-network structure provided for the first embodiment of the bait counting method based on the adaptive scale aggregation network of this application;
[0036] Figure 7 A schematic diagram showing the comparison of the predicted density map, the true density map and the original image provided in Example 1 of the bait counting method based on the adaptive scale aggregation network of this application;
[0037] Figure 8 This is an example diagram of the collected and processed image and density map provided in Example 1 of the bait counting method based on the adaptive scale aggregation network of this application;
[0038] Figure 9 A flow chart of the second embodiment of the bait counting method based on the adaptive scale aggregation network of the present application is provided;
[0039] Figure 10 This is a schematic diagram of the module structure of a bait counting device based on an adaptive scale aggregation network according to an embodiment of the present application;
[0040] Figure 11 Schematic diagram of the device structure of the hardware operating environment involved in the bait counting method based on the adaptive scale aggregation network in the embodiment of the present application.
[0041] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0042] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0043] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0044] With the rapid development of the global fisheries and aquaculture industries, the inefficiency and high costs of traditional artificial aquaculture methods have become increasingly prominent, leading to the emergence of intelligent aquaculture technology, which has become a research hotspot. In factory aquaculture, accurate feeding of bait is crucial to improving efficiency and reducing costs, and effectively calculating the remaining amount of bait in the aquaculture pond is the key to achieving intelligent aquaculture management. Currently, research on bait feeding mainly focuses on counting using computer vision technology, mostly based on visible light image analysis. However, existing methods have many problems in low-light environments. For example, the dynamic changes in the distribution density of bait particles lead to the emergence of multi-scale feature patterns, which increases the difficulty of counting. At the same time, existing methods have difficulty in accurately identifying bait particles and are easily interfered by noise, which reduces the accuracy of counting.
[0045] The main solution of the embodiment of the present application is as follows: First, the near-infrared image is preprocessed to obtain a high-quality enhanced image dataset through hybrid filtering denoising and contrast enhancement. Next, an adaptive scale aggregation network model is constructed, whose feature map encoder processes multi-scale features through dynamic weight allocation. Then, the model is trained using the enhanced image dataset and two loss functions to optimize model performance. Finally, the real-time bait image is input into the model, and the bait density map and quantity are output, thereby improving the accuracy of fish bait counting in low-light environments.
[0046] It should be noted that the execution subject of the embodiments of the present application may be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device or computer system capable of implementing the above functions. The following describes this embodiment and the following embodiments using a computer system as an example.
[0047] Based on this, the embodiment of the present application provides a bait counting method based on an adaptive scale aggregation network, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the bait counting method based on the adaptive scale aggregation network of this application.
[0048] In this embodiment, the bait counting method based on the adaptive scale aggregation network includes steps S10 to S40:
[0049] Step S10 , performing hybrid filtering and dynamic contrast limited adaptive histogram equalization on the near-infrared image dataset to obtain an enhanced image dataset.
[0050] It should be noted that the near-infrared image dataset refers to the image collection used to train and test ASANet (Adaptive Scale Aggregation Network) in this embodiment, which was collected through a specific experimental environment and equipment. These images were taken using a near-infrared night vision camera in a low-light environment, recording the distribution of fish bait in the water. Specifically, the experimental equipment includes an infrared night vision camera, an aerator, a 200L bucket and several red carp, and the camera is located directly above the breeding bucket. Bait is fed at 7:00pm-8:00pm at night. The camera starts recording 20 seconds before feeding and lasts for 30 to 60 minutes. The shooting cycle is two weeks. Images are acquired by extracting one frame every 5 seconds and screening them. Finally, 772 remaining bait images in BMP format were selected. After annotation and preprocessing, these images are used to construct the input dataset of the model. Please refer to Figure 2 , Figure 2 The original image and annotation example diagrams provided for Example 1 of the bait counting method based on the adaptive scale aggregation network of this application. The image on the left is an unprocessed original image, showing a school of fish and bait particles scattered in the water. Due to the lighting conditions and turbid water, the identification and counting of bait are somewhat difficult. The image on the right is a labeled version of the same scene, in which the bait particles are clearly marked with red marks to facilitate subsequent bait counting and analysis. These annotations are usually done manually by professionals and provide accurate ground truth data for model training. Through this comparison, you can intuitively see the distribution of bait in the original image and the precise positioning of bait particles during the annotation process, which is crucial for the development and verification of image-based bait counting algorithms.
[0051] Hybrid filtering is an image preprocessing technique used to improve the contrast between target bait and background in near-infrared images. In this embodiment, hybrid filtering is achieved through a cascade strategy of multiple filters.
[0052] CLAHE (Contrast Limited Adaptive Histogram Equalization) is an image enhancement technique used to further improve image contrast, particularly in low-light environments. In this embodiment, CLAHE is implemented through adaptive adjustment of the Clip Limit (contrast limit threshold) driven by block entropy. Specifically, the image is divided into 32×32 sub-blocks, and the Clip Limit (1.5 to 3.5) is dynamically set based on the local entropy value (range 3 to 6). In low-entropy areas (such as dark areas on the water surface), the Clip Limit is set to 1.5 to suppress noise amplification; in high-entropy areas (such as bait concentrations), the Clip Limit is set to 3.5 to enhance detail. In addition, by limiting the maximum grayscale mapping slope (clipping is triggered when the slope is > 3), overexposure of highly reflective areas (such as metal bait boxes) is prevented, thereby preserving highlight details. Through this dynamic adjustment, CLAHE effectively enhances the overall contrast of the image while avoiding noise amplification and detail loss caused by over-enhancement.
[0053] The enhanced image dataset consists of images processed with hybrid filtering and dynamic contrast-limited adaptive histogram equalization. These images significantly enhance contrast and detail, making it easier to distinguish between target bait and background. This enhanced processing enables the model to more accurately identify and count bait particles, thereby improving bait counting accuracy.
[0054] Please refer to Figure 3 , Figure 3 A schematic diagram of the preprocessing effect provided for Example 1 of the bait counting method based on the adaptive scale aggregation network of this application shows a comparison of the effects of near-infrared images after hybrid filtering and CLAHE preprocessing. The left side is the original image, which shows an image containing bait particles acquired under a low-light environment, in which the contrast between the bait particles and the background is low, and the details are not clear enough. The right side is the preprocessed image, which effectively removes noise and retains image details through hybrid filtering technology, and then applies dynamic CLAHE technology to enhance the contrast of the image, making the distinction between the bait particles and the background more obvious, and the edges and shapes of the particles clearer. This preprocessing step significantly improves the image quality, provides more accurate visual information for subsequent bait identification and counting, and thus improves the counting accuracy and robustness of the entire system.
[0055] It is understood that, first, the computer system performs hybrid filtering on the collected near-infrared image dataset to remove noise from the image while preserving important details. Second, the computer system performs dynamic contrast-limited adaptive histogram equalization on the hybrid-filtered image. By performing block processing and dynamically adjusting the contrast limit based on local entropy, the computer system enhances the contrast between the target bait and the background in the image while avoiding noise amplification and detail loss caused by over-enhancement.
[0056] As an example, the step of performing hybrid filtering and dynamic contrast limited adaptive histogram equalization on the near-infrared image dataset to obtain an enhanced image dataset includes: performing adaptive median filtering on the near-infrared image dataset to obtain a denoised image dataset; performing non-local mean filtering on the denoised image dataset to obtain a smoothed image dataset; and performing dynamic contrast limited adaptive histogram equalization and dark field correction on the smoothed image dataset to obtain an enhanced dataset.
[0057] Adaptive median filtering is a noise removal technique that dynamically adjusts the filter window size. It detects noise pixel by pixel and dynamically expands the filter window based on noise intensity to accurately locate and remove impulse noise. This method can effectively remove hot pixels with a removal rate of up to 98%. The calculation formula for adaptive median filtering is as follows:
[0058]
[0059] Among them, W init is the initial filtering window (3×3), T noise is the impulse noise judgment threshold (30 gray levels), and med() is the window median calculation function.
[0060] The denoised image dataset refers to a collection of images that have been processed with adaptive median filtering. After removing impulse noise, the background noise of these images is significantly reduced, and the signal-to-noise ratio of the images is improved.
[0061] Non-local mean filtering is a denoising method based on image block similarity. It searches for regions in the entire image that are similar to the current pixel block and performs weighted denoising. This method not only removes Gaussian noise but also preserves edges and details while suppressing noise. The calculation formula for non-local mean filtering is as follows:
[0062]
[0063] Among them, P x,y is a pixel block centered at (x, y), P i,j For a block centered at (i, j), G σ is a Gaussian function, d is the Gaussian weighted Euclidean distance between the two, h is the attenuation parameter, Ω is the local window, INLM (x,y) is the weighted average of all similar blocks in the final filter output.
[0064] Smoothed image datasets refer to a collection of images that have been processed with non-local mean filtering. After removing Gaussian noise, the overall smoothness of these images is improved while retaining edge and detail information.
[0065] Dark field correction is a technology used to eliminate image sensor non-uniformity and grayscale baseline drift. By updating the dark field template every 15 frames, it eliminates grayscale baseline drift caused by rising sensor temperature. This method can reduce the image non-uniformity coefficient from 8.7% to 1.2%, thereby improving overall image quality and consistency.
[0066] First, the computer system applies adaptive median filtering to each image in the near-infrared image dataset. The specific operation is as follows: starting with a 3×3 filter window, the noise in the image is detected pixel by pixel. When impulse noise is detected, the filter window size is dynamically expanded to 5×5 or 7×7 until the noise is effectively removed, thereby obtaining a denoised image dataset. This process can effectively remove impulse noise such as hot pixels while retaining important details in the image, providing a clearer image foundation for subsequent processing. Secondly, the system performs non-local mean filtering on the denoised image dataset. Specifically, for each pixel block in the image, the system searches for similar areas in the entire image and calculates the weighted average of these similar areas to replace the value of the current pixel block. In this way, Gaussian noise is removed while enhancing the edge and detail information of the image, resulting in a smooth image dataset and further optimizing image quality. Finally, the system performs CLAHE on the smoothed image dataset to suppress noise amplification in low-entropy areas and enhance details in high-entropy areas. At the same time, dark field correction is performed to eliminate the grayscale baseline drift caused by sensor temperature changes. Ultimately, an enhanced image dataset is obtained, which provides high-quality input data for subsequent bait counting model training and improves the robustness and accuracy of the model in low-light environments.
[0067] As an example, the step of performing dynamic contrast limited adaptive histogram equalization and dark field correction on the smoothed image dataset to obtain an enhanced dataset includes: dividing each frame image in the smoothed image dataset into sub-blocks of a preset size; calculating the local entropy value of the sub-block, and marking the sub-block as a high entropy area or a low entropy area according to the local entropy value and a preset entropy threshold; determining a contrast limiting threshold parameter according to the high entropy area and the low entropy area; performing histogram equalization on the sub-block according to the contrast limiting threshold parameter, and eliminating boundary artifacts between the sub-blocks through bilinear interpolation to obtain a reference dataset; generating a dark field template through a Gaussian blur kernel, and applying the dark field template to the reference dataset to obtain an enhanced dataset.
[0068] The preset size refers to the size of each small block when the image is divided into multiple small blocks.
[0069] Sub-blocks refer to smaller regions into which an image is divided.
[0070] The local entropy value refers to the complexity or information content of the pixel grayscale distribution within a sub-block. Entropy is an indicator of the richness of image information. The higher the local entropy value, the more complex the grayscale changes within the sub-block.
[0071] The preset entropy threshold is a fixed value used to distinguish high entropy regions from low entropy regions.
[0072] The high entropy area refers to the sub-block with a higher local entropy value, which usually indicates that the texture of the area is complex, such as the bait gathering area.
[0073] The low entropy area refers to a sub-block with a low local entropy value, which usually means that the texture of the area is simple, such as the dark area on the water surface.
[0074] The contrast limit threshold parameter refers to a parameter used to control the contrast enhancement degree of each sub-block in CLAHE, and determines the degree of limitation of the contrast enhancement of each sub-block.
[0075] Bilinear interpolation is an image processing technique used to smooth the boundaries between sub-blocks during image processing.
[0076] Boundary artifacts refer to discontinuous or unnatural boundary effects caused by different processing methods between sub-blocks during image processing.
[0077] The reference dataset refers to an image dataset that has been processed with dynamic contrast limited adaptive histogram equalization. The image has been significantly improved in terms of contrast and details, but has not yet been dark field corrected.
[0078] The Gaussian blur kernel is a filter used for image smoothing, and its shape conforms to the Gaussian distribution. In this embodiment, the size of the Gaussian blur kernel is 21×21, and the standard deviation σ=5.
[0079] The dark field template is an image generated using a Gaussian blur kernel and is used to correct for sensor non-uniformity and grayscale baseline drift. In this embodiment, the dark field template is updated every 15 frames to eliminate grayscale baseline drift caused by sensor temperature changes.
[0080] First, the computer system divides each frame in the smoothed image dataset into 32×32 pixel sub-blocks. This allows for local image processing, allowing for targeted contrast adjustments based on the characteristics of each region. Next, the system calculates the local entropy value of each sub-block—the complexity of the pixel grayscale distribution within the sub-block—and compares the local entropy value with a preset entropy threshold (ranging from 3 to 6) to label the sub-block as either a high-entropy or low-entropy region. Specifically, sub-blocks with local entropy values within this range are labeled as high-entropy regions, typically corresponding to areas with complex textures, such as bait concentrations; sub-blocks with local entropy values below this range are labeled as low-entropy regions, typically corresponding to areas with simple textures, such as dark areas on the water surface. The system then determines the contrast limit threshold parameters based on the sub-block labels: the Clip Limit for low-entropy regions is set to 1.5 to suppress noise amplification, and the Clip Limit for high-entropy regions is set to 3.5 to enhance detail. Subsequently, the system performs histogram equalization on each sub-block based on these contrast limit threshold parameters. The specific operation is as follows: for each sub-block, its histogram is calculated and the contrast limit threshold is applied to avoid noise amplification and detail loss caused by over-enhancement. During the histogram equalization process, the system limits the histogram peak of each sub-block to not exceed the set Clip Limit value, and then normalizes the histogram to enhance the contrast within the sub-block. In order to eliminate boundary artifacts between sub-blocks, the system uses bilinear interpolation technology for smooth transition to obtain a reference data set. Finally, the system generates a dark field template through a Gaussian blur kernel and applies the dark field template to the reference data set to correct the sensor non-uniformity and grayscale baseline drift, and finally obtains an enhanced data set, which provides high-quality input data for subsequent bait counting model training.
[0081] Step S20: construct an adaptive scale aggregation network model including a feature map encoder and a density map decoder, wherein the feature map encoder is composed of multiple adaptive scale aggregation modules to perform dynamic weight allocation of multi-scale convolution kernels.
[0082] It should be noted that the feature map encoder is a key component of ASANet. Its main function is to extract useful feature information from the input near-infrared image. Through a series of convolution operations and the ASAM (Adaptive Scale Aggregation Module), it gradually converts the pixel-level information of the input image into high-level feature representations. These feature representations can capture features of different scales and complexities in the image, providing basic data for the subsequent density map decoder. The feature map encoder is designed to effectively extract and aggregate multi-scale features in the image, thereby improving the model's ability to identify and count fish bait.
[0083] The density map decoder is another key component of ASANet. It is responsible for converting the high-level feature information extracted by the feature map encoder into a high-resolution density map of the same size as the input image. A density map is an image that represents the distribution density of objects (such as fish bait) in an image, where the value of each pixel corresponds to the bait density at that location. The density map decoder gradually restores the resolution of the feature map through a series of convolutional layers and transposed convolutional layers (upsampling operations) and refines the details of the feature map, ultimately generating a density map of the same size as the input image, which is used to accurately count bait particles in the image.
[0084] ASANet is a deep learning model specifically designed for counting fish bait in near-infrared images. It consists of two components: a feature map encoder and a density map decoder. ASANet's core advantage lies in its ability to dynamically adjust convolution kernel weights to accommodate image features of varying scale and complexity, thereby improving the model's robustness and accuracy in low-light and complex environments.
[0085] Please refer to Figure 4 , Figure 4 The schematic diagram of the adaptive scale aggregation network structure provided in the first embodiment of the bait counting method based on the adaptive scale aggregation network of the present application includes two main parts: a feature map encoder (FME) and a density map decoder (DME). In the feature map encoder, the input original image first passes through four adaptive scale aggregation modules (ASAM1 to ASAM4), each of which contains convolution layers with convolution kernels of different sizes (such as conv1_32, conv3_32, etc.), as well as a maximum pooling layer for extracting multi-scale features and performing downsampling. Each ASAM module also includes a dynamic convolution selection mechanism to adaptively adjust the weights of different convolution kernels to better capture features of different scales in the image. The feature map processed by the feature map encoder is fed into the density map decoder, which consists of a series of transposed convolution layers (Trans_conv2_1 to Trans_conv2_4) and convolution layers (conv3_16 to conv9_128) to gradually restore the resolution of the feature map and ultimately generate a predicted density map of the same size as the input image. The entire network structure achieves effective aggregation of features of different scales by adaptively adjusting the weights of the convolution kernel, thereby improving the accuracy of fish bait counting in near-infrared images under low-light conditions.
[0086] ASAM is a core component in the feature map encoder, implementing dynamic weight allocation for multi-scale convolution kernels. Each ASAM contains multiple convolution kernels of different sizes, which are used to extract features at different scales in the image. The ASAM module dynamically calculates the weight of each convolution kernel through an attention subnetwork, automatically adjusting the weight allocation based on the feature distribution of the input image.
[0087] Please refer to Figure 5 , Figure 5 The schematic diagram of the adaptive scale aggregation module structure provided in the first embodiment of the bait counting method based on the adaptive scale aggregation network of this application is a key component for feature extraction in ASANet. In ASAM, the input image is first processed by four parallel convolution kernels of different sizes (conv1_64, conv3_64, conv5_64, conv7_64), and each convolution kernel extracts feature information of different scales. Subsequently, the output of these parallel convolution layers is subjected to a dynamic convolution selection mechanism to dynamically adjust the weights of each convolution kernel according to the content of the input image to highlight the most useful features of the current image. Next, the weighted feature map passes through a 2x2 maximum pooling layer to reduce the spatial dimension of the feature map while retaining the most important feature information. This structural design enables ASAM to flexibly adapt to image features of different scales, thereby improving the accuracy and robustness of the network in identifying bait particles in complex scenes.
[0088] Multi-scale convolution kernels refer to a set of convolution kernels of different sizes used in ASAM, including 1×1, 3×3, 5×5, and 7×7. These convolution kernels are used to extract features of different scales in images to accommodate the various levels of aggregation and texture characteristics of bait particles in an image. Small convolution kernels (such as 1×1 and 3×3) are better suited for capturing small targets and detailed information, while large convolution kernels (such as 5×5 and 7×7) are better suited for capturing a wider range of contextual information. By using multi-scale convolution kernels, the model is able to more comprehensively understand and analyze features in images, thereby improving its ability to identify and count fish bait.
[0089] As you can understand, the computer system first constructs the ASANet model, which consists of two components: a feature map encoder and a density map decoder. The feature map encoder's task is to extract rich feature information from the input near-infrared image. To this end, it is composed of multiple ASAMs. Each ASAM module integrates convolution kernels of different sizes to capture features at different scales within the image. The weights of these convolution kernels are not fixed but are dynamically assigned through an attention mechanism. Specifically, the system automatically adjusts the weights of each convolution kernel based on the specific content of the input image. When processing areas with dense bait, the weights of small convolution kernels are increased to accurately capture small targets; while when processing areas with complex backgrounds, the weights of large convolution kernels are increased to better understand contextual information. This dynamic weight allocation mechanism enables the feature map encoder to flexibly adapt to the feature extraction requirements of different scenarios, providing high-quality feature representations for the subsequent density map decoder, thereby improving the accuracy and robustness of the entire model in fish bait counting tasks.
[0090] As an example, the steps of constructing an adaptive scale aggregation network model including a feature map encoder and a density map decoder include: constructing multiple cascaded adaptive scale aggregation modules containing multiple convolution kernels; introducing an attention subnetwork into the adaptive scale aggregation module to obtain a feature map encoder, and the attention subnetwork is used to calculate the weight of the convolution kernel; constructing a density map decoder based on multiple convolution layers of different sizes and multiple transposed convolution layers; connecting the feature map encoder with the density map decoder to form an adaptive scale aggregation network model.
[0091] The attention subnetwork is a key component in ASAM, responsible for dynamically calculating the weights of different convolution kernels. Specifically, the attention subnetwork receives an input feature map and extracts local features at different scales through a series of convolution operations and nonlinear activation functions (such as ReLU). It then fuses these feature maps and normalizes them using the Softmax function to generate a weight map for each convolution kernel, thus achieving dynamic weight allocation.
[0092] Please refer to Figure 6 , Figure 6 The schematic diagram of the attention sub-network structure provided in the first embodiment of the bait counting method based on the adaptive scale aggregation network of this application is a key component of ASAM, which is used to dynamically calculate the weights of different convolution kernels. First, the input feature Figure XFour parallel convolutional branches are convolved using kernels of 1×1, 3×3, 5×5, and 7×7, respectively, to generate initial feature maps F_1×1, F_3×3, F_5×5, and F_7×7. To reduce computational complexity, each branch's feature map undergoes dimensionality reduction, reducing the number of channels to 1 / 4 of the original input. This results in reduced-dimensional feature maps F_1×1_reduced, F_3×3_reduced, F_5×5_reduced, and F_7×7_reduced. These reduced-dimensional features are activated using a ReLU activation function to introduce nonlinearity, generating activated feature maps F_1×1_act, F_3×3_act, F_5×5_act, and F_7×7_act. All activated feature maps are then concatenated and fused along the channel dimension to form a composite feature map that synergizes information at multiple scales. The composite feature map further compresses the number of channels through a 1×1 convolution and is normalized along the spatial-channel dimension using a Softmax function, ultimately generating a dynamic weight map W. The four channels of weight map W correspond to the importance weights of different convolution kernels, ensuring that the sum of the weights at each spatial location is 1. This mechanism enables the network to dynamically allocate the contribution of each convolution kernel based on the adaptive needs of the input features, thereby improving the recognition accuracy of features at different scales.
[0093] Convolutional layers are the fundamental building blocks of deep learning networks, used to extract local features from input data. They generate new feature maps by applying convolutional kernels of varying sizes to the input feature map. Each kernel is responsible for extracting specific local features, such as edges, textures, or shapes. By stacking multiple convolutional layers, the network is able to extract progressively higher-level feature representations.
[0094] The transposed convolution layer (also called the deconvolution layer or upsampling layer) is used to gradually restore the resolution of the feature map to the same size as the input image. The transposed convolution layer expands each pixel of the feature map into a larger area by learning the upsampling weights, thereby increasing the spatial dimension of the feature map.
[0095] First, a computer system constructs multiple cascaded ASAMs, each module integrating convolutional kernels of various sizes. These kernels perform convolution operations on the input feature map in parallel to extract local features at different scales. Second, an attention subnetwork is introduced within each adaptive scale aggregation module. Then, a density map decoder is constructed based on multiple convolutional layers of varying sizes and multiple transposed convolutional layers. The convolutional layers refine the details of the feature map, while the transposed convolutional layers gradually restore the resolution of the feature map, ultimately generating a high-resolution density map of the same size as the input image. Finally, the feature map encoder and density map decoder are connected to form the complete ASANet, which can efficiently extract features from near-infrared images and generate accurate density maps for counting fish bait.
[0096] Step S30: training the adaptive scale aggregation network model according to the enhanced image dataset, Euclidean loss, and local pattern consistency loss to obtain a target model.
[0097] It should be noted that the Euclidean loss (L E ) is a commonly used loss function that measures the estimation error at each pixel level and sums it over the entire sample. The calculation method is as follows:
[0098]
[0099] where N is the number of pixels in the density map, X is the input image, Y is the corresponding ground-truth density map, and F(X;θ) represents the predicted density map and denotes the parameters of the model.
[0100] Local pattern consistency loss (L C ) is a structural similarity loss function used to measure the structural similarity between the predicted density map and the true density map. It is evaluated by calculating the SSIM between the predicted density map and the true density map. This index is designed to capture and strengthen the structural consistency between the predicted results and the true labels in the local area. C The calculation method is as follows:
[0101]
[0102] Among them, C1 and C2 are constants used to prevent the denominator from being zero, μ F (x) is the average value of all pixels in the local window centered at position x in the predicted density map, μ Y (x) is the average value of all pixels in the local window centered at position x of the true density map, is the variance of pixel values in a local window centered at position x in the predicted density map, is the variance of pixel values in a local window centered at position x on the true density map, σYF (x) is the covariance of the pixel values of the predicted density map and the true density map at the same location x.
[0103] Combining Euclidean loss and local pattern consistency loss to construct the adaptive scale aggregation network model loss function:
[0104] L=L E +α(x)·L C
[0105]
[0106] Where α(x) represents L C The weight of , D(x) represents the local density value, μ and σ represent the mean and standard deviation of the density respectively, α max and α min Represent the upper and lower bounds of the weight, namely 0.01 and 0.001 respectively.
[0107] The target model is the trained ASANet. This model is trained on an augmented image dataset and is capable of processing near-infrared images in low-light environments and generating high-resolution density maps for accurate counting of bait particles.
[0108] It can be understood that the computer system first uses the enhanced image dataset as input and feeds these images into the adaptive scale aggregation network model through forward propagation to generate a predicted density map. Next, the system calculates the Euclidean loss between the predicted density map and the true density map. This is done by taking the squared difference between the predicted value and the true value for each pixel and summing them to obtain a numerical measure of overall error. Simultaneously, the system also calculates the local pattern consistency loss, using the structural similarity index to assess the structural similarity between the predicted and true density maps, thereby obtaining another numerical measure of error. The system then combines the Euclidean loss and the local pattern consistency loss to form a comprehensive loss function, which is used to guide model optimization. Finally, through backpropagation and optimization algorithms, the model parameters are adjusted according to the comprehensive loss function, gradually reducing the error. After multiple iterative training, the optimized target model is ultimately obtained, which can more accurately count fish bait in low-light environments.
[0109] As an example, the step of training the adaptive scale aggregation network model according to the enhanced image dataset, Euclidean loss and local pattern consistency loss to obtain a target model includes: iteratively training the adaptive scale aggregation network model according to the enhanced image dataset using an Adam optimizer with a preset learning rate and a preset batch size; in each iterative training, calculating the Euclidean loss and structural similarity index between the predicted density map output by the adaptive scale aggregation network model and the true density map in the enhanced image dataset; calculating the local pattern consistency loss according to the number of pixels in the true density map and the structural similarity index; weightedly fusing the Euclidean loss and the local pattern consistency loss according to the mean and standard deviation of the density distribution of the enhanced image dataset to obtain a joint loss; when the joint loss meets the preset conditions, stopping the training to obtain the target model.
[0110] The Adam optimizer is an adaptive learning rate optimization algorithm widely used in deep learning model training. It can automatically adjust the learning rate to improve training efficiency and stability.
[0111] The preset learning rate is the initial learning rate set before training begins, which controls the step size of the model parameter updates. The learning rate determines the magnitude of the parameter adjustments in each iteration. In this example, the preset learning rate is 1e-6, indicating that the step size of each parameter update is 1e-6.
[0112] The preset batch size refers to the number of samples used to calculate the loss function and update model parameters in each iteration. In this example, the preset batch size is 8, indicating that each iteration uses 8 augmented images and their corresponding true density maps for training. The choice of batch size affects training efficiency and the speed of model convergence.
[0113] The predicted density map is the output of the adaptive scale aggregation network model, where the value of each pixel represents the density of bait at that location. The predicted density map is the result of the model processing the input augmented image and is used to estimate the distribution of bait in the image.
[0114] The ground-truth density map is an annotated image corresponding to an image in the augmented image dataset, where the value of each pixel represents the actual bait density at that location. The ground-truth density map is obtained through manual annotation or calculation and is used to evaluate the accuracy of the model prediction.
[0115] Please refer to Figure 7 , Figure 7The diagram of the comparison of the predicted density map, the true density map and the original image provided in the first embodiment of the bait counting method based on the adaptive scale aggregation network of this application is shown. The left side is the predicted density map generated by ASANet, in which the intensity of the color represents the density of the bait. The red area indicates a higher density of the bait, while the blue area indicates a lower density. The middle is the true density map, which is usually obtained by manual annotation or precise measurement and serves as a benchmark for evaluating the accuracy of the model prediction. The right side is the original image, which shows an image containing bait particles taken by a near-infrared camera in a low-light environment. The figure also marks the predicted total number of baits (54.49) and the actual number of baits (41). By comparison, it can be seen that ASANet can accurately predict the distribution and quantity of baits. Although there is a certain difference between the predicted value and the true value, the overall trend and distribution are consistent with the actual situation, verifying the effectiveness and accuracy of the model in counting baits in a low-light environment.
[0116] SSIM (Structural Similarity Index Measure) is an indicator for measuring the structural similarity between two images, which takes into account the brightness, contrast and structural information of the images.
[0117] The density distribution mean refers to the average value of all pixel values in the true density map, which is used to describe the average level of bait density in the entire image.
[0118] Joint loss refers to a comprehensive loss function that combines Euclidean loss and local pattern consistency loss.
[0119] The precondition refers to the condition used to determine whether to stop training during the training process. In this embodiment, the precondition is that the loss value does not decrease for 10 consecutive rounds. When this condition is met, the model is considered to have converged and the training process stops.
[0120] First, a computer system uses an augmented image dataset as input and iteratively trains an adaptive scale aggregation network model using the Adam optimizer with a preset learning rate and batch size. In each iteration, the system feeds a batch of 8 augmented images into the model, which then outputs a predicted density map. Next, the system calculates the Euclidean loss (the sum of the squared differences between the predicted and true values) between the predicted density map and the corresponding ground-truth density map. The SSIM (Short Sense Mean Squared) is also calculated to measure the structural similarity between the predicted and true density maps. Then, based on the number of pixels in the true density map and the SSIM value, the local pattern consistency loss is calculated, which focuses on the structural consistency between the predicted and true density maps. The system then performs a weighted fusion of the Euclidean and local pattern consistency losses based on the mean and standard deviation of the density distribution of the augmented image dataset to produce a joint loss. Finally, training stops when the joint loss meets a preset condition—i.e., if the loss value does not decrease after 10 consecutive rounds—the resulting model becomes the target model, which can more accurately count fish bait in low-light environments.
[0121] Step S40: input the target bait image into the target model to obtain a bait density map and bait quantity.
[0122] It should be noted that the target bait image refers to the near-infrared image acquired in real time in actual applications, which contains the distribution of fish bait in the water.
[0123] A bait density map is an image output by the target model, where the value of each pixel represents the bait density at that location. The density map is a high-resolution image that visually displays the distribution of bait in the water. By analyzing the density map, the amount and location of bait can be accurately calculated.
[0124] Please refer to Figure 8 , Figure 8 The following is an example diagram of the collected and processed images and density maps provided in Example 1 of the bait counting method based on the adaptive scale aggregation network of this application. The left side of the figure is a collected near-infrared image, which shows the distribution of fish and bait in the aquarium. Due to the lighting conditions and turbid water, the identification and counting of bait particles are somewhat difficult. The right side is a processed density map, in which the intensity of the color represents the density of the bait particles, the red and yellow areas indicate a higher density of bait, and the blue area indicates a lower density. This density map is generated by ASANet, which can analyze near-infrared images and identify the position and number of bait particles, thereby generating an image reflecting the distribution density of the bait. Through this processing, the distribution of bait in the aquarium can be more intuitively understood, providing accurate data support for the intelligent feeding system, and helping to improve breeding efficiency and management accuracy.
[0125] Bait count refers to the total number of bait pellets obtained by analyzing the bait density map. The target model calculates the number of bait pellets based on the pixel values in the density map by adding up all the pixel values in the density map to get the total number of bait pellets.
[0126] As can be understood, the computer system first acquires real-time images of the target bait. These images, captured by a near-infrared camera in a low-light environment, depict the distribution of fish bait in the water. These images are then fed into a trained target model, which processes the input images using its feature map encoder and density map decoder to generate a bait density map. Finally, the system analyzes the generated bait density map to calculate the total number of bait particles, thereby determining the bait count.
[0127] This embodiment provides a bait counting method based on an adaptive scale aggregation network. First, a near-infrared image dataset is subjected to hybrid filtering and dynamic contrast-limited adaptive histogram equalization to remove noise and enhance image contrast, thereby obtaining an enhanced image dataset. This provides high-quality input data for subsequent model training. Next, an adaptive scale aggregation network model including a feature map encoder and a density map decoder is constructed. The feature map encoder dynamically allocates weights of multi-scale convolution kernels through multiple adaptive scale aggregation modules, thereby improving the model's adaptability to features of different scales. Then, the model is trained based on the enhanced image dataset, Euclidean loss, and local pattern consistency loss to obtain a target model. This process optimizes the performance of the model. Finally, the target bait image is input into the target model to obtain a bait density map and bait quantity, thereby improving the accuracy of fish bait counting in near-infrared images under low illumination environments.
[0128] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 9 , Figure 9 This is a flow chart of the second embodiment of the bait counting method based on the adaptive scale aggregation network of the present application. The target model includes a feature map encoder and a density map decoder. Step S40 of the method includes steps S41 to S44:
[0129] Step S41 , extracting features from the target bait image through the feature mapping encoder to obtain multi-scale features.
[0130] It should be noted that multi-scale features refer to feature information extracted from different scales or resolution levels in image processing and analysis. In this embodiment, multi-scale features refer to feature representations at multiple scales extracted from the target bait image using a feature map encoder. These features can capture details of varying sizes and complexities within the image. For example, small-scale features can capture tiny bait particles, while large-scale features can capture the distribution and overall structure of the bait.
[0131] As you can understand, the target bait image is first input into the feature map encoder. The feature map encoder is composed of multiple ASAM modules, each containing convolution kernels of varying sizes. These kernels perform convolution operations on the input image in parallel, extracting local features at different scales. Through this multi-scale convolution operation, the feature map encoder can simultaneously extract both detailed information and contextual information from the image, ultimately generating a multi-scale feature map.
[0132] Step S42: reconstructing the resolution of the multi-scale features through the density mapping decoder to obtain a bait density map.
[0133] As you can understand, the multi-scale features extracted by the feature map encoder are first input into the density map decoder. The density map decoder consists of multiple convolutional and transposed convolutional layers, which work together to gradually restore the resolution of the feature map. The convolutional layers refine the details of the feature map and extract higher-level feature representations; the transposed convolutional layers gradually upscale the low-resolution feature map to the same resolution as the input image through upsampling. In this process, the decoder leverages the details and contextual information in the multi-scale features to generate a high-resolution bait density map, where the value of each pixel represents the bait density at that location.
[0134] Step S43, performing region segmentation on the bait density map using a region growing algorithm to obtain connected regions.
[0135] It should be noted that the region growing algorithm is an image segmentation technology that starts from a set of seed pixels and gradually merges adjacent pixels with similar features (such as grayscale value, color or texture) into the same region to achieve image segmentation.
[0136] A connected region is a set of pixels in an image that are connected by some connectivity rule. In image segmentation, a connected region usually represents an independent object or part in the image.
[0137] It can be understood that, first, the system analyzes the bait density map and selects pixels with grayscale values higher than a preset threshold as seed pixels. These seed pixels are usually located in the central area of the bait particles and can better represent the presence of bait. Secondly, starting from each seed pixel, the system checks its adjacent pixels according to the 4-connected or 8-connected rule. If the difference between the grayscale value of the adjacent pixel and the grayscale value of the seed pixel is less than the preset similarity threshold (for example, the grayscale difference is less than 10), the adjacent pixel is considered to belong to the same area as the seed pixel and is merged into the currently growing connected area. This process will be repeated until all adjacent pixels that meet the similarity conditions are merged into the connected area, thereby completing the growth of a connected area. Finally, the system continues to search for unassigned pixels in the bait density map and repeats the above growth process until all pixels in the entire density map are assigned to the corresponding connected areas, and finally obtains multiple independent connected areas, each of which represents a bait particle or a group of tightly arranged bait particles, thereby achieving regional segmentation of the bait density map.
[0138] Step S44: determining the amount of bait according to the integral density value of the connected area.
[0139] It should be noted that the integrated density value refers to the sum of the density values of all pixels within each connected region in the bait density map. Specifically, each pixel value in the density map represents the bait density at that location, and the sum of all pixel values within a connected region is the integrated density value of that connected region. By calculating the integrated density value of each connected region, the total amount of bait in that region can be quantified, thereby determining the bait quantity.
[0140] It can be understood that first, the system traverses the pixels in each connected area, adds up the density value of each pixel, and obtains the integral density value of the connected area. Then, the system converts the integral density value into a specific amount of bait based on the preset mapping relationship between density and quantity. For example, if the average density value corresponding to each bait particle is determined in advance through experiments, the amount of bait can be obtained by dividing the integral density value of the connected area by the average density value. Finally, the system accumulates the amount of bait in all connected areas to obtain the total amount of bait in the entire bait density map.
[0141] As an example, the step of determining the amount of bait based on the integral density value of the connected area includes: calculating the sum of the density values of all pixels in each connected area to obtain the integral density value; obtaining the number of bait particles corresponding to each connected area based on the ratio of the integral density value and the preset unit density threshold; and accumulating the number of bait particles in all the connected areas to obtain the amount of bait.
[0142] The preset unit density threshold is the average density value corresponding to each bait pellet in the bait density graph. This value is set experimentally or in advance and is used to convert the integrated density value into the number of bait pellets.
[0143] The number of bait particles refers to the number of bait particles calculated based on the integral density value in each connected area.
[0144] First, the system traverses all pixels within each connected region and accumulates the density values of each pixel to obtain the integrated density value of the connected region. This step is to quantify the total amount of bait in each connected region. Second, the system divides the integrated density value of each connected region by the preset unit density threshold to calculate the number of bait particles corresponding to each connected region. Finally, the system adds up the number of bait particles in all connected regions to obtain the total number of bait in the entire bait density map.
[0145] This embodiment first extracts features from the target bait image through a feature mapping encoder to obtain multi-scale features. This step can capture details of different sizes and complexities in the image, providing a rich information basis for subsequent processing, thereby improving the model's ability to recognize bait features. Next, a density mapping decoder is used to reconstruct the resolution of the multi-scale features to generate a bait density map of the same size as the original image, making the density information more intuitive and accurate. Then, a region growing algorithm is used to segment the bait density map to obtain connected regions, which helps to distinguish different bait particles or aggregation areas and provide a basis for accurate counting. Finally, the number of baits is determined based on the integral density value of the connected area. By quantifying the sum of the density of each area and converting it into the number of particles, accurate statistics of the number of baits are achieved, thereby improving the accuracy of fish bait counting in near-infrared images under low-light conditions and providing reliable data support for intelligent breeding.
[0146] This application also provides a bait counting device based on an adaptive scale aggregation network, please refer to Figure 10 , the bait counting device based on the adaptive scale aggregation network includes:
[0147] The data enhancement module 10 is used to perform hybrid filtering and dynamic contrast limited adaptive histogram equalization on the near-infrared image dataset to obtain an enhanced image dataset;
[0148] A model construction module 20 is used to construct an adaptive scale aggregation network model including a feature map encoder and a density map decoder, wherein the feature map encoder is composed of multiple adaptive scale aggregation modules to perform dynamic weight allocation of multi-scale convolution kernels;
[0149] A model training module 30 is used to train the adaptive scale aggregation network model according to the enhanced image dataset, Euclidean loss and local pattern consistency loss to obtain a target model;
[0150] The bait counting module 40 is used to input the target bait image into the target model to obtain the bait density map and the bait quantity.
[0151] In one embodiment, the model construction module 20 is further used to construct multiple cascaded adaptive scale aggregation modules containing multiple convolution kernels; introduce an attention subnetwork into the adaptive scale aggregation module to obtain a feature map encoder, and the attention subnetwork is used to calculate the weights of the convolution kernels; construct a density map decoder based on multiple convolution layers of different sizes and multiple transposed convolution layers; and connect the feature map encoder with the density map decoder to form an adaptive scale aggregation network model.
[0152] In one embodiment, the bait counting module 40,
[0153] The feature map encoder is used to extract features of the target bait image to obtain multi-scale features; the density map decoder is used to reconstruct the resolution of the multi-scale features to obtain a bait density map; the bait density map is segmented into regions using a region growing algorithm to obtain connected regions; and the amount of bait is determined based on the integral density value of the connected region.
[0154] In one embodiment, the bait counting module 40 is also used to calculate the sum of the density values of all pixels in each of the connected areas to obtain an integral density value; based on the ratio of the integral density value and a preset unit density threshold, the number of bait particles corresponding to each of the connected areas is obtained; and the number of bait particles in all the connected areas is accumulated to obtain the amount of bait.
[0155] In one embodiment, the data enhancement module 10 is further used to perform adaptive median filtering on the near-infrared image dataset to obtain a denoised image dataset; perform non-local mean filtering on the denoised image dataset to obtain a smoothed image dataset; and perform dynamic contrast limited adaptive histogram equalization and dark field correction on the smoothed image dataset to obtain an enhanced dataset.
[0156] In one embodiment, the data enhancement module 10 is further used to divide each frame image in the smoothed image dataset into sub-blocks of a preset size; calculate the local entropy value of the sub-block, and mark the sub-block as a high entropy area or a low entropy area according to the local entropy value and a preset entropy threshold; determine a contrast limiting threshold parameter according to the high entropy area and the low entropy area; perform histogram equalization on the sub-block according to the contrast limiting threshold parameter, and eliminate boundary artifacts between the sub-blocks through bilinear interpolation to obtain a reference dataset; generate a dark field template through a Gaussian blur kernel, and apply the dark field template to the reference dataset to obtain an enhanced dataset.
[0157] In one embodiment, the model training module 30 is further used to iteratively train the adaptive scale aggregation network model based on the enhanced image dataset using an Adam optimizer with a preset learning rate and a preset batch size; in each iterative training, the Euclidean loss and the structural similarity index between the predicted density map output by the adaptive scale aggregation network model and the true density map in the enhanced image dataset are calculated; the local pattern consistency loss is calculated based on the number of pixels in the true density map and the structural similarity index; the Euclidean loss and the local pattern consistency loss are weightedly fused based on the mean and standard deviation of the density distribution of the enhanced image dataset to obtain a joint loss; when the joint loss meets a preset condition, the training is stopped to obtain the target model.
[0158] The adaptive scale aggregation network-based bait counting device provided in this application, which utilizes the adaptive scale aggregation network-based bait counting method described in the aforementioned embodiments, can address the technical problem of improving the accuracy of fish bait counting in near-infrared images under low-light conditions. Compared to the prior art, the beneficial effects of the adaptive scale aggregation network-based bait counting device provided in this application are the same as those of the adaptive scale aggregation network-based bait counting method described in the aforementioned embodiments. Other technical features of the adaptive scale aggregation network-based bait counting device are the same as those disclosed in the aforementioned embodiments and are not further elaborated here.
[0159] The present application provides a bait counting device based on an adaptive scale aggregation network. The bait counting device based on the adaptive scale aggregation network includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the bait counting method based on the adaptive scale aggregation network in the above-mentioned embodiment 1.
[0160] Reference below Figure 11, which shows a schematic diagram of the structure of a bait counting device based on an adaptive scalable aggregation network suitable for implementing embodiments of the present application. The bait counting device based on an adaptive scalable aggregation network in the embodiments of the present application can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 11 The bait counting device based on the adaptive scale aggregation network shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0161] like Figure 11 As shown, the bait counting device based on the adaptive scaling aggregation network may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in ROM (Read Only Memory) 1002 or programs loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the bait counting device based on the adaptive scaling aggregation network. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007, such as a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008, such as an LCD (Liquid Crystal Display), speaker, vibrator, etc.; storage devices 1003, such as a magnetic tape, hard disk, etc.; and communication devices 1009. The communication devices 1009 can allow the bait counting device based on the adaptive scaling convergent network to communicate with other devices wirelessly or wired to exchange data. Although the figure shows a bait counting device based on the adaptive scaling convergent network with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented or have alternatively.
[0162] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0163] The adaptive scale aggregation network-based bait counting device provided in this application, which utilizes the adaptive scale aggregation network-based bait counting method described in the aforementioned embodiment, can address the technical problem of improving the accuracy of fish bait counting in near-infrared images under low-light conditions. Compared to the prior art, the beneficial effects of the adaptive scale aggregation network-based bait counting device provided in this application are the same as those of the adaptive scale aggregation network-based bait counting method described in the aforementioned embodiment. Other technical features of this adaptive scale aggregation network-based bait counting device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.
[0164] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0165] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0166] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer program) stored thereon, and the computer-readable program instructions are used to execute the bait counting method based on the adaptive scale aggregation network in the above embodiment.
[0167] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or Flash memory), optical fiber, CD-ROM (CD-Read Only Memory, portable compact disk read-only memory), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0168] The computer-readable storage medium may be included in the bait counting device based on the adaptive scale aggregation network; or may exist independently without being assembled into the bait counting device based on the adaptive scale aggregation network.
[0169] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by a bait counting device based on an adaptive scale aggregation network, the bait counting device based on the adaptive scale aggregation network: performs hybrid filtering and dynamic contrast limited adaptive histogram equalization on a near-infrared image data set to obtain an enhanced image data set; constructs an adaptive scale aggregation network model including a feature map encoder and a density map decoder, wherein the feature map encoder is composed of multiple adaptive scale aggregation modules to perform dynamic weight allocation of multi-scale convolution kernels; trains the adaptive scale aggregation network model according to the enhanced image data set, Euclidean loss, and local pattern consistency loss to obtain a target model; and inputs a target bait image into the target model to obtain a bait density map and bait quantity.
[0170] The computer program code for performing the operations of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).
[0171] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0172] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0173] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned bait counting method based on an adaptive scale aggregation network. This computer-readable storage medium addresses the technical problem of improving the accuracy of fish bait counting in near-infrared images under low-light conditions. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the bait counting method based on an adaptive scale aggregation network provided in the aforementioned embodiments, and are not further elaborated here.
[0174] The present application also provides a computer program product, comprising a computer program, which implements the steps of the above-mentioned bait counting method based on the adaptive scale aggregation network when executed by a processor.
[0175] The computer program product provided in this application can solve the technical problem of improving the accuracy of fish bait counting in near-infrared images in low-light environments. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the bait counting method based on the adaptive scale aggregation network provided in the above-mentioned embodiment, and will not be further elaborated here.
[0176] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A bait counting method based on an adaptive scale aggregation network, characterized in that: The method comprises: The near-infrared image dataset is subjected to hybrid filtering and dynamic contrast limited adaptive histogram equalization to obtain an enhanced image dataset; Constructing an adaptive scale aggregation network model including a feature map encoder and a density map decoder. The feature map encoder is composed of multiple adaptive scale aggregation modules to perform dynamic weight allocation of multi-scale convolution kernels. Training the adaptive scale aggregation network model according to the enhanced image dataset, Euclidean loss, and local pattern consistency loss to obtain a target model; The target bait image is input into the target model to obtain a bait density map and bait quantity.
2. The method according to claim 1, wherein The steps of constructing an adaptive scale aggregation network model including a feature map encoder and a density map decoder include: Construct multiple cascaded adaptive scale aggregation modules containing various convolution kernels; Introducing an attention sub-network into the adaptive scale aggregation module to obtain a feature map encoder, wherein the attention sub-network is used to calculate the weight of the convolution kernel; Construct a density mapping decoder based on multiple convolutional layers of different sizes and multiple transposed convolutional layers; The feature map encoder is connected to the density map decoder to form an adaptive scale aggregation network model.
3. The method according to claim 1, wherein The target model includes a feature map encoder and a density map decoder; The step of inputting the target bait image into the target model to obtain the bait density map and the bait quantity comprises: Extracting features from the target bait image using the feature mapping encoder to obtain multi-scale features; Reconstructing the resolution of the multi-scale features by the density mapping decoder to obtain a bait density map; Performing region segmentation on the bait density map using a region growing algorithm to obtain connected regions; The amount of bait is determined according to the integral density value of the connected area.
4. The method according to claim 3, wherein The step of determining the amount of bait according to the integrated density value of the connected area comprises: Calculating the sum of density values of all pixels in each connected area to obtain an integrated density value; Obtaining the number of bait particles corresponding to each of the connected areas according to the ratio of the integrated density value to a preset unit density threshold; The number of bait particles in all the connected areas is accumulated to obtain the bait quantity.
5. The method according to claim 1, wherein The step of performing hybrid filtering and dynamic contrast limited adaptive histogram equalization on the near-infrared image dataset to obtain an enhanced image dataset includes: Adaptive median filtering is performed on the near-infrared image dataset to obtain a denoised image dataset; Performing non-local mean filtering on the denoised image dataset to obtain a smoothed image dataset; Dynamic contrast limited adaptive histogram equalization and dark field correction are performed on the smoothed image dataset to obtain an enhanced dataset.
6. The method according to claim 5, wherein The step of performing dynamic contrast limited adaptive histogram equalization and dark field correction on the smoothed image data set to obtain an enhanced data set comprises: Dividing each frame of the smoothed image data set into sub-blocks of a preset size; Calculating a local entropy value of the sub-block, and marking the sub-block as a high entropy area or a low entropy area according to the local entropy value and a preset entropy threshold; Determining a contrast limit threshold parameter according to the high entropy region and the low entropy region; performing histogram equalization processing on the sub-blocks according to the contrast limit threshold parameter, and eliminating boundary artifacts between the sub-blocks by bilinear interpolation to obtain a reference data set; A dark field template is generated by using a Gaussian blur kernel, and the dark field template is applied to the reference dataset to obtain an enhanced dataset.
7. The method according to any one of claims 1 to 6, characterized in that The step of training the adaptive scale aggregation network model according to the enhanced image dataset, Euclidean loss, and local pattern consistency loss to obtain a target model comprises: Iteratively training the adaptive scale aggregation network model using an Adam optimizer with a preset learning rate and a preset batch size according to the enhanced image dataset; In each iterative training, calculating the Euclidean loss and the structural similarity index between the predicted density map output by the adaptive scale aggregation network model and the true density map in the enhanced image dataset; Calculating a local pattern consistency loss based on the number of pixels in the true density map and the structural similarity index; Performing weighted fusion on the Euclidean loss and the local pattern consistency loss according to the density distribution mean and standard deviation of the enhanced image dataset to obtain a joint loss; When the combined loss meets the preset conditions, training is stopped and the target model is obtained.
8. A bait counting device based on an adaptive scale aggregation network, characterized in that: The device comprises: A data enhancement module is used to perform hybrid filtering and dynamic contrast limited adaptive histogram equalization on the near-infrared image dataset to obtain an enhanced image dataset; A model building module for building an adaptive scale aggregation network model including a feature map encoder and a density map decoder. The feature map encoder is composed of multiple adaptive scale aggregation modules to perform dynamic weight allocation of multi-scale convolution kernels. A model training module, configured to train the adaptive scale aggregation network model according to the enhanced image dataset, Euclidean loss, and local pattern consistency loss to obtain a target model; The bait counting module is used to input the target bait image into the target model to obtain the bait density map and the bait quantity.
9. A bait counting device based on an adaptive scale aggregation network, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the bait counting method based on an adaptive scale aggregation network according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the bait counting method based on an adaptive scale aggregation network as claimed in any one of claims 1 to 7 are implemented.