A fish school biomass estimation method and system based on a multi-modal large model
Patent Information
- Application Number
- CN202610905817.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]目前,水产养殖领域中常用的鱼群生物量估算方法主要有传统图像处理的估算方法、基于声呐与传感器的估算方法和基于单模态深度学习的估算方法,然而,这些鱼群生物量估算方法普遍准确性较低,无法精准地估算出鱼群的生物量
本申请提供了一种基于多模态大模型的鱼群生物量估算方法及系统,通过采集水下鱼群图像数据、声呐数据和环境数据等多模态数据,并引入多模态大模型,对多模态数据进行预处理与特征提取、跨模态数据融合以及动态目标跟踪与遮挡补全,从而实现多模态数据之间的跨模态数据融合,以及鱼群跟踪和鱼群遮挡区域的补全,使得鱼群生物量估算的基础数据更加全面、完整、清晰,基于此,再采用回归模型进行鱼群生物量估算,可以得到更加准确、可靠的鱼群生物量估算结果,有效提升了鱼群生物量估算的准确性。
Smart Images

Figure CN122594735A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fish biomass estimation technology, and in particular to a method and system for estimating fish biomass based on a multimodal large model. Background Technology
[0002] Accurate estimation of underwater biomass not only improves the production efficiency and economic benefits of aquaculture but also promotes environmental protection and sustainable development. Biomass estimation allows for the rational allocation of stocking densities, optimization of water quality management, and prevention of water quality deterioration and fish diseases caused by over-farming. It also enables precise feed delivery based on actual needs, avoiding waste and reducing feed costs. Furthermore, accurate biomass data helps predict harvest time and yield, allowing for the development of reasonable sales plans using big data analytics and intelligent management systems, thereby enhancing economic efficiency.
[0003] Currently, commonly used methods for estimating fish biomass in aquaculture include traditional image processing methods, sonar and sensor-based methods, and single-modal deep learning-based methods. However, these methods generally have low accuracy and cannot accurately estimate the biomass of fish populations. Therefore, improving the accuracy of fish biomass estimation has become a pressing technical problem in this field. Summary of the Invention
[0004] The purpose of this application is to provide a method and system for estimating fish biomass based on a multimodal large model, which can improve the accuracy of fish biomass estimation.
[0005] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a method for estimating fish biomass based on a multimodal large model, comprising the following steps: Collect multimodal data of the target area, including underwater fish image data, sonar data, and environmental data.
[0006] Based on the multimodal large model, the multimodal data is preprocessed and feature extracted, cross-modal data fused, and dynamic target tracking and occlusion completion are performed sequentially to obtain target tracking and completion information; the target tracking and completion information refers to the information obtained after tracking the fish school and completing the occluded areas of the fish school.
[0007] Based on the target tracking and completion information, a regression model is used to estimate the fish biomass, and the fish biomass estimation results are obtained.
[0008] Secondly, this application provides a fish biomass estimation system based on a multimodal large model, including a multimodal data acquisition module, a data preprocessing and feature extraction module, a cross-modal attention mechanism fusion module, a dynamic target tracking and occlusion completion module, and a biomass regression and result output module connected in sequence. The multimodal data acquisition module is used to collect multimodal data of the target area, including underwater fish image data, sonar data, and environmental data.
[0009] The data preprocessing and feature extraction module, the cross-modal attention mechanism fusion module, and the dynamic target tracking and occlusion completion module are used to perform preprocessing and feature extraction, cross-modal data fusion, and dynamic target tracking and occlusion completion on the multimodal data based on a multimodal large model, to obtain target tracking and completion information; the target tracking and completion information refers to the information obtained after tracking the fish school and completing the occluded areas of the fish school.
[0010] The biomass regression and result output module is used to estimate the biomass of the fish population using a regression model based on the target tracking and completion information, and to obtain the biomass estimation result of the fish population.
[0011] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a method and system for estimating fish biomass based on a multimodal large model. By collecting multimodal data such as underwater fish image data, sonar data, and environmental data, and introducing a multimodal large model, the multimodal data is preprocessed and feature extracted, cross-modal data fusion is performed, and dynamic target tracking and occlusion completion are carried out. This achieves cross-modal data fusion between multimodal data, as well as fish tracking and completion of fish occlusion areas, making the basic data for fish biomass estimation more comprehensive, complete, and clear. Based on this, a regression model is then used for fish biomass estimation, which can yield more accurate and reliable fish biomass estimation results, effectively improving the accuracy of fish biomass estimation. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart illustrating a method for estimating fish biomass based on a multimodal large model, provided as an embodiment of this application.
[0014] Figure 2This is a schematic diagram of the structure of a fish biomass estimation system based on a multimodal large model, provided in an embodiment of this application. Detailed Implementation
[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0016] Currently, the challenges faced by traditional machine learning and deep learning algorithms in aquaculture biomass estimation involve multiple aspects, including image quality, fish movement and occlusion, partial information processing, data acquisition and annotation, and real-time processing. First, underwater lighting conditions are complex and highly variable, with light refraction, reflection, and scattering often affecting image quality. Light intensity varies at different depths, and water turbidity and suspended particulate matter also impact the clarity of fish images. Fish swim relatively fast, easily causing motion blur, which negatively affects image clarity and subsequent recognition. Second, the distance of fish from the underwater camera causes significant visual size variations; distant fish appear smaller, while closer fish appear larger, requiring models to consider scale inconsistencies when estimating biomass. Furthermore, fish often occlude each other, making the identification and estimation of individual fish difficult, especially in dense schools of fish. Accurately estimating the weight of a fish when only part of its body is visible due to occlusion or movement presents a significant challenge. Finally, acquiring and annotating underwater fish images and videos is costly, and high-quality labeled data is required to train the model; insufficient data can affect the model's generalization ability. In practical applications, especially in large-scale fishing grounds, the model needs to be able to quickly process large amounts of video data and output results; real-time performance is crucial.
[0017] Traditional image processing estimation methods involve acquiring images via underwater cameras, extracting fish outlines using thresholding and morphological operations, combining this with statistical methods such as target detection algorithms like YOLO (You Only Look Once) and Faster R-CNN (Fast Region Convolutional Neural Network), and then estimating biomass using a volume-weight model. However, these methods are significantly affected by underwater image quality, occlusion, and lighting variations, resulting in low accuracy. Sonar and sensor-based estimation methods indirectly infer fish biomass through sonar echo signals or environmental sensors (dissolved oxygen, temperature, etc.). While unaffected by water quality, these methods have low accuracy and cannot effectively distinguish between different fish species. Single-modal deep learning-based estimation methods utilize single image or video data and employ convolutional neural networks (CNNs) to predict fish quantity or density. These methods rely on large amounts of labeled data for model training and struggle to cope with complex underwater environments and occlusion issues, thus also resulting in low accuracy.
[0018] In summary, the main drawbacks of the commonly used fish biomass estimation methods are as follows: (1) Image quality sensitivity: Uneven underwater lighting and turbid water quality make traditional image processing algorithms less stable in practical applications and prone to misjudgment. (2) Occlusion and scale issues: Fish swarms are subject to occlusion and changes in scale at different distances, which leads to a decrease in the accuracy of target detection models. (3) Strong data dependence: Many methods rely on a large amount of labeled data to train models, which is cumbersome and costly, and the existing models have poor generalization ability. (4) Insufficient real-time performance: Traditional fish biomass estimation models are computationally complex and slow, which cannot meet the real-time processing needs of large-scale fish farms. (5) Lack of multimodal fusion: Most existing technologies rely on a single data source (images or sonar) and fail to effectively utilize the complementarity of multiple modal data.
[0019] Based on this, this embodiment aims to propose a fish biomass estimation method based on a multimodal large model. Utilizing multimodal large model technology for fish biomass estimation improves the accuracy of biomass estimation and addresses various problems associated with the aforementioned methods. The multimodal large model enables the collection, cleaning, extraction, and integration of multimodal fishery data, providing comprehensive and accurate intelligent support for aquaculture workers, managers, and decision-makers. The multimodal large model is mainly divided into four modules: "Ask Me," "Listen to Me," "See Me," and "Decision Maker," representing four major decision-making scenarios: text, voice, video, and IoT, respectively. Users can query different applications in fisheries. This approach leverages various modern technologies to achieve the digitalization and unmanned operation of aquaculture.
[0020] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] like Figure 1 As shown, this embodiment provides a method for estimating fish biomass based on a multimodal large model, specifically including the following steps: Step S1: Collect multimodal data of the target area. Multimodal data includes underwater fish image data, sonar data, and environmental data, etc.
[0022] In this embodiment, step S1, which involves collecting multimodal data of the target area, specifically includes the following steps: Step S11: Use an underwater camera to collect dynamic underwater fish image data of the target area, record the density, swimming speed and morphological changes of the fish, and obtain underwater fish image data.
[0023] Step S12: Use a sonar array to collect the location, movement and shape information of the underwater fish in the target area to obtain sonar data.
[0024] Step S13: Using environmental sensors such as temperature sensors, dissolved oxygen sensors, and turbidity sensors, parameters such as water temperature, dissolved oxygen, and turbidity are collected to obtain environmental data.
[0025] Step S2: Based on the multimodal large model, the multimodal data is preprocessed and feature extracted sequentially, cross-modal data fusion is performed, and dynamic target tracking and occlusion completion are carried out to obtain target tracking and completion information. Here, target tracking and completion information refers to the information obtained after tracking the fish school and completing the occluded areas of the fish school.
[0026] In this embodiment, step S2, based on a multimodal large model, sequentially performs preprocessing and feature extraction, cross-modal data fusion, and dynamic target tracking and occlusion completion on the multimodal data to obtain target tracking and completion information. Specifically, it includes the following steps: Step S21: Preprocess and extract features from the multimodal data to obtain image features, sonar features, and environmental features.
[0027] In this embodiment, step S21 preprocesses and extracts features from the multimodal data to obtain image features, sonar features, and environmental features, specifically including the following steps: Step S211: Perform illumination compensation and descattering processing on the underwater fish school image data to obtain the underwater fish school image data after illumination compensation and descattering.
[0028] Step S212: Perform edge detection and target segmentation on the underwater fish school image data after illumination compensation and descattering, and extract fish school features, including the morphological features, contour features and movement speed features of the fish school.
[0029] Step S213: Perform scale normalization on the fish school features to obtain scale-normalized fish school features, which are used as image features.
[0030] Step S214: Use the Kalman filter algorithm to remove background noise from the sonar data to obtain denoised sonar data.
[0031] Step S215: Extract features from the denoised sonar data to obtain sonar features.
[0032] Step S216: Standardize the environmental data to obtain standardized environmental data.
[0033] Step S217: Extract features from the standardized environmental data to obtain environmental features.
[0034] Step S22: Using a cross-modal attention mechanism model, perform cross-modal data fusion on the image features, the sonar features, and the environmental features to obtain multimodal fusion features.
[0035] In this embodiment, step S22 utilizes a cross-modal attention mechanism model to perform cross-modal data fusion on the image features, the sonar features, and the environmental features to obtain multimodal fused features. Specifically, this includes the following steps: Step S221: Construct a cross-modal attention mechanism model; the cross-modal attention mechanism model refers to a model with a Transformer architecture used for cross-modal data fusion.
[0036] Step S222: Based on the image features, sonar features, and environmental features in the multimodal data, train the cross-modal attention mechanism model to obtain a trained cross-modal attention mechanism model.
[0037] In this embodiment, step S222 trains the cross-modal attention mechanism model based on the image features, sonar features, and environmental features in the multimodal data to obtain a trained cross-modal attention mechanism model, specifically including the following steps: The image features, sonar features, and environmental features from the multimodal data are input into the cross-modal attention mechanism model for training. The cross-modal attention mechanism model learns the correlation between the image features, sonar features, and environmental features. Combining the temporal information and spatial features in the multimodal data, the model uses a multi-head self-attention mechanism to process the image features, sonar features, and environmental features in parallel, and dynamically adjusts the fusion weights of each modality in the multimodal data. After training, a trained cross-modal attention mechanism model is obtained.
[0038] Step S223: Input the image features, the sonar features, and the environmental features into the trained cross-modal attention mechanism model to obtain multimodal fusion features.
[0039] In this embodiment, the multimodal large model is based on a multimodal data fusion architecture. Through deep integration of a cross-modal attention mechanism and a Transformer architecture, it addresses the heterogeneity problem between multimodal data. Furthermore, precise time synchronization, spatial calibration, and data preprocessing techniques ensure effective alignment of underwater fish image data, sonar data, and environmental data. Building upon this foundation, the model achieves high-precision, multimodal data fusion for underwater biomass estimation tasks through efficient feature extraction, information alignment, and accuracy enhancement. This provides a practical technical path for intelligent monitoring and biomass estimation in large-scale aquaculture environments. The cross-modal attention mechanism and Transformer architecture solve the synchronization problem between multimodal data sources, handle the heterogeneity of different modalities, and perform efficient information alignment and feature fusion.
[0040] In this embodiment, underwater fish image data comes from an underwater camera, recording the fish school and its movement, capturing images of the fish's motion. Sonar data comes from an underwater sonar array, capturing information such as the distance, depth, and shape of the fish school. Environmental data is collected through environmental sensors, including water temperature, dissolved oxygen, and turbidity. For sonar data collected by the underwater sonar array, underwater LiDAR point cloud data can be used for modeling, providing high-precision 3D data to replace sonar data for fish school identification. To ensure temporal consistency of data from different modalities, all sensors and devices have accurate timestamps. The synchronization strategy uses a timestamp synchronization algorithm, aligning time using a global clock or device local timestamps. To align the data from the underwater camera, sonar array, and environmental sensors, the spatial position and viewing angle of each device are calibrated. Spatial consistency of the data source is ensured through sensor position information and viewing angle calibration methods. Each modal multimodal data acquisition module is equipped with a high-precision time synchronization system (based on NTP or PTP protocol synchronization technology) to ensure time alignment of data from different modalities. Hardware-triggered synchronization technology is used to ensure the real-time nature of image and sonar acquisition. All data can be aligned using a unified timeframe, reducing discrepancies between modalities. Due to significant differences between data from different modalities (underwater fish image data, sonar data, and environmental data), each modal's data undergoes independent preprocessing and feature extraction to ensure effective fusion.
[0041] This embodiment employs illumination compensation and descattering techniques to address noise and illumination issues in underwater fish images. A convolutional neural network (CNN) is used for edge detection and target segmentation to extract key fish features. Classical target detection models such as YOLO or Faster R-CNN are used to extract information such as fish contours, size, and speed. The fish and background in the underwater fish image are converted into high-dimensional feature vectors, which will serve as input for subsequent data fusion.
[0042] Sonar data typically contains noise, requiring Kalman filtering to remove background noise. Sonar data provides distance and depth information for fish swarms; point cloud reconstruction techniques convert sonar signals into usable spatial coordinates, extracting features such as distance, intensity, and velocity for each point. High-dimensional vector embedding is performed using these sonar signal features, transforming them into feature vectors with temporal information for subsequent cross-modal data fusion. Environmental data (temperature, dissolved oxygen, etc.) collected by environmental sensors requires standardization to eliminate dimensional differences between different sensors. Environmental data is usually in time-series format; sliding window techniques are used to extract temporal features for alignment with other modal data. An embedding layer converts environmental data into high-dimensional feature vectors to accommodate subsequent feature fusion. The cross-modal attention mechanism in the Transformer architecture enables effective alignment and feature fusion of different modal data. Each modality's data (image features, sonar features, environmental features) is converted into a unified feature vector space after passing through its respective encoder.
[0043] A multi-head self-attention mechanism is employed to process features from different modalities in parallel, ensuring that features from each modality receive sufficient attention during fusion. In this stage, the correlations between cross-modal data (the consistency of fish position in images and sonar) are weighted, thereby enhancing the model's understanding of the commonalities and differences between different modalities. By calculating the similarity matrix between image features, sonar features, and environmental features, similar features are aligned using the attention mechanism. This eliminates information bias caused by heterogeneity between different modalities, ensuring that the influence of each data type is appropriate for the fusion task. In the Transformer model, a bidirectional long short-term memory network (BiLSTM) is used for multimodal data fusion. BiLSTM can capture temporal information, especially in underwater environments where fish movement and behavior patterns often exhibit temporal dependencies. By combining the temporal information of each modality, the system can better capture the dynamic changes in the underwater environment. By enhancing the parameter learning capability of the multi-head attention layer, the fusion accuracy between different modalities can be effectively improved, especially when there is occlusion or different scales between image and sonar modalities. Strengthening feature similarity matching enhances the model's robustness and accuracy in complex environments. Training with a large-scale underwater multimodal dataset enables the model to operate effectively in different waters, with varying water qualities and fish species. Multiple samples of each modality's data input enhance the model's adaptability to various environmental changes. Combining transfer learning techniques allows model parameters trained in a specific water environment to be transferred to other similar environments, reducing training time and improving generalization ability.
[0044] To address the issues of underwater illumination compensation and descattering, this embodiment utilizes multi-scale convolutional networks and generative adversarial networks (GANs) to achieve adaptive illumination compensation and descattering functions for underwater environments. By employing adversarial learning, the generator and discriminator jointly optimize the process, resulting in denoised images that effectively resolve underwater illumination and scattering noise problems. This method can adapt to different aquatic environments, improve underwater image quality, and enhance the accuracy of fish biomass estimation.
[0045] Underwater lighting conditions are complex and highly variable. Factors such as water depth, turbidity, season, and weather can affect the intensity, distribution, and color of underwater light. Turbid water and particulate matter can cause light scattering, severely impacting image clarity and target detection. Generative adversarial networks (GANs) can learn the characteristics of the underwater environment to effectively remove scattering noise from underwater images and restore the original lighting conditions.
[0046] To adapt to varying lighting conditions and water quality in different water environments, this embodiment designs a multi-scale underwater image compensation and denoising generative adversarial network (UW-GAN), mainly comprising a generator and a discriminator, as well as denoising layers and convolutional layers. The generator receives underwater images contaminated by noise and outputs denoised and lighting-compensated images. The generator employs a multi-scale convolutional neural network (CNN) to process local and global information of underwater images at different scales. The generator network structure includes an input layer that takes the original underwater image as input, including water surface reflection, scattering noise, and lighting variations. The convolutional layers are multi-layered convolutional neural networks that extract low-level features of texture and color distribution. Multi-scale feature fusion captures image information at different scales through multiple convolutional kernel sizes, enabling better handling of lighting variations and details in the underwater environment. The denoising layer uses a deep convolutional neural network (Deep CNN) to denoise the input image, especially removing light scattering noise caused by water turbidity. The illumination compensation layer incorporates a physical illumination model (Snell's law of reflection) to simulate the propagation of underwater light within the network, thus compensating for the illumination in underwater images. The discriminator is responsible for judging the quality of the generated images, ensuring the realism of the denoised and illuminated images. The discriminator can employ a traditional Convolutional Neural Network (CNN) structure to judge the quality of the generated images and optimize the generator through adversarial learning. The discriminator network structure includes: an input layer with the denoised image output from the generator and a real, noise-free image as input; multiple convolutional layers to extract local and global features of the image and learn how to determine whether an image is realistic; and an output layer where the generated image is judged for realism by the discriminator. The generator and discriminator are trained together through adversarial learning. The generator continuously improves to output denoised results that are closer to the real image, while the discriminator strives to improve its ability to distinguish between real and denoised images. Ultimately, through adversarial learning between the generator and discriminator, the goals of removing scattering noise and compensating for illumination are achieved.
[0047] To train the UW-GAN model, an underwater fish image dataset was constructed, covering underwater scenes under different water quality, lighting, and depth conditions. Underwater images from various water areas (clear, turbid, containing different particulate matter, etc.) were collected. Lighting variations included underwater images taken under different lighting conditions, such as direct sunlight, cloudy days, and nighttime. Noisy images were generated by simulating or artificially adding noise, resulting in underwater images with scattering and lighting distortion issues. In the pre-training phase, the generator was first unsupervised pre-trained, using a physical model to compensate for lighting distortion in the underwater images, generating preliminary denoised images. The generator and discriminator were trained adversarially, with the generator optimizing to generate higher-quality images and the discriminator optimizing to determine whether the images matched real images. Generative adversarial loss and L2 loss (image reconstruction error) were used to measure the difference between the generated images and real images, promoting improvements in the quality of the generated images. During training, through mixed training with data from different water areas, the generator learned how to generate clear images in waters with different lighting, turbidity, and depths. The generator learns from context, enabling it to adapt to changes in lighting and water quality in different environments. For specific environments, an environment adaptation module can be introduced to adjust the UW-GAN generation strategy based on environmental parameters (water quality, light intensity, etc.) to achieve image optimization under different water conditions.
[0048] The UW-GAN model compensates for the lighting problem in underwater images, restoring the brightness, color, and contrast of the underwater images, making the images clearer. This facilitates subsequent target detection, removes water scattering noise, and improves the image recognition accuracy of fish schools. In particular, it significantly improves the quality of generated images in highly turbid waters.
[0049] Step S23: Perform dynamic target tracking and occlusion completion on the multimodal fusion features to obtain target tracking and completion information.
[0050] In this embodiment, step S23 performs dynamic target tracking and occlusion completion on the multimodal fusion features to obtain target tracking and completion information, specifically including the following steps: Step S231: Input the multimodal fusion features into the spatiotemporal Transformer model, and use the spatiotemporal Transformer model to continuously track the position of the fish school to obtain the fish school position tracking information.
[0051] Step S232: Input the multimodal fusion features into the generative adversarial network model, and use the generative adversarial network model to complete the occluded area of the fish school to obtain the occluded area completion information.
[0052] Step S233: Obtain target tracking and completion information based on the fish school location tracking information and the occlusion area completion information.
[0053] By combining a spatiotemporal Transformer model with a generative adversarial network (GAN) model, the problems of dynamic target tracking and occlusion region completion in underwater fish swarm tracking are effectively solved. Through global temporal information processing by the spatiotemporal Transformer model, combined with contextual inference from image, sonar, and environmental data, the model can accurately track the target's location. Furthermore, the GAN model completes occlusion regions, thereby enhancing the model's accuracy and robustness. This combination not only improves the accuracy of target tracking but also addresses the complex and dynamic changes in the underwater environment, making it suitable for real-time monitoring in large-scale aquaculture.
[0054] When fish are obscured, a combination of a spatiotemporal Transformer model and a multimodal large model offers significant advantages for target tracking and occlusion region completion in underwater environments, addressing the challenges of dynamic fish tracking and partial information completion. The spatiotemporal Transformer captures the movement patterns and contextual information of the fish school, while the multimodal large model generates the probability distribution of the occluded region and completes the missing information.
[0055] The Spatiotemporal Transformer is a deep learning model that processes sequential data through a self-attention mechanism. It can capture global information in both time and space, making it suitable for dynamic target tracking tasks. The Spatiotemporal Transformer can simultaneously process the spatiotemporal characteristics of video frame sequences and target locations, thereby achieving accurate dynamic target tracking.
[0056] The input data consists of consecutive video frames and sonar data. Fish image features in the video frames are extracted using a convolutional neural network (CNN), while sonar data is processed using a point cloud network. A spatiotemporal Transformer encoder is used to learn image features (fish appearance, direction of movement, speed, etc.) and sonar features (fish density, location, etc.) in each frame. By fusing image features, sonar features, and temporal information, the spatiotemporal Transformer can capture the temporal relationships of fish movement and generate a spatiotemporal representation of each target. The spatiotemporal Transformer utilizes a self-attention mechanism to learn the correlation between different video frames, capturing the movement trajectories of the fish at different points in time. Especially when occlusion occurs, it can infer the fish's location based on previous movement trajectories.
[0057] The spatiotemporal Transformer model can infer the position of fish in the current frame based on the image and sonar data from the previous frame. Especially when fish are occluded, the model can make inferences using contextual information (fish texture, swimming patterns, etc.). The spatiotemporal Transformer model can not only track targets in the current frame but also predict the state of fish in future frames based on information from past frames, further enhancing the accuracy of dynamic target tracking.
[0058] In practical applications, fish schools often experience occlusion, especially in dense schools or complex underwater environments. To address this issue, contextual information (fish texture, swimming patterns) is incorporated to complete the occluded areas. Features such as fish appearance, movement trajectories, and textures are extracted from underwater fish school image data to help the model understand the specific morphology and behavioral patterns of the fish school. In situations without a complete field of view, fish texture and movement patterns can provide supplementary information. Sonar data provides information about the fish school's location, density, and depth. When fish are occluded, sonar data can provide auxiliary information to help infer the school's specific location and morphology. A point cloud generation network is used to process sonar data, further enhancing the robustness of target tracking. Environmental information provided by environmental sensors (such as water temperature, dissolved oxygen, and turbidity) can be used to infer the fish school's behavioral patterns. When the water temperature is low, fish may gather in groups, providing clues about the school's location and density.
[0059] A multimodal large model is used to generate the probability distribution of occluded regions, inferring potential fish features in these areas. A self-attention mechanism is employed, combining image, sonar, and environmental data to complete missing image information. UW-GAN is used for image restoration, generating reasonable predicted images of the missing regions during the completion process. A generator predicts underwater fish images of the occluded areas, while a discriminator judges whether the generated completed data matches the real-world fish morphology, ensuring the realism of the completion. Combining contextual information (fish texture, swimming patterns) and existing image features, a Conditional GAN is used to generate underwater fish images of the occluded areas. Conditional generation allows the model to provide more accurate completion results based on the current context.
[0060] The multimodal large model, trained through self-supervised learning, learns the regular movement patterns and occlusion behaviors of fish schools from existing image and sonar data. Based on this learning, the model can generate the probability distribution of occluded areas and assess the fish states that may be contained in different areas. A spatiotemporal Transformer model is used to model the temporal information of video frames, generating predictions of occlusion areas in future frames. This prediction is updated in conjunction with image and sonar data, allowing the model to dynamically complete occlusion areas. By combining the spatiotemporal Transformer model, contextual information, and multimodal data, the model can effectively perform target tracking and information completion, exhibiting higher robustness and accuracy in underwater environments.
[0061] By fusing target location predictions generated by a spatiotemporal Transformer model with contextual completion from multimodal data, more accurate fish school images can be generated, effectively reducing the impact of occlusion on target tracking accuracy. Incorporating sonar data for inference further improves the model's performance in complex environments. Incremental learning and self-supervised learning enable the model to continuously adapt to constantly changing underwater environments. Especially under conditions of significant changes in fish density and water quality, the model continuously optimizes its predictive capabilities through contextual completion. The fusion of multimodal data (images, sonar, and environmental data) allows the model to accurately infer target locations even when a single data source is missing, enhancing the system's robustness.
[0062] This embodiment also employs self-supervised pre-training and few-shot learning. Self-supervised learning is performed using a large amount of unlabeled underwater video data to pre-train a multimodal large model, followed by fine-tuning with a small amount of labeled data. This effectively reduces data labeling costs and improves the model's adaptability and generalization ability. By combining self-supervised learning and few-shot learning techniques, the data scarcity problem in aquaculture can be effectively addressed, improving the model's accuracy and generalization ability. Self-supervised learning trains the model to obtain feature representations using unlabeled data, while few-shot learning enables the model to quickly adapt to new tasks even with insufficient labeled data. The combination of these two techniques significantly reduces dependence on labeled data and enhances the model's performance in diverse environments.
[0063] Since underwater fish image data is often difficult to annotate extensively, self-supervised learning utilizes large amounts of unlabeled data for model pre-training to obtain general feature representations and reduce reliance on labeled data. By comparing the similarity of different modalities (underwater fish image data and sonar data), the model learns the correlation between different modalities of data in an aquaculture environment. The multimodal contrastive loss function maximizes the similarity between underwater fish image data and sonar data at the same time or spatial location, while minimizing the distance between data at different times or spatial locations.
[0064] First, underwater fish school image data and sonar data are used for feature extraction through their respective encoding networks (CNN and PointNet). Then, contrastive learning is used to learn the mapping between the underwater fish school image data and sonar data, so that the representation of the same fish school in the underwater fish school image and sonar data is as close as possible, and the representation of different fish schools is as far apart as possible. Historical video frames are used to predict future video frames, thereby helping the model learn the movement patterns of the fish school and the laws of environmental change. An autoregressive model (autoregressive prediction based on Transformer) is used to process the image sequence. By inputting several historical frames of underwater fish school image data, the model is trained to predict the content of the next frame. This task not only improves the model's understanding of temporal data, but also learns the dynamic changes of the underwater environment. The model is trained to recover occluded or noise-contaminated image or sonar data, promoting the model's understanding of the internal structure of the data. Generative adversarial networks are used to design generators and discriminators to recover missing modal data. If some parts of the image or sonar data are occluded or affected by noise, the model recovers these data through generation tasks, thereby learning the diversity and structural features of the data.
[0065] Pre-training on unlabeled data using self-supervised tasks allows the model to learn feature representations from unlabeled underwater videos, sonar data, and environmental data, establishing preliminary perception capabilities. Unlabeled data can be used for training with routine monitoring data from underwater environments (fish swimming videos, real-time sonar detection data, etc.). When labeled data is scarce, few-shot learning techniques are employed to fine-tune the model. During fine-tuning, supervised training is performed using a small amount of labeled data, allowing the model to be optimized for specific tasks and improve accuracy in diverse environments. Few-shot learning techniques are particularly important when labeled data is scarce; by designing few-shot learning schemes suitable for aquaculture tasks, the dependence on large-scale labeled data can be significantly reduced. Transfer learning reduces the need for new labeled data by applying learned knowledge to new tasks. Through transfer learning, the model can leverage knowledge gained in one environment and quickly apply it to a new one. A base model is trained on a labeled fish dataset (fish data from a specific body of water), and then transferred to a target body of water or a task involving a new fish species. By fine-tuning the transferred model, training time is reduced and accuracy is improved. Transfer learning can help a model learn fish detection knowledge from one body of water (clear water) and then transfer it to other bodies of water (turbid water), and fine-tune it with a small amount of labeled data to optimize the model's performance in the new environment.
[0066] Data augmentation and pseudo-label generation can significantly expand the training set for few-shot learning. Simulation techniques are used to augment existing labeled data, including simulations of fish size, morphological changes, and swimming patterns. Underwater images are enhanced through image rotation, scaling, and flipping to increase sample diversity and help the model learn more universal features. When sufficient labeled data is unavailable, pseudo-labels are generated through self-supervised learning to supplement few-shot learning. The feature representations obtained through self-supervised learning can generate preliminary labels for unlabeled data, effectively increasing the training dataset.
[0067] By combining self-supervised learning and few-shot learning, the model can learn important features on unlabeled data and quickly fine-tune using a small amount of labeled data, thus significantly reducing the need for data labeling. Self-supervised learning enables the model to understand the common characteristics of different aquaculture environments, while few-shot learning helps the model quickly adapt to new environments, thereby improving the model's generalization ability, especially in situations with diverse water quality, fish species, and significant environmental changes. Combining these two techniques allows the model to maintain high accuracy even with scarce data, effectively addressing the problem of insufficient data, particularly in aquaculture environments where obtaining large amounts of labeled data is difficult.
[0068] Step S3: Based on the target tracking and completion information, a regression model is used to estimate the fish biomass, and the fish biomass estimation result is obtained.
[0069] In this embodiment, step S3 uses a regression model to estimate the fish biomass based on the target tracking and completion information to obtain the fish biomass estimation result. Specifically, it includes the following steps: Based on the target tracking and completion information, combined with the multimodal data, a regression model is used to estimate the fish biomass using an ensemble learning method, resulting in the fish biomass estimation results, including the total weight of the fish, density distribution, and the number of fish per unit volume.
[0070] In this embodiment, to ensure efficient deployment of multimodal large models on edge computing devices and meet the needs of real-time monitoring of large-scale fish farms, it is necessary to address the challenges posed by limitations in computing resources (processing power, storage, bandwidth, etc.) while maintaining a balance between high accuracy and real-time performance. By employing a large model distillation technique, the multimodal large model is compressed into a lightweight sub-model suitable for underwater deployment on edge devices, optimizing computational efficiency and enabling real-time estimation of fish biomass to meet the needs of large-scale fish farms.
[0071] By employing techniques such as lightweight model design, computational resource optimization, incremental learning, and edge device hardware acceleration, the resource limitations of edge computing devices can be effectively addressed, ensuring high accuracy while meeting the real-time monitoring needs of large-scale fish farms. Reasonable task scheduling and hierarchical architecture design further improve the model's operational efficiency on edge devices, enabling real-time and efficient biomass estimation and fish swarm monitoring even in complex environments.
[0072] Edge devices have limited computing resources, so lightweight strategies must be adopted when designing models to ensure accuracy without excessive computational consumption. Model distillation is a method that compresses a large "teacher model" into a small "student model." By training the student model under the guidance of the teacher model, the student model's performance approaches that of the teacher model, but with a smaller structure and higher computational efficiency. Using a distillation strategy, a trained multimodal Transformer model is used as the teacher model, and the student model is trained by minimizing the output difference between the teacher and student models. The student model's architecture can be designed to be more compact, reducing computational and storage requirements. Although small in size, the student model inherits the knowledge of the teacher model, maintains high accuracy, and significantly reduces computational and storage resource consumption when running on edge devices.
[0073] To ensure that edge devices can meet the real-time monitoring needs of large-scale fish farms, it is essential to optimize computing resource usage and ensure real-time performance. For large-scale fish farm environments, model sharding or distributed inference can be used to distribute computing tasks across multiple edge devices. This reduces the computational burden on each device, improving the overall system's processing power and response speed. The entire multimodal model is decomposed into multiple sub-models, each handling a specific data modality (underwater fish image data, sonar data, environmental data, etc.). Different edge devices can be responsible for processing data of different modalities and aggregating the results. Using an edge computing grid to connect multiple devices and leveraging collaborative computing among devices for distributed inference ensures efficiency and real-time performance during large-scale deployments.
[0074] Because aquaculture environments are constantly changing, factors such as fish behavior patterns and environmental conditions may evolve. Incremental learning or online learning strategies allow models to be continuously updated on edge devices without retraining the entire model. The model can be designed to be fine-tuned online based on new sensor data or underwater fish image data after deployment. When the edge device receives new data, the model only needs to update some parameters instead of retraining the entire network. Incremental learning methods allow the model to continuously learn from new data without significantly increasing computational burden, ensuring the model adapts to environmental changes. To improve inference speed, edge devices can utilize dedicated hardware acceleration, such as GPUs, TPUs, and FPGAs, to accelerate the inference process of deep learning models. These hardware accelerators can significantly improve model efficiency and support complex multimodal computing tasks. Designing and deploying specifically optimized deep learning models allows them to efficiently utilize the hardware acceleration capabilities of the devices. Federated learning frameworks can also be used for collaborative training across multiple distributed fish farms, effectively protecting data privacy while improving model versatility.
[0075] In large-scale fish farm monitoring, real-time data processing is crucial. By using data caching and preprocessing mechanisms, the computational burden during each inference can be significantly reduced, improving real-time performance. Caching mechanisms store recent calculation results and intermediate feature values, avoiding redundant calculations. Edge devices can cache the current location information of fish swarms and recent environmental data, allowing for direct use of cached data instead of recalculation during subsequent inferences. Preprocessing sensor data, including denoising, standardization, and feature extraction during data acquisition, reduces data transmission and computational burden. The overall system architecture should be optimized based on the resource limitations of edge computing devices, ensuring reasonable resource allocation and enabling real-time processing. A layered architecture is adopted, where more complex computational tasks can be performed in the cloud or on more powerful edge devices, while simpler tasks are handled by edge devices with fewer resources. A more streamlined model is deployed on edge devices for real-time preprocessing and preliminary analysis, while complex computational and data fusion tasks are handled by the cloud or more powerful edge nodes. A flexible task scheduling mechanism is designed based on the computing power and real-time requirements of the devices, prioritizing different tasks to ensure the real-time performance of critical tasks. Deploy multi-level data streams and priority scheduling in the system to ensure that critical data (fish location, significant environmental changes, etc.) can be processed in a timely manner, while non-critical data can be processed with appropriate delay.
[0076] In summary, this embodiment addresses the limitations of traditional techniques in underwater biomass estimation by using a multimodal large model combined with multimodal data fusion, generative adversarial networks, Transformer architecture, and few-shot learning. In particular, the application of cross-modal attention mechanisms, dynamic target tracking, occlusion completion, and self-supervised pre-training improves the robustness and accuracy of the model and overcomes the impact of the underwater environment on image quality.
[0077] In one exemplary embodiment, a fish biomass estimation system based on a multimodal large model is provided. This system mainly includes a data acquisition layer, a preprocessing layer, a multimodal fusion large model, and an output layer. The data acquisition layer consists of an underwater camera, a multi-band sonar array, and multiple environmental sensors, responsible for synchronously acquiring multimodal data. The preprocessing layer includes preprocessing steps such as illumination compensation, sonar signal filtering, and multimodal data alignment. The multimodal fusion large model is based on a Transformer encoder-decoder structure, with inputs being image features, sonar features (sonar point clouds), and environmental feature (environmental parameters) embedding vectors, performing cross-modal fusion. The output layer outputs the estimated biomass (weight of fish per unit volume), density heatmap, and anomaly warning (disease risk) information.
[0078] In one exemplary embodiment, another fish biomass estimation system based on a multimodal large model is provided, such as... Figure 2 As shown, the system includes a multimodal data acquisition module, a data preprocessing and feature extraction module, a cross-modal attention mechanism fusion module, a dynamic target tracking and occlusion completion module, a biomass regression and result output module, and a lightweight edge computing deployment module, all connected in sequence. The multimodal data acquisition module collects multimodal data from the target area, including underwater fish image data, sonar data, and environmental data. The data preprocessing and feature extraction module, the cross-modal attention mechanism fusion module, and the dynamic target tracking and occlusion completion module are used to perform preprocessing and feature extraction, cross-modal data fusion, and dynamic target tracking and occlusion completion on the multimodal data sequentially based on a multimodal large model, obtaining target tracking and completion information; this target tracking and completion information refers to the information obtained after tracking the fish school and completing the occluded areas of the fish school. The biomass regression and result output module is used to estimate the fish school biomass using a regression model based on the target tracking and completion information, obtaining the fish school biomass estimation result. The lightweight edge computing deployment module is used to deploy the multimodal large model to edge devices in a lightweight manner to enable real-time inference of the multimodal large model on the edge devices.
[0079] A multimodal data acquisition module is used to simultaneously acquire multimodal data to support multimodal fusion processing for fish biomass estimation. This module includes an underwater camera, a sonar array, and environmental sensors. The underwater camera acquires underwater images of the fish school, providing visual information and capturing dynamic movement patterns (such as swimming speed, density, and morphological changes), and performs adaptive optimization under varying lighting conditions. The sonar array captures distance, depth, and shape information of the fish school, supplementing the temporal and spatial information in the image and sonar data to aid in estimating dynamic changes in fish biomass. The environmental sensors collect environmental data such as water temperature, dissolved oxygen, and turbidity.
[0080] The multimodal data acquisition module transmits underwater fish image data captured by the underwater camera, sonar data such as distance and depth obtained by the sonar array, and environmental data such as water temperature and dissolved oxygen collected by the environmental sensor to the data preprocessing and feature extraction module. The data preprocessing and feature extraction module transmits the preprocessed and extracted image features, sonar features, and environmental feature vectors to the cross-modal attention mechanism fusion module for feature fusion and information alignment. The cross-modal attention mechanism fusion module transmits the feature vectors of the fused multimodal features to the dynamic target tracking and occlusion completion module for dynamic target tracking and occlusion region completion. The dynamic target tracking and occlusion completion module transmits the fish movement trajectory and density information after target tracking and occlusion completion to the biomass regression and result output module for fish biomass estimation, obtains the fish biomass estimation results, and outputs biomass, density heatmaps, and abnormal warning information.
[0081] The multimodal data acquisition module transmits raw data to the data preprocessing and feature extraction module for necessary illumination compensation, noise reduction, and feature extraction. The data preprocessing and feature extraction module then sends the processed data to the cross-modal attention mechanism fusion module for multimodal data fusion. The cross-modal attention mechanism fusion module transmits the fused multimodal features to the dynamic target tracking and occlusion completion module for accurate target tracking and occlusion completion. The dynamic target tracking and occlusion completion module then transmits the completed fish movement trajectory and other information to the biomass regression and result output module for final biomass estimation and result output.
[0082] The data preprocessing and feature extraction module is used to process the collected multimodal data and provide high-quality features for subsequent fusion and estimation.
[0083] The data preprocessing and feature extraction module processes the raw data from the multimodal data acquisition module, extracts useful features, and optimizes data quality. The processed data is then transmitted to the cross-modal attention mechanism fusion module. The transmitted data includes underwater fish school images after illumination compensation and descattering, clear image features after removing illumination issues and water quality interference; sonar data features such as depth, distance, and shape after Kalman filtering; and standardized environmental data with time-series features of water quality information such as temperature and dissolved oxygen.
[0084] The cross-modal attention mechanism fusion module is based on the Transformer architecture and combines the temporal information of multimodal data. It dynamically fuses underwater camera, sonar array and environmental data in real time through the cross-modal attention mechanism to facilitate accurate estimation of fish biomass.
[0085] The main task of the cross-modal attention mechanism fusion module is to fuse processed underwater fish school image data, sonar data, and environmental data using a deep learning model to achieve efficient integration of different data modalities and spatiotemporal feature modeling. After adjusting the fusion weights of the multimodal data, the generated fused features are transmitted to the dynamic target tracking and occlusion completion module for further analysis of the fish school's underwater movement trajectory, especially for completion when occlusion occurs. By adaptively adjusting the fusion weights of image, sonar, and environmental data, the contribution of each modality is optimized based on the characteristics and temporal information of different modalities, thereby improving the accuracy of fish school biomass estimation. The transmitted content includes fused features such as fish school location, density, and depth information from image, sonar, and environmental sensors.
[0086] A cross-modal attention fusion model dynamically fuses underwater camera, sonar array, and environmental data in real time through a cross-modal attention mechanism to accurately estimate fish biomass. Based on the Transformer architecture, this model dynamically fuses data from underwater cameras, sonar arrays, and the environment in real time, accurately estimating fish biomass through joint modeling of temporal information and spatial features. By utilizing time-series modeling and combining the temporal information and spatial features of underwater camera, sonar array, and environmental data, the fusion weights of each modality are dynamically adjusted, improving the accuracy and robustness of fish biomass estimation. The cross-modal attention mechanism dynamically adjusts the fusion weights of image, sonar, and environmental data, automatically and adaptively adjusting attention allocation based on the characteristics of different modalities (illuminance variations in the image modality, distance information in the sonar modality, and time series data from the sensors). Furthermore, self-supervised pre-training and few-shot learning can be used. Self-supervised learning is performed using a large amount of unlabeled underwater video data, followed by fine-tuning with a small amount of labeled data, reducing data labeling costs and improving the model's adaptability and generalization ability.
[0087] The dynamic target tracking and occlusion completion module addresses the dynamic tracking problem of underwater fish schools in complex environments, particularly in cases of occlusion and missing information. Based on the transmitted fused feature data, this module utilizes a spatiotemporal Transformer model to track dynamic targets, resolving information loss issues caused by occlusion and movement within underwater fish schools. This module then transmits the completed fish school position and movement information to the biomass regression and result output module. The latter uses this complete information to perform biomass regression estimation and outputs the final biomass estimate. The transmitted content includes the fish school's dynamic position, velocity, density, and other completed fish movement trajectory and density information underwater. This embodiment combines image features, sonar features, and contextual information from environmental data through the dynamic target tracking and occlusion completion module to predict the position of fish schools under occlusion or data loss conditions, enhancing the model's robustness in complex underwater environments.
[0088] The lightweight edge computing deployment module enables real-time inference of large multimodal models on edge devices, supporting the real-time monitoring needs of large-scale aquaculture farms. Through model distillation technology, large-scale multimodal models are compressed into lightweight sub-models suitable for edge device deployment, adapting to the limited computing resources of edge devices. Hardware acceleration and incremental learning strategies optimize computational efficiency, supporting an inference speed of 30 frames per second (FPS) in real-time monitoring of large-scale fish farms, meeting real-time processing requirements. This allows the system to perform real-time inference on edge devices, achieving efficient and stable operation of multimodal models on edge devices and supporting real-time monitoring of large-scale fish farms.
[0089] The biomass regression and results output module receives data from the dynamic target tracking and occlusion completion module, uses regression analysis to calculate the total biomass of the fish population, and outputs the estimated biomass results. These results can be presented to farm managers in the form of heatmaps, data tables, etc., and provide data support for subsequent feed allocation and stocking density adjustment decisions. The transmitted content includes the estimated total weight of the fish population and density distribution maps.
[0090] This embodiment presents a method for estimating fish biomass based on a multimodal large model. The method relies on the aforementioned system, and its specific implementation includes the following steps: Step 1: The underwater camera in the multimodal data acquisition module captures dynamic underwater fish school images under different water quality conditions, recording data such as fish density, swimming speed, and morphological changes. This underwater fish school image data can be further used to estimate the number and distribution of fish. A sonar array collects distance, depth, and shape information of the fish school. In situations with poor underwater visibility and image quality, sonar data assists in fish school identification by providing spatial information from the underwater fish school images, offering the three-dimensional position and movement status of the fish. To further enrich the environmental background data for fish biomass estimation, environmental sensors such as temperature, dissolved oxygen, and turbidity sensors are used to collect water quality parameters such as temperature, dissolved oxygen, and turbidity. This provides background support for dynamically adjusting the model and helps in estimating fish biomass. The multimodal data acquisition module ensures synchronous data acquisition from different sensors and includes timestamp recording to ensure temporal consistency.
[0091] Step 2: Preprocessing and feature extraction based on the multimodal data from Step 1. Specifically, illumination compensation and descattering are performed on the underwater fish image data using multi-scale convolutional networks and generative adversarial networks. This improves image quality, enhances target detection accuracy, and addresses issues related to lighting variations, water turbidity, and scattering in the underwater environment. Convolutional neural networks are then used for edge detection and target segmentation, extracting key features such as the fish's morphology, contours, and movement speed.
[0092] To address the challenge of varying fish size across different underwater dimensions, scale normalization was performed to ensure a consistent representation of fish size at different depths and viewpoints. Kalman filtering was applied to the sonar data to remove background noise, improving data quality and providing a more accurate basis for analyzing the depth, shape, and location of the fish. Environmental data was standardized to eliminate dimensional differences between different sensors. Key features were extracted from underwater fish image data, sonar data, and environmental data, including fish morphology, spatial point cloud features from sonar, and time-series data from environmental sensors, to prepare for subsequent model input.
[0093] Step 3: Based on the preprocessed and feature-extracted data from Step 2, perform cross-modal attention mechanism fusion. Specifically, the cross-modal attention mechanism in the Transformer architecture is used to jointly model and fuse the image features, sonar features, and environmental features extracted in Step 2. By learning the correlation between image, sonar, and environmental data, and combining temporal information and spatial features, a multi-head self-attention mechanism is used to process the features of each modality in parallel, dynamically adjusting the data weights of each modality to ensure that the features of each data modality receive sufficient attention during the fusion process. Temporal information modeling analyzes the movement trajectory of fish schools, combining the temporal information of underwater fish school image data, sonar data, and environmental data. When fish schools experience occlusion, the model can effectively capture the dynamic changes in fish school behavior. Spatial feature modeling fuses the spatial features of sonar and environmental data, improving the model's adaptability and accuracy under different water depths and environmental conditions, overcoming the problems of data scarcity and insufficient annotation.
[0094] Step 4: Based on the multimodal fusion features from Step 3, perform dynamic target tracking and occlusion completion. The specific process includes: Dynamic multi-target tracking utilizes a spatiotemporal Transformer model to continuously track the position of the fish school based on its movement trajectory, image, and sonar data spatiotemporal characteristics, ensuring accurate target tracking even in dense fish schools.
[0095] When fish schools lose information due to mutual occlusion or other factors, generative adversarial networks are used to fill in the occluded areas by utilizing contextual information such as fish body texture, head features, and movement patterns, thereby restoring the occluded fish school image and morphology and ensuring the accuracy of biomass estimation.
[0096] Step 5: Based on the target tracking and completion information from Step 4, estimate the fish biomass and output the results. Using the fish location and movement trajectory obtained in Step 4, combined with the fish volume, density characteristics, and depth information provided by sonar in the image, a regression model is used to estimate the fish biomass. The regression model uses an ensemble learning method to calculate the total biomass of the underwater fish population based on the weight-volume relationship of different fish species. Fish biomass includes the estimated total weight of the fish population, density distribution, and the number of fish per unit volume. The regression calculation results are output in the form of heatmaps showing fish distribution and data tables, facilitating real-time monitoring of fish growth and prediction of key data such as the future harvest period by aquaculture managers.
[0097] While multimodal learning and cross-modal attention mechanisms have been applied in image processing and object detection, their application in aquaculture and underwater biomass estimation still faces unique challenges. These challenges include variations in underwater ambient lighting, unstable data quality, and fish occlusion, and existing technologies are rarely able to effectively address these specific issues.
[0098] The underwater environment is characterized by intense, non-uniform lighting, complex water quality variations (turbidity, dissolved oxygen levels, etc.), and dynamic changes caused by water flow and fish density. These factors result in more severe quality issues for underwater fish image data compared to those used in terrestrial applications. Current multimodal technologies are mostly used in static or relatively clear environments; applications in dynamic, highly complex underwater scenarios are still in the realm of innovation.
[0099] This embodiment utilizes a cross-modal attention mechanism to automatically adjust attention allocation based on the characteristics of different modalities. Sonar data and underwater fish school image data differ significantly in terms of time synchronization, spatial resolution, data scarcity, and quality. By introducing adaptive learning capabilities, the system can dynamically adjust the attention focus even in low-quality images and with missing information, ensuring high accuracy in biomass estimation. Utilizing dynamic modal alignment technology, through time-series modeling, spatial location calibration, and multi-scale fusion methods of different modalities, the alignment accuracy is further improved, effectively handling the heterogeneity between sonar, image, and environmental data, enabling more precise integration of image and sonar data of fish schools in complex underwater environments.
[0100] Traditional Transformer architectures are primarily used for processing sequential data (text, speech, etc.), while sonar data typically presents as two-dimensional or three-dimensional point cloud data with noise and complex depth information. To better utilize Transformer for sonar data processing, a sonar embedding method is designed to convert sonar data into a format suitable for Transformer processing, transforming it into a high-dimensional feature vector or point cloud encoded representation. This method can handle the spatial and temporal information of sonar data, thus enabling effective integration with underwater fish swarm image data.
[0101] Fish in underwater environments often exhibit occlusion, especially in dense schools, making it difficult for traditional target detection methods to accurately track their movement. By incorporating a spatiotemporal Transformer architecture, this system enables dynamic target tracking in underwater environments and uses contextual information (fish texture, swimming patterns, etc.) to fill in occluded information. Compared to traditional convolutional neural networks or simple RNNs, the Transformer's self-attention mechanism can capture global dependencies over long time sequences, thus exhibiting greater adaptability to dynamic scenes.
[0102] Existing multimodal learning techniques typically rely on large amounts of labeled data to train models, but acquiring labeled data is very expensive and time-consuming in underwater biomass estimation. This embodiment significantly reduces the dependence on labeled data by introducing self-supervised learning and few-shot learning techniques.
[0103] Self-supervised learning enables pre-training on unlabeled data, allowing for fine-tuning of the model. By applying unsupervised learning to image, sonar, and environmental data from underwater videos, the system can automatically extract features and relationships from the data, reducing the cost of data annotation. Unlike traditional methods that rely on large amounts of manually labeled data, self-supervised learning allows the model to effectively capture patterns in the underwater environment even with limited data.
[0104] In aquaculture, labeled data is scarce and difficult to obtain, especially under certain special environments. Through few-shot learning techniques, the system can quickly adapt and optimize on limited labeled data, improving the model's accuracy under different environmental conditions. Compared to traditional fully supervised learning methods, few-shot learning effectively improves the model's generalization ability and data adaptability.
[0105] Existing deep learning models often require powerful computing resources, especially in real-time monitoring applications in large-scale fish farms where real-time performance is crucial. To deploy the multimodal model of this invention on edge devices, model distillation and lightweight design methods are employed, enabling the model to run on low-power devices while maintaining accuracy and supporting real-time inference. Larger pre-trained models are "distilled" into smaller, more computationally efficient models, ensuring efficient biomass estimation even with limited device computing power. By optimizing the model into a lightweight sub-model adapted for edge devices, low-latency and high-performance real-time data processing is achieved, meeting the needs of real-time monitoring in large-scale fish farms.
[0106] By fusing multimodal data (images, sonar, and environmental data), more comprehensive information about fish populations is provided, including density, distribution, and movement patterns. A cross-modal attention mechanism leverages the strengths of different modalities, dynamically adjusting the weights of each modality to ensure spatiotemporal consistency and accurate fusion. This provides fundamental data support for biomass estimation. This module addresses the limitations of single data sources (such as images or sonar), such as occlusion, illumination variations, and insufficient depth information, improving the accuracy of biomass estimation through the complementarity of multimodal data.
[0107] The data preprocessing and feature extraction module addresses the impact of underwater lighting variations on image quality, improving image clarity and target detection accuracy. This module uses generative adversarial networks (GANs) and multi-scale convolutional networks to perform illumination compensation and descattering on underwater images. This addresses illumination and scattering issues in underwater images, improving image quality and target detection accuracy, ensuring the quality of underwater images is suitable for subsequent target detection and biomass estimation. Core problem solved: This module addresses issues such as insufficient lighting and image blurring by compensating for illumination and descattering, providing high-quality underwater fish school image data for target detection and object segmentation, thereby optimizing the accuracy of biomass estimation.
[0108] The dynamic target tracking and occlusion completion module addresses the dynamic tracking of fish schools underwater, resolving target loss due to occlusion or movement. This module employs a spatiotemporal Transformer model, inferring the location of occluded fish schools based on contextual information and using sonar and environmental data to complete missing portions of the image. This module effectively solves the occlusion problem of fish schools in dynamic environments, ensuring continuous tracking of fish schools in complex underwater conditions and providing accurate movement trajectories and density estimation data for biomass estimation.
[0109] This lightweight edge computing deployment module enables multimodal data processing and real-time inference on edge devices with limited computing resources, meeting the real-time monitoring needs of large-scale fish farms. Through model distillation technology, large models are compressed into smaller sub-models suitable for edge devices, optimizing computational efficiency and ensuring real-time inference on these devices. This module addresses the real-time processing requirements of aquaculture farms, enabling rapid and accurate biomass estimation on edge devices and improving the overall system performance and stability through distributed inference.
[0110] The synergistic effect of all modules ensures high accuracy and efficiency in underwater biomass estimation. Multimodal data fusion provides comprehensive fish swarm information, illumination compensation optimizes image quality, target tracking solves occlusion problems, and the lightweight edge computing deployment module ensures the system's real-time performance and large-scale deployment capabilities. Through this modular and collaborative design, the task of underwater fish biomass estimation can be completed efficiently and accurately in practical applications.
[0111] This embodiment utilizes multimodal fusion technology to improve estimation accuracy. Sonar data effectively supplements image information, and environmental parameter data helps correct biomass estimation, thereby improving the overall model's accuracy. Self-supervised learning technology significantly reduces the need for labeled data, lowering the cost of data labeling. A lightweight model enables edge computing devices to achieve an inference speed of 30 FPS with low power consumption, meeting the real-time monitoring needs of large-scale fish farms. The contextual learning capability of the large multimodal model allows it to adapt to environments with different fish species and stocking densities, enhancing the model's generalization and adaptability.
[0112] This embodiment introduces a multimodal large model into the field of aquaculture biomass estimation for the first time. It adopts a cross-modal attention mechanism and self-supervised learning technology, which effectively solves the problems of image quality, occlusion and data scarcity in the existing technology, while ensuring high accuracy and real-time performance. This system achieves accurate estimation of underwater fish biomass by fusing multimodal data (underwater fish image data, sonar data, and environmental data). It includes a multimodal data acquisition module, a data preprocessing and feature extraction module, a cross-modal attention mechanism fusion module, a dynamic target tracking and occlusion completion module, a biomass regression and result output module, and a lightweight edge computing deployment module. Specifically, through a cross-modal attention mechanism, it dynamically fuses multimodal data collected by underwater cameras, sonar arrays, and environmental sensors in real time. Employing self-supervised learning and few-shot learning techniques significantly reduces data annotation requirements. Furthermore, through lightweight model optimization, it enables real-time inference on edge devices, effectively improving the accuracy of fish biomass estimation and solving the problems of image quality, occlusion, and data scarcity in traditional technologies. This meets the real-time monitoring needs of large-scale fish farms, improving the efficiency and economic benefits of aquaculture. In addition, the system has strong generalization capabilities, adapting to different fish species and environmental conditions. Combining deep learning and multimodal data fusion technology, it provides an intelligent solution with broad application potential in the aquaculture field.
[0113] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0114] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for fish school biomass estimation based on a multi-modal large model, characterized in that, include: Collect multimodal data of the target area, including underwater fish image data, sonar data, and environmental data; Based on the multimodal large model, the multimodal data is preprocessed and feature extracted, cross-modal data fused, and dynamic target tracking and occlusion completion are performed sequentially to obtain target tracking and completion information; The target tracking and completion information refers to the information obtained after tracking the fish school and completing the areas obscured by the fish school; Based on the target tracking and completion information, a regression model is used to estimate the fish biomass, and the fish biomass estimation results are obtained.
2. The fish biomass estimation method based on a multimodal large model according to claim 1, characterized in that, Collect multimodal data of the target area, specifically including: Using an underwater camera, dynamic underwater fish images of the target area are collected, and the density, swimming speed and morphological changes of the fish are recorded to obtain underwater fish image data. A sonar array is used to collect the location, movement, and shape information of underwater fish schools in the target area to obtain sonar data. Temperature, dissolved oxygen, and turbidity sensors were used to collect water temperature, dissolved oxygen, and turbidity data to obtain environmental data.
3. The fish biomass estimation method based on a multimodal large model according to claim 1, characterized in that, Based on a multimodal large model, the multimodal data is preprocessed and feature extracted sequentially, cross-modal data fusion is performed, and dynamic target tracking and occlusion completion are conducted to obtain target tracking and completion information, specifically including: The multimodal data is preprocessed and features are extracted to obtain image features, sonar features, and environmental features; A cross-modal attention mechanism model is used to perform cross-modal data fusion on the image features, the sonar features, and the environmental features to obtain multimodal fused features; Dynamic target tracking and occlusion completion are performed on the multimodal fusion features to obtain target tracking and completion information.
4. The fish biomass estimation method based on a multimodal large model according to claim 3, characterized in that, The multimodal data is preprocessed and feature extracted to obtain image features, sonar features, and environmental features, specifically including: The underwater fish image data is subjected to illumination compensation and descattering processing to obtain underwater fish image data after illumination compensation and descattering. Edge detection and target segmentation are performed on the underwater fish school image data after illumination compensation and descattering to extract fish school features, including fish school morphological features, contour features and movement speed features; The fish school features are scale-normalized to obtain scale-normalized fish school features, which are used as image features. The Kalman filter algorithm is used to remove background noise from the sonar data to obtain denoised sonar data. Feature extraction is performed on the denoised sonar data to obtain sonar features; The environmental data is standardized to obtain standardized environmental data. Feature extraction is performed on the standardized environmental data to obtain environmental features.
5. The fish biomass estimation method based on a multimodal large model according to claim 3, characterized in that, Using a cross-modal attention mechanism model, cross-modal data fusion is performed on the image features, sonar features, and environmental features to obtain multimodal fused features, specifically including: Construct a cross-modal attention mechanism model; the cross-modal attention mechanism model refers to a model with a Transformer architecture used for cross-modal data fusion; Based on the image features, sonar features, and environmental features in the multimodal data, the cross-modal attention mechanism model is trained to obtain a trained cross-modal attention mechanism model; The image features, sonar features, and environmental features are input into the trained cross-modal attention mechanism model to obtain multimodal fusion features.
6. The fish biomass estimation method based on a multimodal large model according to claim 5, characterized in that, Based on the image features, sonar features, and environmental features in the multimodal data, the cross-modal attention mechanism model is trained to obtain a trained cross-modal attention mechanism model, specifically including: The image features, sonar features, and environmental features from the multimodal data are input into the cross-modal attention mechanism model for training. The cross-modal attention mechanism model learns the correlation between the image features, sonar features, and environmental features. Combining the temporal information and spatial features in the multimodal data, the model uses a multi-head self-attention mechanism to process the image features, sonar features, and environmental features in parallel, and dynamically adjusts the fusion weights of each modality in the multimodal data. After training, a trained cross-modal attention mechanism model is obtained.
7. The fish biomass estimation method based on a multimodal large model according to claim 3, characterized in that, Dynamic target tracking and occlusion completion are performed on the multimodal fusion features to obtain target tracking and completion information, specifically including: The multimodal fusion features are input into the spatiotemporal Transformer model, and the spatiotemporal Transformer model is used to continuously track the position of the fish school to obtain the fish school position tracking information. The multimodal fusion features are input into a generative adversarial network model, and the generative adversarial network model is used to complete the occluded area of the fish school to obtain the occluded area completion information. Based on the fish school location tracking information and the occlusion area completion information, target tracking and completion information are obtained.
8. The fish biomass estimation method based on a multimodal large model according to claim 1, characterized in that, Based on the target tracking and completion information, a regression model is used to estimate the fish biomass, yielding the fish biomass estimation results, specifically including: Based on the target tracking and completion information, combined with the multimodal data, a regression model is used to estimate the fish biomass using an ensemble learning method, resulting in the fish biomass estimation results, including the total weight of the fish, density distribution, and the number of fish per unit volume.
9. A fish biomass estimation system based on a multimodal large model, characterized in that, It includes a multimodal data acquisition module, a data preprocessing and feature extraction module, a cross-modal attention mechanism fusion module, a dynamic target tracking and occlusion completion module, and a biomass regression and result output module, which are connected in sequence. The multimodal data acquisition module is used to acquire multimodal data of the target area, including underwater fish image data, sonar data, and environmental data; The data preprocessing and feature extraction module, the cross-modal attention mechanism fusion module, and the dynamic target tracking and occlusion completion module are used to perform preprocessing and feature extraction, cross-modal data fusion, and dynamic target tracking and occlusion completion on the multimodal data in sequence based on a multimodal large model, so as to obtain target tracking and completion information. The target tracking and completion information refers to the information obtained after tracking the fish school and completing the areas obscured by the fish school; The biomass regression and result output module is used to estimate the biomass of the fish population using a regression model based on the target tracking and completion information, and to obtain the biomass estimation result of the fish population.
10. The fish biomass estimation system based on a multimodal large model according to claim 9, characterized in that, It also includes a lightweight edge computing deployment module, which is connected to the biomass regression and result output module. The lightweight edge computing deployment module is used to deploy the multimodal large model to an edge device in a lightweight manner to realize real-time inference of the multimodal large model on the edge device.