Schooling fish precision feeding method and device based on multi-modal perception and uncertainty weighted multi-task learning
By employing multimodal perception and uncertainty-weighted multi-task learning, the problems of limited data and insufficient robustness of decision-making models in fish feeding have been solved, enabling efficient and precise fish feeding and improving aquaculture efficiency and environmental sustainability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG OCEAN UNIVERSITY
- Filing Date
- 2026-04-08
- Publication Date
- 2026-07-03
AI Technical Summary
Existing fish feeding methods rely on single-modal data, which makes it difficult to fully reflect the feeding needs of fish and environmental changes. They also lack multimodal fusion mechanisms, and the decision-making models have insufficient robustness and generalization ability in complex environments.
We employ a multimodal perception and uncertainty-weighted multitask learning approach to obtain a unified multimodal joint feature vector through heterogeneous data processing. This vector is then combined with a dual attention mechanism and an adaptive loss balancing mechanism to generate optimized feeding decisions.
It improves the accuracy of feeding decisions and feed utilization, optimizes the sustainability of the aquaculture environment, and promotes fish health.
Smart Images

Figure CN122319977A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent aquaculture technology, and in particular to a method and device for precise feeding of fish swarms based on multimodal perception and uncertainty-weighted multi-task learning. Background Technology
[0002] Aquaculture is a vital industry for ensuring human food supply, and its efficiency and sustainability are crucial. Traditional fish feeding methods rely mainly on manual experience or simple timed and quantitative models. This extensive management approach is no longer suitable for the high-efficiency and environmentally friendly requirements of modern fisheries. Research on precision fish feeding aims to accurately assess the feeding needs of fish schools and the state of the aquaculture environment in real time by applying advanced sensing technologies and intelligent decision-making models, thereby achieving refined control over the amount, speed, and area of feed. This has significant economic and ecological benefits for improving feed utilization, reducing aquaculture costs, minimizing the pollution load of feed residue on water bodies, and improving fish growth. Therefore, developing high-precision fish feeding methods is key to promoting the intelligent and sustainable development of aquaculture.
[0003] While existing methods have made some progress in precise fish feeding, they generally suffer from the following shortcomings: First, most methods rely on single-modal data, making it difficult to comprehensively and accurately reflect the immediate feeding needs of fish and complex environmental changes. Second, even when using multi-source data, they often involve simple feature splicing, lacking effective multimodal fusion mechanisms to extract high-value joint features. More importantly, the decision outputs of existing models are usually single-objective or lack consideration of system uncertainties. In scenarios involving multi-task collaborative optimization such as feeding intensity prediction and residual erbium density quantification, the models lack adaptive balancing mechanisms for differences in the magnitude of losses across different tasks and inherent data noise, resulting in decision models exhibiting poor robustness and generalization ability in complex, dynamic, and aquaculture environments.
[0004] Chinese patent CN120694207A discloses a "Vision-Based Adaptive Fish Feeding Method and Device," which uses machine vision technology to analyze the real-time feeding behavior of fish during the feeding process and uses this as the basis for formulating feeding strategies. This type of method solves the problem that traditional timed and quantitative feeding does not take into account the actual feeding needs of fish. However, this type of method has limitations: First, it is highly dependent on data, and its performance is easily affected by complex underwater environments such as turbid water and changes in lighting, leading to a decrease in the accuracy of fish behavior and residual feed detection. Second, the decision indicators are relatively singular, mainly focusing on the behavior of fish groups, and lacking systematic quantification and processing of the fine spatial distribution of residual feed, feed waste rate, and the inherent uncertainty of multi-source data. Summary of the Invention
[0005] The purpose of this invention is to at least address one of the shortcomings of the prior art and provide a precise feeding method for fish swarms based on multimodal perception and uncertainty-weighted multi-task learning.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] Specifically, a precise feeding method for fish swarms based on multimodal perception and uncertainty-weighted multi-task learning is proposed, including the following: Acquire pre-selected multi-source heterogeneous data as model input data; The input data of the model is processed through a preset feature engineering process to obtain a unified multimodal joint feature vector; The unified multimodal joint feature vector is processed by an uncertainty-weighted multitask model that integrates dual attention mechanism and adaptive loss balance mechanism to obtain optimized feeding decision; The optimized feeding strategy is analyzed to generate feeding control instructions, and the fish are accurately fed based on the feeding control instructions.
[0008] Furthermore, specifically, acquiring pre-selected multi-source heterogeneous data as model input data includes, The system synchronously collects heterogeneous environmental and biological data, gathers real-time data from underwater space, and uses underwater sensors to perform non-contact real-time monitoring of fish populations in the target water area. Key indicators of the aquaculture water body are detected using environmental sensors. High-resolution real-time image sequences for constructing residual erbium density maps are periodically acquired. After processing, these image sequences are used to quantitatively analyze the spatial distribution of residual erbium, the proportion of residual area, and the decay rate of erbium particles in the underwater or suspended state. Key performance indicators of feeding were extracted, and the feeding intensity was calculated based on the feeding behavior of the fish as a quantitative basis for the fish's immediate demand for feed. The feed waste rate was calculated based on the total amount of residual feed to reflect the degree of loss of input.
[0009] Furthermore, specifically, the process of obtaining a unified multimodal joint feature vector by processing the model input data through preset feature engineering includes, Based on the input data of the model, the residual bait density spectrum is constructed and the average speed feature of the fish school is extracted; The preset key performance indicators in the model input data are aligned with the residual bait density spectrum and the average speed characteristics of the fish school in the spatiotemporal dimension to obtain a unified multimodal joint feature vector.
[0010] Furthermore, specifically, the process of constructing the residual bait density spectrum based on the model input data includes, The real-time image sequence is subjected to target recognition and pixel segmentation. By calculating the pixel proportion and spatial distribution characteristics of the residual erbium target region, a residual erbium density quantization map reflecting the abundance and concentration of residual erbium is constructed. If there are N residual erbium in the image, the density matrix of the j-th image with N residual erbium is represented as: ; in, Let represent the Gaussian kernel size of the i-th residual bait, and s represent the two-dimensional coordinates of any pixel in the image. Let the values be represented as the actual two-dimensional coordinates of the i-th residual erbium in the image, when obtaining the density matrix. Then, an adaptive Gaussian density kernel function is used to perform Gaussian kernel blurring on the density matrix. Represented as: ; Where d represents the offset relative to the center point of the Gaussian kernel function, the final density map generation function is expressed as: ; Where s represents the two-dimensional coordinates of any pixel in the image.
[0011] Furthermore, specifically, the process of extracting the average speed feature of the fish swarm based on the model input data includes, The optical flow calculation first requires identifying high-quality feature points within the image frames. Specifically, based on the real-time image sequence acquired by the underwater camera, two consecutive frames within an adjacent time interval are selected as input image frames for optical flow calculation. These image frames are real-time image sequences that have undergone denoising, enhancement, and grayscale preprocessing. Feature points satisfying preset feature criteria are identified within these image frames. These feature criteria include at least one of the following: pixel gradient magnitude, corner response value, local texture intensity, and stability index. For each feature point, the pixel displacement between the two consecutive image frames is calculated by minimizing the error function of the optical flow constraint variance, thereby obtaining the corresponding optical flow velocity. ; in, Represents the velocity of the feature point. Represents the displacement vector of the feature point. Indicates time interval, This represents the x-coordinate of feature point i at time t. Let represent the y-coordinate of feature point i at time t. This represents the x-coordinate of feature point i at time t+1. Let represent the y-coordinate of feature point i+1 at time t, where t represents the time in the previous frame and t+1 represents the time in the next frame. Its velocity component is represented as: ; ; in, This represents the horizontal offset of feature point i. The vertical offset of feature point i is represented; the average velocity of the fish swarm is obtained by arithmetically averaging the velocities of all feature points in the current frame. : ; Where Q represents the total number of detected feature points.
[0012] Furthermore, specifically, the process of obtaining optimized feeding decisions includes: preprocessing the unified multimodal joint feature vector, i.e., performing preliminary feature extraction and batch normalization through a feature sharing network to obtain processed features; then, adaptively weighting the processed features through channel attention and spatial attention mechanisms to obtain a fused feature vector enhanced by the dual attention mechanism; subsequently, inputting the fused feature vector into the feeding intensity classification network and the uneaten food quantity estimation network respectively to obtain the corresponding feeding intensity prediction results and uneaten food quantity prediction results; finally, determining the optimized feeding decision based on the feeding intensity prediction results and uneaten food quantity prediction results. The channel attention mechanism obtains contextual statistics for different channel dimensions by performing global average pooling and max pooling on the processed features, and generates channel weights based on the contextual statistics. These weights are used to weight the processed features along the channel dimension. The formula for calculating the channel attention weights is as follows: ; in, and These represent the average pooling and max pooling of the processed features, respectively. , For the MLP parameters, σ represents the Sigmoid function; The formula for calculating the spatial attention weights is as follows: ; in and These represent the average pooling and max pooling features of the output features processed by the channel attention mechanism, respectively. The comprehensive prediction result is obtained by weighting and fusing the prediction results of feeding intensity and uneaten food quantity. The specific formula is as follows: ; in, This represents the output of the feeding intensity classification network in the i-th dimension. This represents the output of the bait quantity estimation network in the i-th dimension. Z represents the learnable fusion parameters used to achieve the fusion of output results, C represents the comprehensive prediction result, and c represents the total dimension of features involved in the fusion process. The optimal feeding strategy is determined based on Z.
[0013] Furthermore, the feeding intensity classification loss required for training the model is defined. Regression loss from uneaten bait count Binary cross-entropy loss is used to measure the difference between the model's predicted fish feeding categories and the labels in the training data. Feeding intensity classification loss The calculation formula is: ; Where N represents the number of training data points. The parentheses represent the predicted category, and p() represents the probability of the event occurring. Regression loss from bait count The calculation formula is as follows: ; ; ; Where w is a preset weighting coefficient, This represents the mean squared error loss, used to measure the predicted residual bait density map. Labels for residual bait density map Pixel-level differences between them SSIM represents the structural similarity loss, used to measure the similarity between the predicted and actual bait density maps at the pixel level. The expression for SSIM is: ; in, This represents the mean of the predicted residual bait density map. This represents the variance of the predicted residual bait density map. This represents the mean value of the residual bait density map labels. This represents the variance of the residual bait density map labels. This represents the covariance between the predicted bait density map and the bait density map label. and This is a preset constant used to prevent the denominator from being zero.
[0014] Furthermore, to address the issues of inconsistent loss magnitudes and varying convergence difficulties among different tasks in multi-task learning, an uncertainty-based log-likelihood maximization method is employed to construct the total loss function. This achieves adaptive balancing of task weights, and the total loss function formula is derived from the above. and The weighted combination, with the introduction of a regularization term, is calculated using the following formula: ; and These represent the learnable noise variance parameters that the parameters obey.
[0015] Furthermore, specifically, the optimized feeding strategy is analyzed, feeding control instructions are generated, and precise feeding of the fish is achieved based on the feeding control instructions, including: The optimized feeding strategy is analyzed to obtain quantitative indicators; Based on the analyzed quantitative indicators and combined with the preset feeding strategy rules, a machine-executable feeding control instruction is generated. The feeding control instruction is used to transform the abstract decisions of feeding amount, feeding speed and feeding area into specific feeding machine parameter settings. The generated feeding control command is sent to the feeding execution device, which precisely controls the feeding action according to the feeding control command, and at the same time sends back the actual feeding status and execution feedback to complete closed-loop control.
[0016] This invention also proposes a precise fish feeding device based on multimodal perception and uncertainty-weighted multi-task learning, comprising the following: The data acquisition module is used to acquire pre-selected multi-source heterogeneous data as model input data; The data processing module is used to process the model input data through a preset feature engineering process to obtain a unified multimodal joint feature vector; The optimized feeding decision calculation module is used to process the unified multimodal joint feature vector through an uncertainty-weighted multi-task model that integrates dual attention mechanism and adaptive loss balance mechanism to obtain optimized feeding decision; The feeding control module is used to parse the optimized feeding strategy, generate feeding control instructions, and implement precise feeding of the fish based on the feeding control instructions.
[0017] The beneficial effects of this invention are as follows: This invention proposes a method and apparatus for precise fish feeding based on multimodal perception and uncertainty-weighted multi-task learning, comprising: acquiring pre-selected multi-source heterogeneous data as model input data; processing the model input data through pre-defined feature engineering to obtain a unified multimodal joint feature vector; processing the unified multimodal joint feature vector through an uncertainty-weighted multi-task model integrating a dual attention mechanism and an adaptive loss balancing mechanism to obtain an optimized feeding decision; parsing the optimized feeding strategy to generate feeding control instructions, and achieving precise fish feeding based on the feeding control instructions. This invention employs an uncertainty-weighted multi-task model to solve the difficulties of single data and multi-task balancing. Through optimization using a dual attention mechanism, it improves decision accuracy and feed utilization, achieving high-efficiency aquaculture. This invention optimizes the sustainability of the aquaculture environment and promotes water quality and fish health. Attached Figure Description
[0018] The above and other features of this disclosure will become more apparent from the detailed description of the embodiments illustrated in conjunction with the accompanying drawings. In the accompanying drawings, the same reference numerals denote the same or similar elements. Obviously, the drawings described below are merely some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without any creative effort. In the drawings: Figure 1 The flowchart shown is a process for the precise feeding method for fish swarms based on multimodal perception and uncertainty-weighted multi-task learning according to the present invention. Figure 2 The diagram shows a flowchart of obtaining a unified multimodal joint feature vector in this invention. Figure 3 The diagram shown is a schematic of the uncertainty-weighted multi-task model proposed in this invention. Detailed Implementation
[0019] The following will provide a clear and complete description of the concept, specific structure, and technical effects of the present invention in conjunction with embodiments and accompanying drawings, so as to fully understand the purpose, solution, and effects of the present invention. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The same reference numerals used throughout the accompanying drawings indicate the same or similar parts.
[0020] Example 1, referring to Figure 1 This invention proposes a precise feeding method for fish swarms based on multimodal perception and uncertainty-weighted multi-task learning, including the following: Step S1: Obtain pre-selected multi-source heterogeneous data as model input data; Step S2: The model input data is processed through a preset feature engineering process to obtain a unified multimodal joint feature vector; Step S3: The unified multimodal joint feature vector is processed by an uncertainty-weighted multi-task model that integrates dual attention mechanism and adaptive loss balance mechanism to obtain optimized feeding decision; Step S4: Analyze the optimized feeding strategy, generate feeding control instructions, and implement precise feeding of the fish based on the feeding control instructions.
[0021] In a preferred embodiment of the present invention, specifically, acquiring pre-selected multi-source heterogeneous data as model input data includes... The system synchronously collects heterogeneous environmental and biological data, gathers real-time data from underwater space, and uses underwater sensors to perform non-contact real-time monitoring of fish populations in the target water area. Key indicators of the aquaculture water body are detected using environmental sensors. High-resolution real-time image sequences for constructing residual erbium density maps are periodically acquired. After processing, these image sequences are used to quantitatively analyze the spatial distribution of residual erbium, the proportion of residual area, and the decay rate of erbium particles in the underwater or suspended state. Key performance indicators of feeding were extracted, and the feeding intensity was calculated based on the feeding behavior of the fish as a quantitative basis for the fish's immediate demand for feed. The feed waste rate was calculated based on the total amount of residual feed to reflect the degree of loss of input.
[0022] Reference Figure 2 In a preferred embodiment of the present invention, specifically, the process of obtaining a unified multimodal joint feature vector by processing the model input data through a preset feature engineering process includes, Based on the input data of the model, the residual bait density spectrum is constructed and the average speed feature of the fish school is extracted; The preset key performance indicators in the model input data are aligned with the residual bait density spectrum and the average speed characteristics of the fish school in the spatiotemporal dimension to obtain a unified multimodal joint feature vector.
[0023] In a preferred embodiment of the present invention, specifically, the process of constructing the residual bait density spectrum based on the model input data includes, The real-time image sequence is subjected to target recognition and pixel segmentation. By calculating the pixel proportion and spatial distribution characteristics of the residual erbium target region, a residual erbium density quantization map reflecting the abundance and concentration of residual erbium is constructed. If there are N residual erbium in the image, the density matrix of the j-th image with N residual erbium is represented as: ; in, Let represent the Gaussian kernel size of the i-th residual bait, and s represent the two-dimensional coordinates of any pixel in the image. Let the values be represented as the actual two-dimensional coordinates of the i-th residual erbium in the image, when obtaining the density matrix. Then, an adaptive Gaussian density kernel function is used to perform Gaussian kernel blurring on the density matrix. Represented as: ; Where d represents the offset relative to the center point of the Gaussian kernel function, the final density map generation function is expressed as: ; Where s represents the two-dimensional coordinates of any pixel in the image.
[0024] In a preferred embodiment of the present invention, specifically, the process of extracting the average speed feature of the fish school based on the model input data includes: The optical flow calculation first requires identifying high-quality feature points within the image frames. Specifically, based on the real-time image sequence acquired by the underwater camera, two consecutive frames within an adjacent time interval are selected as input image frames for optical flow calculation. These image frames are real-time image sequences that have undergone denoising, enhancement, and grayscale preprocessing. Feature points satisfying preset feature criteria are identified within these image frames. These feature criteria include at least one of the following: pixel gradient magnitude, corner response value, local texture intensity, and stability index. For each feature point, the pixel displacement between the two consecutive image frames is calculated by minimizing the error function of the optical flow constraint variance, thereby obtaining the corresponding optical flow velocity. ; in, Represents the velocity of the feature point. Represents the displacement vector of the feature point. Indicates time interval, This represents the x-coordinate of feature point i at time t. Let represent the y-coordinate of feature point i at time t. This represents the x-coordinate of feature point i at time t+1. Let represent the y-coordinate of feature point i+1 at time t, where t represents the time in the previous frame and t+1 represents the time in the next frame. Its velocity component is represented as: ; ; in, This represents the horizontal offset of feature point i. The vertical offset of feature point i is represented; the average velocity of the fish swarm is obtained by arithmetically averaging the velocities of all feature points in the current frame. : ; Where Q represents the total number of detected feature points.
[0025] Reference Figure 3 As a preferred embodiment of the present invention, the process of obtaining optimized feeding decisions specifically includes: preprocessing a unified multimodal joint feature vector, i.e., performing preliminary feature extraction and batch normalization through a feature sharing network to obtain processed features; then, adaptively weighting the processed features through channel attention and spatial attention mechanisms to obtain a fused feature vector enhanced by the dual attention mechanism; then, inputting the fused feature vector into a feeding intensity classification network and a residual bait quantity estimation network respectively to obtain corresponding feeding intensity prediction results and residual bait quantity prediction results; finally, determining the optimized feeding decision based on the feeding intensity prediction results and residual bait quantity prediction results. The channel attention mechanism obtains contextual statistics for different channel dimensions by performing global average pooling and max pooling on the processed features, and generates channel weights based on the contextual statistics. These weights are used to weight the processed features along the channel dimension. The formula for calculating the channel attention weights is as follows: ; in, and These represent the average pooling and max pooling of the processed features, respectively. , For the MLP parameters, σ represents the Sigmoid function; The formula for calculating the spatial attention weights is as follows: ; in and These represent the average pooling and max pooling features of the output features processed by the channel attention mechanism, respectively. The comprehensive prediction result is obtained by weighting and fusing the prediction results of feeding intensity and uneaten food quantity. The specific formula is as follows: ; in, This represents the output of the feeding intensity classification network in the i-th dimension. This represents the output of the bait quantity estimation network in the i-th dimension. Z represents the learnable fusion parameters used to achieve the fusion of output results, C represents the comprehensive prediction result, and c represents the total dimension of features involved in the fusion process. The optimal feeding strategy is determined based on Z.
[0026] The "uncertainty weighted fusion" is a decision-level fusion performed at the model output layer. Its purpose is to adaptively adjust the contribution of each task's prediction results to the final result based on their uncertainty, thereby reducing the interference of high-uncertainty results on feeding decisions and improving the overall stability and reliability of the decision. The fused prediction results are used to comprehensively assess the current feeding needs and uneaten food status of the fish population, and serve as the basis for generating optimized feeding decisions.
[0027] In this preferred embodiment, the unified joint feature vector is preprocessed, i.e., preliminary feature extraction and batch normalization are performed through a feature sharing network to obtain processed features. Based on this, channel attention and spatial attention mechanisms are introduced to adaptively weight the processed features to enhance the model's ability to identify key features, resulting in a fused feature vector enhanced by the dual attention mechanism. This fused feature vector is a feature representation after joint weighting by channel attention and spatial attention. The channel attention mechanism obtains contextual statistics under different channel dimensions by performing global average pooling and max pooling on the processed features, and generates channel weights based on these contextual statistics. These weights are used to weight the processed features along the channel dimension to highlight feature channels that contribute significantly to task discrimination. The spatial attention mechanism is based on… Spatial weights are generated based on the saliency distribution of features in spatial location. These weights are used to enhance the channel-weighted features in the spatial dimension, thereby strengthening key spatial regions related to fish behavior and uneaten food distribution. The channel weights and spatial weights act on the corresponding channel and spatial dimensions, respectively, jointly weighting the same processed features to achieve feature enhancement fusion based on a dual-attention mechanism. The fused feature vector enhanced by the dual-attention mechanism is then input into the feeding intensity classification network and the uneaten food quantity estimation network, respectively, to obtain the corresponding feeding intensity prediction results and uneaten food quantity prediction results. Further, based on the feeding intensity prediction results and uneaten food quantity prediction results, the current feeding state is comprehensively evaluated, and the feeding device is controlled to perform the feeding operation. Z is input into the pre-trained feeding decision model to output the corresponding optimized feeding scheme. The channel attention mechanism and spatial attention mechanism belong to the feature-level weighting enhancement process. The subsequent weighted fusion based on uncertainty weights belongs to the model output level fusion processing. Both are independent in function and hierarchical level, and together they are used to improve the stability and reliability of feeding decisions.
[0028] In a preferred embodiment of the present invention, the feeding intensity classification loss required for model training is defined. Regression loss from uneaten bait count Binary cross-entropy loss is used to measure the difference between the model's predicted fish feeding categories and the labels in the training data. Feeding intensity classification loss The calculation formula is: ; Where N represents the number of training data points. The parentheses represent the predicted category, and p() represents the probability of the event occurring. Regression loss from bait count The calculation formula is as follows: ; ; ; Where w is a preset weighting coefficient, This represents the mean squared error loss, used to measure the predicted residual bait density map. Labels for residual bait density map Pixel-level differences between them SSIM represents the structural similarity loss, used to measure the similarity between the predicted and actual bait density maps at the pixel level. The expression for SSIM is: ; in, This represents the mean of the predicted residual bait density map. This represents the variance of the predicted residual bait density map. This represents the mean value of the residual bait density map labels. This represents the variance of the residual bait density map labels. This represents the covariance between the predicted bait density map and the bait density map label. and This is a preset constant used to prevent the denominator from being zero.
[0029] As a preferred embodiment of the present invention, to address the issues of inconsistent loss magnitudes and varying convergence difficulties among different tasks in multi-task learning, a log-likelihood maximization method based on uncertainty is used to construct the total loss function. This achieves adaptive balancing of task weights, and the total loss function formula is derived from the above. and The weighted combination, with the introduction of a regularization term, is calculated using the following formula: ; and These represent the learnable noise variance parameters that the parameters obey.
[0030] In this preferred embodiment, the loss function is based on the following probability model, assuming... The output of a neural network with weights W after processing the input x is represented as the probabilistic model for a regression task: ; in, This indicates the probability of the event within the parentheses occurring. This represents the actual output value of the regression task. This represents forecast uncertainty. This indicates that the random variable follows a normal distribution, where the first parameter is the mean of the distribution and the second parameter is the variance of the distribution. The variance of the Gaussian distribution represents the inherent noise level of the regression task. For classification tasks, it can be expressed as: ; Among them, this place This represents the actual output category of the classification task. This represents a commonly used activation function that transforms the raw output of a neural network into a probability distribution.
[0031] In a preferred embodiment of the present invention, specifically, the optimized feeding strategy is parsed, feeding control instructions are generated, and precise feeding of the fish is achieved based on the feeding control instructions, including: The optimized feeding strategy is analyzed to obtain quantitative indicators; Based on the analyzed quantitative indicators and combined with the preset feeding strategy rules, a machine-executable feeding control instruction is generated. The feeding control instruction is used to transform the abstract decisions of feeding amount, feeding speed and feeding area into specific feeding machine parameter settings. The generated feeding control command is sent to the feeding execution device, which precisely controls the feeding action according to the feeding control command, and at the same time sends back the actual feeding status and execution feedback to complete closed-loop control.
[0032] Example 2: This invention also proposes a precise fish feeding device based on multimodal perception and uncertainty-weighted multi-task learning, comprising the following: The data acquisition module is used to acquire pre-selected multi-source heterogeneous data as model input data; The data processing module is used to process the model input data through a preset feature engineering process to obtain a unified multimodal joint feature vector; The optimized feeding decision calculation module is used to process the unified multimodal joint feature vector through an uncertainty-weighted multi-task model that integrates dual attention mechanism and adaptive loss balance mechanism to obtain optimized feeding decision; The feeding control module is used to parse the optimized feeding strategy, generate feeding control instructions, and implement precise feeding of the fish based on the feeding control instructions.
[0033] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0034] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or system capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0035] Although the description of the invention has been quite detailed and particularly of several described embodiments, it is not intended to limit it to any of these details or embodiments or any particular embodiment, but should be considered as providing a broad possible interpretation of the claims by referring to the appended claims and taking into account the prior art, thereby effectively covering the intended scope of the invention. Furthermore, the invention has been described above with respect to embodiments foreseeable by the inventors in order to provide a useful description, and non-substantial modifications to the invention that have not yet been foreseen may still represent equivalent modifications.
[0036] The above description is merely a preferred embodiment of the present invention. The present invention is not limited to the above-described embodiments. Any embodiment that achieves the technical effects of the present invention using the same means should fall within the protection scope of the present invention. Within the protection scope of the present invention, various modifications and variations can be made to the technical solutions and / or implementation methods.
Claims
1. A fish precise feeding method based on multi-modal perception and uncertainty weighted multi-task learning, characterized in that, Including the following: Acquire pre-selected multi-source heterogeneous data as model input data; The input data of the model is processed through a preset feature engineering process to obtain a unified multimodal joint feature vector; The unified multimodal joint feature vector is processed by an uncertainty-weighted multitask model that integrates dual attention mechanism and adaptive loss balance mechanism to obtain optimized feeding decision; The optimized feeding strategy is analyzed to generate feeding control instructions, and the fish are accurately fed based on the feeding control instructions.
2. The method for precise feeding of fish swarms based on multimodal perception and uncertainty-weighted multi-task learning according to claim 1, characterized in that, Specifically, pre-selected multi-source heterogeneous data is acquired as model input data, including: The system synchronously collects heterogeneous environmental and biological data, gathers real-time data from underwater space, and uses underwater sensors to perform non-contact real-time monitoring of fish populations in the target water area. Key indicators of the aquaculture water body are detected using environmental sensors. High-resolution real-time image sequences for constructing residual erbium density maps are periodically acquired. After processing, these image sequences are used to quantitatively analyze the spatial distribution of residual erbium, the proportion of residual area, and the decay rate of erbium particles in the underwater or suspended state. Key performance indicators of feeding were extracted, and the feeding intensity was calculated based on the feeding behavior of the fish as a quantitative basis for the fish's immediate demand for feed. The feed waste rate was calculated based on the total amount of residual feed to reflect the degree of loss of input.
3. The method for precise fish feeding based on multimodal perception and uncertainty-weighted multi-task learning according to claim 2, characterized in that, Specifically, the process of obtaining a unified multimodal joint feature vector by processing the model input data through preset feature engineering includes: Based on the input data of the model, the residual bait density spectrum is constructed and the average speed feature of the fish school is extracted; The preset key performance indicators in the model input data are aligned with the residual bait density spectrum and the average speed characteristics of the fish school in the spatiotemporal dimension to obtain a unified multimodal joint feature vector.
4. The method for precise feeding of fish swarms based on multimodal perception and uncertainty-weighted multi-task learning according to claim 3, characterized in that, Specifically, the process of constructing the residual bait density spectrum based on the model input data includes: The real-time image sequence is subjected to target recognition and pixel segmentation. By calculating the pixel proportion and spatial distribution characteristics of the residual erbium target region, a residual erbium density quantization map reflecting the abundance and concentration of residual erbium is constructed. If there are N residual erbium in the image, the density matrix of the j-th image with N residual erbium is represented as: ; in, Let represent the Gaussian kernel size of the i-th residual bait, and s represent the two-dimensional coordinates of any pixel in the image. Let the values be represented as the actual two-dimensional coordinates of the i-th residual erbium in the image, when obtaining the density matrix. Then, an adaptive Gaussian density kernel function is used to perform Gaussian kernel blurring on the density matrix. Represented as: ; Where d represents the offset relative to the center point of the Gaussian kernel function, the final density map generation function is expressed as: ; Where s represents the two-dimensional coordinates of any pixel in the image.
5. The method for precise feeding of fish swarms based on multimodal perception and uncertainty-weighted multi-task learning according to claim 3, characterized in that, Specifically, the process of extracting the average speed feature of the fish swarm based on the model input data includes: The optical flow calculation first requires identifying high-quality feature points within the image frames. Specifically, based on the real-time image sequence acquired by the underwater camera, two consecutive frames within an adjacent time interval are selected as input image frames for optical flow calculation. These image frames are real-time image sequences that have undergone denoising, enhancement, and grayscale preprocessing. Feature points satisfying preset feature criteria are identified within these image frames. These feature criteria include at least one of the following: pixel gradient magnitude, corner response value, local texture intensity, and stability index. For each feature point, the pixel displacement between the two consecutive image frames is calculated by minimizing the error function of the optical flow constraint variance, thereby obtaining the corresponding optical flow velocity. ; in, Represents the velocity of the feature point. Represents the displacement vector of the feature point. Indicates time interval, This represents the x-coordinate of feature point i at time t. This represents the y-coordinate of feature point i at time t. This represents the x-coordinate of feature point i at time t+1. Let represent the y-coordinate of feature point i+1 at time t, where t represents the time in the previous frame and t+1 represents the time in the next frame. Its velocity component is expressed as: ; ; in, This represents the horizontal offset of feature point i. The vertical offset of feature point i is represented; the average velocity of the fish swarm is obtained by arithmetically averaging the velocities of all feature points in the current frame. : ; Where Q represents the total number of detected feature points.
6. The method for precise fish feeding based on multimodal perception and uncertainty-weighted multi-task learning according to claim 1, characterized in that, Specifically, the process of obtaining optimized feeding decisions includes: preprocessing the unified multimodal joint feature vector, i.e., performing preliminary feature extraction and batch normalization through a feature sharing network to obtain processed features; then, adaptively weighting the processed features through channel attention and spatial attention mechanisms to obtain a fused feature vector enhanced by the dual attention mechanism; then, inputting the fused feature vector into the feeding intensity classification network and the uneaten food quantity estimation network respectively to obtain the corresponding feeding intensity prediction results and uneaten food quantity prediction results; finally, determining the optimized feeding decision based on the feeding intensity prediction results and uneaten food quantity prediction results. The channel attention mechanism obtains contextual statistics for different channel dimensions by performing global average pooling and max pooling on the processed features, and generates channel weights based on the contextual statistics. These weights are used to weight the processed features along the channel dimension. The formula for calculating the channel attention weights is as follows: ; in, and These represent the average pooling and max pooling of the processed features, respectively. , For the MLP parameters, σ represents the Sigmoid function; The formula for calculating the spatial attention weights is as follows: ; in and These represent the average pooling and max pooling features of the output features processed by the channel attention mechanism, respectively. The comprehensive prediction result is obtained by weighting and fusing the prediction results of feeding intensity and uneaten food quantity. The specific formula is as follows: ; in, This represents the output of the feeding intensity classification network in the i-th dimension. This represents the output of the bait quantity estimation network in the i-th dimension. Z represents the learnable fusion parameters used to achieve the fusion of output results, C represents the comprehensive prediction result, and c represents the total dimension of features involved in the fusion process. The optimal feeding strategy is determined based on Z.
7. The method for precise feeding of fish swarms based on multimodal perception and uncertainty-weighted multi-task learning according to claim 6, characterized in that, Define the feeding intensity classification loss required for training the model. Regression loss from uneaten bait count Binary cross-entropy loss is used to measure the difference between the model's predicted fish feeding categories and the labels in the training data. Feeding intensity classification loss The calculation formula is: ; Where N represents the number of training data points. The parentheses represent the predicted category, and p() represents the probability of the event occurring. Regression loss from bait count The calculation formula is as follows: ; ; ; Where w is a preset weighting coefficient. This represents the mean squared error loss, used to measure the predicted residual bait density map. Labels with residual bait density map Pixel-level differences between them SSIM represents the structural similarity loss, used to measure the similarity between the predicted and actual bait density maps at the pixel level. The expression for SSIM is: ; in, This represents the mean of the predicted residual bait density map. This represents the variance of the predicted residual bait density map. This represents the mean value of the residual bait density map labels. This represents the variance of the residual bait density map labels. This represents the covariance between the predicted bait density map and the bait density map label. and This is a preset constant used to prevent the denominator from being zero.
8. The method for precise fish feeding based on multimodal perception and uncertainty-weighted multi-task learning according to claim 7, characterized in that, To address the issues of inconsistent loss magnitudes and varying convergence difficulties among different tasks in multi-task learning, an uncertainty-based log-likelihood maximization method is used to construct the total loss function. This achieves adaptive balancing of task weights, and the total loss function formula is derived from the above. and The weighted combination, with the introduction of a regularization term, is calculated using the following formula: ; and These represent the learnable noise variance parameters that the parameters obey.
9. The method for precise feeding of fish swarms based on multimodal perception and uncertainty-weighted multi-task learning according to claim 1, characterized in that, Specifically, the optimized feeding strategy is analyzed, feeding control instructions are generated, and precise feeding of the fish is achieved based on the feeding control instructions, including: The optimized feeding strategy is analyzed to obtain quantitative indicators; Based on the analyzed quantitative indicators and combined with the preset feeding strategy rules, a machine-executable feeding control instruction is generated. The feeding control instruction is used to transform the abstract decisions of feeding amount, feeding speed and feeding area into specific feeding machine parameter settings. The generated feeding control command is sent to the feeding execution device, which precisely controls the feeding action according to the feeding control command, and at the same time sends back the actual feeding status and execution feedback to complete closed-loop control.
10. A precise fish feeding device based on multimodal perception and uncertainty-weighted multi-task learning, characterized in that, Including the following: The data acquisition module is used to acquire pre-selected multi-source heterogeneous data as model input data; The data processing module is used to process the model input data through a preset feature engineering process to obtain a unified multimodal joint feature vector; The optimized feeding decision calculation module is used to process the unified multimodal joint feature vector through an uncertainty-weighted multi-task model that integrates dual attention mechanism and adaptive loss balance mechanism to obtain optimized feeding decision; The feeding control module is used to parse the optimized feeding strategy, generate feeding control instructions, and implement precise feeding of the fish based on the feeding control instructions.
Citation Information
Patent Citations
Fish self-adaptive feeding method and device based on vision
CN120694207A