A non-invasive underwater organism feeding monitoring system integrating passive acoustics and image recognition

The underwater biological feeding monitoring system, which integrates passive acoustics and image recognition, solves the problems of low efficiency, decreased accuracy, and high cost in aquaculture monitoring technology. It enables all-weather, multi-dimensional biological feeding identification and precise feeding decisions, thereby reducing economic losses.

CN122490197APending Publication Date: 2026-07-31SHANGHAI OCEAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI OCEAN UNIV
Filing Date
2026-05-06
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing aquaculture monitoring technologies suffer from problems such as low efficiency of manual monitoring, decreased accuracy of single monitoring technologies, insufficient integration depth, lack of causal relationship modeling, and high costs, making it difficult to achieve precise feeding and reduce economic losses.

Method used

A non-invasive underwater biological feeding monitoring system that integrates passive acoustics and image recognition achieves all-weather, high-precision, and multi-dimensional biological feeding identification and decision support through multimodal data acquisition, enhanced single-modal feature extraction, multimodal feature interaction and fusion, joint learning of feeding causal chains, and perception of state uncertainty.

Benefits of technology

It achieves all-weather, high-precision biological feeding recognition, reduces feed waste and breeding costs, adapts to complex environments, supports precise feeding decisions, and has promotional value among small and medium-sized farmers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490197A_ABST
    Figure CN122490197A_ABST
Patent Text Reader

Abstract

This invention discloses a non-invasive underwater organism feeding monitoring system that integrates passive acoustics and image recognition. The system includes: a multimodal data acquisition module that simultaneously acquires acoustic and image signals; an enhanced single-modal feature extraction module, wherein the image feature extraction network includes a fish school clustering effect perception module that outputs the fish school clustering area and compactness index; a multimodal feature interaction and fusion module that adaptively fuses acoustic and image feature maps at the feeding stage; a feeding causal chain joint learning module that outputs feeding intensity level, uneaten food quantity, and fish school clustering degree in parallel under causal consistency constraints; a feeding state uncertainty perception module that calculates stage boundaries, modal conflicts, and temporal anomalies, and selects high-uncertainty samples for incremental training; and a control and decision output module that generates feeding instructions and warning signals based on the feeding intensity index. This invention achieves high-precision, multi-dimensional feeding monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent monitoring technology for aquaculture, specifically to a non-invasive underwater organism feeding monitoring system that integrates passive acoustics and image recognition. Background Technology

[0002] my country is the world's largest aquaculture producer, but its aquaculture monitoring technology still faces the following bottlenecks:

[0003] 1. Low efficiency of manual monitoring: Traditional feeding relies on the experience and judgment of farmers, making nighttime monitoring difficult and delaying the detection of abnormal feeding behavior, which can easily lead to feed waste or insufficient feeding and economic losses.

[0004] 2. Single monitoring technologies have inherent defects: Optical monitoring technologies (images / videos) are easily affected by factors such as water turbidity, changes in light, and fish obstruction, resulting in a significant decrease in recognition accuracy in complex aquaculture environments; Although acoustic monitoring technologies can penetrate complex water bodies, they are easily affected by environmental noise from aerators, water pumps, etc., and lack intuitive verification methods, making it difficult to ensure the reliability of recognition results.

[0005] 3. Existing fusion technologies lack sufficient fusion depth and are devoid of biological mechanism-driven approaches: Most existing acoustic-optical fusion solutions employ a "post-fusion" strategy, which involves simple weighting of the results after separate identification, failing to achieve deep interaction at the feature level. More importantly, existing technologies all adopt a "data-driven" paradigm—collecting large amounts of labeled data to train general models. These models learn statistical correlations in the data rather than causal patterns in biological behavior, resulting in poor performance in edge scenarios (such as overfeeding phases and bimodal information conflicts).

[0006] 4. Limited monitoring dimensions and lack of causal relationship modeling: Existing methods typically only output binary classification results of "feeding / non-feeding", failing to provide richer decision-making basis such as the amount of uneaten food and the degree of fish aggregation, and failing to consider the causal time sequence relationship between these indicators (uneaten food presence → fish aggregation → feeding behavior → uneaten food reduction → decreased feeding intensity → fish dispersal), making it difficult to meet the demand for multidimensional information for precise feeding.

[0007] 5. High cost and deployment threshold: The unit price of imported similar equipment exceeds 100,000 yuan, making it difficult to popularize among small and medium-sized farmers.

[0008] Therefore, there is an urgent need for a non-invasive underwater biological feeding monitoring system that is driven by biological feeding mechanisms, has a deeper integration depth, richer output dimensions, and causal reasoning capabilities, so as to promote the intelligent upgrading of aquaculture. Summary of the Invention

[0009] This invention provides a non-invasive underwater organism feeding monitoring system that integrates passive acoustics and image recognition to achieve all-weather, high-precision, and multi-dimensional identification of underwater organism feeding behavior, reduce feed waste and aquaculture costs, and provide comprehensive decision support for precise feeding.

[0010] This invention provides a non-invasive underwater organism feeding monitoring system that integrates passive acoustics and image recognition. The system includes:

[0011] A multimodal data acquisition module is used to simultaneously acquire acoustic and image signals from underwater organism feeding scenarios;

[0012] An enhanced single-modal feature extraction module is used to extract features from acoustic signals and image signals respectively, and obtain acoustic feature maps and image feature maps accordingly;

[0013] The multimodal feature interaction and fusion module is used to perform multi-level fusion of acoustic feature maps and image feature maps. Multi-level fusion includes feature-level interaction and feeding stage adaptive fusion. Feature-level interaction is carried out through bidirectional parallel interaction units to exchange information multiple times between acoustic features and image features. Feeding stage adaptive fusion is to build a feeding stage perceptron, infer the current feeding stage in real time based on the rate of change of fish aggregation degree and the rate of uneaten food consumption, and dynamically adjust the fusion weights of multi-scale features according to the feeding stage.

[0014] The feeding causal chain joint learning module is used to output multiple task results with causal relationships in parallel based on the fused features. The multiple task results include at least the biological feeding intensity level classification results, the estimated results of the number or distribution of uneaten food, and the fish school aggregation degree assessment results.

[0015] The feeding state uncertainty perception module is used to calculate the feeding stage boundary uncertainty, modal conflict uncertainty, and time series anomaly uncertainty of the sample, and selects the sample with the highest uncertainty score for incremental training.

[0016] The control and decision output module is used to generate precise feeding control commands and abnormal warning signals based on a segmented control strategy based on the feeding intensity index, according to the results of multiple tasks.

[0017] In some embodiments of the present invention, the multimodal data acquisition module includes:

[0018] The passive acoustic acquisition unit includes four high-sensitivity piezoelectric hydrophones, which are placed at a depth of 30cm-80cm below the surface of the aquaculture pond.

[0019] The underwater image acquisition unit includes an industrial-grade waterproof low-light camera with a resolution of 1280×720 and a frame rate of no less than 15fps.

[0020] The clock synchronization component uses the internal clock of the embedded processor or an external high-precision clock source to synchronously acquire sound and image signals, with a timestamp accuracy of ≤1ms.

[0021] In some embodiments of the present invention, the enhanced single-modal feature extraction module includes an acoustic feature extraction network;

[0022] Acoustic feature extraction networks include:

[0023] The signal preprocessing unit is used to perform framing, windowing, pre-emphasis, and noise suppression on the acoustic signal to obtain the preprocessed acoustic signal.

[0024] The multi-scale convolutional unit adopts a hierarchical residual structure to divide the preprocessed acoustic signal into multiple sub-feature groups along the channel dimension. For each sub-feature group, a convolutional kernel of different size is used to extract features to obtain a feature map.

[0025] The coordinate attention unit is used to perform average pooling along the height and width directions of the feature map, embedding positional information into the channel attention, generating attention weights that simultaneously contain channel and spatial positional information, and multiplying them with the input feature map to obtain the acoustic feature map.

[0026] In some embodiments of the present invention, the enhanced single-modal feature extraction module includes an image feature extraction network, the image feature extraction network includes a fish swarming effect perception module, and the fish swarming effect perception module includes:

[0027] The cascaded large kernel attention unit is used to decompose the two-dimensional large convolution kernel into cascaded vertical one-dimensional convolution and horizontal one-dimensional convolution to simulate the perception of information in front of and to the sides of a school of fish.

[0028] The aggregation compactness calculation unit is used to calculate a compactness index that quantifies the degree of coordination among fish groups based on the spatial distribution density and movement direction consistency of individual fish. When the compactness index exceeds a first threshold and continues to rise, it is determined that the feeding period has begun. When the compactness index decreases but there is still uneaten food remaining, it is determined that the satiated period has begun.

[0029] In some embodiments of the present invention, adaptive fusion during the feeding phase includes:

[0030] The feeding phase sensor is used to infer the current feeding phase based on the currently detected rate of change in fish aggregation ΔA / Δt and the rate of uneaten food consumption ΔB / Δt, according to the following rules:

[0031] When ΔA / Δt > 0 and ΔB / Δt ≈ 0, it is determined to be the search period;

[0032] When ΔA / Δt > 0 and ΔB / Δt < 0, it is determined to be the aggregation period;

[0033] When ΔA / Δt ≈ 0 and ΔB / Δt < 0, it is determined to be the feeding period;

[0034] When ΔA / Δt < 0 and ΔB / Δt > 0, it is determined to be the satiated period;

[0035] When ΔA / Δt < 0 and ΔB / Δt ≈ 0, it is determined to be a period of fasting;

[0036] The weighted dynamic adjustment unit is used to dynamically adjust the fusion weights of multi-scale features according to the feeding stage. It enhances small-scale features during the aggregation and feeding periods to capture individual feeding details, and enhances large-scale features during the searching and satiated periods to capture group movement trends.

[0037] In some embodiments of the present invention, the bidirectional parallel interaction unit includes:

[0038] The first interactive subunit is used to use the acoustic feature map as the query vector, the image feature map as the key vector and the value vector, and to calculate the acoustic features enhanced by the image features through an attention mechanism to enhance the global semantic information of the acoustic features.

[0039] The second interactive subunit is used to use the image feature map as the query vector and the acoustic feature map as the key vector and value vector, and to calculate the image features enhanced by acoustic features through an attention mechanism to enhance the local detail information of the image features.

[0040] The feature fusion subunit is used to add the enhanced acoustic feature map with the enhanced image feature map by residual addition or channel splicing to obtain the fused feature map.

[0041] In some embodiments of the present invention, the feeding causal chain joint learning module adopts a hard parameter shared network structure, and the feeding causal chain joint learning module includes a shared feature extraction network and a task-specific branch network;

[0042] A shared feature extraction network is used to receive the fused feature map and extract a general semantic representation from the fused feature map;

[0043] Task-specific branch networks include:

[0044] A classification branching network is used to output the classification results of feeding intensity levels, which include at least high active feeding, moderate active feeding, low active feeding, and no feeding.

[0045] A counting branch network is used to output a residual bait density map or an estimate of the number of residual baits.

[0046] The detection branch network is used to output the location, area, and cluster compactness index of the fish gathering area;

[0047] The loss function of the feeding causal chain joint learning module during training includes the following components:

[0048] The target detection loss uses a size-adaptive penalty factor, with the denominator of the penalty factor being the width and height of the target bounding box;

[0049] Temporal consistency loss constrains the smooth variation of feeding intensity prediction results in adjacent time frames;

[0050] The causal consistency loss constrains the correlation coefficients among the fish school aggregation rate, uneaten food consumption rate, and feeding intensity level to meet the preset sign constraint.

[0051] Modal consistency loss is used to minimize the difference between the feeding intensity predicted by acoustic features and the feeding intensity predicted by image features.

[0052] In some embodiments of the present invention, the feeding state uncertainty perception module includes:

[0053] The stage boundary uncertainty calculation unit is used to calculate the probability that a sample is at the boundary of the feeding stage based on the information entropy of the probability distribution of feeding intensity level;

[0054] The modal conflict uncertainty calculation unit is used to calculate the degree of bimodal information conflict based on the difference between acoustic prediction results and image prediction results;

[0055] The temporal anomaly uncertainty calculation unit is used to calculate the degree of temporal anomaly based on the degree of deviation between the predicted temporal changes in feeding intensity and the preset causal law;

[0056] The sample selection unit is used to select a preset number of samples with the highest scores based on the weighted sum of uncertainty scores, which are then labeled by experts and added to the training set for incremental training.

[0057] In some embodiments of the present invention, the control and decision output module calculates the feeding intensity index using the following formula:

[0058]

[0059] in, To determine the confidence level for fish aggregation. This represents the area where the fish congregate. To determine the confidence level for residual bait detection, The area of ​​the uneaten bait. This is the adjustment coefficient;

[0060] Based on the feeding intensity index I, a segmented control strategy based on biological feeding patterns is adopted:

[0061] When I ≥ 0.7, it is determined to be a period of high feeding demand, and the amount of feed should be increased;

[0062] When 0.3 < I < 0.7, it is determined to be a normal feeding period, and normal feeding should be maintained;

[0063] When I ≤ 0.3, it is determined to be a period of satiation or fasting, and feeding should be stopped.

[0064] In some embodiments of the present invention, the control and decision output module includes:

[0065] The abnormal warning unit will push an alarm signal locally or remotely when the feeding intensity index drops more than a threshold within a preset time window, or when there is equipment failure or a sudden increase in environmental noise.

[0066] The data storage unit is used to store the original audio signals, image keyframes, and historical statistical results of each task, and supports data export and report generation.

[0067] In the non-invasive underwater biological feeding monitoring system that integrates passive acoustics and image recognition provided by this invention, a swarm sensing mechanism is introduced to simulate fish school aggregation behavior. A feeding stage sensor dynamically adjusts the fusion strategy, and the prediction process is constrained by a feeding causal chain to ensure that the prediction results conform to the biological laws of fish feeding behavior. Through feature-level bidirectional interaction and joint learning of the feeding causal chain, not only can the feeding intensity level be output, but also the quantity of uneaten food, fish school aggregation degree, and aggregation compactness index can be simultaneously provided, thus achieving a multi-dimensional representation of the feeding state. Furthermore, the acoustic modality can penetrate turbid water to obtain target information, while the image modality provides intuitive visual verification. The two are deeply fused at the intermediate layer through feature-level interaction, effectively improving the system's adaptability to complex environmental factors such as water turbidity, light changes, and background noise. Simultaneously, this invention includes a feeding state uncertainty perception module to automatically filter stage boundary samples, modal conflict samples, and temporal anomaly samples, thereby achieving continuous model optimization with low annotation costs and enabling the system to adapt to the characteristic changes of different farming stages and different farmed species. Attached Figure Description

[0068] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0069] Figure 1 This is a schematic diagram of the structure of the non-invasive underwater organism feeding monitoring system that integrates passive acoustics and image recognition, provided in an embodiment of the present invention.

[0070] Figure 2This is a flowchart illustrating the fish swarming effect sensing module provided in an embodiment of the present invention. Detailed Implementation

[0071] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0072] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0073] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.

[0074] The use of "applies to" or "configured to" in this invention implies an open and inclusive language, which does not exclude the applicability to or configuration to devices performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more conditions or values ​​may in practice be based on additional conditions or values ​​beyond those conditions.

[0075] In this invention, the term "exemplary" is used to mean "serving as an example, illustration, or description." Any embodiment described as "exemplary" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0076] The following description, in conjunction with the accompanying drawings, introduces a non-invasive underwater organism feeding monitoring system that integrates passive acoustics and image recognition, based on an embodiment of the present invention.

[0077] like Figure 1As shown, this embodiment of the invention provides a non-invasive underwater organism feeding monitoring system that integrates passive acoustics and image recognition. The system includes a multimodal data acquisition module, an enhanced single-modal feature extraction module, a multimodal feature interaction and fusion module, a feeding causal chain joint learning module, a feeding state uncertainty perception module, and a control and decision output module.

[0078] The multimodal data acquisition module is used to simultaneously acquire acoustic and image signals from underwater organisms feeding, and a clock synchronization component binds a unified timestamp to the acquired data. The acoustic and image signals correspond to each other. Underwater organisms include bass, whiteleg shrimp, and goldfish, among others.

[0079] The enhanced single-modal feature extraction module is used to extract features from acoustic signals and image signals respectively, and obtain acoustic feature maps and image feature maps accordingly.

[0080] In some examples, acoustic feature extraction networks are used to extract features from acoustic signals to obtain acoustic feature maps. The acoustic feature extraction network adopts a multi-layer residual structure, segments features along the channel dimension, and processes them in parallel with convolutional kernels of different sizes to achieve multi-scale receptive field capture; a coordinate attention mechanism is introduced to encode positional information along the height and width directions to enhance the expression of key frequency band features.

[0081] An image feature extraction network is used to extract features from image signals, obtaining image feature maps. This network includes a fish swarming effect perception module, which is based on the biological swarm sensing mechanism (individual fish perceive the movement of neighboring individuals through vision and lateral line systems, forming cooperative group behavior). A cascaded large-kernel attention structure is designed, decomposing a two-dimensional large convolutional kernel into cascaded one-dimensional convolutions in the vertical and horizontal directions to simulate the perception of information about the frontal and lateral neighbors by individual fish. The output of the fish swarming effect perception module includes not only the location and area of ​​the fish swarming region but also a swarm compactness index, used to quantify the degree of cooperation within the fish group.

[0082] The multimodal feature interaction and fusion module is used to perform multi-level fusion of acoustic feature maps and image feature maps. Multi-level fusion includes feature-level interaction and adaptive fusion during the feeding stage.

[0083] Feature-level interaction involves multiple information exchanges between acoustic and image features through bidirectional parallel interaction units. Image features are used as queries to enhance the global semantics of acoustic features, while acoustic features are used as queries to enhance the local details of image features. Residual connections preserve the original information, thus achieving mutual enhancement of local details and global semantics.

[0084] The adaptive fusion of feeding phases involves constructing a feeding phase perceiver that infers the current feeding phase (including searching, gathering, feeding, satiation, and cessation of feeding) in real time based on the rate of change in fish aggregation density and the rate of uneaten food consumption. The fusion weights of multi-scale features are dynamically adjusted according to the feeding phase. During the gathering and feeding phases, fish are spatially densely distributed, requiring enhancement of small-scale features to capture individual feeding details; during the searching and satiation phases, fish are spatially dispersed, requiring enhancement of large-scale features to capture group movement trends.

[0085] The feeding causal chain joint learning module is used to output multiple task results with causal relationships in parallel based on the fused features. The multiple task results include at least the biological feeding intensity level classification results, the estimated results of the number or distribution of uneaten food, and the fish school aggregation degree assessment results.

[0086] It is understandable that causal association refers to the causal temporal relationship of feeding behavior, namely, uneaten food → fish gathering → feeding behavior occurring → uneaten food decreasing → feeding intensity decreasing → fish dispersing. The feeding causal chain joint learning module adopts a hard parameter shared network structure, and a shared feature extraction network extracts general semantic representations.

[0087] The feeding state uncertainty perception module is used to calculate the feeding stage boundary uncertainty, modal conflict uncertainty, and time series anomaly uncertainty of the sample, and selects the sample with the highest uncertainty score for incremental training.

[0088] Understandably, stage boundary uncertainty refers to the increased uncertainty when the model's predicted results fall within the boundaries of feeding stages; such samples are crucial for distinguishing adjacent stages. Modal conflict uncertainty refers to the increased uncertainty when acoustic and image predictions are inconsistent; such samples are crucial for improving model robustness. Temporal anomaly uncertainty refers to the increased uncertainty when the model's predicted temporal changes in feeding intensity violate causal laws; such samples reflect data labeling errors or extreme scenarios. Automatically selecting high-uncertainty samples for expert labeling enables continuous model evolution.

[0089] The control and decision output module is used to generate precise feeding control commands and abnormal warning signals based on a segmented control strategy based on the feeding intensity index, according to the results of multiple tasks.

[0090] This invention provides a non-invasive underwater biological feeding monitoring system that integrates passive acoustics and image recognition. It introduces a swarm sensing mechanism to simulate fish schooling behavior, dynamically adjusts the fusion strategy through a feeding stage sensor, and constrains the prediction process through a feeding causal chain to ensure the prediction results conform to the biological laws of fish feeding behavior. Through feature-level bidirectional interaction and joint learning of the feeding causal chain, it can output not only the feeding intensity level but also simultaneously provide the number of uneaten baits, fish school aggregation degree, and aggregation compactness index, thus achieving a multi-dimensional representation of the feeding state. Experimental results show that the accuracy of this invention in feeding intensity classification tasks can reach over 93%, the mean absolute error of uneaten bait counting can be controlled within 23, and the mAP50 of fish school aggregation detection can reach over 95%, demonstrating high detection accuracy and recognition reliability. Furthermore, the acoustic modality can penetrate turbid water to acquire target information, while the image modality provides intuitive visual verification. The two are deeply fused at the intermediate layer through feature-level interaction, effectively improving the system's adaptability to complex environmental factors such as water turbidity, lighting changes, and background noise. Meanwhile, this invention incorporates a feeding state uncertainty perception module to automatically filter stage boundary samples, modal conflict samples, and temporal anomaly samples, thereby achieving continuous model optimization at a lower annotation cost. This enables the system to adapt to changes in the characteristics of different farming stages and different farmed species. In practical applications, the hardware cost of this invention can be controlled within 5,000 yuan. By generating precise feeding decisions, it can reduce feed waste by 15% to 20%, thereby reducing economic losses in farming. It has good economic viability and promotional value, making it suitable for large-scale deployment in small and medium-sized farming scenarios.

[0091] In some embodiments of the present invention, the multimodal data acquisition module includes a passive acoustic acquisition unit, an underwater image acquisition unit, and a clock synchronization component.

[0092] The passive acoustic acquisition unit includes four high-sensitivity piezoelectric hydrophones, which are placed at different depths of 30cm-80cm below the surface of the aquaculture pond to cover the main feeding area.

[0093] The underwater image acquisition unit includes an industrial-grade waterproof low-light camera with a resolution of 1280×720 and a frame rate of no less than 15fps, which, together with infrared illumination, enables all-weather imaging.

[0094] The clock synchronization component uses the internal clock of the embedded processor or an external high-precision clock source to synchronously acquire sound and image signals, with a timestamp accuracy of ≤1ms.

[0095] In some embodiments of the present invention, the enhanced single-modal feature extraction module includes an acoustic feature extraction network.

[0096] The acoustic feature extraction network includes a signal preprocessing unit, a multi-scale convolutional unit, and a coordinate attention unit.

[0097] The signal preprocessing unit is used to perform framing, windowing, pre-emphasis and noise suppression on the acoustic signal to obtain the preprocessed acoustic signal.

[0098] The multi-scale convolutional unit adopts a hierarchical residual structure to divide the preprocessed acoustic signal into multiple sub-feature groups along the channel dimension. For each sub-feature group, a convolutional kernel of different size is used to extract features to obtain a feature map.

[0099] The coordinate attention unit is used to perform average pooling along the height and width directions of the feature map, embedding positional information into the channel attention, generating attention weights that simultaneously contain channel and spatial positional information, and multiplying them with the input feature map to obtain the acoustic feature map.

[0100] In some embodiments of the present invention, the enhanced single-modal feature extraction module includes an image feature extraction network, and the image feature extraction network includes a fish swarming effect perception module, such as... Figure 2 As shown, the fish swarming effect perception module includes a cascaded large-kernel attention unit and a cluster compactness calculation unit.

[0101] The cascaded large kernel attention unit is used to decompose a two-dimensional large convolution kernel into cascaded vertical one-dimensional convolutions and horizontal one-dimensional convolutions to simulate the perception of information in front of and to the sides of a school of fish.

[0102] The aggregation compactness calculation unit is used to calculate a compactness index that quantifies the degree of coordination among fish groups based on the spatial distribution density and movement direction consistency of individual fish. When the compactness index exceeds a first threshold and continues to rise, it is determined that the feeding period has begun. When the compactness index decreases but there is still uneaten food remaining, it is determined that the satiated period has begun.

[0103] In some examples, vertical orientation is perceived. (That is, simulating the perception of the lateral neighborhood by an individual fish in a school) is as follows:

[0104]

[0105] Horizontal perception (That is, simulating the perception of the surrounding neighborhood by an individual fish in a school) is as follows:

[0106]

[0107] Extended perception (Simulating the perception of an individual fish in a school of fish towards an individual at a greater distance) is as follows:

[0108]

[0109]

[0110] Attention weight generation and cluster compactness calculation:

[0111]

[0112]

[0113]

[0114] The compactness index quantifies the spatial clustering of individual fish in a school; a higher value indicates a denser cluster. F represents the input feature map, d is the dilation rate, and k is the size of the large kernel convolution. This is for rounding down. This represents a depth-separable convolution in the vertical direction, with a kernel size of [value missing]. expansion rate ; This represents a depth-separable convolution in the horizontal direction; The hole depth in the vertical direction represents the separable convolution, and the hole rate is... In the formula for calculating the cluster compactness index, The total number of individual fish in the school. For the first The location coordinates of each individual The center of the fish school This refers to the area of ​​the region where the fish congregate.

[0115] In some embodiments of the present invention, the feeding phase adaptive fusion includes a feeding phase sensor and a weight dynamic adjustment unit.

[0116] The feeding phase sensor is used to infer the current feeding phase based on the currently detected rate of change in fish aggregation ΔA / Δt and the rate of uneaten food consumption ΔB / Δt, according to the following rules:

[0117] When ΔA / Δt > 0 and ΔB / Δt ≈ 0, it is determined to be the search period;

[0118] When ΔA / Δt > 0 and ΔB / Δt < 0, it is determined to be the aggregation period;

[0119] When ΔA / Δt ≈ 0 and ΔB / Δt < 0, it is determined to be the feeding period;

[0120] When ΔA / Δt < 0 and ΔB / Δt > 0, it is determined to be the satiated period;

[0121] When ΔA / Δt < 0 and ΔB / Δt ≈ 0, it is determined to be a period of fasting.

[0122] Table 1. Correspondence between feeding stages and biological behavioral characteristics

[0123] Search period , As the bait was thrown in, the fish began to move toward the feeding area. Gathering period , The fish quickly gathered and began to feed. Feeding period , The fish are clustered together, and the bait is consumed rapidly. period of fullness , The fish gradually dispersed, and the remaining food was no longer being consumed. Fasting period , The fish dispersed completely, having finished feeding.

[0124] The weighted dynamic adjustment unit is used to dynamically adjust the fusion weights of multi-scale features according to the feeding stage. It enhances small-scale features during the aggregation and feeding periods to capture individual feeding details, and enhances large-scale features during the searching and satiated periods to capture group movement trends.

[0125] In the process of multi-scale feature fusion, the multi-scale feature extraction network outputs feature maps of different scales. To simplify the description, the feature maps are divided into small-scale feature groups. and large-scale feature groups The small-scale feature map has high spatial resolution and is used to capture individual feeding details; the large-scale feature map has a large receptive field and is used to capture group movement trends. The fused feature map... We obtain the result through weighted summation: in, and The fusion weights are dynamically adjustable and satisfy the following conditions: .

[0126] In some examples, the weights are adaptively adjusted based on the feeding stage:

[0127] During the aggregation and feeding periods, fish schools are densely distributed spatially, requiring enhanced small-scale features to capture details of individual feeding activities. , ;

[0128] During the searching and feeding periods, fish schools are spatially dispersed, requiring enhanced large-scale features to capture school movement trends. , ;

[0129] During the fasting period, take , .

[0130] The weight values ​​mentioned above can be fine-tuned in practical applications based on the specific aquaculture species and environment.

[0131] In some embodiments of the present invention, the bidirectional parallel interaction unit includes a first interaction subunit, a second interaction subunit, and a feature fusion subunit.

[0132] The first interactive subunit is used to use the acoustic feature map as the query vector, the image feature map as the key vector and the value vector, and to calculate the acoustic features enhanced by the image features through an attention mechanism, so as to enhance the global semantic information of the acoustic features.

[0133] The second interactive subunit is used to use the image feature map as the query vector, the acoustic feature map as the key vector and the value vector, and to calculate the image features enhanced by acoustic features through an attention mechanism, so as to enhance the local detail information of the image features.

[0134] The feature fusion subunit is used to add the enhanced acoustic feature map with the enhanced image feature map by residual addition or channel splicing to obtain the fused feature map.

[0135] In some embodiments of the present invention, the feeding causal chain joint learning module adopts a hard parameter shared network structure, and the feeding causal chain joint learning module includes a shared feature extraction network and a task-specific branch network.

[0136] A shared feature extraction network is used to receive fused feature maps and extract general semantic representations from the fused feature maps.

[0137] Task-specific branch networks include classification branch networks, counting branch networks, and detection branch networks.

[0138] The classification branch network is used to output the classification results of feeding intensity levels, which include at least high active feeding, medium active feeding, low active feeding, and no feeding.

[0139] The counting branch network is used to output a residual bait density map or an estimate of the number of residual baits.

[0140] The detection branch network is used to output the location, area, and cluster compactness index of the fish gathering area.

[0141] The loss function of the feeding causal chain joint learning module during training includes the following components:

[0142] The target detection loss uses a size-adaptive penalty factor, specifically: Wherein, IoU is the intersection-union ratio between the predicted bounding box and the ground truth bounding box; and These are the width and height of the actual bounding box, respectively; , These are the lateral offsets of the predicted bounding box's left and right boundaries relative to the true bounding box, respectively. , These are the vertical offsets of the upper and lower boundaries of the predicted bounding box relative to the true bounding box, respectively.

[0143] Temporal consistency loss This constrains the smooth variation of feeding intensity predictions across adjacent time frames. In some examples, In the formula, T represents the total number of frames in a time segment. Let be the predicted value of the feeding intensity index for frame t.

[0144] Causal consistency loss is constrained by ensuring that the correlation coefficients among the rate of change in fish aggregation, the rate of uneaten food consumption, and the feeding intensity level meet the preset sign constraints. In some examples, ,in The Pearson correlation coefficient is used. This represents the feeding intensity level (which can be quantified as a continuous value or a discrete level between 0 and 1). The first constraint is that the increase in aggregation should be positively correlated with the consumption of uneaten food, and the second constraint is that the consumption of uneaten food should be positively correlated with the feeding intensity.

[0145] Modality consistency loss is used to minimize the difference between the feeding intensity predicted by acoustic features and the feeding intensity predicted by image features. In some examples, In the formula, This refers to the probability distribution (or feeding intensity index) of feeding intensity predicted based on acoustic features. This is the probability distribution (or feeding intensity index) of feeding intensity predicted based on image features.

[0146] Total loss for:

[0147] in, , , To balance the weighting coefficients of each loss term, in one example of this invention, the following values ​​are taken respectively: , , In actual training, these coefficients can be adjusted based on the performance of the validation set.

[0148] In some embodiments of the present invention, the feeding state uncertainty perception module includes a stage boundary uncertainty calculation unit, a modal conflict uncertainty calculation unit, a temporal anomaly uncertainty calculation unit, and a sample selection unit.

[0149] The stage boundary uncertainty calculation unit is used to calculate the probability that a sample is at the stage boundary of feeding based on the information entropy of the feeding intensity level probability distribution. In some examples, stage boundary uncertainty exists. In the formula, Let be the probability that the feeding intensity predicted by the model belongs to the i-th level. There are four feeding intensity levels: high-activity feeding, medium-activity feeding, low-activity feeding, and no feeding. The higher the information entropy, the more likely the sample is to be at the stage boundary.

[0150] The modal conflict uncertainty calculation unit is used to calculate the degree of bimodal information conflict based on the difference between acoustic prediction results and image prediction results. In some examples, modal conflict uncertainty... .

[0151] The temporal anomaly uncertainty calculation unit is used to calculate the degree of temporal anomaly based on the deviation of predicted temporal changes in feeding intensity from a pre-defined causal law. In some examples, temporal anomaly uncertainty... In the formula, The rate of change of the feeding intensity index over time (the difference between adjacent frames); To predict the direction of change, based on biological feeding patterns, the feeding intensity should increase after feeding. The intensity of food intake should decrease after a full meal. When the actual change is opposite to the expected direction, this item is positive, resulting in a penalty.

[0152] Total uncertainty The weighted sum of the three uncertainties mentioned above: in, , , As weighting coefficients, in one example of this invention, respectively take... , , It can be adjusted according to the actual application scenario.

[0153] The sample selection unit is used to select a preset number of samples with the highest scores based on the weighted sum of uncertainty scores, which are then labeled by experts and added to the training set for incremental training.

[0154] In some embodiments of the present invention, the control and decision output module calculates the feeding intensity index using the following formula:

[0155]

[0156] in, To determine the confidence level for fish aggregation. This represents the area where the fish congregate. To determine the confidence level for residual bait detection, The area of ​​the uneaten bait. This is the adjustment coefficient.

[0157] Based on the feeding intensity index I, a segmented control strategy based on biological feeding patterns is adopted:

[0158] When I ≥ 0.7, it is determined to be a period of high feeding demand, and the amount of feed should be increased;

[0159] When 0.3 < I < 0.7, it is determined to be a normal feeding period, and normal feeding should be maintained;

[0160] When I ≤ 0.3, it is determined to be a period of satiation or fasting, and feeding should be stopped.

[0161] In some embodiments of the present invention, the control and decision output module includes an anomaly warning unit and a data storage unit.

[0162] The abnormal warning unit is used to push alarm signals locally or remotely when the feeding intensity index drops more than a threshold within a preset time window, equipment malfunctions, or environmental noise suddenly increases.

[0163] The data storage unit is used to store the original sound signals, image keyframes, and historical statistical results of each task, and supports data export and report generation.

[0164] In some embodiments of the present invention, the application scenario of the non-invasive underwater biological feeding monitoring system that integrates passive acoustics and image recognition is used as an example of the industrialized recirculating aquaculture system for sea bass. The specific configuration is as follows:

[0165] Culture pond specifications: 4m in diameter, 1m in water depth, and a stocking density of approximately 100 fish / m³. Equipped with a circulating water treatment system (microfilter, protein skimmer, biofilter, UV sterilizer) and an aeration system.

[0166] Acoustic acquisition unit: 4-channel high-sensitivity piezoelectric hydrophones (sensitivity -165dB re 1V / μPa, frequency response range 20Hz-20kHz), deployed 50cm below the water surface, evenly distributed in the four corners of the aquaculture pond.

[0167] An industrial-grade waterproof low-light camera (1920×1080 resolution, 25fps frame rate, minimum illumination 0.01Lux) is installed on a pillar by the pool, about 1.2m above the water surface, and tilted downwards at a 30° angle towards the feeding area; it is equipped with two 850nm infrared lights that start and stop synchronously with the camera.

[0168] Embedded processing unit: responsible for data acquisition synchronization and preliminary preprocessing.

[0169] Edge computing devices: deploy model inference and decision-making algorithms, and connect to the automatic feeder via RS485 communication protocol.

[0170] In addition, regarding dataset construction and annotation of biological feeding patterns, two weeks of continuous aquaculture videos and synchronous acoustic signals were collected. Keyframes were extracted from the videos at 25fps, and the acoustic signals were sampled at 16kHz. After annotation by aquaculture experts, a dataset containing 5000 sets of image-acoustic paired samples was constructed and divided into training, validation, and test sets in a 7:2:1 ratio.

[0171] In addition to the usual targets, the annotations specifically included the following information related to the feeding patterns of the organisms:

[0172] Feeding intensity level (high / medium / low / non-feeding, four categories); location of uneaten food (point annotation, used to generate density map); fish gathering area (boundary box annotation); feeding phase annotation (searching period / gathering period / feeding period / satisfaction period / cessation of feeding period), used to train the feeding phase perceptron; temporal correlation annotation: the temporal relationship of the time of uneaten food appearance, the time of fish gathering, the time of peak feeding, and the time of uneaten food depletion in the same feeding event.

[0173] Table 2 Ablation Experiment

[0174] Model Configuration Accuracy (%) MAE (leftover bait) Fish swarm detection mAP50(%) Timing consistency violation rate (%) Baseline model (decision-level fusion) 86.0 35.2 90.3 18.5 +Perception of fish schooling effect 88.5 32.1 92.8 15.2 +Adaptive fusion during feeding phase 90.2 28.5 94.1 12.1 +Feature-level bidirectional interaction 91.8 25.3 95.3 9.8 +Joint learning of the causal chain of food intake 93.0 22.4 96.2 4.2 +Uncertainty perception and active learning 93.8 21.0 96.8 3.5

[0175] The ablation experiment results are shown in Table 2. The temporal consistency violation rate refers to the proportion of the model's predicted temporal change in feeding intensity that contradicts the biological feeding pattern (feeding intensity should increase when uneaten food decreases). This indicator further decreased from 4.2% to 3.5%, verifying the effectiveness of the causal chain constraint. The ablation experiment results show that the core technological innovations of this invention, such as the fish swarming effect perception module, adaptive fusion of feeding stages, feature-level bidirectional interaction, joint learning of feeding causal chains, and active learning for uncertainty perception, all bring quantifiable performance improvements. Field deployment verification shows that the system can reduce feed consumption by more than 16% in a factory aquaculture environment, and can operate stably under daytime, nighttime, and slightly turbid water conditions. The accuracy rate for determining the feeding stage reaches 89%, and the anomaly prediction rate is reduced to 2.1%.

[0176] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0177] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.

[0178] The foregoing has provided a detailed description of a non-invasive underwater biological feeding monitoring system that integrates passive acoustics and image recognition, as provided in the embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A passive acoustic and image recognition fusion based non-invasive underwater organism feeding monitoring system, characterized in that, The system includes: A multimodal data acquisition module is used to simultaneously acquire acoustic and image signals from underwater organism feeding scenarios; An enhanced single-modal feature extraction module is used to extract features from the acoustic signal and the image signal respectively, and obtain acoustic feature maps and image feature maps accordingly; The multimodal feature interaction and fusion module is used to perform multi-level fusion of the acoustic feature map and the image feature map. The multi-level fusion includes feature-level interaction and feeding stage adaptive fusion. The feature-level interaction is carried out by a bidirectional parallel interaction unit to exchange information multiple times between acoustic features and image features. The feeding stage adaptive fusion is to construct a feeding stage perceptron, infer the current feeding stage in real time based on the rate of change of fish aggregation degree and the rate of uneaten food consumption, and dynamically adjust the fusion weight of multi-scale features according to the feeding stage. The feeding causal chain joint learning module is used to output multiple task results with causal relationships in parallel based on the fused features. The multiple task results include at least the biological feeding intensity level classification result, the estimated result of the number or distribution of uneaten food, and the fish school aggregation degree assessment result. The feeding state uncertainty perception module is used to calculate the feeding stage boundary uncertainty, modal conflict uncertainty, and time series anomaly uncertainty of the sample, and selects the sample with the highest uncertainty score for incremental training. The control and decision output module is used to generate precise feeding control commands and abnormal warning signals based on a segmented control strategy based on the feeding intensity index, according to the results of the multiple tasks.

2. The passive acoustic and image recognition fusion based non-intrusive underwater organism feeding monitoring system as claimed in claim 1, wherein, The multimodal data acquisition module includes: The passive acoustic acquisition unit includes four high-sensitivity piezoelectric hydrophones, which are placed at a depth of 30cm-80cm below the surface of the aquaculture pond. The underwater image acquisition unit includes an industrial-grade waterproof low-light camera with a resolution of 1280×720 and a frame rate of no less than 15fps. The clock synchronization component uses the internal clock of the embedded processor or an external high-precision clock source to synchronously acquire sound and image signals, with a timestamp accuracy of ≤1ms.

3. The non-invasive underwater biological feeding monitoring system fusion of passive acoustics and image recognition according to claim 1, characterized in that, The enhanced single-modal feature extraction module includes an acoustic feature extraction network; The acoustic feature extraction network includes: The signal preprocessing unit is used to perform framing, windowing, pre-emphasis and noise suppression on the acoustic signal to obtain the preprocessed acoustic signal. The multi-scale convolutional unit adopts a hierarchical residual structure to divide the preprocessed acoustic signal into multiple sub-feature groups along the channel dimension. For each sub-feature group, a convolutional kernel of different size is used to extract features to obtain a feature map. The coordinate attention unit is used to perform average pooling along the height and width directions of the feature map, embedding position information into the channel attention, generating attention weights that simultaneously contain channel and spatial position information, and multiplying them with the input feature map to obtain the acoustic feature map.

4. The non-invasive underwater organism feeding monitoring system fusion of passive acoustics and image recognition according to claim 1, characterized in that, The enhanced single-modal feature extraction module includes an image feature extraction network, which in turn includes a fish swarming effect perception module. The fish swarming effect perception module includes: The cascaded large kernel attention unit is used to decompose the two-dimensional large convolution kernel into cascaded vertical one-dimensional convolution and horizontal one-dimensional convolution to simulate the perception of information in front of and to the sides of a school of fish. The aggregation compactness calculation unit is used to calculate a compactness index that quantifies the degree of fish group coordination based on the spatial distribution density and movement direction consistency of individual fish in the school; wherein, when the compactness index exceeds a first threshold and continues to rise, it is determined that the feeding period has begun; when the compactness index decreases but there is still uneaten food remaining, it is determined that the satiated period has begun.

5. The non-invasive underwater organism feeding monitoring system fusion of passive acoustics and image recognition according to claim 1, characterized in that, The adaptive fusion during the feeding phase includes: The feeding phase sensor is used to infer the current feeding phase based on the currently detected rate of change in fish aggregation ΔA / Δt and the rate of uneaten food consumption ΔB / Δt, according to the following rules: When ΔA / Δt > 0 and ΔB / Δt ≈ 0, it is determined to be the search period; When ΔA / Δt > 0 and ΔB / Δt < 0, it is determined to be the aggregation period; When ΔA / Δt ≈ 0 and ΔB / Δt < 0, it is determined to be the feeding period; When ΔA / Δt < 0 and ΔB / Δt > 0, it is determined to be the satiated period; When ΔA / Δt < 0 and ΔB / Δt ≈ 0, it is determined to be a period of fasting; The weighted dynamic adjustment unit is used to dynamically adjust the fusion weights of multi-scale features according to the feeding stage. It enhances small-scale features during the aggregation and feeding periods to capture individual feeding details, and enhances large-scale features during the searching and satiated periods to capture group movement trends.

6. The non-invasive underwater organism feeding monitoring system fusion of passive acoustics and image recognition according to claim 1, characterized in that, The bidirectional parallel interaction unit includes: The first interactive subunit is used to use the acoustic feature map as the query vector, the image feature map as the key vector and the value vector, and to calculate the acoustic features enhanced by the image features through an attention mechanism to enhance the global semantic information of the acoustic features. The second interactive subunit is used to use the image feature map as the query vector and the acoustic feature map as the key vector and value vector, and to calculate the image features enhanced by acoustic features through an attention mechanism to enhance the local detail information of the image features. The feature fusion subunit is used to add the enhanced acoustic feature map with the enhanced image feature map by residual addition or channel splicing to obtain the fused feature map.

7. The non-invasive underwater organism feeding monitoring system fusion of passive acoustics and image recognition according to claim 1, characterized in that, The feeding causal chain joint learning module adopts a hard parameter shared network structure, and the feeding causal chain joint learning module includes a shared feature extraction network and a task-specific branch network; The shared feature extraction network is used to receive the fused feature map and extract a general semantic representation from the fused feature map; The task-specific branch network includes: A classification branch network is used to output the classification results of feeding intensity levels, wherein the levels include at least high active feeding, moderate active feeding, low active feeding, and no feeding; A counting branch network is used to output a residual bait density map or an estimate of the number of residual baits. The detection branch network is used to output the location, area, and cluster compactness index of the fish gathering area; The loss function of the feeding causal chain joint learning module during training includes the following components: The target detection loss employs a size-adaptive penalty factor, where the denominator of the penalty factor is the width and height of the target bounding box. Temporal consistency loss constrains the smooth variation of feeding intensity prediction results in adjacent time frames; The causal consistency loss constrains the correlation coefficients among the fish school aggregation rate, uneaten food consumption rate, and feeding intensity level to meet the preset sign constraint. Modal consistency loss is used to minimize the difference between the feeding intensity predicted by acoustic features and the feeding intensity predicted by image features.

8. The non-invasive underwater organism feeding monitoring system fusion of passive acoustics and image recognition according to claim 1, characterized in that, The feeding state uncertainty perception module includes: The stage boundary uncertainty calculation unit is used to calculate the probability that a sample is at the boundary of the feeding stage based on the information entropy of the probability distribution of feeding intensity level; The modal conflict uncertainty calculation unit is used to calculate the degree of bimodal information conflict based on the difference between acoustic prediction results and image prediction results; The temporal anomaly uncertainty calculation unit is used to calculate the degree of temporal anomaly based on the degree of deviation between the predicted temporal changes in feeding intensity and the preset causal law; The sample selection unit is used to select a preset number of samples with the highest scores based on the weighted sum of uncertainty scores, which are then labeled by experts and added to the training set for incremental training.

9. The non-invasive underwater organism feeding monitoring system fusion of passive acoustics and image recognition according to claim 1, characterized in that, The control and decision output module calculates the feeding intensity index using the following formula: in, To determine the confidence level for fish aggregation. This represents the area where the fish congregate. To determine the confidence level for residual bait detection, The area of ​​the uneaten bait. This is the adjustment coefficient; Based on the feeding intensity index I, a segmented control strategy based on biological feeding patterns is adopted: When I ≥ 0.7, it is determined to be a period of high feeding demand, and the amount of feed should be increased; When 0.3 < I < 0.7, it is determined to be a normal feeding period, and normal feeding should be maintained; When I ≤ 0.3, it is determined to be a period of satiation or fasting, and feeding should be stopped.

10. The non-invasive underwater organism feeding monitoring system fusion of passive acoustics and image recognition according to claim 1, characterized in that, The control and decision output module includes: The abnormal warning unit pushes an alarm signal locally or remotely when the feeding intensity index drops more than a threshold within a preset time window, or when there is equipment failure or a sudden increase in environmental noise. The data storage unit is used to store the original audio signals, image keyframes, and historical statistical results of each task, and supports data export and report generation.