A Precision Feeding Decision-Making Method for Shrimp Based on Multimodal Sound and Image Fusion
By combining multimodal sound and shadow fusion technology with deep learning and reinforcement learning, the feeding status of shrimp can be accurately identified and the feeding strategy optimized. This solves the problems of adaptability and accuracy of feeding methods in existing technologies, and improves farming efficiency and economic benefits.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- OCEAN UNIV OF CHINA
- Filing Date
- 2026-03-06
- Publication Date
- 2026-05-26
AI Technical Summary
Current shrimp farming feeding methods mainly rely on human experience or fixed-time and fixed-quantity feeding, which are difficult to adapt to different farming stages and environmental changes, resulting in low feed utilization and reduced farming efficiency. Existing multimodal information fusion methods have failed to fully explore the inherent correlations and lack self-learning ability, making it difficult to achieve precise feeding.
By employing multimodal sound and image fusion technology, combined with deep learning and reinforcement learning, and through the fusion of information from acoustic perception and video perception, a cross-modal alignment network is constructed to achieve accurate identification of shrimp feeding status and optimization of feeding strategies. A self-learning feeding control mechanism is established, forming an end-to-end intelligent feeding system.
It improved the accuracy of shrimp feeding status identification and the system's adaptability, reduced data acquisition costs, realized closed-loop control and full-process automation of shrimp feeding, and improved aquaculture efficiency and economic benefits.
Smart Images

Figure CN121789147B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent control technology for aquaculture based on reinforcement learning, and particularly relates to a method for precise feeding decision-making for shrimp based on multimodal sound and shadow fusion. Background Technology
[0002] Shrimp farming is an important component of my country's aquaculture industry, and its yield and quality directly affect the economic benefits and sustainable development of the industry. Feeding management is a crucial aspect of shrimp farming, influencing farming efficiency, feed utilization, and the stability of the aquatic environment. Overfeeding leads to feed waste, increased ammonia nitrogen and nitrite concentrations, and the induction of diseases, while underfeeding results in slow shrimp growth, uneven size, and even competition for food, reducing overall farming efficiency. Therefore, achieving precise and intelligent control of shrimp feeding is a critical technical problem that urgently needs to be solved in the aquaculture sector.
[0003] Currently, shrimp farming still relies primarily on manual, experience-based feeding or timed, quantitative automated feeding. Manual feeding depends on farmers judging the amount of feed based on water color, shrimp activity, and experience; this is highly subjective, lacks repeatability, and is difficult to adapt to changes in different farming stages and environmental conditions. While timed, quantitative automated feeding reduces labor costs, it cannot dynamically adjust feeding strategies based on the shrimp's real-time feeding status, easily leading to low feed utilization and failing to meet the demands of refined farming. With the development of sensor and artificial intelligence technologies, some research and applications have begun to explore intelligent shrimp feeding decisions using sensor data. For example, existing technologies use environmental parameters such as water temperature, dissolved oxygen, and turbidity to establish empirical models or threshold rules to adjust feeding amounts. However, these methods only reflect the state of the farming environment and cannot directly depict the actual feeding behavior of shrimp, thus limiting their guiding role in feeding decisions. Other technologies attempt to use video monitoring to analyze shrimp activity, but due to factors such as water turbidity, changes in lighting, and obstruction, single visual information is prone to problems such as unstable recognition and high misjudgment rates, making long-term stable operation difficult.
[0004] In recent years, acoustic sensing technology has been increasingly applied to aquaculture. Studies have shown that shrimp generate characteristic acoustic signals during feeding, reflecting feeding intensity and stage. However, existing acoustic signal-based feeding methods often employ simple energy thresholds or traditional feature classification models, which have limited adaptability to complex aquaculture environments. Furthermore, acoustic information is easily affected by water flow, equipment noise, and other disturbances, resulting in insufficient stability when used alone. Therefore, relying solely on single-modal information for feeding decisions makes it difficult to balance accuracy and robustness. To improve the reliability of shrimp feeding status recognition, multimodal information fusion has become a research hotspot. Some existing technologies simply splice acoustic and video information or analyze them independently before making a comprehensive judgment. However, these methods fail to fully explore the inherent relationships between different modalities and lack a unified semantic expression, limiting the fusion effect. Simultaneously, most methods rely on manually labeled feeding status samples for supervised learning, resulting in high data acquisition costs and difficulty adapting to changes in aquaculture species, density, and environment, making engineering promotion challenging. Furthermore, at the feeding decision-making level, existing technologies mostly determine the feeding amount based on rules or static models, lacking feedback-based self-learning capabilities. This prevents the formation of a closed-loop control mechanism of "perception, analysis, decision-making, and feedback," making it difficult to achieve truly precise feeding. Especially given the significant temporal evolutionary characteristics of shrimp feeding behavior, the lack of modeling capabilities for dynamic changes in feeding states becomes a key bottleneck limiting the performance improvement of intelligent feeding systems. Summary of the Invention
[0005] To address the aforementioned problems, this invention discloses a method for precise shrimp feeding decisions based on multimodal acoustic-video fusion. The aim is to achieve precise identification of shrimp feeding status and intelligent optimization of feeding strategies by introducing multimodal information fusion technology of acoustic perception and video perception, combined with deep learning and reinforcement learning methods.
[0006] This invention provides a method for precise feeding decision-making for shrimp based on multimodal sound and image fusion, comprising the following processes:
[0007] S1, Acoustic signals and video image data generated during shrimp feeding are collected; the collected data are preprocessed and synchronized to obtain acoustic time-frequency diagrams and video time window segments;
[0008] S2, the acoustic time-frequency map and video time window segment are input into the acoustic coding sub-network and the visual coding sub-network respectively to extract acoustic feature vectors and visual feature vectors. Then, the obtained acoustic feature vectors and visual feature vectors are input into the cross-modal alignment module to generate a unified multimodal joint feature representation.
[0009] S3. The multimodal joint feature representation is input into the feeding state evolution modeling network. The joint features within a continuous time window are dynamically modeled through state recursion to obtain the feeding state sequence of shrimp. Then, the feeding state sequence is quantified to obtain the feeding intensity index sequence and the feeding saturation index sequence.
[0010] S4 uses the feeding intensity index sequence and the feeding saturation index sequence as the state input of the reinforcement learning environment, and uses the feeding amount increment and feeding time interval as action variables to establish a continuously adjustable feeding control space; it learns the strategy by constructing a reward function that takes into account both the feeding improvement effect and the feed residue penalty, updates the feeding strategy through real-time data feedback, and outputs the optimal feeding control command.
[0011] Preferably, the preprocessing includes performing denoising, frame segmentation, and time-frequency transformation on the acquired acoustic signal to obtain an acoustic time-frequency map. The video image data is subjected to denoising, target enhancement, and temporal window segmentation to transform it into structurally stable short video temporal window segments. .
[0012] Preferably, the acoustic coding subnetwork uses a two-dimensional convolutional neural network as its main structure, and maps the original time-frequency representation into a low-dimensional acoustic feature vector layer by layer;
[0013] First, the first convolutional feature extraction layer, Conv1, is used to process the input acoustic time-frequency map. The algorithm slides across the time-frequency plane to calculate the response of acoustic modes in local frequency bands and short time intervals. The convolutional output is batch normalized to stabilize the feature distribution, and then activated by the ReLU function to obtain the primary acoustic feature map. This indicates a localized change in the sound energy of feeding;
[0014] Secondly, the primary acoustic feature map Downsampling is performed to reduce the feature space resolution while preserving the main acoustic response regions; the resulting acoustic feature map is obtained after compression. Next, feature extraction is performed using residual convolutional modules. Each residual convolutional module includes a convolutional layer, a batch normalization layer, and an activation layer. The original information is preserved through residual connections to obtain high-level acoustic feature maps. ;
[0015] Finally, in the global average pooling layer, the high-level acoustic feature map is processed. Averaging is performed along the time-frequency dimension to compress the two-dimensional feature map into a one-dimensional feature vector, thus obtaining the acoustic feature vector. .
[0016] Preferably, the visual coding subnetwork adopts a three-dimensional convolutional structure to simultaneously model spatial and temporal information; firstly, it uses video time window segments... Using this as input, a 3D convolutional kernel is simultaneously slid across the temporal and spatial dimensions to extract features of shrimp population movement direction, velocity changes, and aggregation behavior, resulting in a primary spatiotemporal feature tensor. ;
[0017] Secondly, regarding the primary spatiotemporal feature tensor Downsampling is performed to reduce the temporal and spatial resolution, resulting in a compressed spatiotemporal feature tensor. Furthermore, high-level behavioral features are extracted using a 3D residual convolution module, with residual connections used to preserve the original spatiotemporal information, resulting in a high-level spatiotemporal feature tensor. ;
[0018] Finally, in the global spatiotemporal pooling layer, the high-level spatiotemporal feature tensors are processed. Pooling operations are performed in both time and space dimensions to compress the data into a one-dimensional vector, yielding the video feature vector. .
[0019] Preferably, the input layer of the cross-modal alignment module receives acoustic feature vectors. With visual feature vectors First, a feature space transformation is performed, and the two feature vectors are passed to the subsequent shared projection layer; then, the acoustic feature vectors are transformed in the shared projection layer. and visual feature vectors Each feature is mapped to an embedding space of the same dimension and the same metric to eliminate scale differences and statistical biases between different modal features. After shared projection, a contrastive constraint learning mechanism is introduced. By constructing a contrastive loss function, the distance between acoustic and visual embedding features in the same time window is minimized in the shared space, while the feature distance in different time windows is maximized, thus completing cross-modal semantic alignment.
[0020] Finally, the two types of embedded features are jointly constructed to form a unified multimodal feature representation.
[0021] Preferably, the overall structure of the feeding state evolution modeling network includes an input mapping layer, a state recursion layer, and a state output layer;
[0022] The input mapping layer takes input from the multimodal joint feature representation and performs a linear mapping on the joint feature representation to obtain a low-dimensional input vector suitable for state modeling. ;
[0023] Secondly, the continuous change of feeding state over time is modeled in the state recursion layer. The gated recursive unit (GRU) structure is adopted, and the gradual evolution of feeding state is realized through an explicit state update mechanism, that is, the current features are combined with the past states and updated into a new feeding state.
[0024] set up For time windows Given the hidden state vector, the state update process is as follows:
[0025]
[0026] in, Indicates the first Latent variables representing feeding status within a time window include information on shrimp feeding activity and behavioral trends. For gated loop unit network functions;
[0027] Through the Feature projection and statistical analysis are performed to map them to an interpretable feeding behavior indicator space; when features When exhibiting a combination of high acoustic energy and high amplitude of motion, it indicates a state of high feeding activity; when the characteristics... When the overall trend is downward and the magnitude of the change in state gradually decreases, it is determined that the shrimp have entered a state of reduced or saturated feeding.
[0028] Through the state recursion layer, features within consecutive time windows are gradually accumulated, enabling the network to recognize the enhancement, stabilization, or decay processes of the feeding state. Finally, the hidden state vector output by the state recursion layer is mapped to an interpretable representation of the feeding state evolution.
[0029]
[0030] in, This is a sequence of feeding states evolution. and These are represented as the output matrix and the output bias term, respectively.
[0031] Preferably, the feeding intensity index sequence is obtained by linearly mapping and normalizing the state vector. The output is mapped to the interval [0,1], where 0 represents no obvious feeding behavior and 1 represents extremely strong feeding behavior.
[0032] The feeding saturation index sequence is obtained by accumulating the changes in the feeding state evolution sequence. When the state change gradually decreases, the cumulative value tends to stabilize, indicating that the shrimp gradually enters the state of feeding saturation. The index value ranges from [0,1], where close to 0 indicates that feeding has just begun and close to 1 indicates that feeding has approached saturation.
[0033] Preferably, the state space construction and feeding action space definition of the reinforcement learning environment are as follows:
[0034] Construct a reinforcement learning environment state vector based on feeding state information, and in the... Within a time window, the environment state is defined as follows: ; Indicating the current level of feeding activity, This indicates whether food intake is close to saturation. For the most recent The state mean vector within each time window reflects the feeding trend;
[0035] Secondly, the action space of the reinforcement learning model is defined as a continuous action space, and the actions are used to directly control the feeding device; the action vector is defined as: ;in, This indicates an increase in the amount of feed given. This indicates the adjustment amount for the next feeding interval; the feeding execution amount is obtained from the following formula:
[0036]
[0037] This represents the amount of food to be fed during the current feeding cycle. This is the time interval after the current feeding cycle is updated.
[0038] Preferably, the reinforcement learning model adopts an Actor-Critic dual-network structure, which includes four sub-networks: the current policy network (Actor); the current value network (Critic); the target policy network (Target Actor); and the target value network (TargetCritic).
[0039] In the Within a time window, the system obtains the environmental state vector. And input it into the Actor network to obtain the action output. The action, once performed, creates a new environmental state. and rewards Interaction data is stored in the experience replay pool.
[0040] The input layer of the Actor network is the environment state vector. The environment state vector is subjected to nonlinear feature mapping using the ReLU activation function through a two-layer fully connected network in the feature mapping layer.
[0041] After the feature mapping layer, the network outputs a continuous action vector. Its dimension is 2, corresponding to the increase in feeding amount respectively. and feeding interval increment The Tanh activation function is used to limit the output to [-1, 1] to prevent abrupt changes in the feeding action; finally, the action is converted into a real feeding control quantity through a proportional mapping function.
[0042]
[0043] in, This is the action scaling factor; through the policy network, the system can adaptively output feeding adjustment instructions based on the feeding status. Indicates the increase in feeding amount; Indicates the increment of the feeding time interval;
[0044] In a value network structure, the Critic network simultaneously receives both state and action as input. The input is fused and mapped through a fully connected layer. ,in, and These are the weight matrices for the state and the action, respectively. To incorporate the fusion mapping bias term, the value network outputs the Q-value of the state-action pair. : , The value is used to assess the combined contribution of current feeding decisions to future feed intake improvement and feed utilization efficiency.
[0045] Preferably, the reward function is as follows:
[0046]
[0047] in, Indicates the degree of improvement in appetite; This indicates the level of feed residue detected through audio-visual feedback; , and These are weighting coefficients used to balance different objectives.
[0048] Compared with the prior art, the present invention has the following innovative features and beneficial effects:
[0049] (1) A cross-modal aligned acoustic-visual joint representation learning method is proposed. This invention constructs a cross-modal aligned network to map acoustic and visual features to a unified feature space, ensuring that acoustic and visual information within the same time window remains semantically consistent, thereby achieving joint representation of shrimp feeding behavior. This method effectively solves the problem of low fusion efficiency caused by semantic inconsistency in acoustic and visual information, and improves the accuracy of feeding state recognition.
[0050] (2) A self-supervised learning-based modeling mechanism for feeding state evolution is proposed. This invention introduces self-supervised learning into shrimp feeding state modeling. By predicting future short-term joint features to construct training objectives, it achieves automatic learning of the temporal evolution law of feeding behavior without the need for manual labeling of feeding state samples. This method significantly reduces data acquisition costs and enables the model to adapt to different farming densities, environmental conditions, and growth stages, exhibiting good generalization ability and engineering feasibility.
[0051] (3) Constructing a reinforcement learning feeding decision model based on feeding state feedback. This invention directly uses the modeling results of shrimp feeding state as the state of the reinforcement learning environment, establishes a self-learning mechanism for feeding strategies, and enables the feeding system to continuously optimize the feeding amount and feeding sequence based on real-time audio-visual feedback, thereby achieving closed-loop control of the feeding process.
[0052] (4) Forming an end-to-end closed-loop intelligent feeding decision system. This invention organically combines multimodal perception, feature learning, state modeling and decision optimization to form an end-to-end intelligent feeding system of perception, understanding, decision-making and feedback, realizing full-process automated control from data collection to feeding execution. The overall technical architecture has systematic innovation. Attached Figure Description
[0053] Figure 1 This is a flowchart illustrating the overall implementation process of the present invention.
[0054] Figure 2 This is a diagram of the cross-modal alignment network structure.
[0055] Figure 3 The diagram shows a reinforcement learning model for feeding decisions based on deep deterministic policy gradients.
[0056] Figure 4 This is a heatmap of the cross-modal alignment effect in an embodiment of the present invention.
[0057] Figure 5 This is a three-dimensional evolution diagram of the feeding state in an embodiment of the present invention.
[0058] Figure 6 This is a comparison diagram between the intelligent feeding strategy and the traditional fixed feeding strategy in an embodiment of the present invention. Detailed Implementation
[0059] This invention proposes a precise feeding decision-making method for shrimp based on multimodal sound and shadow fusion. The overall process is as follows: Figure 1 As shown, the method includes:
[0060] S1. First, underwater acoustic sensors and video acquisition devices are deployed in the shrimp farming water to synchronously collect acoustic signals and video image data generated during shrimp feeding, and the acquisition time of each modality is recorded using a unified timestamp. Then, the acquired acoustic signals are subjected to denoising, frame segmentation, and time-frequency transformation processing, while the video image data is subjected to denoising, target enhancement, and time window segmentation processing, transforming different modal data into structurally stable short segments. Finally, based on the timestamp information, the acoustic and video data are matched one-to-one and synchronized to construct a time-consistent acoustic-video joint sample set. .
[0061] S2, Sound-shadow joint feature learning based on cross-modal alignment network. First, the sound-shadow joint sample set... Input the acoustic coding subnetwork and the visual coding subnetwork respectively, and extract the acoustic feature vector. and visual feature vectors Subsequently, the obtained acoustic feature vectors With visual feature vectors The input cross-modal alignment module, through shared projection space and contrast constraint loss function, ensures that the sound and shadow features within the same time window correspond one-to-one in the feature space, ultimately generating a unified multimodal joint feature representation. This serves as input for subsequent modeling of the evolution of feeding states.
[0062] S3, Multimodal fusion discrimination based on self-supervised feeding state evolution modeling. First, the multimodal joint feature set... The feeding state evolution modeling network is input, and the joint features within a continuous time window are dynamically modeled through a state recursion method to obtain the state sequence of shrimp feeding states. Secondly, through the evolutionary sequence of feeding states... Quantification yielded two key indicators of feeding status: the feeding intensity index series. and food saturation index Finally, a self-supervised learning objective is constructed based on future feature prediction, outputting a quantified feeding intensity index and a feeding saturation index to characterize the current feeding status of shrimp.
[0063] S4, Precise feeding decisions based on feeding state feedback reinforcement learning. Using the aforementioned feeding intensity index sequence... and food saturation index The system takes the following steps: First, it uses the environmental state as input. Second, it uses the increase in feeding amount and the feeding time interval as action variables to establish a continuously adjustable feeding control space. Third, it constructs a reward function that balances the improvement in feeding and the penalty for feed residue to drive the agent to learn the strategy. Finally, it continuously updates the feeding strategy through real-time data feedback and outputs the optimal feeding control command. This enables adaptive and precise control of the shrimp feeding process.
[0064] The specific implementation process of the present invention will be further described below with reference to specific embodiments.
[0065] S1. Synchronous acquisition and preprocessing of multimodal sound and shadow data
[0066] S1-1 Multimodal Acoustic and Video Data Synchronous Acoustic ... The video acquisition device is positioned on the side or surface of the water body to collect continuous video sequences of shrimp feeding behavior. ,in, Indicates the first Frame video images, ,and This represents the total number of video frames.
[0067] S1-2 Acoustic Signal Processing: Since the original acoustic signal contains a large amount of interference signals unrelated to feeding, in order to improve the quality of the acoustic data, the original acoustic data is processed... Bandpass filtering is performed. Specifically, by setting the filter band, signal components within the frequency range of shrimp feeding sounds are preserved, while low-frequency water flow noise and high-frequency mechanical noise are suppressed, resulting in a denoised acoustic signal. .
[0068] Furthermore, to extract the short-time stable features of the acoustic signal, the denoised acoustic signal is... Framing is performed using fixed-length time windows. The continuous acoustic signal is divided into multiple short-time stable acoustic data frame sets. .
[0069] Finally, for each frame of acoustic data Performing time-frequency transformation processing, namely short-time Fourier transform, converts the one-dimensional time series signal into a two-dimensional time-frequency representation, thus obtaining the corresponding time-frequency representation. Through the above time-frequency transformation operation, the acoustic time-frequency diagram corresponding to each frame of acoustic data is obtained. .
[0070] S1-3 Video Signal Processing: For the original video sequence Image denoising and enhancement are performed frame by frame.
[0071] For the original video sequence The image denoising and enhancement operations are performed frame by frame. First, spatial filtering methods are used to remove noise such as suspended matter in the water and lighting jitter, resulting in denoised video frames. Secondly, contrast enhancement and background suppression processing are performed on the denoised video frames to highlight the shrimp population and feeding behavior areas, thus obtaining enhanced video frames. An enhanced video frame sequence with clear shrimp group outlines and prominent feeding behavior characteristics was obtained. To ensure temporal consistency between the video and acoustic data, the enhanced video sequence was... The video is divided into multiple video action segments based on a fixed time window length. .
[0072] S1-4 Acoustic Data Synchronization Alignment and Joint Sample Construction: Using a Uniform Timestamp to Perform Acoustic Time-Frequency Map With video time window segments Perform synchronous matching to construct a joint audio-visual sample set. This audio-visual joint sample set This serves as the joint feature learning for subsequent cross-modal alignment networks.
[0073] S2. Sound-shadow joint feature learning based on cross-modal alignment network
[0074] This step is used to obtain the time-synchronized audio-visual joint sample set from S1. In this invention, high-level semantic features characterizing shrimp feeding behavior are extracted, and acoustic and visual information are mapped to a unified feature space through a cross-modal alignment mechanism, thereby constructing a joint feature representation that is consistent with sound and visual information. This invention constructs a cross-modal alignment network (CMAN). The CMAN network consists of an acoustic coding sub-network, a visual coding sub-network, and a cross-modal alignment module, and its input is the joint sound and visual sample set output by S1. The output is a set of multimodal joint feature representations. .
[0075] Specifically, firstly, features are extracted from the acoustic time-frequency map and video segments using acoustic and visual coding sub-networks, respectively. Then, the obtained acoustic and visual feature vectors are input into a cross-modal alignment module. By using a shared projection space and a contrast constraint loss function, the acoustic and visual features within the same time window are mapped one-to-one in the feature space, ultimately generating a unified multimodal joint feature representation, which serves as input for subsequent feeding state evolution modeling.
[0076] The structure and feature extraction process of the S2-1 acoustic coding subnetwork: using acoustic time-frequency maps As input, the network uses a two-dimensional convolutional neural network as its main structure, mapping the original time-frequency representation into low-dimensional acoustic feature vectors layer by layer. Specifically:
[0077] First, the first convolutional feature extraction layer, Conv1, is used to process the input acoustic time-frequency map. The algorithm slides across the time-frequency plane to calculate the response of acoustic modes in local frequency bands and short time intervals. The convolutional output is batch normalized to stabilize the feature distribution, and then activated by the ReLU function to obtain the primary acoustic feature map. This feature map can represent local changes in the sound energy of feeding.
[0078] Secondly, the primary acoustic feature map Downsampling is performed to reduce the feature space resolution while preserving the main acoustic response regions, thus reducing subsequent computational load. The resulting acoustic feature map is obtained after compression. Next, feature extraction is performed using residual convolutional modules. Each residual module includes a convolutional layer, a batch normalization layer, and an activation layer. Residual connections are used to preserve the original information and prevent degradation during deep network training. This yields high-level acoustic feature maps. .
[0079] Finally, in the global average pooling layer, the high-level acoustic feature map is processed. Averaging is performed along the time-frequency dimension to compress the two-dimensional feature map into a one-dimensional feature vector, thus obtaining the acoustic feature vector. .
[0080] S2-2 Structure and Feature Extraction of the Visual Encoding Subnetwork: The visual encoding subnetwork is used to extract spatiotemporal behavioral features of shrimp populations from video time windows. The network adopts a three-dimensional convolutional structure to simultaneously model spatial and temporal information. Specifically:
[0081] First, using video time window segments Using this as input, a 3D convolutional kernel is simultaneously slid across the temporal and spatial dimensions to extract features of shrimp population movement direction, velocity changes, and aggregation behavior. This yields a primary spatiotemporal feature tensor. .
[0082] Secondly, regarding the primary spatiotemporal feature tensor Downsampling is performed to reduce the temporal and spatial resolution, resulting in a compressed spatiotemporal feature tensor. Furthermore, high-level behavioral features are extracted using a 3D residual convolution module, with residual connections used to preserve the original spatiotemporal information, resulting in a high-level spatiotemporal feature tensor. .
[0083] Finally, in the global spatiotemporal pooling layer, the high-level spatiotemporal feature tensors are processed. Pooling operations are performed in both time and space dimensions to compress the data into a one-dimensional vector, resulting in the video feature vector. .
[0084] S2-3 Cross-modal Alignment Module: Structure and Feature Alignment
[0085] In this invention, the cross-modal alignment module receives acoustic feature vectors output from the acoustic coding subnetwork. and visual feature vectors from the output of the visual encoding subnetwork. Furthermore, through a shared projection layer and a contrast constraint learning mechanism, a stable one-to-one correspondence is established between sound and shadow features within the same time window in the feature space. This stage is divided into: a cross-modal alignment module input layer, a shared projection layer, and a contrast constraint layer. Specifically:
[0086] Acoustic feature vectors are received at the input layer of the cross-modal alignment module. With visual feature vectors These two feature vectors differ in feature distribution, amplitude range, and semantic level, necessitating a feature space transformation beforehand. The two feature vectors are then passed to the subsequent shared projection layer.
[0087] Next, the acoustic feature vectors are projected onto the shared projection layer. and visual feature vectors Each feature is mapped to an embedding space of the same dimension and the same metric, eliminating scale differences and statistical biases between different modal features.
[0088] Specifically, by analyzing acoustic feature vectors and visual feature vectors By performing linear mapping operations, the acoustic features are converted into embedding feature vectors of consistent dimension and directly comparable characteristics, thus obtaining the acoustic embedding features. and visual embedding features These two features reside in a unified semantic space.
[0089] Finally, in the contrast constraint layer, to further enhance the semantic consistency between acoustic and visual features, this invention introduces a contrast constraint learning mechanism after shared projection. By constructing a contrast loss function, the distance between acoustic and visual embedding features within the same time window is minimized in the shared space, while the feature distance between different time windows is maximized. In a training batch, there are a total of Therefore, the contrastive loss function is defined as follows: (The number of time windows is specified in the original text.)
[0090]
[0091] in, Represents the cosine similarity function. The temperature coefficient is used to adjust the smoothness of the feature distribution. By minimizing the contrastive loss function described above, the model is forced to learn that the semantics of the sound-shadow embedding features within the same time window are consistent, and the semantics of features in different time windows are distinguishable, thereby achieving cross-modal semantic alignment.
[0092] Through the above alignment process, acoustic and visual embedding features that are strictly aligned in semantic space are finally obtained.
[0093] Multimodal joint feature construction: After completing cross-modal alignment, acoustic embedding features With visual embedding features They are mapped to the same semantic space. To fully utilize the complementary information of acoustic and visual modalities during shrimp feeding, this invention jointly constructs the two embedded features to form a unified multimodal feature representation.
[0094] Specifically, firstly, the aligned acoustic embedding features are... With visual embedding features By concatenating the features along the feature dimension, a joint feature vector is obtained. .
[0095] To avoid the redundant feature problem caused by simple concatenation, this invention further performs nonlinear mapping on the concatenated features through a fusion layer to obtain the final multimodal joint feature vector. .in To fuse the weight matrix, To merge the bias term, This is the activation function.
[0096] Furthermore, the joint features from all time windows are summarized to obtain a multimodal feature representation set: This joint feature set fully describes the acoustics of shrimp during continuous feeding. The sequence of visual behavior changes is used as input data for S3 to achieve temporal modeling and dynamic adjustment of feeding decisions. Figure 2 This is a diagram of the cross-modal alignment network structure.
[0097] S3. Multimodal fusion discrimination based on self-supervised feeding state evolution modeling
[0098] After completing the construction of multimodal joint features, S2 outputs a set of multimodal joint feature representations. The acoustic and visual joint behavioral characteristics of shrimp within a continuous time window are described. However, individual feature vectors can only reflect instantaneous feeding behavior and cannot characterize the dynamic evolution of the feeding process over time. Therefore, this invention constructs a self-supervised feeding state evolution modeling network (SSEMN) in S3, which dynamically models multimodal feature sequences through state recursion and outputs quantified feeding state indicators. This process specifically includes:
[0099] S3-1 Feeding State Evolution Modeling Network Structure (SSEMN) Design: The SSEMN network adopts a gated recursive state modeling structure to characterize the gradual change process of feeding state over time. Its overall structure includes an input mapping layer, a state recursion layer, and a state output layer, forming an end-to-end state evolution modeling network.
[0100] Specifically, firstly, in the input mapping layer, the input comes from the multimodal joint feature representation. Because the joint features have a high dimensionality and the numerical ranges of different dimensions vary significantly, an input mapping layer is used to linearly map the joint features, resulting in a low-dimensional input vector suitable for state modeling. .
[0101] Secondly, the continuous change of the feeding state over time is modeled in the state recursion layer. This invention adopts a gated recursive unit (GRU) structure and realizes the gradual evolution of the feeding state through an explicit state update mechanism, that is, combining the current features with the past states and updating them into a new feeding state.
[0102] set up For time windows Given the hidden state vector, the state update process is as follows:
[0103]
[0104] in, Indicates the first Latent variables representing feeding status within a time window include information on shrimp feeding activity and behavioral trends. Among them, This is a gated loop unit network function.
[0105] The hidden state vector As a state representation in a high-dimensional feature space, the energy release and group movement characteristics during shrimp feeding are indirectly reflected through the joint encoding of acoustic and visual features. Specifically:
[0106] Through the By performing feature projection and statistical analysis, it can be mapped to an interpretable space of feeding behavior indicators. When features When it exhibits the combined characteristics of high acoustic energy and high amplitude of motion, it indicates a state of high feeding activity.
[0107] When features When the overall trend is downward and the magnitude of the change in state gradually decreases, it indicates that the intensity of acoustic activity and group movement are gradually weakening, thus indicating that the shrimp have entered a state of reduced or saturated feeding.
[0108] Through the state recursion layer, features within consecutive time windows are gradually accumulated, enabling the network to identify the enhancement, stabilization, or decay of feeding states, rather than relying solely on the instantaneous features of a single time window.
[0109] Finally, the hidden state vector output by the state recursion layer is mapped to an interpretable representation of the feeding state evolution: .in, This is a sequence vector of feeding state evolution, which serves as the basic input for subsequent calculations of feeding state indices. and These are represented as the output matrix and the output bias term, respectively.
[0110] After completing the evolutionary modeling of feeding states and obtaining the state sequence Subsequently, the present invention further performs the following quantification processing on the state sequence to generate feeding state indicators that can be used for feeding decisions.
[0111] S3-2 Calculation method of feeding intensity index and saturation index: After completing the state evolution modeling, the feeding state evolution sequence is analyzed. Quantification yielded two key indicators of feeding status.
[0112] (1) Feeding intensity index sequence. The feeding intensity index is obtained by linearly mapping and normalizing the feeding state evolution sequence:
[0113]
[0114] This is a weight vector used to extract state components related to feeding intensity. The normalization function maps the output to the interval [0,1], where 0 represents no obvious feeding behavior and 1 represents extremely strong feeding behavior. This index reflects the feeding activity level of shrimp within the current time window; a higher value indicates stronger feeding behavior. This index indicates whether feeding should continue at this time.
[0115] (2) Feed saturation index sequence. The feed saturation index is obtained by accumulating the changes in the evolutionary sequence of feeding states:
[0116]
[0117] This index is used to accumulate the range of changes in feeding status. As the changes gradually decrease, the cumulative value tends to stabilize, indicating that the shrimp are gradually entering a state of feeding saturation. The index value ranges from [0,1], where close to 0 indicates that feeding has just begun, and close to 1 indicates that feeding is approaching saturation. This index is used to determine whether it is necessary to reduce or stop feeding.
[0118] Finally, the output of S3 includes a feeding state evolution sequence. Feeding intensity index sequence and food saturation index .
[0119] S3-3 Self-supervised learning mechanism based on future feature prediction: In order to avoid the high cost of manually labeling feeding states, this invention introduces a self-supervised learning mechanism to construct training targets by predicting future features.
[0120] First, we construct a task for predicting future features. This is done within the current time window. Hidden state The prediction network generates predictions of future joint features: .in, For the predicted multimodal joint features of the next time window, and These are the weights and biases of the prediction layer, respectively.
[0121] Secondly, the predicted features This reflects the model's comprehensive judgment on possible changes in feeding sound intensity, group movement amplitude, and behavioral persistence in the next time window, based on the current feeding status.
[0122] Finally, during the training phase, the predicted features will be... Joint features with the actual next time window By comparison, a self-supervised loss function is constructed:
[0123]
[0124] in, The total number of time windows for the feature sequence is used, and this loss is calculated using the L2 norm. By minimizing this loss, the network is forced to learn the hidden states. It can contain the key state information that determines future behavior changes to the greatest extent, thus avoiding the hidden state from encoding only transient noise features.
[0125] S4. Precise feeding decisions based on reinforcement learning of feeding status feedback
[0126] S4-1 Reinforcement Learning Environment State Space Construction and Feeding Action Space Definition:
[0127] First, a reinforcement learning environment state vector is constructed based on the feeding state information output by S3. In the... Within a time window, the environment state is defined as follows: . Indicating the current level of feeding activity, This indicates whether food intake is close to saturation. For the most recent The mean vector of states within each time window reflects the feeding trend.
[0128] Secondly, the action space of the reinforcement learning model is defined as a continuous action space, where actions are used to directly control the feeding device. The action vector is defined as follows: .in, This indicates an increase in the amount of feed given. This indicates the adjustment amount for the next feeding interval. Through this action definition, the model can achieve fine-grained control over the feeding amount.
[0129] The feeding execution amount is obtained by the following formula, which enables continuous adjustment of the feeding strategy.
[0130]
[0131] This represents the amount of food to be fed during the current feeding cycle. This is the time interval after the current feeding cycle is updated.
[0132] S4-2 Reinforcement Learning Model Structure and Policy Network Design: To achieve continuous regulation and long-term optimal control of shrimp feeding behavior, this invention constructs a feeding decision reinforcement learning model based on deep deterministic policy gradients. This model employs an Actor-Critic dual-network structure. The policy network (Actor) outputs feeding actions, and the value network (Critic) evaluates the long-term benefits of these actions, thereby achieving stable and efficient feeding policy learning. Specifically:
[0133] The reinforcement learning model consists of four sub-networks: 1) Current policy network (Actor); 2) Current value network (Critic); 3) Target policy network (Target Actor); and 4) Target value network (Target Critic).
[0134] The target network is used to improve training stability by slowly tracking the parameters of the main network through soft updates.
[0135] In the Within a time window, the system obtains the environmental state vector through S3. And input it into the Actor network to obtain the action output. This action, once executed, creates a new environmental state. and rewards The aforementioned interaction data is stored in the experience replay pool for subsequent network training, thereby breaking temporal correlation and improving the model's generalization ability.
[0136] First, the input layer of the Actor network is the environment state vector. The environmental state vector undergoes nonlinear feature mapping using the ReLU activation function in the feature mapping layer through a two-layer fully connected network. This enables the network to learn the complex mapping relationship between feeding state and feeding behavior.
[0137] After the feature mapping layer, the network outputs a continuous action vector. Its dimension is 2, corresponding to the increase in feeding amount respectively. and feeding interval increment The Tanh activation function is used to limit the output to [-1, 1] to prevent abrupt changes in feeding actions and improve system stability. Finally, the actions are converted into actual feeding control quantities through a proportional mapping function.
[0138]
[0139] in, This represents the action scaling factor. Through the policy network, the system can adaptively output feeding adjustment commands based on the feeding status. Indicates the increase in feeding amount; This indicates the increment of the feeding time interval.
[0140] Secondly, a benefit evaluation mechanism is implemented within the value network structure. The Critic network simultaneously receives both state and action as input. The input vectors are fused and mapped through a fully connected layer. .in, and These are the weight matrices for the state and the action, respectively. This is the fusion mapping bias term. The Q-value of the value network output state-action pair. : .Should The value is used to evaluate the combined contribution of current feeding decisions to future improvements in feed intake and feed utilization efficiency. The value network guides the optimization of the strategy network, causing feeding behavior to evolve towards long-term optimality.
[0141] S4-3 Reward Function Design and Feeding Effect Constraint Mechanism: To guide the agent to achieve a high-cost-performance feeding objective, the following reward function is constructed:
[0142]
[0143] in, Indicates the degree of improvement in appetite; This indicates the level of feed residue detected through audio-visual feedback; , and The weighting coefficients are used to balance different objectives. This reward function ensures that the feeding strategy promotes feeding while avoiding overfeeding and water pollution.
[0144] S4-4 Real-time Acoustic and Visual Feedback Strategy Update and Feeding Command Output: After completing the feeding execution, the system re-acquires acoustic signals and video image data through S1–S3, and recalculates the new feeding state evolution sequence. Intensity index and food saturation index Thus, the environmental state at the next moment can be obtained. The reinforcement learning model optimizes the feeding strategy in real time based on a closed-loop mechanism of execution, feedback, and update.
[0145] In the Within each decision cycle, the value network is updated according to the following loss function:
[0146]
[0147] in, This is a discount factor used to balance the importance of current returns with future returns; The target policy network is used to generate reference actions for the next state; This is a target value network used to calculate a stable target Q value; and These are the parameters of the current value network and the target value network, respectively. By minimizing the above loss function, the value network continuously approximates the true long-term feeding return assessment, thus providing a reliable basis for strategy updates.
[0148] Furthermore, after the value network completes its update, the policy network updates its parameters based on the gradient information output by the value network, with the optimization objective being to maximize the long-term cumulative return. Through this update mechanism, the strategy network gradually learns the optimal feeding action to take under different feeding states, enabling the feeding strategy to adapt to the dynamic changes in shrimp feeding behavior.
[0149] After the policy network converges, the system will output the action vector. Mapped to specific feeding control commands: . This refers to the actual amount of feed given during this feeding session. This is the time interval for the next feeding. The feeding command is sent to the automatic feeding device via a communication interface, enabling the feeding device to dynamically adjust the feeding rhythm according to the current feeding status of the shrimp, achieving precise, energy-saving, and environmentally friendly feeding control. Figure 3 This is a reinforcement learning model for feeding decisions based on deep deterministic policy gradients.
[0150] Simulation results explanation:
[0151] Figure 4 A heatmap visually illustrates the cross-modal alignment effect. Before alignment, the similarity between acoustic and visual features exhibits an irregular distribution and cannot form a clear diagonal structure, indicating that the features of the two modalities are difficult to correlate directly. After alignment, however, a clear high-value band appears near the main diagonal of the similarity matrix, indicating that the acoustic and visual features within the same time window are effectively aligned, and the cross-modal semantic consistency is significantly improved, thus verifying the effectiveness of the cross-modal joint feature learning model.
[0152] Figure 5 The three-dimensional feeding state evolution diagram visually illustrates the complete evolutionary path of shrimp feeding from initiation, enhancement to saturation: in the early stage, feeding intensity rises rapidly while saturation is low; in the middle stage, intensity remains high while saturation continues to accumulate; in the later stage, intensity decreases while saturation approaches 1, indicating that feeding tends to terminate. This diagram visually verifies the ability of the constructed feeding state evolution model to distinguish different feeding stages.
[0153] like Figure 6 As shown in the results, the reinforcement learning-based intelligent feeding method exhibits significant advantages in several aspects compared to the traditional fixed feeding strategy. Regarding feed quantity control, the intelligent feeding strategy can gradually adjust the feed quantity according to the feeding status, avoiding overfeeding caused by fixed feeding. In terms of feeding response, the reinforcement learning strategy maintains the shrimp's feeding intensity at a consistently high level, indicating that the feeding decisions are more in line with the shrimp's actual feeding needs. Regarding feed utilization, the amount of uneaten feed is significantly reduced under the intelligent feeding strategy, demonstrating that the proposed method can effectively reduce feed waste and improve the aquatic environment.
[0154] The comparative results show that the feeding decision-making method based on multimodal sound and shadow fusion and reinforcement learning proposed in this invention is superior to the traditional feeding method in terms of feeding accuracy and breeding efficiency, and has significant practical application value.
[0155] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0156] While the specific embodiments of the present invention have been described above, they are not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for precise feeding decision-making for shrimp based on multimodal sound and image fusion, characterized in that, Includes the following processes: S1, Acoustic signals and video image data generated during shrimp feeding are collected; the collected data are preprocessed and synchronized to obtain acoustic time-frequency diagrams and video time window segments; S2, the acoustic time-frequency map and video time window segment are input into the acoustic coding sub-network and the visual coding sub-network respectively to extract acoustic feature vectors and visual feature vectors. Then, the obtained acoustic feature vectors and visual feature vectors are input into the cross-modal alignment module to generate a unified multimodal joint feature representation. The input layer of the cross-modal alignment module receives acoustic feature vectors. With visual feature vectors First, a feature space transformation is performed, and the two feature vectors are passed to the subsequent shared projection layer; then, the acoustic feature vectors are transformed in the shared projection layer. and visual feature vectors Each feature is mapped to an embedding space of the same dimension and the same metric to eliminate scale differences and statistical biases between different modal features. After shared projection, a contrastive constraint learning mechanism is introduced. By constructing a contrastive loss function, the distance between acoustic and visual embedding features in the same time window is minimized in the shared space, while the feature distance in different time windows is maximized, thus completing cross-modal semantic alignment. Finally, the two types of embedded features are jointly constructed to form a unified multimodal feature representation; S3. The multimodal joint feature representation is input into the feeding state evolution modeling network. The joint features within a continuous time window are dynamically modeled through state recursion to obtain the shrimp feeding state sequence. Then, the feeding state sequence is quantified to obtain the feeding intensity index sequence and the feeding saturation index sequence. The overall structure of the feeding state evolution modeling network includes an input mapping layer, a state recursion layer, and a state output layer. The input mapping layer takes input from the multimodal joint feature representation and performs a linear mapping on the joint feature representation to obtain a low-dimensional input vector suitable for state modeling. ; Secondly, the continuous change of feeding state over time is modeled in the state recursion layer. The gated recursive unit (GRU) structure is adopted, and the gradual evolution of feeding state is realized through an explicit state update mechanism, that is, the current features are combined with the past states and updated into a new feeding state. set up For time windows Given the hidden state vector, the state update process is as follows: in, Indicates the first Latent variables representing feeding status within a time window include information on shrimp feeding activity and behavioral trends. For gated loop unit network functions; Through the Feature projection and statistical analysis are performed to map them to an interpretable feeding behavior indicator space; when features When exhibiting a combination of high acoustic energy and high amplitude of motion, it indicates a state of high feeding activity; when the characteristics... When the overall trend is downward and the magnitude of the change in state gradually decreases, it is determined that the shrimp have entered a state of reduced or saturated feeding. Through the state recursion layer, features within consecutive time windows are gradually accumulated, enabling the network to recognize the enhancement, stabilization, or decay processes of the feeding state. Finally, the hidden state vector output by the state recursion layer is mapped to an interpretable representation of the feeding state evolution. in, This is a sequence of feeding states evolution. and These are represented as the output matrix and the output bias term, respectively. The feeding intensity index sequence is obtained by linearly mapping and normalizing the feeding state evolution sequence. The output is mapped to the interval [0,1], where 0 represents no obvious feeding behavior and 1 represents extremely strong feeding behavior. The feeding saturation index sequence is obtained by accumulating the changes in the feeding state evolution sequence. When the state change gradually decreases, the cumulative value tends to stabilize, indicating that the shrimp gradually enters the feeding saturation state. The index value ranges from [0,1], where close to 0 indicates that feeding has just begun, and close to 1 indicates that feeding has approached saturation. S4 uses the feeding intensity index sequence and the feeding saturation index sequence as the state input of the reinforcement learning environment, and uses the feeding amount increment and feeding time interval as action variables to establish a continuously adjustable feeding control space; it learns the strategy by constructing a reward function that takes into account both the feeding improvement effect and the feed residue penalty, updates the feeding strategy through real-time data feedback, and outputs the optimal feeding control command.
2. The shrimp precision feeding decision-making method based on multimodal sound and image fusion as described in claim 1, characterized in that: The preprocessing includes performing noise reduction, frame segmentation, and time-frequency transformation on the acquired acoustic signals to obtain an acoustic time-frequency map. The video image data is subjected to denoising, target enhancement, and temporal window segmentation to transform it into structurally stable short video temporal window segments. .
3. The shrimp precision feeding decision-making method based on multimodal sound and image fusion as described in claim 1, characterized in that: The acoustic coding subnetwork uses a two-dimensional convolutional neural network as its main structure, and maps the original time-frequency representation into a low-dimensional acoustic feature vector layer by layer. First, the first convolutional feature extraction layer, Conv1, is used to process the input acoustic time-frequency map. The algorithm slides across the time-frequency plane to calculate the response of acoustic modes in local frequency bands and short time intervals. The convolutional output is batch normalized to stabilize the feature distribution, and then activated by the ReLU function to obtain the primary acoustic feature map. This indicates a localized change in the sound energy of feeding; Secondly, the primary acoustic feature map Downsampling is performed to reduce the feature space resolution while preserving the main acoustic response regions; the resulting acoustic feature map is obtained after compression. Next, feature extraction is performed using residual convolutional modules. Each residual convolutional module includes a convolutional layer, a batch normalization layer, and an activation layer. The original information is preserved through residual connections to obtain high-level acoustic feature maps. ; Finally, in the global average pooling layer, the high-level acoustic feature map is processed. Averaging is performed along the time-frequency dimension to compress the two-dimensional feature map into a one-dimensional feature vector, thus obtaining the acoustic feature vector. .
4. The shrimp precision feeding decision-making method based on multimodal sound and image fusion as described in claim 1, characterized in that: The visual coding subnetwork employs a three-dimensional convolutional structure to simultaneously model spatial and temporal information; firstly, it uses video time windows... Using this as input, a 3D convolutional kernel is simultaneously slid across the temporal and spatial dimensions to extract features of shrimp population movement direction, velocity changes, and aggregation behavior, resulting in a primary spatiotemporal feature tensor. ; Secondly, regarding the primary spatiotemporal feature tensor Downsampling is performed to reduce the temporal and spatial resolution, resulting in a compressed spatiotemporal feature tensor. ; High-level behavioral features are extracted using a 3D residual convolution module, and residual connections are used to preserve the original spatiotemporal information to obtain the high-level spatiotemporal feature tensor. ; Finally, in the global spatiotemporal pooling layer, the high-level spatiotemporal feature tensors are processed. Pooling operations are performed in both time and space dimensions to compress the data into a one-dimensional vector, yielding the video feature vector. .
5. The shrimp precision feeding decision-making method based on multimodal sound and image fusion as described in claim 1, characterized in that: The state space and feeding action space of the reinforcement learning environment are defined as follows: Construct a reinforcement learning environment state vector based on feeding state information, and in the... Within a time window, the environment state is defined as follows: ; Indicating the current level of feeding activity, This indicates whether food intake is close to saturation. For the most recent The state mean vector within each time window reflects the feeding trend; Secondly, the action space of the reinforcement learning model is defined as a continuous action space, and the actions are used to directly control the feeding device; the action vector is defined as: ;in, This indicates an increase in the amount of feed given. This indicates the adjustment amount for the next feeding interval; the feeding execution amount is obtained from the following formula: This represents the amount of food to be fed during the current feeding cycle. This is the time interval after the current feeding cycle is updated.
6. The shrimp precision feeding decision-making method based on multimodal sound and image fusion as described in claim 5, characterized in that: The reinforcement learning model employs an Actor-Critic dual-network structure, comprising four sub-networks: the current policy network (Actor); the current value network (Critic); the target policy network (Target Actor); and the target value network (Target Critic). In the Within a time window, the system obtains the environmental state vector. And input it into the Actor network to obtain the action output. The action, once performed, creates a new environmental state. and rewards ; Interaction data is stored in the experience replay pool; The input layer of the Actor network is the environment state vector. ; The environment state vector is non-linearly mapped using the ReLU activation function in the feature mapping layer through a two-layer fully connected network. After the feature mapping layer, the network outputs a continuous action vector. Its dimension is 2, corresponding to the increase in feeding amount respectively. and feeding interval increment The Tanh activation function is used to limit the output to [-1, 1] to prevent abrupt changes in the feeding action; finally, the action is converted into a real feeding control quantity through a proportional mapping function. in, This is the action scaling factor; through the policy network, the system can adaptively output feeding adjustment instructions based on the feeding status. Indicates the increase in feeding amount; Indicates the increment of the feeding time interval; In a value network structure, the Critic network simultaneously receives both state and action as input. The input is fused and mapped through a fully connected layer. ,in, and These are the weight matrices for the state and the action, respectively. To incorporate the fusion mapping bias term, the value network outputs the Q-value of the state-action pair. : , The value is used to assess the combined contribution of current feeding decisions to future feed intake improvement and feed utilization efficiency.
7. The shrimp precision feeding decision-making method based on multimodal sound and image fusion as described in claim 6, characterized in that: The reward function is as follows: in, Indicates the degree of improvement in appetite; This indicates the level of feed residue detected through audio-visual feedback; , and These are weighting coefficients used to balance different objectives.