Breeding animal behavior tracking system and method based on artificial intelligence

By using a multimodal data acquisition and dynamic contour anchoring Mamba network processing module, combined with a behavioral dynamic map construction and deviation assessment module, the problems of insufficient multimodal data processing and single behavioral assessment dimensions in existing technologies are solved. This enables efficient and accurate tracking and in-depth analysis of animal behavior, providing reliable data support for smart farming.

CN120995126APending Publication Date: 2025-11-21BAODING FENGRAO AGRI TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511417562.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies lack efficient time-series modeling capabilities for multimodal data processing in animal behavior tracking, making it difficult to achieve effective correlation between cross-frame data. The behavioral assessment dimensions are limited, and there is a lack of quantitative assessment of dynamic behavioral deviations, which fails to meet the efficiency and accuracy requirements of smart farming for animal behavior tracking.

Method used

An AI-based animal behavior tracking system is adopted. The system integrates video, audio, and infrared thermal imaging data through a multimodal data acquisition module. Combined with a dynamic contour anchoring Mamba network processing module, it achieves accurate extraction of animal contours and cross-frame temporal modeling. The behavior dynamic map construction module and the deviation evaluation module work together to construct a map based on the coordinates of key nodes and calculate the feature deviation between the real-time map and the standard map. The smart farming multimodal spatiotemporal engine module integrates multimodal spatiotemporal information and temporal features to construct spatiotemporally correlated behavior data.

Benefits of technology

It improves the accuracy of behavioral feature extraction, enables precise reflection of the continuous changes in animal behavior, makes up for the shortcomings of the assessment system that is single-dimensional and lacks dynamic deviation quantitative assessment, and constructs comprehensive spatiotemporal correlated behavioral data. This meets the needs of smart farming for high efficiency, accuracy and in-depth analysis of animal behavior tracking, and provides reliable data support for farming management decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995126A_ABST
    Figure CN120995126A_ABST
Patent Text Reader

Abstract

The invention discloses a bred animal behavior tracking system and method based on artificial intelligence, and the system comprises six modules which are sequentially connected to achieve data transmission and processing. Video streams, audio streams and infrared thermal imaging data of animals in a breeding scene are obtained through a multi-modal data acquisition device, animal contours are extracted through a dynamic contour anchoring Mamba network processing module, time sequence modeling is carried out, and a behavior dynamic map construction module constructs a behavior dynamic map. The behavior dynamic map deviation evaluation module calculates a characteristic deviation value between real time and a standard map, the intelligent breeding multi-mode space-time engine module fuses space-time information and time sequence characteristics to generate space-time correlation behavior data, and finally the behavior tracking result output module outputs animal real-time behavior categories and movement tracks. According to the invention, the behavior feature extraction precision is improved to reflect the continuous change process of animal behaviors; and the requirements of intelligent breeding on high efficiency, accuracy and deep analysis of animal behavior tracking are met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of farmed animal behavior tracking, and in particular to a farmed animal behavior tracking system and method based on artificial intelligence. BACKGROUND

[0002] In the current development process of intelligent farming industry, the contradiction between the continuous expansion of farming scale and the demand for fine management is increasingly prominent. The traditional animal behavior monitoring method relying on manual observation is difficult to realize real-time and accurate tracking of individual animal behavior in large-scale farming scenarios. With the penetration of artificial intelligence technology in the field of agriculture, it has become a trend to use multi-modal data collection and intelligent algorithms to analyze animal behavior. In the farming scenario, animal behavior characteristics need to be captured through video, audio, infrared and other multi-dimensional data to build a time-series and spatially related behavior analysis system to realize dynamic monitoring of animal feeding, activity, health status and other behaviors and provide data support for farming management decisions. However, the existing technology still has deficiencies in multi-modal data fusion processing, accurate behavior feature modeling and spatio-temporal correlation analysis, which makes it difficult to meet the efficiency and accuracy requirements of animal behavior tracking in intelligent farming.

[0003] The existing technology has two significant shortcomings in the tracking of farmed animal behavior. On the one hand, the existing system lacks efficient time-series modeling capability in processing multi-modal data, often failing to fully combine the dynamic change characteristics of animal behavior. When extracting animal contours and key node information, it is difficult to effectively associate cross-frame data, resulting in insufficient behavior feature extraction accuracy and inability to accurately reflect the continuous change process of animal behavior. On the other hand, the behavior evaluation system of the existing technology is single-dimensional, only capable of simple classification of animal behavior, lacking quantitative evaluation of behavior dynamic deviation, and failing to fuse spatio-temporal information to build a multi-dimensional evaluation model, making it difficult to fully judge the deviation degree of animal behavior from the standard behavior and meet the needs of deep analysis and accurate management of animal behavior in the intelligent farming scenario. SUMMARY

[0004] In order to overcome the shortcomings and deficiencies of the existing technology, the present application provides a farmed animal behavior tracking system and method based on artificial intelligence.

[0005] The technical scheme adopted by the present application is a breeding animal behavior tracking system based on artificial intelligence, comprising: a multi-modal data acquisition module for acquiring video stream, audio stream and infrared thermal imaging data of animals in a breeding scene; a dynamic contour anchoring Mamba network processing module receiving output data of the multi-modal data acquisition module, performing contour extraction on animal targets through a dynamic contour anchoring unit, and using a linear attention mechanism of the Mamba network to perform time series modeling on the extracted contour sequence; a behavior dynamic atlas construction module receiving time series contour data output by the dynamic contour anchoring Mamba network processing module, and constructing a behavior dynamic atlas based on animal body calibration node coordinates; a behavior dynamic atlas deviation evaluation module receiving atlas data output by the behavior dynamic atlas construction module, and calculating feature deviation values of real-time atlas and preset standard behavior atlas; a smart breeding multi-modal spatio-temporal engine module receiving deviation values output by the behavior dynamic atlas deviation evaluation module, fusing spatio-temporal information of the multi-modal data acquisition module and time series features of the dynamic contour anchoring Mamba network processing module, and generating spatio-temporally correlated behavior data; and a behavior tracking result output module receiving spatio-temporally correlated behavior data output by the smart breeding multi-modal spatio-temporal engine module, and outputting real-time behavior categories and motion trajectory data of animal individuals, wherein the multi-modal data acquisition module and the dynamic contour anchoring Mamba network processing module are connected through a data transmission interface, the dynamic contour anchoring Mamba network processing module and the behavior dynamic atlas construction module are connected through a time series data bus, the behavior dynamic atlas construction module and the behavior dynamic atlas deviation evaluation module are connected through an atlas feature channel, the behavior dynamic atlas deviation evaluation module and the smart breeding multi-modal spatio-temporal engine module are connected through a deviation value transmission link, and the smart breeding multi-modal spatio-temporal engine module and the behavior tracking result output module are connected through a result output channel.

[0006] Further, in the behavior dynamic atlas deviation evaluation module, a contour time series feature extracted by the dynamic contour anchoring Mamba network and a behavior dynamic atlas feature are used to construct a deviation evaluation model, and the model formula is: wherein, is a total deviation value of the behavior dynamic atlas, is a time series frame number processed by the dynamic contour anchoring Mamba network, is a calibration node number in the behavior dynamic atlas, is a weight coefficient of the th calibration node in the th frame, is a feature vector of the th calibration node in the th frame output by the dynamic contour anchoring Mamba network, is a feature vector of the th calibration node in the th frame in the preset standard behavior atlas, the first frame of the output of the intelligent breeding multi-modal spatio-temporal engine the spatio-temporal correlation coefficient of the frame of the calibration node.

[0007] Further, in the intelligent breeding multi-modal spatio-temporal engine module, the time sequence features output by the dynamic contour anchored Mamba network and the spatio-temporal information of the multi-modal data are fused to construct a spatio-temporal feature fusion model, and the formula is: wherein, is the fused spatio-temporal feature matrix, is the fusion weight coefficient, is the time sequence feature matrix output by the dynamic contour anchored Mamba network, is the spatio-temporal feature matrix output by the multi-modal data acquisition module, is the matrix Hadamard product operation, is the gradient matrix of the time sequence feature of the dynamic contour anchored Mamba network, is the gradient matrix of the multi-modal spatio-temporal feature, is the matrix multiplication operation.

[0008] Further, in the dynamic contour anchored Mamba network processing module, when modeling the contour time sequence of the animal, the improved Mamba network state update formula is used: wherein, is the network hidden state vector of the frame, is the state transition matrix, is the network hidden state vector of the frame, is the input weight matrix, is the input feature transformation matrix, is the contour feature vector of the frame, is the bias vector, is the dynamic contour anchoring term weight coefficient, is the dynamic contour anchoring function, is the contour anchor point coordinate set of the frame.

[0009] Further, in the behavior dynamic atlas construction module, the contour data output by the dynamic contour anchored Mamba network is used to calculate the calibration node coordinates, and the formula is: wherein, is the coordinate of the frame of the calibration node, is the number of contour pixel neighborhoods output by the dynamic contour anchored Mamba network, is the contour feature vector of the frame. Frame number The coordinates of each contour pixel. The contour feature enhancement coefficient, To use dynamic contour anchoring to generate convolutional kernels for Mamba networks right The result of the convolution operation.

[0010] Furthermore, in the behavior tracking result output module, the output data of the intelligent breeding multimodal spatiotemporal engine and the behavior dynamic map deviation assessment results are combined to calculate the probability of animal behavior categories, using the following formula: ,in, Animals belong to the category The probability, This refers to the spatiotemporal feature vector output by the multimodal spatiotemporal engine for intelligent aquaculture. The deviation value output by the behavioral dynamic graph deviation assessment module. For category The feature weight vector, For category The bias term, The deviation value influence coefficient. This represents the total number of behavior categories.

[0011] Furthermore, the behavior dynamic atlas construction module includes: a calibration node candidate region generation unit, which receives temporal contour data output by the dynamic contour anchoring Mamba network processing module, traverses the contour image through a sliding window, and selects regions with a grayscale value change rate greater than a preset threshold as calibration node candidate regions, annotates the candidate regions with bounding boxes, and records the center coordinates and size information of each candidate region; and a calibration node feature extraction unit, which extracts features from the image regions within the bounding boxes annotated by the calibration node candidate region generation unit, uses a 3×3 convolution kernel to perform feature mapping on the region image, extracts texture features and edge features, and concatenates the extracted features into a feature vector. The system correlates the temporal features of the dynamic contour anchoring Mamba network output with the data to obtain the temporal feature vector of each candidate region. The calibration node matching unit receives the temporal feature vector output by the calibration node feature extraction unit, calculates the cosine similarity of the candidate region feature vectors between adjacent frames, determines the candidate regions with similarity greater than a preset threshold as the same calibration node, and updates the calibration node coordinates. The graph topology construction unit calculates the Euclidean distance between adjacent calibration nodes based on the calibration node coordinates determined by the calibration node matching unit, constructs an adjacency matrix based on the distance value, and combines the adjacency matrix with the temporal features of the calibration nodes to form the topology of the behavioral dynamic graph.

[0012] Further, the behavior dynamic graph deviation evaluation module comprises: a standard graph feature library construction unit, which collects multi-modal data under a normal behavior state of the farmed animal, processes the multi-modal data through a dynamic contour anchored Mamba network, generates a standard behavior graph through a behavior dynamic graph construction module, extracts standard graph calibration node features, adjacency matrix features and time sequence change features, establishes a standard graph feature library and stores feature threshold ranges; a real-time graph feature extraction unit, which receives a real-time behavior dynamic graph output by the behavior dynamic graph construction module, extracts real-time graph calibration node features, adjacency matrix features and time sequence change features according to a feature extraction mode of the standard graph feature library construction unit, and forms a real-time graph feature vector; a feature deviation calculation unit, which compares the real-time graph feature vector output by the real-time graph feature extraction unit with standard feature vectors in the standard graph feature library dimension by dimension, calculates absolute deviation values and relative deviation values of each dimension, and obtains a feature deviation total sum by using a weighted summation method; and a deviation level determination unit, which determines a deviation level of the real-time behavior dynamic graph by comparing the feature deviation total sum output by the feature deviation calculation unit with deviation threshold values of different levels according to a preset deviation level division rule, and outputs a deviation level result.

[0013] Further, the intelligent farming multi-modal spatio-temporal engine module comprises: a multi-modal data spatio-temporal alignment unit, which receives video streams, audio streams and infrared thermal imaging data output by the multi-modal data collection module, extracts timestamp information and spatial coordinate information of each modal data, synchronizes the audio streams and the infrared thermal imaging data in time based on the timestamp of the video streams, calibrates the spatial coordinates of each modal data based on a spatial coordinate system of the farming scene, and performs spatio-temporal alignment of the multi-modal data; a time sequence feature fusion unit, which receives time sequence contour features output by the dynamic contour anchored Mamba network processing module and the aligned multi-modal data output by the multi-modal data spatio-temporal alignment unit, assigns weights to time sequence features of each modal data by using an attention mechanism, multiplies the weights by the time sequence features of the corresponding modal, and sums the results to obtain fused time sequence features; a spatial feature enhancement unit, which processes spatial features output by the multi-modal data spatio-temporal alignment unit, expands a receptive field by using a hollow convolution, extracts spatial features of different scales, splices the spatial features of different scales, enhances the spliced spatial features in combination with contour spatial information output by the dynamic contour anchored Mamba network, and obtains enhanced spatial features; and a spatio-temporal correlation modeling unit, which receives the fused time sequence features output by the time sequence feature fusion unit and the enhanced spatial features output by the spatial feature enhancement unit, correlates and models the spatio-temporal features by using a graph neural network, takes the time sequence features as node features and the spatial features as edge features, constructs a spatio-temporal correlation graph, updates the node features through graph convolution operation, and generates behavior data correlated in space and time.

[0014] The artificial intelligence-based farmed animal behavior tracking method comprises the following steps: S1, a multi-modal data acquisition device is used to acquire video stream, audio stream and infrared thermal imaging data of animals in a farming scene, encapsulate the acquired data according to a preset format, and send the encapsulated data to a dynamic contour anchoring Mamba network processing module through a data transmission link; S2, the dynamic contour anchoring Mamba network processing module receives the encapsulated data, decapsulates the data, extracts frame images in the video stream, performs contour extraction on animal targets in the frame images through a dynamic contour anchoring unit, obtains an animal contour pixel coordinate set, inputs the contour pixel coordinate set into a Mamba network, uses a linear attention mechanism of the Mamba network to perform time series modeling on the contour pixel coordinate set, and outputs a time series contour feature sequence; S3, a behavior dynamic atlas construction module receives the time series contour feature sequence, identifies animal body calibration nodes based on the time series contour feature sequence, calculates the coordinates of each calibration node in different frames, calculates the distance and angle between adjacent calibration nodes according to the calibration node coordinates, and constructs a behavior dynamic atlas; S4, a behavior dynamic atlas deviation evaluation module receives the behavior dynamic atlas, calls a preset standard behavior atlas feature library, calculates the feature deviation value of the real-time behavior dynamic atlas and the standard behavior atlas, normalizes the feature deviation value, and obtains a deviation evaluation result; S5, a smart farming multi-modal spatio-temporal engine module receives the deviation evaluation result, the time series contour feature sequence and the multi-modal data, fuses the spatio-temporal information of the multi-modal data and the time series contour feature sequence, constructs a spatio-temporal correlation model, inputs the deviation evaluation result into the spatio-temporal correlation model, and generates spatio-temporally correlated behavior data; S6, a behavior tracking result output module receives the spatio-temporally correlated behavior data, analyzes the behavior data, identifies the behavior category of an animal individual, draws the motion trajectory of the animal based on the coordinate information in the behavior data, arranges the behavior category and the motion trajectory according to a preset format, and outputs the behavior tracking result to a display terminal and a storage device.

[0015] Beneficial Effects: This invention proposes an artificial intelligence-based animal behavior tracking system and method. It integrates video, audio, and infrared thermal imaging data through a multimodal data acquisition module, and combines this with a dynamic contour anchoring Mamba network processing module to achieve accurate extraction of animal contours and cross-frame temporal modeling. This effectively solves the problems of insufficient temporal modeling capability in multimodal data processing and weak cross-frame data correlation, improving the accuracy of behavioral feature extraction to reflect the continuous changes in animal behavior. The behavioral dynamic map construction module works in conjunction with the deviation assessment module to construct a map based on key node coordinates and calculate the feature deviation between the real-time map and the standard map, compensating for the single dimension of the assessment system. The lack of dynamic deviation quantification assessment is a deficiency that hinders the comprehensive judgment of the degree of behavioral deviation. The smart farming multimodal spatiotemporal engine module integrates multimodal spatiotemporal information and temporal characteristics to construct spatiotemporally correlated behavioral data, further enhancing data correlation and analysis depth. Meanwhile, the behavior tracking result output module accurately outputs behavior categories and movement trajectories, comprehensively meeting the needs of smart farming for efficient, accurate, and in-depth analysis of animal behavior tracking. This provides more reliable data support for farming management decisions and comprehensively overcomes the shortcomings of poor multimodal data fusion and processing, inaccurate behavioral feature modeling, insufficient spatiotemporal correlation analysis, and single evaluation dimensions in animal behavior. Attached Figure Description

[0016] Figure 1 This is a diagram showing the system module composition of the present invention; Figure 2 This is a flowchart of the method steps of the present invention. Detailed Implementation

[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0018] like Figure 1 As shown, the artificial intelligence-based animal behavior tracking system includes: The multimodal data acquisition module is used to collect video streams, audio streams, and infrared thermal imaging data of animals in the breeding scene; Specifically, the multi-modal data acquisition module collects video stream, audio stream and infrared thermal imaging data of animals in the breeding scene. The implementation process needs to configure acquisition equipment that meets the needs of the breeding scene. The video acquisition equipment uses a network camera with a resolution of 1920x1080 pixels and a frame rate of 30 frames per second. The lens focal length is set to 4mm-12mm to cover a collection range of 5m-20m, ensuring clear capture of animal limb movement details. The audio acquisition equipment uses an omnidirectional microphone with a sampling rate of 44.1kHz and a bit depth of 16bit. The sensitivity is controlled at -38dB±2dB, which can effectively collect audio information such as animal calls and activity friction sounds within a range of 30 meters. The infrared thermal imaging equipment uses a detector with a resolution of 640x512 pixels. The temperature measurement range is set to -20℃-150℃, and the thermal sensitivity is ≤50mK. It can accurately capture the temperature distribution and position information of animals in low light or night environments. This module achieves timestamp alignment of the three modal data through the device synchronization control unit, with a time synchronization error controlled within ±10ms. The collected data is stored in H.264 encoding format for video stream, WAV format for audio stream, and TIFF format for infrared thermal imaging data. It is sent to the subsequent processing module through the Ethernet interface at a transmission rate of 100Mbps, providing comprehensive and accurate raw data support for the entire system. It solves the problem of incomplete information of single modal data, ensures that the subsequent behavior analysis can combine visual, auditory and temperature information, and improves the comprehensiveness and accuracy of behavior recognition.

[0019] The dynamic contour anchored Mamba network processing module receives the output data of the multi-modal data acquisition module, extracts the contour of the animal target through the dynamic contour anchoring unit, and uses the linear attention mechanism of the Mamba network to model the extracted contour sequence in time sequence. Specifically, the dynamic contour anchoring Mamba network processing module receives the data output by the multi-modal data acquisition module. In the implementation process, the input video stream is first frame extracted, and key frames are obtained at a frequency of extracting 1 frame every 1 frame, reducing the data processing amount while ensuring the behavior timing continuity. Then, the dynamic contour anchoring unit is started, and an edge detection algorithm is used to preliminarily extract the animal target in the key frame. The edge gradient threshold is set to 15-25 to filter environmental noise interference. Then, the preliminary contour is optimized through a contour growth algorithm. The growth step is set to 1 pixel, and the iteration number is controlled within 5-8 times to ensure that the contour completely covers the animal body area and obtains the animal contour pixel coordinate set. Subsequently, the contour pixel coordinate set is input into the Mamba network. The network hidden layer dimension is set to 512 dimensions, and the window size of the linear attention mechanism is set to 10-15 frames. The contour sequence is time series modeled through a sliding window, the similarity of adjacent frame contours is calculated, and the network state is updated. The state update step is set to 0.01-0.05, the model training iteration number is 1000-1500 times, the training learning rate initial value is set to 0.001, and every 200 iterations is attenuated to 0.5 times of the original value, to ensure that the network can accurately learn the timing change rule of the animal contour. Through the dynamic contour anchoring, the animal target positioning accuracy is ensured, the linear attention mechanism of the Mamba network is used to improve the timing data processing efficiency, precise timing contour features are provided for subsequent behavior dynamic atlas construction, and the problems of traditional contour extraction being easily affected by environmental interference and low timing modeling efficiency are solved.

[0020] The behavior dynamic atlas construction module receives the timing contour data output by the dynamic contour anchoring Mamba network processing module, and constructs a behavior dynamic atlas based on the animal body calibration node coordinates. Specifically, the behavior dynamic graph construction module receives the time series contour data output by the dynamic contour anchored Mamba network processing module. In the implementation process, the time series contour data is first subjected to key node identification. Based on the gray value change and curvature characteristics of the contour pixel coordinates, the key nodes such as the head, torso, and limb ends of the animal are determined. Each animal target is provided with 12-16 key nodes. The confidence threshold for key node identification is set to be above 0.85. Nodes below the threshold need to be supplemented by the interpolation method of adjacent frame key nodes. The interpolation step is set to be 0.5-1 pixel unit according to the frame interval. Then, the coordinates of each key node in different frames are calculated. The coordinate calculation accuracy is controlled within ±0.5 pixels. Then, based on the key node coordinates, the Euclidean distance and angle between adjacent key nodes are calculated. The distance calculation error is controlled within ±1%. The angle calculation range is 0°-360°, and the accuracy is controlled within ±0.5°. Subsequently, the key nodes are taken as graph nodes, and the distance, angle, and time series change rate of adjacent key nodes are taken as edge weights to construct a behavior dynamic graph. The node feature dimension of the graph is set to be 20-30 dimensions, including key node coordinates, gray value, and motion speed information. The edge feature dimension is set to be 5-8 dimensions, including distance, angle, and change frequency information. The graph data is stored in the form of an adjacency matrix. The matrix dimension is consistent with the number of key nodes. The abstract contour time series data is converted into structured graph data, which directly reflects the spatial structure and time series change of animal behavior, provides a clear feature carrier for subsequent deviation evaluation, and solves the problems of difficult quantification and unclear structure of traditional behavior features.

[0021] The behavior dynamic graph deviation evaluation module receives the graph data output by the behavior dynamic graph construction module, and calculates the feature deviation value of the real-time graph and the preset standard behavior graph. Specifically, the behavior dynamic atlas deviation evaluation module receives the atlas data output by the behavior dynamic atlas construction module. In the implementation process, a preset standard behavior atlas library is first established, behavior data of the farmed animals in normal feeding, activity, rest and other states is collected, and a standard atlas is generated through the same process as the behavior dynamic atlas construction module. Each standard behavior category stores 500-1000 sample atlases, the mean and standard deviation of the characteristics of the sample atlases are calculated, and the standard characteristic range is determined. Then, feature extraction is performed on the real-time input atlas data. The extracted features are consistent with the feature types of the standard atlas library, including the mean and variance of the node features, the change trend of the edge features, etc., and a total of 30-40 feature dimensions are extracted. Then, the deviation value of the real-time atlas features and the standard atlas features is calculated, and the total deviation is calculated by weighted summation, in which the weight proportion of the node features is 60%-70%, and the weight proportion of the edge features is 30%-40%. The weights are determined by the analytic hierarchy process, and the deviation value is calculated by combining the absolute deviation and the relative deviation. The absolute deviation threshold is set to 1.5-2 times the standard deviation of the standard features, and the relative deviation threshold is set to 5%-10%. Finally, the calculated deviation value is normalized, and the normalization range is 0-1. 0 represents no deviation, and 1 represents the maximum threshold deviation. The deviation evaluation result is output in numerical form, and the output frequency is consistent with the atlas generation frequency of the behavior dynamic atlas construction module, which is 1-2 times per second. By quantifying the deviation of real-time behavior and standard behavior, animal behavior abnormalities are accurately identified, and key deviation basis is provided for the subsequent spatio-temporal engine module, solving the problem of traditional behavior evaluation relying on subjective judgment and lacking of quantitative standard.

[0022] The smart farming multi-modal spatio-temporal engine module receives the deviation value output by the behavior dynamic atlas deviation evaluation module, fuses the spatio-temporal information of the multi-modal data acquisition module and the time sequence features of the dynamic contour anchoring Mamba network processing module, and generates behavior data associated with space and time. Specifically, the smart breeding multi-modal spatio-temporal engine module receives the deviation value output by the behavior dynamic graph deviation evaluation module. In the implementation process, the spatio-temporal information of the multi-modal data acquisition module is first extracted. The time information is based on the timestamp of data acquisition, accurate to millisecond level. The spatial information is based on the coordinate system preset for the breeding scene. The origin of the coordinate system is set at the upper left corner of the breeding area. The X-axis is along the horizontal direction, and the Y-axis is along the vertical direction. The coordinate unit is meter. The pixel coordinates are converted into actual spatial coordinates through image calibration. The calibration error is controlled within ±0.1 meters. Then the time sequence features of the dynamic contour anchored Mamba network processing module are fused. The time sequence features include contour change rate, key node motion period, etc., a total of 15-20 feature dimensions. In the fusion process, the attention fusion mechanism is adopted. Higher attention weight is given to the period with larger deviation value. The weight adjustment range is 1.2-1.5 times, ensuring that the features of the abnormal behavior period are given special attention. Then the spatio-temporal correlation model is constructed. The model input is multi-modal spatio-temporal information, time sequence features and deviation value, a total of 60-80 feature dimensions. The model is constructed using a deep learning framework, including 3-5 fully connected layers, with 128-256 neurons per layer. The activation function uses the ReLU function. In the model training process, the cross-entropy loss function is used. The training batch size is set to 32-64, and the iteration number is 800-1200 times, ensuring that the model can accurately establish the correlation between spatio-temporal features and behavior. Finally, the spatio-temporal correlated behavior data is generated. The data includes the position of the animal in the actual spatial coordinate system, behavior category, deviation level and time sequence number. The data update frequency is 1 time / second. Integrating multi-source information, a comprehensive spatio-temporal correlation model is constructed, improving the integrity and correlation of behavior data, providing accurate comprehensive data for subsequent result output, and solving the problem of separation of spatio-temporal information and weak data correlation in traditional technology.

[0023] The behavior tracking result output module receives the spatio-temporal correlated behavior data output by the smart breeding multi-modal spatio-temporal engine module, and outputs the real-time behavior category and motion trajectory data of the animal individual.

[0024] Specifically, the behavior tracking result output module receives spatiotemporally correlated behavioral data output by the smart aquaculture multimodal spatiotemporal engine module. During implementation, the input data is first parsed. The parsing process extracts information such as animal individual identifiers, real-time behavior categories, spatial coordinates, timestamps, and deviation levels according to a preset data format. The data parsing accuracy must reach over 99%. Data with parsing errors is marked and fed back to the smart aquaculture multimodal spatiotemporal engine module for reprocessing. Next, the behavior categories are standardized and output. Behavior categories are divided into 8-10 categories based on common farmed animal behaviors, such as eating, drinking, activity, resting, and abnormal stress. The output is in Chinese text format, along with the confidence level of the behavior category. The confidence threshold is set to above 0.7; behavior categories below the threshold are marked as "pending confirmation." Then, animal movement trajectories are plotted based on spatial coordinate information. The trajectories are plotted as line graphs, with the horizontal axis representing time and the vertical axis representing spatial coordinates. The time interval is set to 10-30 seconds. The trajectory lines are distinguished by individual animal identifiers, with different colors used for different individuals. The line width is set to 2-3 pixels. Key time points are marked with the behavior categories at critical time points, selected when the behavior category changes or the deviation level exceeds a threshold. Finally, the behavior categories and movement trajectories are organized according to preset formats, including Excel spreadsheets and image files. The Excel spreadsheets contain fields such as individual identifier, time, behavior category, confidence level, spatial coordinates, and deviation level. The image files are movement trajectory graphs with a resolution of 1920×1080 pixels, sent to the display terminal via Ethernet interface, and simultaneously stored in a local database for 30-90 days for subsequent querying and analysis. This process transforms complex spatiotemporal behavioral data into intuitive and easy-to-understand output results, providing clear behavioral tracking information for livestock managers, assisting in rapid management decision-making, and solving the problems of traditional output formats being monotonous and difficult to read.

[0025] The multimodal data acquisition module is connected to the dynamic contour anchoring Mamba network processing module through a data transmission interface. The dynamic contour anchoring Mamba network processing module is connected to the behavior dynamic map construction module through a time-series data bus. The behavior dynamic map construction module is connected to the behavior dynamic map deviation evaluation module through a map feature channel. The behavior dynamic map deviation evaluation module is connected to the smart aquaculture multimodal spatiotemporal engine module through a deviation value transmission link. The smart aquaculture multimodal spatiotemporal engine module is connected to the behavior tracking result output module through a result output channel.

[0026] Preferably, in the behavioral dynamic graph deviation assessment module, the deviation assessment model is constructed using the contour temporal features extracted by the dynamic contour pinning Mamba network and the behavioral dynamic graph features. The model formula is as follows: ,in, This represents the total deviation value of the behavioral dynamic spectrum. To determine the number of time frames processed by the Mamba network for dynamic profile pins, This represents the number of labeled nodes in the behavioral dynamic graph. For the first Frame number The weight coefficients of each calibration node, Anchoring the output of the Mamba network for dynamic contours Frame number Each calibrated node feature vector For the first in the preset standard behavior map Frame number Each calibrated node feature vector The first output of the multimodal spatiotemporal engine for smart aquaculture Frame number The spatiotemporal correlation coefficient of each calibration node.

[0027] Specifically, during the implementation of the behavioral dynamic map deviation assessment module, the number of time-series frames for dynamic contour anchoring Mamba network processing is first determined. Based on the frequency of animal behavior changes in the breeding scenario, the number of time-series frames is set to 30-60 frames to ensure coverage of the complete behavioral cycle. The number of key nodes in the behavioral dynamic map is adjusted according to the animal's body size, with 12-14 for small and medium-sized animals and 15-16 for large animals, ensuring that key nodes can comprehensively reflect the body's movement state. The weight coefficient of the i-th key node in the t-th frame is determined through animal behavioral feature importance analysis. The weight of key nodes in the head and trunk is set to 0.8-1.0, and the weight of key nodes in the extremities is set to 0.5-0.7, highlighting the impact of core part behavioral changes on the overall behavioral assessment. The spatiotemporal correlation coefficient output by the smart breeding multimodal spatiotemporal engine is dynamically adjusted according to the reliability of the multimodal data. The correlation coefficient corresponding to video stream data is set to 0.6-0.8, and the correlation coefficient corresponding to audio stream and infrared thermal imaging data is set to 0.3-0.5, ensuring data reliability and weight matching. When calculating the total deviation value, the L2 norm of the feature vector difference is first calculated frame-by-frame and node-by-node, and then accumulated by combining the weighting coefficient and the spatiotemporal correlation coefficient. The deviation value calculation frequency is consistent with the map generation frequency, which is 1-2 times / second. By integrating multi-dimensional parameters, the accuracy of behavioral dynamic map deviation assessment is improved, avoiding errors caused by single feature assessment. This allows the deviation value to more realistically reflect the difference between animal behavior and standard behavior, providing a more reliable quantitative basis for subsequent behavioral anomaly identification.

[0028] Preferably, in the intelligent aquaculture multimodal spatiotemporal engine module, the temporal features of the dynamic contour-guided Mamba network are integrated with the spatiotemporal information of multimodal data to construct a spatiotemporal feature fusion model, the formula of which is: ,in, The fused spatiotemporal feature matrix, To integrate the weighting coefficients, To anchor the temporal feature matrix output by the Mamba network for dynamic contours, This is the spatiotemporal feature matrix output by the multimodal data acquisition module. For the Hadamard product operation of matrices, The gradient matrix of the temporal features of the Mamba network is determined for the dynamic contour needle. The gradient matrix of multimodal spatiotemporal features. This refers to matrix multiplication.

[0029] Specifically, the spatiotemporal feature fusion of the smart aquaculture multimodal spatiotemporal engine module requires dynamic adjustment of the fusion weight coefficient based on the data characteristics of the aquaculture scenario. During periods of stable lighting and frequent animal activity, the fusion weight coefficient is set to 0.6-0.7, emphasizing the temporal features of the dynamic contour anchoring Mamba network. During periods of significant lighting changes and when audio information is more critical, the fusion weight coefficient is set to 0.3-0.4, emphasizing the spatiotemporal features of the multimodal data. The dimension of the temporal feature matrix output by the dynamic contour anchoring Mamba network is determined based on the number of key nodes. A feature matrix with 12-14 key nodes corresponds to a dimension of 512×30, and one with 15-16 key nodes corresponds to a dimension of 512×40, ensuring that the feature dimension matches the number of key nodes. The spatiotemporal feature matrix output by the multimodal data acquisition module includes timestamps, spatial coordinates, and modality identification information. The matrix dimension is set to 256×30, with timestamp accuracy at the millisecond level and spatial coordinate accuracy at ±0.1 meters. When calculating the Hadamard product and gradient matrix multiplication, 32-bit floating-point operations are used to ensure computational accuracy. The gradient matrix is ​​obtained by first-order differencing of the temporal and spatiotemporal feature matrices, with a differencing step size of one frame. The fused spatiotemporal feature matrix needs to be normalized, with a normalization range of 0-1. The processed data is used for subsequent behavioral data generation. By combining flexible weight adjustments with multi-matrix operations, deep fusion of temporal and spatiotemporal features is achieved, solving the problems of fixed feature weights and poor fusion effects in traditional fusion methods, and improving the ability of spatiotemporal features to represent animal behavior.

[0030] Preferably, in the dynamic contour anchoring Mamba network processing module, when modeling the time series sequence of animal contours, an improved Mamba network state update formula is used: ,in, For the first The network hidden state vector of a frame. Here is the state transition matrix. For the first The network hidden state vector of a frame. For the input weight matrix, Given the input feature transformation matrix, For the first The contour feature vector of the frame. For bias vectors, For dynamic contour needle fixed term weight coefficient, For dynamic contour anchoring function, For the first The set of coordinates of the frame's outline anchor points.

[0031] Specifically, the state update mechanism of the dynamic contour anchoring Mamba network processing module is improved. During implementation, the state transition matrix is ​​set according to the temporal correlation of animal behavior, with a matrix dimension consistent with the network's hidden state vector dimension (512×512). Matrix element values ​​range from 0.1 to 0.3, ensuring that state transitions reflect behavioral continuity. The input weight matrix dimension is set to 512×256, matching the contour feature vector dimension, while the bias vector dimension is 512×1, with initial values ​​set to 0.01-0.05, optimized iteratively through model training. The contour anchor point coordinate set in the dynamic contour anchoring function is extracted from contour data of 3-5 historical frames, with the number of anchor points matching the number of key nodes, ensuring that the anchoring function can optimize the contour features of the current frame based on historical information. The GELU activation function is used to perform a non-linear mapping on the input feature transformation results, enhancing the network's ability to fit complex behavioral features. During network training, 100-200 sets of farmed animal behavior samples are used in each iteration, covering different behavioral states. During training, parameters such as the state transition matrix and input weight matrix are adjusted in real time using the validation set. Training stops when the validation set loss function value is less than 0.001 for 10 consecutive iterations. By improving the network state update formula and introducing a dynamic contour anchoring function and the GELU activation function, the modeling accuracy of the Mamba network for the temporal features of animal contours is improved. This addresses the problem of traditional networks relying on a single input for state updates and having poor adaptability to dynamic contour changes, ensuring that the network can accurately capture subtle changes in animal behavior.

[0032] Preferably, in the behavior dynamic map construction module, the coordinates of the calibration nodes are calculated based on the contour data output by the dynamic contour anchoring Mamba network, using the following formula: ,in, For the first Frame number The coordinates of each calibration node, To anchor the number of contour pixel neighbors in the output of the Mamba network for dynamic contours, For the first Frame number The coordinates of each contour pixel. The contour feature enhancement coefficient, To use dynamic contour anchoring to generate convolutional kernels for Mamba networks right The result of the convolution operation.

[0033] Specifically, the key node coordinate calculation of the behavior dynamic atlas construction module, the number of contour pixel neighborhoods output by the dynamic contour anchoring Mamba network is determined according to the size of the animal contour during implementation, 20-30 neighborhoods are set for small and medium-sized animals, and 35-45 neighborhoods are set for large animals, each neighborhood contains 8-12 contour pixels, and the neighborhood can cover the feature information around the key node. The contour feature enhancement coefficient is dynamically adjusted according to the contour sharpness, and is set to 0.3-0.5 when the contour edge is clear, and is set to 0.6-0.8 when the contour edge is blurred, and the feature recognition degree of the blurred contour is improved through the enhancement coefficient. The size of the convolution kernel generated by the dynamic contour anchoring Mamba network is set to 3x3, the number of convolution kernels is consistent with the contour feature dimension, which is 256, and the step size is set to 1 and the padding is set to 1 during convolution operation, which ensures that the feature map size after convolution is consistent with the input. When calculating the key node coordinates, first, the contour pixel coordinates in each neighborhood are averaged, then the product of the convolution operation result and the contour feature enhancement coefficient is added, and after obtaining the preliminary coordinates, the abnormal coordinate values are corrected by comparing with the historical 2-3 frame key node coordinates, and the coordinate correction error is controlled within ±0.3 pixels. The key node coordinates are calculated by combining multi-neighborhood averaging and convolution enhancement, which improves the accuracy and stability of the coordinate calculation, and solves the problem that the traditional key node coordinate calculation depends on a single pixel and is easily disturbed by noise, providing accurate node coordinate data for subsequent behavior dynamic atlas construction.

[0034] Preferably, in the behavior tracking result output module, the animal behavior category probability is calculated by combining the output data of the smart breeding multi-modal spatio-temporal engine and the behavior dynamic atlas deviation evaluation result, and the formula is: wherein, is the probability that the animal belongs to category , is the spatio-temporal feature vector output by the smart breeding multi-modal spatio-temporal engine, is the deviation value output by the behavior dynamic atlas deviation evaluation module, is the feature weight vector of category , is the bias term of category , is the deviation value influence coefficient, is the total number of behavior categories.

[0035] Specifically, the behavior tracking result output module calculates the behavior category probability. In the implementation process, the total number of behavior categories is determined according to common behavior types of farmed animals, and is set to 8-10 categories, covering states such as eating, drinking, moving, resting, abnormal stress, etc. The dimension of the category feature weight vector is consistent with the dimension of the spatiotemporal feature vector output by the intelligent farming multi-modal spatiotemporal engine, which is 256x1. The initial value of the bias term is set to 0.01-0.03, which is optimized based on farmed animal behavior sample data through model training. The number of samples of each behavior category in the training sample is kept balanced, and the difference in the number of samples of each category is not more than 10%. The bias value influence coefficient is determined according to the discrimination degree of the bias value to the behavior category. When the bias value has strong correlation with the behavior category, it is set to 0.8-1.0, and when the correlation is weak, it is set to 0.4-0.6, so as to ensure that the bias value can reasonably affect the calculation of the behavior category probability. When calculating the behavior category probability, the product of the spatiotemporal feature vector and the category feature weight vector, the bias term, and the product of the bias value and the influence coefficient are first linearly combined, and then normalized by the Softmax function to obtain the probability value of each behavior category. The probability value calculation accuracy is controlled to 4 decimal places. The spatiotemporal feature and the bias evaluation result are combined to calculate the behavior category probability, which improves the accuracy of behavior category recognition, solves the problem of low recognition accuracy of traditional behavior recognition which only relies on a single feature, and ensures that the output behavior category can truly reflect the actual behavior state of the animal.

[0036] Preferably, the behavior dynamic atlas construction module comprises: a calibration node candidate region generation unit that receives the time series contour data output by the dynamic contour anchor Mamba network processing module, traverses the contour image through a sliding window, filters out regions with a gray value change rate greater than a preset threshold as calibration node candidate regions, and labels the boundary boxes of the candidate regions and records the center coordinates and size information of each candidate region; a calibration node feature extraction unit that extracts features from the image region within the boundary box labeled by the calibration node candidate region generation unit, maps the region image using a 3x3 convolution kernel to extract texture features and edge features, concatenates the extracted features into a feature vector, and associates the feature vector with the time series features output by the dynamic contour anchor Mamba network to obtain a time series feature vector for each candidate region; a calibration node matching unit that receives the time series feature vector output by the calibration node feature extraction unit, calculates the cosine similarity of the candidate region feature vectors between adjacent frames, determines candidate regions with a similarity greater than a preset threshold as the same calibration node, and updates the calibration node coordinates; and a graph topology structure construction unit that calculates the Euclidean distance between adjacent calibration nodes according to the calibration node coordinates determined by the calibration node matching unit, constructs an adjacency matrix based on the distance value, combines the adjacency matrix with the calibration node time series features, and forms the topology structure of the behavior dynamic atlas.

[0037] Specifically, the unit splitting and refinement of the behavior dynamic graph construction module is implemented. After the key node candidate region generation unit receives the time sequence contour data output by the dynamic contour anchor Mamba network, a 16x16 pixel sliding window is used to traverse the contour image, and the sliding step is set to 4 pixels. The candidate region is selected by calculating the gray value change rate in the window, and the region with a gray value change rate greater than 15% is marked as a candidate region. The boundary box labeling accuracy is controlled within ±1 pixel, and the center coordinates (accurate to 1 decimal place) and width and height size (range 20-50 pixels) of each candidate region are recorded. The key node feature extraction unit uses a 3x3 convolution kernel to map the features of the image region in the boundary box, with a convolution step of 1 pixel and padding of 1 pixel. The 128-dimensional texture features and 64-dimensional edge features are extracted and concatenated into a 192-dimensional feature vector. Then, the 64-dimensional time sequence features output by the dynamic contour anchor Mamba network are associated to form a 256-dimensional time sequence feature vector. The association process uses feature concatenation rather than weighted fusion. The key node matching unit calculates the cosine similarity of the feature vectors of adjacent frame candidate regions, and determines that the similarity greater than 0.85 is the same key node. The key node coordinate history data of the previous 3 frames is retained when updating the key node coordinates of each frame, and if the deviation between the current coordinates and the historical average coordinates exceeds 8 pixels, a re-matching is triggered. The graph topology structure construction unit calculates the Euclidean distance (accuracy ±0.2 pixels) between adjacent key nodes, and constructs an adjacency matrix based on the distance value (less than 30 pixels is considered adjacent). The matrix element value is the distance normalization result (range 0-1). Then, the adjacency matrix is combined with the 256-dimensional key node time sequence features to form a graph topology structure data with a dimension of (key node number x key node number + 256). By splitting the units, the technical parameters of each link are determined to ensure the accuracy and operability of the behavior dynamic graph construction, and to solve the problem of fuzzy implementation details at the module level.

[0038] Preferably, the behavior dynamic graph deviation evaluation module comprises: a standard graph feature library construction unit, which collects multi-modal data under normal behavior state of the farmed animals, processes the multi-modal data through a dynamic contour anchored Mamba network, generates a standard behavior graph with a behavior dynamic graph construction module, extracts calibrated node features, adjacency matrix features and time sequence change features of the standard graph, establishes a standard graph feature library and stores feature threshold ranges; a real-time graph feature extraction unit, which receives a real-time behavior dynamic graph output by the behavior dynamic graph construction module, extracts calibrated node features, adjacency matrix features and time sequence change features of the real-time graph according to the feature extraction mode of the standard graph feature library construction unit, and forms a real-time graph feature vector; a feature deviation calculation unit, which compares the real-time graph feature vector output by the real-time graph feature extraction unit with standard feature vectors in the standard graph feature library dimension by dimension, calculates absolute deviation values and relative deviation values of each dimension, and obtains a feature deviation total sum by using a weighted summation method; and a deviation level determination unit, which compares the feature deviation total sum output by the feature deviation calculation unit with deviation threshold values of different levels according to a preset deviation level division rule, determines a deviation level of the real-time behavior dynamic graph, and outputs a deviation level result.

[0039] Specifically, the unit functions of the behavior dynamic atlas deviation evaluation module are refined. When the standard atlas feature library construction unit collects normal behavior data of farmed animals, 1000 groups of samples are collected for each normal behavior (eating, resting, etc.), and each group of samples contains 60 frames of time series contour data. After the standard atlas is generated by the behavior dynamic atlas construction module, 80-dimensional key node features (such as coordinate mean, change rate), 40-dimensional adjacency matrix features (such as average distance, connection density), and 30-dimensional time series change features (such as inter-frame similarity) are extracted, a total of 150-dimensional features. At the same time, set the threshold range of each feature (such as the key node coordinate change rate threshold 0.5-2.0 pixels / frame), and the feature library storage capacity is not less than 50GB to meet the multi-category storage demand. The real-time atlas feature extraction unit strictly follows the extraction method of the standard atlas feature library, extracts 150-dimensional features for each frame of atlas data input in real time, and the extraction time is controlled within 50 milliseconds to ensure real-time performance. The feature deviation calculation unit compares the real-time feature vector with the standard feature vector dimension by dimension, calculates the absolute deviation (|real-time value-standard value|) and the relative deviation (|real-time value-standard value| / standard value x 100%), in which the key node feature weight accounts for 65%, the adjacency matrix feature accounts for 25%, and the time series change feature accounts for 10%. The weighted sum formula (deviation total sum = Σ (absolute deviation x weight)) is used to calculate the feature deviation total sum, and the weight is determined by the analytic hierarchy process and optimized through 5 rounds of verification. The deviation level determination unit presets 3 deviation levels, the deviation total sum 0-5 is normal (level 1), 5-15 is slightly abnormal (level 2), and above 15 is seriously abnormal (level 3), and the output delay of the determination result is not more than 100 milliseconds. Through unitization and parameter quantization, the standardization and accuracy of deviation evaluation are improved, and the subjective problem of traditional evaluation is avoided.

[0040] Preferably, the smart farming multi-modal spatio-temporal engine module comprises: a multi-modal data spatio-temporal alignment unit receiving the video stream, the audio stream and the infrared thermal imaging data output by the multi-modal data acquisition module, extracting the timestamp information and the spatial coordinate information of each modal data, taking the timestamp of the video stream as the reference, synchronizing the time of the audio stream and the infrared thermal imaging data, calibrating the spatial coordinates of each modal data based on the spatial coordinate system of the farming scene, and performing spatio-temporal alignment of the multi-modal data; a time sequence feature fusion unit receiving the time sequence contour features output by the dynamic contour anchoring Mamba network processing module and the aligned multi-modal data output by the multi-modal data spatio-temporal alignment unit, using an attention mechanism to assign weights to the time sequence features of each modal data, multiplying the weights by the time sequence features of the corresponding modal and summing them up to obtain the fused time sequence features; a spatial feature enhancement unit processing the spatial features output by the multi-modal data spatio-temporal alignment unit, using a hollow convolution to expand the receptive field, extracting spatial features of different scales, splicing the spatial features of different scales, combining the contour spatial information output by the dynamic contour anchoring Mamba network, enhancing the spliced spatial features to obtain enhanced spatial features; a spatio-temporal correlation modeling unit receiving the fused time sequence features output by the time sequence feature fusion unit and the enhanced spatial features output by the spatial feature enhancement unit, using a graph neural network to model the spatio-temporal correlation, taking the time sequence features as node features and the spatial features as edge features, constructing a spatio-temporal correlation graph, updating the node features through graph convolution operation, and generating behavior data with spatio-temporal correlation.

[0041] Specifically, the unit technology of the smart breeding multi-modal spatio-temporal engine module is implemented. After the multi-modal spatio-temporal alignment unit receives the multi-modal data, it extracts the timestamps of each modality (the timestamp precisions of video stream, audio stream, and infrared thermal imaging data are 10 milliseconds, 5 milliseconds, and 15 milliseconds, respectively). Taking the video stream timestamp as the reference, the linear interpolation method is used for the audio stream, and the nearest-neighbor method is used for the infrared data to achieve time synchronization, with a synchronization error controlled within ±8 milliseconds. The spatial calibration is based on the breeding scene coordinate system (with the origin set at the southwest corner of the breeding area, the X-axis pointing east, and the Y-axis pointing north, with the unit in meters). The mapping relationship between pixel coordinates and spatial coordinates is established through four calibration boards (coordinates known), with a mapping error ≤0.1 meters, ensuring the spatial position unity of multi-modal data. The time sequence feature fusion unit uses the attention mechanism to allocate weights. The video stream time sequence feature (64 dimensions), audio stream time sequence feature (32 dimensions), and infrared time sequence feature (32 dimensions) are calculated for attention weights, with the weight sum being 1. When the animal is active, the video stream weight is set to 0.6-0.7, the audio stream weight is 0.2-0.3, and the infrared weight is 0.1-0.2. Then, the weights are multiplied by the corresponding modality time sequence features and summed to obtain 64-dimensional fused time sequence features. The spatial feature enhancement unit uses a 3x3 hollow convolution with a hollow rate of 2 to expand the receptive field, extracts 32-dimensional, 64-dimensional, and 128-dimensional spatial features, and concatenates them into 224-dimensional features. Combined with the 64-dimensional contour spatial information output by the dynamic contour anchored Mamba network, the concatenated features are enhanced through residual connection, and the enhanced feature dimension remains 224-dimensional. The spatio-temporal correlation modeling unit takes the 64-dimensional fused time sequence features as node features and the 224-dimensional enhanced spatial features as edge features to construct a spatio-temporal correlation graph (the number of nodes = the number of animals, and the number of edges = the number of connections between individuals with a distance less than 5 meters). Two-layer graph convolution operations (each layer outputs 128-dimensional features) are used to update the node features, and finally generate 128-dimensional spatio-temporal correlation behavior data with a data update frequency of 1 time / second. Through the explicit unit-level technical parameters and process refinement, the efficiency and correlation of multi-modal spatio-temporal fusion are ensured, and the traditional engine spatio-temporal separation problem is solved.

[0042] The dynamic contour anchoring Mamba network is a core technical component for processing multi-modal data of farmed animals, extracting and modeling the contour time sequence features of the animals in the present application. The implementation process needs to first receive the video stream data output by the multi-modal data acquisition module, extract a key frame at a frequency of 1 frame per interval, preliminarily extract the animal contour through an edge detection algorithm (the edge gradient threshold is set to 15-25), and then optimize the contour through a contour growth algorithm (the growth step is 1 pixel, and the iteration is 5-8 times) to obtain a complete contour pixel coordinate set. Then the coordinate set is input into the Mamba network, the network hidden layer dimension is set to 512 dimensions, a linear attention mechanism with a window size of 10-15 frames is used to model the contour sequence in time sequence, and a dynamic contour anchoring function is introduced to optimize the contour features of the current frame based on the contour anchor point coordinates of the last 3-5 frames (the number of anchor points is consistent with the number of key nodes), and the network is trained for 1000-1500 iterations (the initial learning rate is 0.001, and the learning rate is attenuated by 0.5 times every 200 iterations) to ensure the modeling accuracy. Its role is to accurately extract the animal contour and capture the time sequence variation law of the contour, and to provide high-quality time sequence contour data for subsequent behavior dynamic atlas construction. It solves the problems of traditional contour extraction being easily disturbed by the environment and low efficiency of time sequence modeling, and through the combination of dynamic anchoring and Mamba network, the contour extraction accuracy and time sequence processing efficiency are considered, and the core feature basis for behavior tracking of farmed animals is laid.

[0043] The behavior dynamic atlas deviation evaluation model of the present application is an evaluation component for quantifying the difference between the real-time behavior of farmed animals and the standard behavior, which is realized by relying on the behavior dynamic atlas deviation evaluation module. The implementation needs to first construct a standard atlas feature library, collect 1000 groups of samples (each group contains 60 frames of time sequence contour data) for each normal behavior (eating, resting, etc.), generate a standard atlas, extract 150-dimensional features (80-dimensional key node features, 40-dimensional adjacency matrix features, and 30-dimensional time sequence change features), and set a threshold range (such as the key node coordinate change rate of 0.5-2.0 pixels / frame); then extract 150-dimensional features from the real-time atlas in the same way, calculate the absolute deviation and relative deviation of the real-time and standard features, and weight the sum of the deviations by 65% for key node features, 25% for adjacency matrix features, and 10% for time sequence change features (determined by AHP and optimized for 5 rounds); finally, 3 deviation levels (0-5 for normal, 5-15 for slight abnormal, and above 15 for serious abnormal) are preset, and the delay of the output of the determination result is controlled within 100 milliseconds. Its role is to accurately identify animal behavior abnormalities and provide quantitative deviation basis for subsequent modules. This model breaks through the limitations of traditional behavior evaluation relying on subjective judgment, improves the accuracy and reliability of behavior abnormality identification through standardized and quantitative evaluation, and helps to discover animal health or behavior problems in time for breeding management.

[0044] The intelligent breeding multi-modal space-time engine of the present application is a core engine that fuses multi-modal data space-time information and behavior characteristics, generates space-time associated behavior data, and is realized through the cooperation of four units of the intelligent breeding multi-modal space-time engine module. In the implementation process, the multi-modal data space-time alignment unit takes the video stream timestamp (precision 10 milliseconds) as the reference, uses the linear interpolation method (audio stream) and the nearest-neighbor method (infrared data) to realize time synchronization (error ±8 milliseconds), and completes space calibration based on the breeding scene coordinate system (origin set in the southwest corner, X axis east, Y axis north) and four calibration boards (mapping error ≤0.1 meters); the time sequence feature fusion unit uses the attention mechanism to allocate weights (video stream 0.6-0.7, audio stream 0.2-0.3, infrared 0.1-0.2 for frequent activity period), and weighted summation is used to obtain 64-dimensional fused time sequence features; the spatial feature enhancement unit uses a 3x3 hollow convolution with a hollow rate of 2 to extract three scale features (32-dimensional, 64-dimensional, 128-dimensional), and after splicing, the enhanced features are combined with contour space information to enhance the features to 224-dimensional features; the space-time association modeling unit constructs a space-time association graph with the fused time sequence features as node features and the enhanced spatial features as edge features, and generates 128-dimensional space-time associated behavior data (update frequency 1 time / second) through 2-layer graph convolution. Its role is to integrate multi-source information and improve the association and completeness of behavior data. The significance of the engine lies in solving the problem of separation of space-time information in traditional technology, providing accurate comprehensive data for behavior tracking result output through deep fusion, and promoting the development of breeding animal behavior analysis to be more detailed and comprehensive.

[0045] As Figure 2As shown, the artificial intelligence-based farmed animal behavior tracking method comprises the following steps: S1, a multi-modal data acquisition device is used to acquire video stream, audio stream and infrared thermal imaging data of animals in a farming scene, the acquired data is packaged according to a preset format, and the packaged data is sent to a dynamic contour anchoring Mamba network processing module through a data transmission link; S2, the dynamic contour anchoring Mamba network processing module receives the packaged data, unpackages the data, extracts frame images in the video stream, performs contour extraction on animal targets in the frame images through a dynamic contour anchoring unit, obtains an animal contour pixel coordinate set, inputs the contour pixel coordinate set into a Mamba network, uses a linear attention mechanism of the Mamba network to perform time series modeling on the contour pixel coordinate set, and outputs a time series contour feature sequence; S3, a behavior dynamic graph construction module receives the time series contour feature sequence, identifies animal body calibration nodes based on the time series contour feature sequence, calculates coordinates of each calibration node in different frames, calculates distances and angles between adjacent calibration nodes according to the calibration node coordinates, and constructs a behavior dynamic graph; S4, a behavior dynamic graph deviation evaluation module receives the behavior dynamic graph, calls a preset standard behavior graph feature library, calculates a feature deviation value of a real-time behavior dynamic graph and a standard behavior graph, normalizes the feature deviation value, and obtains a deviation evaluation result; S5, a smart farming multi-modal spatio-temporal engine module receives the deviation evaluation result, the time series contour feature sequence and the multi-modal data, fuses spatio-temporal information of the multi-modal data and the time series contour feature sequence, constructs a spatio-temporal correlation model, inputs the deviation evaluation result into the spatio-temporal correlation model, and generates spatio-temporally correlated behavior data; and S6, a behavior tracking result output module receives the spatio-temporally correlated behavior data, analyzes the behavior data, identifies a behavior category of an animal individual, draws a motion trajectory of the animal based on coordinate information in the behavior data, arranges the behavior category and the motion trajectory according to a preset format, and outputs the behavior category and the motion trajectory to a display terminal and a storage device.

[0046] The artificial intelligence-based farmed animal behavior tracking system and method can integrate video, audio and infrared thermal imaging data through a multi-modal data acquisition module, and can use a dynamic contour anchoring Mamba network processing module to accurately extract animal contours and use a linear attention mechanism of the network to perform time series modeling on contour sequences, thereby effectively solving the problems of insufficient time series modeling capability and weak cross-frame data correlation in the prior art when processing multi-modal data. The module combination can realize coherent connection of animal behavior features between different frames, greatly improve the behavior feature extraction accuracy, clearly reflect the continuous change process of animal behavior from one time to the next, provide more accurate and complete basic data support for subsequent behavior analysis, and make up for the shortcomings of traditional technologies in dynamic behavior capture.

[0047] The behavior dynamic graph construction module and the behavior dynamic graph deviation evaluation module in the system form a synergy, the former constructs a behavior dynamic graph based on the coordinates of key nodes of the animal body, and the latter calculates the feature deviation value of the real-time graph and the preset standard behavior graph. This design breaks the limitation of single dimension of the existing behavior evaluation system, and no longer simply classifies animal behavior, but realizes detailed evaluation of behavior dynamic deviation through quantitative deviation value, can comprehensively judge the deviation degree of real-time behavior and standard behavior of animals, and overcomes the shortcomings of traditional technology lacking quantitative evaluation of dynamic deviation and being unable to deeply analyze behavior abnormalities, and provides more scientific judgment basis for identifying animal abnormal behavior.

[0048] The smart breeding multi-modal spatio-temporal engine module fuses the spatio-temporal information of multi-modal data and the time sequence features of the dynamic contour anchored Mamba network processing module, generates behavior data associated with time and space, strengthens the correlation and analysis depth between data, solves the problem of insufficient spatio-temporal correlation analysis in the prior art; and the behavior tracking result output module accurately outputs the real-time behavior category and motion trajectory data of the animal individual. The overall system meets the needs of smart breeding for efficient, accurate and deep analysis of animal behavior tracking through the ordered connection and cooperative work of the modules, provides more reliable data support for breeding management decision-making, and fully covers the deficiencies of traditional technology in multi-modal data fusion, behavior feature modeling, evaluation dimension, etc., and promotes the change of breeding animal behavior tracking from extensive to fine.

[0049] In the description of the present application, it should be noted that, unless otherwise explicitly specified and limited, the terms "arrangement", "installation", "connection", "link", "fixation" should be understood broadly, for example, can be fixedly connected, can also be detachably connected, or integrally connected; can be mechanically connected, can also be electrically connected; can be directly connected, can also be indirectly connected through an intermediate medium, can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present application can be understood through specific circumstances.

[0050] Although the embodiments of the present application have been shown and described, those skilled in the art can understand that various equivalent changes, modifications, replacements and variations of the embodiments can be made without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalent ranges.

Claims

1. An artificial intelligence based farmed animal behavior tracking system characterized by, The application relates to an intelligent breeding multi-modal spatio-temporal engine system, which comprises the following modules: a multi-modal data acquisition module for acquiring video streams, audio streams and infrared thermal imaging data of animals in a breeding scene; a dynamic contour anchoring Mamba network processing module for receiving data output by the multi-modal data acquisition module, extracting contours of animal targets through a dynamic contour anchoring unit and modeling contour sequences extracted by using a linear attention mechanism of a Mamba network; a behavior dynamic atlas construction module for receiving time-series contour data output by the dynamic contour anchoring Mamba network processing module and constructing a behavior dynamic atlas based on animal body calibration node coordinates; a behavior dynamic atlas deviation evaluation module for receiving atlas data output by the behavior dynamic atlas construction module, calculating feature deviation values of real-time atlases and preset standard behavior atlases; an intelligent breeding multi-modal spatio-temporal engine module for receiving deviation values output by the behavior dynamic atlas deviation evaluation module, fusing spatio-temporal information of the multi-modal data acquisition module and time-series features of the dynamic contour anchoring Mamba network processing module and generating behavior data associated with space and time; a behavior tracking result output module for receiving spatio-temporal associated behavior data output by the intelligent breeding multi-modal spatio-temporal engine module and outputting real-time behavior categories and motion trajectory data of animal individuals, wherein the multi-modal data acquisition module and the dynamic contour anchoring Mamba network processing module are connected through a data transmission interface, the dynamic contour anchoring Mamba network processing module and the behavior dynamic atlas construction module are connected through a time-series data bus, the behavior dynamic atlas construction module and the behavior dynamic atlas deviation evaluation module are connected through an atlas feature channel, the behavior dynamic atlas deviation evaluation module and the intelligent breeding multi-modal spatio-temporal engine module are connected through a deviation value transmission link, and the intelligent breeding multi-modal spatio-temporal engine module and the behavior tracking result output module are connected through a result output channel.

2. The artificial intelligence-based farmed animal behavior tracking system of claim 1, wherein, In the behavior dynamic atlas deviation evaluation module, the contour time sequence features extracted by the dynamic contour anchor Mamba network and the behavior dynamic atlas features are used to construct a deviation evaluation model, and the model formula is: wherein, is the total deviation value of the behavior dynamic atlas, is the number of time sequence frames processed by the dynamic contour anchor Mamba network, is the number of calibration nodes in the behavior dynamic atlas, is the weight coefficient of the th calibration node in the th frame, is the feature vector of the th calibration node in the th frame output by the dynamic contour anchor Mamba network, is the feature vector of the th calibration node in the th frame in the preset standard behavior atlas, is the spatiotemporal correlation coefficient of the th calibration node in the th frame output by the intelligent breeding multi-modal spatiotemporal engine.

3. The artificial intelligence-based farmed animal behavior tracking system of claim 1, wherein, In the multi-modal spatio-temporal engine module of the smart breeding, the spatio-temporal information of the multi-modal data and the time sequence characteristics of the dynamic contour anchored Mamba network are fused to construct a spatio-temporal feature fusion model, and the formula is: wherein, is a fused spatio-temporal feature matrix, is a fusion weight coefficient, is a time sequence feature matrix output by the dynamic contour anchored Mamba network, is a spatio-temporal feature matrix output by the multi-modal data acquisition module, is a matrix Hadamard product operation, is a gradient matrix of the time sequence characteristics of the dynamic contour anchored Mamba network, is a gradient matrix of the multi-modal spatio-temporal characteristics, is a matrix multiplication operation.

4. The artificial intelligence-based farmed animal behavior tracking system of claim 1, wherein, In the dynamic contour anchoring Mamba network processing module, when modeling the time series of animal contours, an improved Mamba network state update formula is used: ,in, For the first The network hidden state vector of a frame. Here is the state transition matrix. For the first The network hidden state vector of a frame. For the input weight matrix, Given the input feature transformation matrix, For the first The contour feature vector of the frame. For bias vectors, For dynamic contour needle fixed term weight coefficient, For dynamic contour anchoring function, For the first The set of coordinates of the frame's outline anchor points.

5. The artificial intelligence-based farmed animal behavior tracking system of claim 1, wherein, In the behavioral dynamic graph construction module, the coordinates of the calibration nodes are calculated based on the contour data output by the dynamic contour anchoring Mamba network. The formula is as follows: ,in, For the first Frame number The coordinates of each calibration node, To anchor the number of contour pixel neighbors in the output of the Mamba network for dynamic contours, For the first Frame number The coordinates of each contour pixel. The contour feature enhancement coefficient, To use dynamic contour anchoring to generate convolutional kernels for Mamba networks right The result of the convolution operation.

6. The artificial intelligence-based farmed animal behavior tracking system of claim 1, wherein, In the behavior tracking results output module, the probability of animal behavior categories is calculated by combining the output data of the smart farming multimodal spatiotemporal engine with the deviation assessment results of the behavior dynamic map. The formula is as follows: ,in, Animals belong to the category The probability, This refers to the spatiotemporal feature vector output by the multimodal spatiotemporal engine for intelligent aquaculture. The deviation value output by the behavioral dynamic graph deviation assessment module. For category The feature weight vector, For category The bias term, The deviation value influence coefficient. This represents the total number of behavior categories.

7. The artificial intelligence-based farmed animal behavior tracking system of claim 1, wherein, The behavior dynamic graph construction module comprises: a calibration node candidate region generation unit receiving time sequence contour data output by the dynamic contour anchoring Mamba network processing module, screening a region with a gray value change rate greater than a preset threshold as a calibration node candidate region by traversing the contour image through a sliding window, labeling a bounding box of the candidate region, and recording center coordinates and size information of each candidate region; a calibration node feature extraction unit extracting features of an image region in the bounding box labeled by the calibration node candidate region generation unit, performing feature mapping on the region image using a 3*3 convolution kernel, extracting texture features and edge features, concatenating the extracted features into a feature vector, associating the time sequence features output by the dynamic contour anchoring Mamba network, and obtaining a time sequence feature vector of each candidate region; a calibration node matching unit receiving the time sequence feature vector output by the calibration node feature extraction unit, calculating the cosine similarity of the candidate region feature vectors between adjacent frames, determining the candidate regions with a similarity greater than a preset threshold as the same calibration node, and updating the calibration node coordinates; and a graph topology structure construction unit calculating the Euclidean distance between adjacent calibration nodes according to the calibration node coordinates determined by the calibration node matching unit, constructing an adjacency matrix based on the distance value, combining the adjacency matrix and the calibration node time sequence features, and forming a topology structure of the behavior dynamic graph.

8. The artificial intelligence-based farmed animal behavior tracking system of claim 1, wherein, The behavior dynamic graph deviation evaluation module comprises: a standard graph feature library construction unit collecting multi-modal data under a normal behavior state of the farmed animals, generating a standard behavior graph through the dynamic contour anchoring Mamba network processing and the behavior dynamic graph construction module, extracting calibration node features, adjacency matrix features and time sequence change features of the standard graph, establishing a standard graph feature library and storing a feature threshold range; a real-time graph feature extraction unit receiving a real-time behavior dynamic graph output by the behavior dynamic graph construction module, extracting calibration node features, adjacency matrix features and time sequence change features of the real-time graph according to the feature extraction mode of the standard graph feature library construction unit, and forming a real-time graph feature vector; a feature deviation calculation unit comparing the real-time graph feature vector output by the real-time graph feature extraction unit with the standard feature vector in the standard graph feature library dimension by dimension, calculating the absolute deviation value and the relative deviation value of each dimension, and obtaining a feature deviation total sum by using a weighted summation method; and a deviation level determination unit determining the deviation level of the real-time behavior dynamic graph by comparing the feature deviation total sum output by the feature deviation calculation unit with the deviation threshold values of different levels according to a preset deviation level division rule, and outputting a deviation level result.

9. The artificial intelligence-based farmed animal behavior tracking system of claim 1, wherein, The intelligent breeding multi-modal space-time engine module comprises: a multi-modal data space-time alignment unit receiving video stream, audio stream and infrared thermal imaging data output by a multi-modal data acquisition module, extracting timestamp information and spatial coordinate information of each modal data, taking the timestamp of the video stream as a reference, performing time synchronization on the audio stream and the infrared thermal imaging data, calibrating the spatial coordinates of each modal data based on the spatial coordinate system of the breeding scene, and performing space-time alignment of the multi-modal data; a time sequence feature fusion unit receiving time sequence contour features output by a dynamic contour anchoring Mamba network processing module and aligned multi-modal data output by the multi-modal data space-time alignment unit, using an attention mechanism to distribute weights to time sequence features of each modal data, multiplying the weights and the time sequence features of the corresponding modal, and summing to obtain fused time sequence features; a spatial feature enhancement unit processing spatial features output by the multi-modal data space-time alignment unit, using a hollow convolution to expand the receptive field, extracting spatial features of different scales, splicing the spatial features of different scales, combining contour spatial information output by the dynamic contour anchoring Mamba network, enhancing the spliced spatial features, and obtaining enhanced spatial features; and a space-time correlation modeling unit receiving fused time sequence features output by the time sequence feature fusion unit and enhanced spatial features output by the spatial feature enhancement unit, using a graph neural network to correlate and model the space-time features, taking the time sequence features as node features and the spatial features as edge features, constructing a space-time correlation graph, updating the node features through graph convolution operation, and generating behavior data correlated in space-time.

10. A method for tracking behavior of farmed animals based on artificial intelligence, characterized by, Comprise: Step S1, a dynamic multi-modal data acquisition device acquires video stream, audio stream and infrared thermal imaging data of animals in a breeding scene, encapsulates the acquired data according to a preset format, and sends the data to a dynamic contour anchoring Mamba network processing module through a data transmission link; step S2, the dynamic contour anchoring Mamba network processing module receives the encapsulated data, decapsulates the data, extracts frame images in the video stream, performs contour extraction on animal targets in the frame images through a dynamic contour anchoring unit, obtains an animal contour pixel coordinate set, inputs the contour pixel coordinate set into a Mamba network, models the contour pixel coordinate set in time sequence by using a linear attention mechanism of the Mamba network, and outputs a time sequence contour feature sequence; Step S3, the behavior dynamic atlas construction module receives the time sequence contour feature sequence, identifies the animal body calibration nodes based on the time sequence contour feature sequence, calculates the coordinates of each calibration node in different frames, calculates the distance and angle between adjacent calibration nodes according to the calibration node coordinates, and constructs the behavior dynamic atlas; Step S4, the behavior dynamic atlas deviation evaluation module receives the behavior dynamic atlas, calls the preset standard behavior atlas feature library, calculates the feature deviation value of the real-time behavior dynamic atlas and the standard behavior atlas, normalizes the feature deviation value, and obtains the deviation evaluation result; Step S5, the intelligent breeding multi-modal space-time engine module receives the deviation evaluation result, the time sequence contour feature sequence and the multi-modal data, fuses the space-time information of the multi-modal data and the time sequence contour feature sequence, constructs a space-time correlation model, inputs the deviation evaluation result into the space-time correlation model, and generates the space-time correlated behavior data; Step S6, the behavior tracking result output module receives the space-time correlated behavior data, analyzes the behavior data, identifies the behavior category of the animal individual, draws the animal motion trajectory based on the coordinate information in the behavior data, and outputs the behavior category and the motion trajectory to the display terminal and the storage device according to the preset format.

Citation Information

Cited By

  • Livestock and poultry breeding whole-process quality information management method based on multi-mode AI

    CN121458241A

  • Rapid zebra fish abnormal behavior recognition system based on computer vision

    CN121505695A

  • A rapid identification system for abnormal behavior in zebrafish based on computer vision

    CN121505695B