A hu sheep breeding management system based on multi-modal data fusion
The multimodal data fusion system solves the problems of lagging health status monitoring and data silos in traditional Hu sheep breeding and management, and realizes accurate assessment and dynamic optimization of the health status of Hu sheep, thereby improving the traceability of the breeding process and the scientific nature of decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YIMU DIGITAL (HANGZHOU) TECHNOLOGY CO LTD
- Filing Date
- 2026-04-20
- Publication Date
- 2026-06-23
Smart Images

Figure CN122264978A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of smart agriculture and animal husbandry, and in particular to a lake sheep breeding management system based on multimodal data fusion. Background Technology
[0002] Hu sheep are an important sheep breed in my country, characterized by high fertility, tolerance to heat and humidity, and suitability for pen rearing. In recent years, the proportion of large-scale farming has been continuously increasing. Traditional Hu sheep farming management mainly relies on manual inspection and experience-based judgment, which has the following drawbacks: 1. Lagging monitoring of individual health status, making it difficult to detect diseases in their early stages; 2. Inefficient decision-making regarding feeding and environmental control, relying on fixed patterns and failing to dynamically optimize based on individual differences and real-time environmental conditions; 3. Scattered and easily tampered production data, making it difficult to achieve reliable traceability throughout the entire process, affecting product quality certification and brand value enhancement. Existing technologies include some livestock management systems using single sensors (such as temperature and humidity sensors) or video monitoring. However, these systems typically suffer from single-modality and data silos, failing to comprehensively and accurately reflect the overall health status of livestock. Furthermore, existing systems have simple decision-making logic, lacking deep analysis capabilities based on multi-dimensional data fusion, and their fixed models make them difficult to adapt to dynamic changes in the farm environment and emerging anomalies. Data storage is also mostly centralized, posing a risk of tampering. Therefore, there is an urgent need for a lake sheep breeding management system that can integrate multimodal information, achieve intelligent analysis and decision-making, and ensure the authenticity and reliability of data. Summary of the Invention
[0003] The purpose of this invention is to provide a Hu sheep breeding management system based on multimodal data fusion. This invention can achieve precise, real-time, and dynamic management of the health status of individual Hu sheep, improving the traceability and scientific basis of the breeding process.
[0004] The technical solution of this invention: A lake sheep breeding management system based on multimodal data fusion, comprising:
[0005] The multimodal data acquisition unit is used to simultaneously acquire video image sequences of Hu sheep, audio of the breeding environment, time-series data of temperature, humidity and ammonia concentration in the barn, and individual identification data of electronic ear tags;
[0006] An adaptive preprocessing unit, connected to the multimodal data acquisition unit, is used to align the acquired data from different modalities to a unified time axis and perform weighted standardization.
[0007] A cross-modal attention fusion analysis unit, connected to the adaptive preprocessing unit, is used to perform fusion analysis on the standardized multimodal data and output a fusion feature vector and health index representing the health status of Hu sheep.
[0008] The dynamic optimization decision unit, connected to the cross-modal attention fusion analysis unit, is used to generate feeding control instructions, disease early warning information, and environmental adjustment instructions based on the health index and environmental parameters.
[0009] The aforementioned multimodal data fusion-based lake sheep breeding management system includes an adaptive preprocessing unit comprising:
[0010] The timestamp alignment module is used to align data from different modalities to a unified time axis using an interpolation method based on the sampling frequency of each acquisition device. The alignment formula is as follows:
[0011] ;
[0012] In the formula, For the first The original sampling time points of each modality, The sampling frequency for the corresponding mode. The basic sampling frequency set for the system;
[0013] The modal normalization module is used to standardize data by assigning adaptive weights to each modality and performing weighted fusion. The standardization formula is as follows:
[0014] ;
[0015] In the formula, , , They are Time-aligned visual, auditory, and sensor modal data. , and These are the weights of each modal data, determined by the signal-to-noise ratio, and the sum of the three is 1.
[0016] The aforementioned multimodal data fusion-based sheep farming management system includes a cross-modal attention fusion analysis unit comprising a multimodal Transformer network based on an attention mechanism, which internally contains:
[0017] Visual encoders, auditory encoders, and sensor encoders are used to extract features from video images, ambient audio, temperature and humidity, and ammonia concentration data, respectively.
[0018] A cross-modal attention fusion layer is used to compute and fuse features from different encoders, where for the ... Features of each modality Its attention weight The calculation satisfies:
[0019] ;
[0020] in, , For the first to participate in the normalization calculation Feature vectors of each modality For the number of modes, For global context features, The weight matrix is a learnable matrix. For bias terms; It is the transpose symbol;
[0021] The feature vector output by fusion ;
[0022] And, a health index quantification model, used to calculate the health index based on the fused feature vector, the calculation formula being:
[0023] ;
[0024] Among them, constraints ; , , and All of these are coefficients from a pre-defined health index quantification model, and their sum is 1. For behavioral scores calculated based on visual data, The audio anomaly score is calculated based on the audio data. Environmental suitability is calculated based on temperature and humidity data. This represents the real-time ammonia concentration. This is the preset safe concentration threshold for ammonia.
[0025] The aforementioned multimodal data fusion-based lake sheep breeding management system includes a dynamic optimization decision-making unit comprising:
[0026] The feeding control subunit is used to dynamically calculate and issue feed dosage control instructions based on the health index, indoor temperature, and daily weight gain deviation. The formula for calculating the feed dosage is:
[0027] ;
[0028] In the formula, Based on the amount of raw materials, This represents the deviation between the actual daily weight gain and the target daily weight gain. This refers to the real-time temperature inside the building. The optimal temperature for the growth of Hu sheep and All are control coefficients;
[0029] The disease early warning subunit is used to determine the degree of behavioral abnormality. Voiceprint anomaly Environmental risk level Weighted composite risk value:
[0030] ;
[0031] In the formula, , and All are risk weighting coefficients;
[0032] Based on weighted composite risk value Does it exceed the preset threshold? To generate and push out epidemic early warning information;
[0033] The environmental control subunit is used to generate environmental control instructions based on the temperature, humidity and ammonia concentration data inside the building.
[0034] The aforementioned multimodal data fusion-based lake sheep breeding management system also includes:
[0035] The temporal causal correlation analysis module is used to extract causal graph features of abnormal events from aligned multimodal temporal data using Granger causality tests or structural causality models. The causal graph features include the causal delay and propagation direction of the abnormality between different modalities. These features are fed into the cross-modal attention fusion analysis unit as auxiliary inputs, or used to locate the source modality of abnormal propagation when the epidemic early warning subunit triggers an early warning.
[0036] The aforementioned multimodal data fusion-based lake sheep breeding management system also includes:
[0037] The model self-iterative update unit, connected to the cross-modal attention fusion analysis unit and the time-series causal correlation analysis module, is used to adjust the attention weight matrix in the cross-modal attention fusion layer based on the feedback from zookeepers, actual disease occurrences, and time window data of the abnormal source modality located by the time-series causal correlation analysis module. This adjustment employs a combination of small-sample fine-tuning and experience replay. Bias terms and the coefficients in the aforementioned health index quantification model , , and Perform incremental learning and updates.
[0038] The aforementioned multimodal data fusion-based lake sheep breeding management system also includes:
[0039] The blockchain storage traceability unit is used to package individual electronic ear tag identifiers, health index sequences, early warning records, feeding records, and environmental data into data blocks, which are then connected and stored using a hash chain to form an immutable traceability chain for the entire breeding process.
[0040] The aforementioned multimodal data fusion-based lake sheep breeding management system also includes:
[0041] The digital twin simulation subunit is connected to the blockchain storage and traceability unit and the model self-iterative update unit. It is used to construct a virtual sheep breeding environment based on historical data and current model parameters, simulate the impact of different feeding strategies and environmental control schemes on the health index, and feed the simulation results back to the dynamic optimization decision-making unit as a decision reference.
[0042] The aforementioned multimodal data fusion-based sheep farming management system further includes, in its cross-modal attention fusion analysis unit:
[0043] The individual record generation module is used to bind the electronic ear tag of each Hu sheep to the trajectory of its corresponding health index change, the fused feature vector, and the associated causal graph feature snapshot, to generate a lifelong digital health record for the Hu sheep from birth to slaughter.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] 1. This invention effectively integrates multi-source heterogeneous data from visual, auditory, and sensor sources through a cross-modal attention mechanism, achieving a comprehensive and accurate assessment of the health status of Hu sheep and overcoming the limitations of single-modal information. Based on real-time calculated health indices and environmental parameters, this invention can dynamically adjust feeding strategies, issue precise disease warnings, and regulate the barn environment, realizing closed-loop intelligent management from perception to decision-making. Furthermore, through temporal causal correlation analysis, this invention not only provides early warnings of anomalies but also reveals the paths of anomaly generation and propagation, enhancing the interpretability of system decisions and helping farmers gain a deeper understanding of the root causes of problems.
[0046] 2. This invention utilizes blockchain technology to store key data, ensuring the authenticity and immutability of data throughout the entire breeding process, providing a reliable data foundation for product traceability and quality certification. The system of this invention possesses self-iterative model updates and digital twin simulation capabilities, enabling it to continuously optimize the model using new data and feedback. By simulating and rehearsing decision-making effects, the system has the ability to continuously learn and optimize, adapting to complex actual breeding environments. Attached Figure Description
[0047] Figure 1 This is a schematic diagram of the system structure of the present invention;
[0048] Figure 2 This is a schematic diagram of the adaptive preprocessing unit;
[0049] Figure 3 This is a schematic diagram of the cross-modal attention fusion analysis unit;
[0050] Figure 4 This is a schematic diagram of the structure of a dynamic optimization decision-making unit. Detailed Implementation
[0051] The following detailed description of specific embodiments, in conjunction with the technical solution of the present invention, is provided. These embodiments are only used to illustrate the present invention and are not intended to limit the scope of protection of the present invention.
[0052] Example: A lake sheep breeding management system based on multimodal data fusion, such as... Figure 1 As shown, the system consists of a multimodal data acquisition unit, an adaptive preprocessing unit, a cross-modal attention fusion analysis unit, a dynamic optimization decision-making unit, a temporal causal correlation analysis module, a model self-iterative update unit, a blockchain storage and traceability unit, and a digital twin simulation subunit.
[0053] The system workflow is as follows: A multimodal acquisition unit simultaneously collects four types of data: video, audio, environmental sensor data, and electronic ear tags; an adaptive preprocessing unit completes timestamp alignment and weighted standardization; a cross-modal attention fusion analysis unit performs feature extraction, cross-modal fusion, and health index quantification; a dynamic optimization decision-making unit outputs feeding, early warning, and environmental control instructions based on health indices and environmental parameters; a time-series causal correlation analysis module locates abnormal causal relationships and assists in decision-making; a model self-iterative update unit continuously optimizes model parameters based on on-site data; a blockchain storage and traceability unit achieves full-process data storage and traceability; and a digital twin simulation subunit completes strategy simulation and provides feedback for optimization decisions. Detailed descriptions of each unit are as follows:
[0054] 1. Multimodal data acquisition unit, used to achieve real-time perception of individual sheep and their breeding environment, including four types of acquisition channels:
[0055] 1.1 Video Image Sequence Acquisition
[0056] The system employs 1080P infrared dome network cameras with a horizontal field of view ≥90°, a frame rate of 15fps, and supports day / night mode switching. One camera is deployed per 30㎡ sheepfold to ensure comprehensive coverage. Data collected includes the sheep's postures and physical characteristics, such as standing, lying down, walking, feeding, drinking, rumination, fighting, lameness, and lethargy.
[0057] 1.2 Audio Collection of the Aquaculture Environment
[0058] A high-sensitivity microphone with a frequency response of 20Hz-20kHz, sensitivity of -38dB, sampling rate of 44.1kHz, and mono 16-bit acquisition is employed. The microphone is used to capture health-related vocal characteristics of sheep, such as coughing, panting, abnormal calls, and respiratory noises. It also collects noise from equipment such as fans and feed lines in the sheepfold for interference removal.
[0059] 1.3 Time-series data collection of the building's internal environment
[0060] Employing an integrated temperature and humidity sensor and an electrochemical ammonia gas sensor:
[0061] Temperature measurement range: 0–50℃, accuracy ±0.5℃;
[0062] Humidity measurement range: 0–95% RH, accuracy ±3% RH;
[0063] Ammonia sensor range: 0–100ppm, accuracy ±1ppm, response time ≤3s;
[0064] The sensor sampling frequency is uniformly set to 1 time / minute, arranged diagonally within the cell, and the average of multiple data points is taken before entering the system.
[0065] 1.4. Collection of Individual Identification Data for Electronic Ear Tags
[0066] RFID low-frequency electronic ear tags are used, operating at a frequency of 125kHz with a reading distance of 0-30cm. Each sheep wears a unique ear tag with a specific ID number. Readers are installed at feeding areas and passageways to automatically read individual tags and link them to basic information such as date of birth, sex, parity, weight, entry time, and immunization records, achieving precise identification with a "one sheep, one file" system. All data collection devices are synchronized via an NTP network time server, with a time error ≤1s, ensuring strict spatiotemporal alignment of multimodal data.
[0067] 2. An adaptive preprocessing unit, connected to the multimodal data acquisition unit, is used to align the acquired data from different modalities to a unified time axis and perform weighted standardization. For example... Figure 2 As shown, it specifically includes:
[0068] 2.1 Timestamp Alignment Module
[0069] Based on the differences in sampling frequencies across modes, a linear interpolation method is used to unify all data onto the same time axis. The system sets a basic sampling frequency. =1Hz, the alignment formula is as follows:
[0070] ;
[0071] In the formula, For the first The original sampling time points of each modality, The sampling frequency for the corresponding mode. The basic sampling frequency set for the system;
[0072] Through the above processing, high-frequency data such as video and audio are downsampled to 1 frame / second, while low-frequency data from the sensor are upsampled to 1 point / second, achieving full-modal timing uniformity.
[0073] 2.2 Modal Normalization Module
[0074] The aligned data is adaptively weighted and standardized, with the weights determined by the real-time signal-to-noise ratio of each modality, ensuring that highly reliable data dominates.
[0075] Standardized formula:
[0076] ;
[0077] In the formula, , , They are Time-aligned visual, auditory, and sensor modal data. , and These are the weights of each modal data, determined by the signal-to-noise ratio, and the sum of the three is 1.
[0078] In this embodiment, when there is insufficient light at night, the visual weight is automatically reduced; when there is a lot of noise in the building, the audio weight is reduced; when the sensor drifts, the sensor weight is reduced, thereby improving the robustness of the system in complex field environments.
[0079] 3. Cross-modal attention fusion analysis unit
[0080] The cross-modal attention fusion analysis unit is the core of the system of this invention. It is connected to the adaptive preprocessing unit and is used to perform fusion analysis on the standardized multimodal data, outputting a fusion feature vector and health index characterizing the health status of the sheep. For example... Figure 3 As shown, it specifically includes:
[0081] 3.1 Multimodal encoder
[0082] The visual encoder extracts spatial and temporal behavioral features from video image sequences. The process is as follows: The spatial feature extraction backbone uses a ResNet50 network. The input is a 224×224×3 standardized video frame, which passes through conv1 (7×7, stride 2), bn1, ReLU, maxpool layers, and four residual blocks (3, 4, 6, 3 layers respectively), outputting a 7×7×2048 spatial feature map. After global average pooling, a 2048-dimensional spatial feature vector is obtained. Temporal feature modeling uses a two-layer bidirectional LSTM with a hidden layer dimension of 512. The input is a sequence of 16 consecutive ResNet50 feature sequences. The bidirectional outputs are concatenated to a dimension of 1024, and a 0.3 Dropout is used to prevent overfitting. The dimension is mapped to 512 through a fully connected layer, and the GELU activation function is used, finally outputting a 512-dimensional visual feature vector. .
[0083] Auditory encoder: Used to extract voiceprint anomaly features from environmental audio. The specific process is as follows: 1. Audio preprocessing: Input a mono 16kHz, 16-bit audio segment, after 512-point framing and 256-point frame shifting, a 64×T Mel spectrogram is generated through a 64-dimensional Mel filter bank; 2. Voiceprint feature extraction CNN: Conv1 (32×3×3, stride 1, padding=1, BN+ReLU), 2×2 max pooling, Conv2 (64×3×3, stride 1, padding=1, BN+ReLU), 2×2 max pooling, Conv3 (128×3×3, stride 1, padding=1, BN+ReLU), and output 128-dimensional local features after global average pooling; 3. Temporal context modeling uses a 1-layer GRU with 256 hidden dimensions, outputting 256-dimensional temporal acoustic features; 4. Mapped to 512 dimensions through a fully connected layer, using GELU. The activation function ultimately outputs a 512-dimensional auditory feature vector. .
[0084] Sensor encoder: Used for feature embedding of environmental time-series data such as temperature, humidity, and ammonia concentration. The specific process is as follows: 1. Temporal location embedding: Input an environmental time-series sequence of length 32, coupled with a 32×64 learnable location code; 2. Environmental feature encoding MLP: Linear1 (4→64, BN+GELU), Linear2 (64→128, BN+GELU), Linear3 (128→256, BN+GELU); 3. Temporal aggregation: Global average pooling and max pooling are used to concatenate the data to obtain 512-dimensional features; 4. Output a 512-dimensional sensor feature vector through a fully connected layer. .
[0085] 3.2 Cross-modal attention fusion layer
[0086] The cross-modal attention fusion layer is used to compute and fuse features from different encoders. The specific steps are as follows:
[0087] Global context feature calculation: , It is a 512-dimensional global context feature;
[0088] For the Features of each modality Its attention weight The calculation satisfies:
[0089] ;
[0090] in, , For the first to participate in the normalization calculation Feature vectors of each modality For the number of modes, For global context features, The weight matrix is a learnable matrix. For bias terms; It is the transpose symbol;
[0091] The feature vector output by fusion ;
[0092] 3.3 Health Index Quantification Model
[0093] The health index is calculated based on the fused feature vector, and the formula is as follows:
[0094] ;
[0095] Among them, constraints ; , , and All of these are coefficients of a preset health index quantification model, and their sum is 1, preferably 0.4, 0.3, 0.2 and 0.1; The behavioral score is calculated based on visual data. The calculation method is to extract behavioral features from video images, such as standing, eating, ruminating, walking, lying down, mental state, limping, coughing and wheezing posture, etc. Normal behaviors are assigned high scores and abnormal behaviors are assigned low scores. A weighted scoring method is used to obtain the comprehensive behavioral score. The audio anomaly score is calculated based on audio data. The calculation method is to extract Mel-spectrum features from the audio to identify cough, wheezing, abnormal noises, and respiratory murmurs. Each type of anomaly is scored, and the higher the degree of anomaly, the lower the score. The environmental suitability score is calculated based on temperature and humidity data. The further the temperature and humidity deviate from the suitable temperature and humidity data, the lower the score. This represents the real-time ammonia concentration. This is the preset safe concentration threshold for ammonia.
[0096] When the health index is ≥80, the person is in good health; when 0.6 ≤ the health index < 0.8, attention is needed; when the health index < 0.6, it indicates an abnormal warning.
[0097] 3.4 Individual Profile Generation Module
[0098] The individual record generation module uniquely identifies each Hu sheep with its electronic ear tag and binds it to real-time health index, historical change trajectory, multimodal fusion feature vector, causal characteristics of abnormal events, environmental data, feeding records, and early warning records to generate a lifelong digital health record from birth to slaughter. It supports single sheep retrieval, time-series backtracking, report export, and integration with third-party platforms.
[0099] 4. Dynamically optimize decision-making unit
[0100] The dynamic optimization decision-making unit is connected to the cross-modal attention fusion analysis unit and is used to generate feeding control instructions, disease early warning information, and environmental adjustment instructions based on the health index and environmental parameters. Figure 4 As shown, it specifically includes:
[0101] 4.1 Feeding Regulation Subunit
[0102] The feeding control subunit dynamically calculates and issues feed dosage control instructions based on the health index, indoor temperature, and daily weight gain deviation. These instructions are then sent to the automatic feeder to achieve individualized, time-segmented, and on-demand feeding, reducing feed waste and improving daily weight gain stability. The formula for calculating the feed dosage is as follows:
[0103] ;
[0104] In the formula, The basic feed amount is set at 2.5% of body weight. This represents the deviation between the actual daily weight gain and the target daily weight gain (i.e., actual daily weight gain - target daily weight gain). This refers to the real-time temperature inside the building. The optimal temperature for the growth of Hu sheep is set at 20℃. and Both are control coefficients, with values of 0.8 and 0.5 respectively.
[0105] 4.2 Disease Early Warning Subunit
[0106] The disease early warning subunit is used to determine the degree of behavioral abnormality. Voiceprint anomaly Environmental risk level Weighted composite risk value:
[0107] ;
[0108] In the formula, , and All are risk weighting coefficients, with values of 0.5, 0.3, and 0.2 respectively;
[0109] Based on weighted composite risk value Does it exceed the preset threshold? To generate and push out epidemic early warning information; for example, when When the value is greater than 0.7, the system will push an alert to the app / SMS. When the value is greater than 0.9, an emergency red alert is triggered, prompting keepers to quickly isolate and examine the animal.
[0110] 4.3 Environmental Regulation Subunit
[0111] The environmental control subunit is used to generate environmental control instructions based on indoor temperature, humidity, and ammonia concentration data.
[0112] Ammonia concentration > 10 ppm: Activate negative pressure fan + spray deodorization;
[0113] Temperature > 28℃: Turn on the cooling water curtain;
[0114] Temperature <10℃: Start the heating and insulation equipment;
[0115] Relative humidity > 85%: Increase ventilation and dehumidification.
[0116] 5. Time-series causal relationship analysis module
[0117] The temporal causal correlation analysis module is used to extract causal graph features of anomalous events from aligned multimodal temporal data using Granger causality tests or structural causality models. These causal graph features include the causal delay and propagation direction of the anomalous event across different modalities. These features are fed as auxiliary input into the cross-modal attention fusion analysis unit, or used to locate the source modality of anomalous propagation when the disease early warning subunit triggers an alert. For example: elevated ammonia levels → lethargic behavior → appearance of coughing voiceprints; when an alert is triggered, the starting point of the anomalous event is directly located, helping zookeepers to respond quickly.
[0118] 6. Model self-iterative update unit
[0119] The model self-iterative update unit is connected to the cross-modal attention fusion analysis unit and the time-series causal correlation analysis module. It is used to adjust the attention weight matrix in the cross-modal attention fusion layer based on the feeder's operational feedback, actual disease occurrences, and the time window data of the abnormal source modality located by the time-series causal correlation analysis module. This adjustment employs a combination of small-sample fine-tuning and experience replay. Bias terms and the coefficients in the aforementioned health index quantification model , , and Incremental learning and updates are performed. For example, input data includes zookeeper feedback, actual disease records, and abnormal causal time window data. Updates are conducted using small-sample fine-tuning and experience playback, with weekly incremental updates and monthly full optimization. This allows for continuous adaptation to new diseases, new environments, and new feeding methods, maintaining high accuracy over the long term.
[0120] 7. Blockchain storage and traceability unit
[0121] The blockchain storage traceability unit is used to package individual electronic ear tag identifiers, health index sequences, early warning records, feeding records, and environmental data into data blocks, which are then connected and stored using a hash chain to form an immutable traceability chain for the entire breeding process, meeting the requirements for pollution-free and green product certification and market traceability. For example, in this embodiment, the following data will be packaged and uploaded to the blockchain periodically: electronic ear tag ID, individual basic information, health index time series, abnormal early warning records, feeding amount, feeding time, feed batch, temperature and humidity, and historical ammonia concentration curves.
[0122] 8. Digital Twin Simulation Subunit
[0123] The digital twin simulation subunit, connected to the blockchain storage and traceability unit and the model self-iterative update unit, is used to construct a virtual sheep farming environment based on historical data and current model parameters. It simulates the impact of different feeding strategies and environmental control schemes on the health index and feeds the simulation results back to the dynamic optimization decision-making unit as a decision-making reference. This is used for strategy pre-playing and optimization, reducing on-site trial-and-error costs. For example, managers can simulate strategies such as "increasing feed intake by 10%" or "controlling ammonia concentration below 15 ppm" and observe the distribution changes of the health index H in the virtual environment. Simulation results show that reducing ammonia concentration can increase the average health index value of the herd by 5.2. This result is fed back to the dynamic optimization decision-making unit as a decision-making reference to assist in actual production control.
[0124] In summary, this system can reliably achieve precise health monitoring, intelligent feeding, early warning of diseases, automatic environmental control, and reliable data traceability for Hu sheep, significantly improving breeding efficiency and economic benefits.
[0125] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A lake sheep breeding management system based on multimodal data fusion, characterized in that, include: The multimodal data acquisition unit is used to simultaneously acquire video image sequences of Hu sheep, audio of the breeding environment, time-series data of temperature, humidity and ammonia concentration in the barn, and individual identification data of electronic ear tags; An adaptive preprocessing unit, connected to the multimodal data acquisition unit, is used to align the acquired different modal data to a unified time axis and perform weighted standardization. A cross-modal attention fusion analysis unit, connected to the adaptive preprocessing unit, is used to perform fusion analysis on the standardized multimodal data and output a fusion feature vector and health index representing the health status of Hu sheep. The dynamic optimization decision unit, connected to the cross-modal attention fusion analysis unit, is used to generate feeding control instructions, disease early warning information, and environmental adjustment instructions based on the health index and environmental parameters.
2. The lake sheep breeding management system based on multimodal data fusion according to claim 1, characterized in that, The adaptive preprocessing unit includes: The timestamp alignment module is used to align data from different modalities to a unified time axis using an interpolation method based on the sampling frequency of each acquisition device. The alignment formula is as follows: ; In the formula, For the first The original sampling time points of each modality, The sampling frequency for the corresponding mode. The basic sampling frequency set for the system; The modal normalization module is used to standardize data by assigning adaptive weights to each modality and performing weighted fusion. The standardization formula is as follows: ; In the formula, , , They are Time-aligned visual, auditory, and sensor modal data. , and These are the weights of each modal data, determined by the signal-to-noise ratio, and the sum of the three is 1.
3. The lake sheep breeding management system based on multimodal data fusion according to claim 1, characterized in that, The cross-modal attention fusion analysis unit includes a multimodal Transformer network based on an attention mechanism, which internally contains: Visual encoders, auditory encoders, and sensor encoders are used to extract features from video images, ambient audio, temperature and humidity, and ammonia concentration data, respectively. A cross-modal attention fusion layer is used to compute and fuse features from different encoders, where for the ... Features of each modality Its attention weight The calculation satisfies: ; in, , For the first to participate in the normalization calculation Feature vectors of each modality For the number of modes, For global context features, The weight matrix is a learnable matrix. For bias terms; It is the transpose symbol; The feature vector output by fusion ; And, a health index quantification model, used to calculate the health index based on the fused feature vector, the calculation formula being: ; Among them, constraints ; , , and All of these are coefficients from a pre-defined health index quantification model, and their sum is 1. For behavioral scores calculated based on visual data, The audio anomaly score is calculated based on the audio data. Environmental suitability is calculated based on temperature and humidity data. This represents the real-time ammonia concentration. This is the preset safe concentration threshold for ammonia.
4. The lake sheep breeding management system based on multimodal data fusion according to claim 1, characterized in that, The dynamic optimization decision-making unit includes: The feeding control subunit is used to dynamically calculate and issue feed dosage control instructions based on the health index, indoor temperature, and daily weight gain deviation. The formula for calculating the feed dosage is: ; In the formula, Based on the amount of raw materials, This represents the deviation between the actual daily weight gain and the target daily weight gain. This refers to the real-time temperature inside the building. The optimal temperature for the growth of Hu sheep and All are control coefficients; The disease early warning subunit is used to determine the degree of behavioral abnormality. Voiceprint anomaly Environmental risk level Weighted composite risk value: ; In the formula, , and All are risk weighting coefficients; Based on weighted composite risk value Does it exceed the preset threshold? To generate and push out epidemic early warning information; The environmental control subunit is used to generate environmental control instructions based on the temperature, humidity and ammonia concentration data inside the building.
5. The lake sheep breeding management system based on multimodal data fusion according to claim 3 or 4, characterized in that, Also includes: The temporal causal correlation analysis module is used to extract causal graph features of abnormal events from aligned multimodal temporal data using Granger causality tests or structural causality models. The causal graph features include the causal delay and propagation direction of the abnormality between different modalities. These features are fed into the cross-modal attention fusion analysis unit as auxiliary inputs, or used to locate the source modality of abnormal propagation when the epidemic early warning subunit triggers an early warning.
6. The lake sheep breeding management system based on multimodal data fusion according to claim 5, characterized in that, Also includes: The model self-iterative update unit, connected to the cross-modal attention fusion analysis unit and the time-series causal correlation analysis module, is used to adjust the attention weight matrix in the cross-modal attention fusion layer based on the feedback from zookeepers, actual disease occurrences, and time window data of the abnormal source modality located by the time-series causal correlation analysis module. This adjustment employs a combination of small-sample fine-tuning and experience replay. Bias terms and the coefficients in the aforementioned health index quantification model , , and Perform incremental learning and updates.
7. The lake sheep breeding management system based on multimodal data fusion according to claim 6, characterized in that, Also includes: The blockchain storage traceability unit is used to package individual electronic ear tag identifiers, health index sequences, early warning records, feeding records, and environmental data into data blocks, which are then connected and stored using a hash chain to form an immutable traceability chain for the entire breeding process.
8. The lake sheep breeding management system based on multimodal data fusion according to claim 7, characterized in that, Also includes: The digital twin simulation subunit is connected to the blockchain storage and traceability unit and the model self-iterative update unit. It is used to construct a virtual sheep breeding environment based on historical data and current model parameters, simulate the impact of different feeding strategies and environmental control schemes on the health index, and feed the simulation results back to the dynamic optimization decision-making unit as a decision reference.
9. The lake sheep breeding management system based on multimodal data fusion according to claim 8, characterized in that, The cross-modal attention fusion analysis unit also includes: The individual record generation module is used to bind the electronic ear tag of each Hu sheep to the trajectory of its corresponding health index change, the fused feature vector, and the associated causal graph feature snapshot, to generate a lifelong digital health record for the Hu sheep from birth to slaughter.