Cattle type state analysis method and device, computer equipment and storage medium

By building a multimodal decision model and an edge-cloud collaborative decision-making mechanism, integrating voiceprint, image and temperature data, the problem of isolated analysis of single modal data is solved, and efficient and accurate monitoring of cattle anomaly behavior is achieved to meet the real-time monitoring needs of large-scale farms and ranches.

CN120493012APending Publication Date: 2025-08-15HEFEI NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510597430.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, cattle health monitoring in farms relies on single modal data analysis, which makes it difficult to meet the requirements of large-scale farms and pastures inadequate utilization of modal features.

Method used

A multimodal decision-making model is built, and a voiceprint analysis model, image behavior recognition model, temperature anomaly detection model and historical data analysis model are integrated through LoRA technology. A multi-head cross-modal attention mechanism and dynamic weight adjustment are adopted, and data processing and resource scheduling are carried out in combination with an edge-cloud collaborative decision-making mechanism.

Benefits of technology

The accuracy and real-time monitoring of cattle abnormal behavior has been improved, the accuracy of multimodal data fusion has been increased to 92.7%, the false alarm rate has been reduced to below 8%, and the utilization rate of computing resources has been increased by 40%, significantly reducing hardware costs and adapting to changes in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493012A_ABST
    Figure CN120493012A_ABST
Patent Text Reader

Abstract

The invention provides a cattle type state analysis method and device, computer equipment and a storage medium, and belongs to the field of farming and pasture breeding, and the method comprises the steps: constructing a plurality of different pre-training models and a multi-feature fusion decision model, fusing the parameters of the plurality of different pre-training models in the multi-feature fusion decision model through an LoRA technology, and obtaining a multi-feature fusion decision model; obtaining a target model; training the target model to obtain a cattle type state analysis model; obtaining the temperature, image, voiceprint and historical record data of cattle in a farm and pasture field; and respectively inputting the temperature, the image, the voiceprint and the historical record data into a cattle state analysis model, and outputting a cattle prediction state. Thus, it is ensured that the state information of the cattle is comprehensively captured from multiple dimensions, features from different modes are rapidly fused, the problem of isolated analysis of single-mode data is solved, and the accuracy and real-time performance of recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of agriculture and animal husbandry, and in particular relates to a cattle status analysis method, device, computer equipment and storage medium. Background Art

[0002] Current cattle health monitoring on farms and ranches primarily relies on technologies such as voiceprint detection, image recognition, and temperature sensing to provide early warnings for abnormal behavior. Voiceprint detection primarily analyzes the spectral characteristics of cattle calls (such as fundamental frequency and resonance peaks) to determine health status, but is susceptible to interference from environmental noise and cannot distinguish voiceprint variations caused by individual differences. Image recognition primarily uses cameras to capture cattle's behavioral posture (such as lying time and abnormal gait), but is affected by lighting conditions, shooting angles, and occlusion, resulting in a false detection rate of up to 30%. Temperature sensing primarily monitors body temperature through ear tags or wearable devices, but a single physiological indicator cannot fully reflect early symptoms of disease (such as decreased appetite or abnormal excretion).

[0003] A relatively similar solution in existing technology is the "Multi-Sensor Cattle Health Monitoring System," which collects data by deploying voiceprint sensors, cameras, and temperature sensors, and uses a threshold alarm mechanism to provide health warnings. Specifically, voiceprint data is matched against a predefined pathological pattern library, and an alarm is triggered when the set decibel threshold is exceeded; image data is identified through target detection algorithms (such as YOLO) and combined with temperature sensor data for comprehensive judgment; and each modal data is independently processed and analyzed. Therefore, existing technologies analyze modal features in isolation, failing to reflect data correlations between different modalities. This insufficiently utilizes data features, resulting in anomaly detection that is difficult to accurately and timely meet the needs of large-scale farms and ranches. Summary of the Invention

[0004] In order to solve the problem of insufficient utilization of data features caused by isolated analysis of multimodal data, the present invention provides a cattle status analysis method, apparatus, computer equipment and storage medium.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] A method for analyzing cattle status, comprising:

[0007] Build pre-trained voiceprint analysis models, image behavior recognition models, temperature anomaly detection models, and historical data analysis models;

[0008] The weight matrices of the four pre-trained models are decomposed using LoRA low-rank adaptation technology to generate adapter matrices for each model, which are then fused through channel dimension splicing to obtain the fusion layer of the multimodal decision model.

[0009] The voiceprint signals, video streams, body temperature data and breeder log text of cattle on farms are input into a multimodal decision model, and the voiceprint feature vector, image feature vector, temperature feature vector and text feature vector are extracted respectively through the fusion layer; the multi-head cross-modal attention mechanism is used to calculate the joint weight matrix of voiceprint features, image features, temperature features and historical text feature vectors; the modal weight ratio is dynamically adjusted according to environmental parameters; fusion features are generated based on the features and joint weight matrix of four different modalities and the modal weight ratio, and the output abnormal categories and corresponding probabilities are identified through a fully connected classifier.

[0010] Optionally, before inputting the acquired cattle data into the multimodal decision model, a two-stage strategy is required to train a fully connected classifier, including:

[0011] In the first stage, the backbone parameters of the multimodal decision model are frozen, and only the fully connected layers of the fully connected classifier are trained;

[0012] In the second stage, all parameters are fine-tuned with a preset learning rate, and the loss function uses FocalLoss to balance the category samples.

[0013] Optionally, before inputting the acquired cattle data into the multimodal decision model, historical baseline modeling is performed, including:

[0014] Construct behavioral baseline curves for individual cattle based on a bidirectional LSTM network;

[0015] The deviation between the real-time monitoring data and the baseline is calculated. When the deviation is greater than the deviation threshold, an alarm message is issued and the multimodal decision model is activated to analyze the current monitoring data.

[0016] Optionally, the deviation threshold is dynamically adjusted according to environmental data, cattle category and age level.

[0017] Optionally, the method further includes an edge-cloud collaborative decision-making mechanism:

[0018] Deploy a hybrid architecture consisting of edge computing nodes and cloud servers. Edge nodes perform real-time data collection, feature extraction, and lightweight model inference, while the cloud is responsible for multimodal fusion and global optimization.

[0019] When the network is interrupted, it switches to the edge local decision-making mode, calls the cloud analysis results of the last 7 days stored in the edge node, combines the real-time output of the edge lightweight model, and generates the final warning result through weighted voting.

[0020] Optionally, before inputting the acquired farm cattle data into the multimodal decision model, the acquired data is preprocessed, including:

[0021] The combination of spectral subtraction and wavelet transform is used to eliminate environmental noise from voiceprint data;

[0022] Image data denoising is performed based on keyframe extraction and optical flow analysis.

[0023] A cattle status analysis device, comprising:

[0024] Construction modules for building pre-trained voiceprint analysis models, image behavior recognition models, temperature anomaly detection models, and historical data analysis models;

[0025] A fusion module is used to decompose the weight matrices of the four pre-trained models using LoRA low-rank adaptation technology to generate adapter matrices for each model, and fuse the adapter matrices through channel dimension splicing to obtain the fusion layer of the multimodal decision model;

[0026] The recognition module is used to input the acquired voiceprint signals, video streams, body temperature data and breeder log text of farm cattle into a multimodal decision model, and respectively extract the voiceprint feature vector, image feature vector, temperature feature vector and text feature vector through the fusion layer; utilize the multi-head cross-modal attention mechanism to calculate the joint weight matrix of the voiceprint features, image features, temperature features and historical text feature vectors; dynamically adjust the modal weight ratio according to environmental parameters; generate fusion features based on the features and joint weight matrix of four different modalities and the modal weight ratio, and identify the output abnormal category and corresponding probability through a fully connected classifier.

[0027] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for analyzing the state of cattle is implemented.

[0028] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned cattle status analysis method when executing the program.

[0029] The cattle status analysis method provided by the present invention has the following beneficial effects:

[0030] First, multiple different pre-trained models are constructed. These models target temperature, image, voiceprint, and historical record data, respectively, and can extract features related to cattle status from different perspectives. This allows for separate training of models corresponding to different modal information, which reduces the amount of training while improving the accuracy of the training results. Secondly, a multi-feature fusion decision model is introduced, using LoRA technology to fuse the parameters of multiple different pre-trained models into the multi-feature fusion decision model. This not only reduces the model size but also achieves effective integration of cross-modal parameters. The multi-head attention mechanism in the multi-feature fusion decision model can dynamically adjust the weights of temperature, image, voiceprint, and historical text features, performing feature fusion based on the real-time and importance of the data, avoiding the limitations of single-modal data and improving the efficiency and accuracy of feature utilization. This ensures that cattle status information is fully captured from multiple dimensions, quickly fusing features from different modalities, solving the problem of isolated analysis of single-modal data, and improving the accuracy and real-time performance of recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] To more clearly illustrate the embodiments of the present invention and its design, the following briefly introduces the drawings required for this embodiment. The drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be derived from these drawings without inventive effort.

[0032] Figure 1 The figure is a flow chart of a method for analyzing cattle status according to an exemplary embodiment of the present invention.

[0033] Figure 2 A schematic diagram of a model for cattle status analysis provided by the present invention according to an exemplary embodiment.

[0034] Figure 3 A schematic diagram of a cloud-edge deployment scheduling system provided according to an exemplary embodiment of the present invention.

[0035] Figure 4 This is a block diagram of a cattle status analysis device provided according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0036] In order to enable those skilled in the art to better understand the technical solution of the present invention and to be able to implement it, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and are not intended to limit the scope of protection of the present invention.

[0037] This paper proposes an innovative method for monitoring abnormal cattle behavior on farms based on multimodal large-scale model fusion. It aims to address existing issues such as modal isolation, poor static threshold adaptability, and low historical data utilization. Through a series of carefully designed key steps, this method achieves efficient and accurate identification of abnormal cattle behavior while optimizing the use and allocation of computing resources. The technical solution of this invention primarily comprises three key modules: a multimodal data preprocessing module, a multi-task model joint fusion training framework, and an intelligent deployment and scheduling system.

[0038] The technical solutions provided by various embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0039] First, the present invention provides a method for analyzing cattle status, specifically Figure 1 As shown, the following steps are included:

[0040] S101. Build a pre-trained voiceprint analysis model, image behavior recognition model, temperature anomaly detection model, and historical data analysis model.

[0041] Specifically, the voiceprint analysis model, image behavior recognition model, temperature anomaly detection model and historical data analysis model can be pre-trained based on WaveNet, YOLOv8s, LSTM and BERT architectures, respectively.

[0042] S102. Use LoRA low-rank adaptation technology to decompose the weight matrices of the four pre-trained models to generate adapter matrices for each model, and fuse the adapter matrices through channel dimension splicing to obtain the fusion layer of the multimodal decision model.

[0043] LoRA, short for Low-Rank Adaptation of Large Language Models, is a low-rank adaptation technique for fine-tuning large language models. It fine-tunes the model by inserting pluggable low-rank matrices into the dense neural network layers of the original model. This approach not only reduces computational requirements but also requires significantly less training resources than directly training the original model.

[0044] In this step, the parameters are fused through LoRA technology to achieve the merging of models and establish a unified multi-task fusion model. The multi-feature fusion decision model may include a data layer, a fusion layer and a fully connected classifier, wherein the data layer is used to receive input multimodal data, the fusion layer is used to generate fusion features based on the multimodal data, and finally the fully connected classifier is used to classify the fusion features. LoRA adjusts the model parameters by performing low-rank decomposition on the weights or parameter matrices of the pre-trained model, thereby achieving adaptive adjustment to multiple tasks without significantly increasing the model size. Based on the LoRA parameter fusion, the present invention constructs a unified multi-task fusion model. This fusion model not only retains the expertise of each individual task model, but also can learn common features across tasks through shared representation. Then, the present invention performs further joint fine-tuning on this fusion model to optimize the overall performance of the model when processing cross-task monitoring.

[0045] Specifically, multiple pre-trained models are first used as base models. LoRa matrices are inserted into specific layers of each base model. For example, the parameter matrix of each pre-trained model is subjected to low-rank decomposition to obtain adapter matrices, which are then inserted into each base model. LoRa low-rank matrices can be injected into key layers of each pre-trained model (such as the last Transformer layer of BERT or the last residual block of ResNet). These low-rank matrices have a rank much smaller than the parameter matrix of the base model, and each model will have one or more LoRa matrices. In the multi-feature fusion decision model, the adapter matrices of all pre-trained models are loaded and fused. Specifically, these LoRa matrices can be fused using a method (such as weighted averaging) to obtain a unified parameter set, thereby achieving multi-model parameter fusion. For example, singular value decomposition (SVD) can be performed on the shared LoRa matrix across modalities to extract the common low-rank components and retain the top k singular values. Finally, the fused model is used for inference and its performance on downstream tasks is evaluated.

[0046] Among them, it is usually chosen to insert the LoRA matrix into certain key layers of the model (such as the attention layer or the fully connected layer). The size of the LoRA matrix (that is, its rank) can be adjusted according to the complexity of the downstream task and the computing resources. When fusing multiple LoRA matrices, it is necessary to select a suitable fusion strategy to ensure the performance of the fused model.

[0047] In addition, when selecting pre-trained models, you can choose multiple pre-trained models of different modalities (such as audio classification models, video action recognition models, text semantic models, etc.) to ensure that their output feature dimensions are aligned. In the multi-feature fusion decision model, the output features of each model are mapped to a unified dimensional space (for example, to 512 dimensions) through a fully connected layer, providing a compatible foundation for feature fusion.

[0048] In this way, LoRA technology enables the model to quickly adapt to different tasks and features by adjusting the parameters of the low-rank matrix instead of the entire model, thereby reducing the complexity and computational cost of fine-tuning. In addition, through parameter fusion, multiple task models can share some parameters, reducing the waste of computing resources and improving the utilization of computing resources. Parameter fusion can also comprehensively utilize the advantages of multiple models, optimize the overall performance through joint fine-tuning, and improve the accuracy and efficiency of abnormal behavior monitoring. Moreover, because LoRA technology allows the model to be fine-tuned for specific tasks without changing the core structure of the model, it enhances the generalization ability of the model, enabling it to better adapt to different environments and scenarios.

[0049] S103: Input the acquired voiceprint signals, video streams, body temperature data, and breeder log text of the cattle on the farm into the multimodal decision model, and output the abnormal category and corresponding probability.

[0050] In this step, before the relevant data of farm cattle are input into the multimodal decision model, a two-stage strategy is required to train the fully connected classifier.

[0051] In the first stage, the backbone parameters of the multimodal decision model are frozen, and only the fully connected layer of the fully connected classifier is trained; in the second stage, all parameters are fine-tuned with a preset learning rate, and the loss function uses FocalLoss to balance the category samples.

[0052] In one embodiment, the multi-task model joint fusion training framework of the present invention is designed to address the problems of modality isolation and resource waste in farmland monitoring systems. The framework uses an innovative phased training method, such as Figure 2 As shown in the figure, in the first stage, real-time data including temperature, images and sounds and historical record data (i.e., keeper log text) are obtained and input into the corresponding preset models for training to obtain multiple pre-trained models. In the second stage, the parameters of multiple pre-trained models are fused and fine-tuned based on LoRA technology. This effectively improves the performance of the model in the abnormal behavior monitoring task, while enhancing the synergy between models and reducing resource consumption.

[0053] In the first stage, the present invention fine-tunes the large model individually for each specific monitoring task in the farm, such as respiratory disease detection, digestive system abnormality identification, behavioral abnormality warning, etc. This process involves fine-tuning the parameters of the model so that they can more accurately capture the data patterns and features related to the specific task. Fine-tuning techniques include but are not limited to transfer learning, in which the pre-trained model parameters are used as a starting point to optimize the model performance through further training on a specific task dataset. For example, WaveNet is used to extract time-frequency features to build a pathological call classification model, an LSTM time series prediction model is trained based on OpenPose skeletal key point detection, and a Bi-LSTM+CRF model is used to extract structured health indicators.

[0054] The second stage is the joint fine-tuning of the multi-task model. In this stage, the present invention fuses and adjusts the parameters of each task model fine-tuned in the first stage, and specifically adopts LoRA (Low-Rank Adaptation) technology to achieve model parameter fusion fine-tuning.

[0055] In addition, for the fused model, the original parameters of all pre-trained models (such as ResNet's convolution kernel weights, BERT's attention parameters, and other pre-trained model parameters) can be frozen, and only the fully connected layer can be trained, reducing the trainable parameters by about 90%. By freezing the backbone, the parameters of the single-task model fused through LoRA can be kept unchanged (that is, its optimized feature extraction capabilities are retained), and only the classifier can be trained, allowing the fully connected layer to quickly adapt to the multimodal fusion features and learn how to map the joint features to abnormal behaviors. This avoids the destruction of pre-trained features due to fluctuations in the backbone parameters during initial training, ensuring stable convergence of the classifier.

[0056] In terms of fusion methods, the present invention adopts a hierarchical fusion strategy. First, at the feature level, a joint representation space (dimension 1024) is constructed through a multi-head attention mechanism to integrate the relevant features of each task to enhance the model's understanding of the monitoring task. Secondly, at the model level, the model parameters are adjusted through LoRA technology to achieve effective collaboration between models. Finally, at the decision-making level, a multi-task decision-making mechanism is designed, which can dynamically adjust the output weights of different task models according to the complexity and urgency of the monitoring task to generate the optimal monitoring results. Through this multi-level and multi-dimensional fusion method, the multi-task model joint fusion training framework of the present invention not only improves the efficiency and accuracy of abnormal behavior monitoring, but also significantly reduces the waste of resources through intelligent resource management and scheduling, achieving significant improvements to the existing technology.

[0057] For example, for the first stage of training, an appropriate model can be selected as a preset model based on the monitoring data, such as an audio classification model, a video action recognition model, or a text semantic model. The preset model is then trained based on pre-acquired training samples. The training samples include pre-acquired sample test data of cattle from a farm and their true feature labels. The sample test data is input into the preset model to obtain a prediction result. The preset model is then trained with the goal of minimizing the difference between the prediction result and the true feature label, resulting in multiple pre-trained models.

[0058] For the second stage of training, a multi-task loss can be designed, using the output of the pre-trained model as the task. Different tasks are weighted and contrastive learning loss is added to enforce alignment of features from different modalities in the latent space. Furthermore, a gradient accumulation approach is used to sequentially calculate the gradients of each modality within a single training batch. Dynamic weight averaging is used to combine gradient updates from different modalities to avoid gradient conflicts between modalities.

[0059] Compared to full parameter fine-tuning, LoRA fusion only requires training 0.5%-2% of the parameters, and the multimodal features complement each other, enhancing the anti-interference ability in various environments and improving accuracy. In the second stage, the backbone is unfrozen. After the classifier has initially converged, the LoRA adapter matrix and classifier parameters are fine-tuned at a lower learning rate (0.0002). This joint optimization further enhances the synergistic effect of multimodal feature fusion, such as enhancing the cross-modal attention mechanism to model the correlation between voiceprint and image features.

[0060] In addition, the trained cattle status analysis model needs to be deployed. Therefore, the present invention also includes an edge-cloud collaborative decision-making mechanism: deploying a hybrid architecture consisting of edge computing nodes and cloud servers. The edge nodes perform real-time data acquisition, feature extraction, and lightweight model inference, while the cloud is responsible for multimodal fusion and global optimization. In the event of a network interruption, the system switches to edge local decision-making mode, calling the cloud analysis results of the last seven days stored by the edge node, combining them with the real-time output of the edge lightweight model, and generating the final warning result through weighted voting.

[0061] The intelligent deployment and scheduling system of the present invention is another key component, which is responsible for intelligently allocating computing resources to achieve optimal utilization of resources, such as Figure 3As shown in the figure, it includes two parts: cloud (On cloud) and edge (At edge). In the cloud (On cloud) part, the core component of the system is the Kubernetes (K8s) cluster, which is uniformly managed and resource scheduled by the K8s master (master node). It includes several computing nodes, such as Node1 and Node2, and multiple edge nodes, such as EdgeNode1, Edge Node2, and Edge Node3. These nodes are used to deploy monitoring models and perform computing tasks. In addition, the CloudCore module, as the core component for communication between the cloud and the edge, is responsible for implementing communication management between the cloud and the edge. In the edge (At edge) part, it contains multiple edge nodes, each of which runs the Edge Core module to manage local devices and implement edge-cloud communication. Each Edge Node is connected to multiple specific monitoring devices (such as Device1 to Device4), which are responsible for actual data collection and front-end processing tasks. Communication between the cloud and edge uses a WebSocket-based messaging mechanism. The CloudCore component enables efficient data transmission and command distribution, enabling real-time and efficient interaction between edge device data and cloud services. This architecture enables centralized cloud management and distributed deployment at the edge, effectively ensuring resource utilization efficiency and the real-time and stability of monitoring services. Built on Kubernetes and Docker technologies, the system dynamically adjusts model deployment and resource allocation based on the current monitoring load and resource usage. The system works by first using monitoring components to collect system performance metrics and resource usage, including 12 core indicators such as GPU utilization, data throughput (>1TB / day), and response latency (<500ms). The scheduler then uses this information and pre-set policies to determine how to allocate resources to meet monitoring needs. For example, when the load of a monitoring task increases, the system can automatically scale up the corresponding computing resources or reschedule the model to balance the load. The system is characterized by a high degree of automation and intelligence, enabling rapid responses based on real-time data, ensuring the stability and responsiveness of the monitoring service. In its implementation, the system deployed Kubernetes cluster management nodes (three Master servers) and edge computing nodes (20 Jetson Xavier servers), and designed a hierarchical degradation strategy to switch to local edge decision-making mode (maintaining over 85% accuracy) in the event of a network outage. This system, primarily composed of software and relying on server hardware and network infrastructure, significantly reduced resource waste through intelligent resource management and scheduling, achieving a significant improvement over existing technologies.

[0062] In one embodiment, pre-installed cameras and temperature sensors can be used to capture temperature, image, voiceprint, and historical data of cattle on farms and pastures. The captured data can also be pre-processed, including removing ambient noise from voiceprint data using a combination of spectral subtraction and wavelet transform, and performing denoising on image data using keyframe extraction and optical flow analysis.

[0063] For example, the multimodal data preprocessing module of the present invention is the foundation of the entire monitoring system. It is responsible for receiving raw data from farms and ranches, including multi-source heterogeneous data such as voiceprints, videos, and manual recordings. This data may contain noise, inconsistencies, or format differences, requiring cleaning, alignment, and feature extraction to ensure data quality. The cleaning process uses a spectral subtraction + wavelet transform noise reduction technique to effectively eliminate environmental noise (improving the signal-to-noise ratio by 15dB). An improved FABF-YOLOv8s lightweight model (with 4.51M parameters) is used for video object detection and ROI region cropping, achieving a detection accuracy of 96.3%. Standardization involves converting data from different sources into a unified format. For example, the BERT model is used to encode unstructured text (such as "decreased appetite") into a 256-dimensional semantic vector. Feature extraction is the core function of this module. By applying statistical analysis, pattern recognition, and other techniques, it extracts from the raw data features useful for abnormal behavior monitoring tasks, including 23 core indicators such as fundamental frequency (F0), proportion of lying time, and appetite index. These features are then used to generate single-task supervised training data and multi-task joint supervised training data, providing input for subsequent model training. This module's design features a high degree of automation and configurability, allowing the preprocessing process to be tailored to different data characteristics and monitoring requirements. The module primarily consists of software, combined with GPU hardware acceleration components, significantly improving data processing efficiency.

[0064] In one embodiment, for the fusion of different features, the features output by each modality LoRA adapter can use a multi-head cross-modal attention mechanism to calculate the joint weight matrix of different feature vectors.

[0065] Specifically, the voiceprint signals, video streams, body temperature data and breeder log texts of cattle on farms are input into the multimodal decision model, and the voiceprint feature vector, image feature vector, temperature feature vector and text feature vector are extracted respectively through the fusion layer; the multi-head cross-modal attention mechanism is used to calculate the joint weight matrix of voiceprint features, image features, temperature features and historical text feature vectors; the modal weight ratio is dynamically adjusted according to environmental parameters; the fusion features are generated according to the features and joint weight matrix of the four different modalities and the modal weight ratio, and the category and corresponding probability of the output abnormal behavior are identified through a fully connected classifier.

[0066] Among them, the adjustment of the modal weight ratio can be dynamically adjusted according to the ambient temperature, ambient light and noise intensity.

[0067] For example, the confidence score of each modal output is calculated in real time. When the confidence score of a modality falls below a threshold, its fusion weight is reduced to 0.1. For example, in a cow monitoring scenario, if the video sensor is blocked and the confidence score drops, the weights of the voiceprint and body temperature modalities are automatically increased to 0.7 and 0.2 respectively.

[0068] For example, to provide early warning for mastitis in dairy cows, the inputs include:

[0069] Audio: MFCC features of cow calls (sampling rate 16kHz, 25ms frame length); Video: cow behavior video; Temperature: udder area infrared temperature time series data (sampling rate 1Hz); Text: milk production log (daily record). The LoRA adapter of the audio model (VGGish) learns abnormal patterns in the fundamental frequency of calls.

[0070] The LoRA fusion process includes:

[0071] Video model (SlowFast) of the LoRA adapter capturing abnormal fluctuations in behavior.

[0072] The LoRA adapter with temperature model (SVM) captures abnormal temperature fluctuations.

[0073] LoRA adapter for text model (LSTM) to analyze declining milk production trends.

[0074] Features are fused via gated attention (weights: audio 0.3, video 0.2, temperature 0.4, text 0.1).

[0075] In another embodiment, a learnable coefficient is used to weight and integrate multimodal prediction results. Specifically, the prediction results can be determined based on different pre-trained models, and the confidence of each modality can be calculated according to the actual situation. The prediction results of the different pre-trained models are weighted by the confidence, and the weighted results are then combined to determine the final prediction result. The specific implementation method can refer to the feature fusion case.

[0076] This invention utilizes multimodal large-scale model fusion and dynamic resource scheduling technology, using a cross-modal attention mechanism to dynamically weight voiceprint, video, and text data. This approach overcomes the limitations of single modality and significantly improves detection comprehensiveness, achieving a significant technological breakthrough in the field of abnormal behavior monitoring for cattle on farms. Multimodal data fusion improves the accuracy of abnormal behavior detection from 75.3% to 92.7% with the existing method, while reducing the false alarm rate to below 8%, significantly outperforming traditional single-modality methods. Real-time performance is enhanced through an edge-cloud collaborative architecture, which reduces system response latency from >5 seconds to <500ms, meeting the real-time monitoring requirements of large-scale ranches. Resource efficiency is optimized through a Kubernetes-based intelligent scheduling system, which increases computing resource utilization by 40%, reduces edge node computing power requirements by 59.48%, and significantly reduces hardware investment costs. The economic benefits are significant, with the early warning mechanism reducing cattle morbidity by 23%, and is expected to save approximately 450,000 yuan annually for ranches with 1,000 head of cattle. Adaptive expansion, dynamic baseline modeling and LoRA fine-tuning technology enable the system to adapt to different varieties, ages and seasonal changes, and has wide applicability.

[0077] In addition, the present invention also constructs a behavioral baseline curve for individual cattle based on a bidirectional LSTM network. The deviation between real-time monitoring data and the baseline is calculated. If the deviation exceeds a deviation threshold, an alarm is issued and a multimodal decision-making model is activated to analyze the current detection data to suppress errors caused by seasonal changes or temporary stress. The calculation of this deviation threshold incorporates an adaptive adjustment mechanism, which dynamically adjusts based on environmental data, cattle type, and age level. This constructs a long-term health baseline model based on the LSTM network and combines it with real-time data for joint analysis, effectively distinguishing temporary anomalies from true abnormal behavior.

[0078] Using the above method, we first construct multiple different pre-trained models. These models target temperature, image, voiceprint, and historical record data, respectively, and can extract features related to cattle status from different perspectives. A multi-feature fusion decision model is then introduced, and the parameters of multiple different pre-trained models are fused in the multi-feature fusion decision model using LoRA technology. This not only reduces the model size but also achieves effective integration of cross-modal parameters. The multi-head attention mechanism in the multi-feature fusion decision model can then dynamically adjust the weights of temperature, image, voiceprint, and historical text features, performing feature fusion based on the real-time and importance of the data, avoiding the limitations of single-modal data and improving the efficiency and accuracy of feature utilization. This ensures that cattle status information is fully captured from multiple dimensions, quickly integrating features from different modalities, solving the problem of isolated analysis of single-modal data, and improving recognition accuracy and real-time performance.

[0079] Secondly, the present invention also provides a cattle status analysis device, such as Figure 4 Shown, including:

[0080] The construction module 401 is used to construct a pre-trained voiceprint analysis model, an image behavior recognition model, a temperature anomaly detection model and a historical data analysis model.

[0081] The fusion module 402 is used to decompose the weight matrices of the four pre-trained models using the LoRA low-rank adaptation technology, generate the adapter matrix of each model, and fuse the adapter matrices through channel dimension splicing to obtain the fusion layer of the multimodal decision model.

[0082] Identification module 403 is used to input the acquired voiceprint signals, video streams, body temperature data and breeder log text of farm cattle into the multimodal decision model, and respectively extract the voiceprint feature vector, image feature vector, temperature feature vector and text feature vector through the fusion layer; use the multi-head cross-modal attention mechanism to calculate the joint weight matrix of the voiceprint features, image features, temperature features and historical text feature vectors; dynamically adjust the modal weight ratio according to environmental parameters; generate fusion features based on the features and joint weight matrix of the four different modalities and the modal weight ratio, and identify the category and corresponding probability of the output abnormal behavior through a fully connected classifier.

[0083] Using the above device, we first construct multiple different pre-trained models. These models target temperature, image, voiceprint, and historical record data respectively, and can extract features related to cattle status from different angles. Then, we introduce a multi-feature fusion decision model and use LoRA technology to fuse the parameters of multiple different pre-trained models in the multi-feature fusion decision model. This not only reduces the model size but also achieves effective integration of cross-modal parameters. Then, the multi-head attention mechanism in the multi-feature fusion decision model can dynamically adjust the weights of temperature, image, voiceprint, and historical text features. Feature fusion is performed based on the real-time and importance of the data, avoiding the limitations of single-modal data and improving the efficiency and accuracy of feature utilization. In this way, we ensure that the status information of cattle is fully captured from multiple dimensions, quickly fuse features from different modalities, solve the problem of isolated analysis of single-modal data, and improve the accuracy and real-time performance of recognition.

[0084] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 Steps for a cattle condition analysis method are provided.

[0085] The present invention also provides a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Steps for a cattle condition analysis method are provided.

[0086] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0087] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0088] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0089] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0090] It should be noted that the above specific embodiments can enable those skilled in the art to more fully understand the present invention, but do not limit the present invention in any way. Therefore, although this specification has described the present invention in detail, those skilled in the art should understand that the present invention can still be modified or replaced with equivalents; and all technical solutions and improvements that do not depart from the spirit and scope of the present invention are included in the scope of protection of the patent for the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. A method for analyzing cattle status, characterized in that: The method comprises: Build pre-trained voiceprint analysis models, image behavior recognition models, temperature anomaly detection models, and historical data analysis models; The weight matrices of the four pre-trained models are decomposed using LoRA low-rank adaptation technology to generate adapter matrices for each model, which are then fused through channel dimension splicing to obtain the fusion layer of the multimodal decision model. The voiceprint signals, video streams, body temperature data and breeder log text of cattle on farms are input into a multimodal decision model, and the voiceprint feature vector, image feature vector, temperature feature vector and text feature vector are extracted respectively through the fusion layer; the multi-head cross-modal attention mechanism is used to calculate the joint weight matrix of voiceprint features, image features, temperature features and historical text feature vectors; the modal weight ratio is dynamically adjusted according to environmental parameters; fusion features are generated based on the features and joint weight matrix of four different modalities and the modal weight ratio, and the output abnormal categories and corresponding probabilities are identified through a fully connected classifier.

2. A cattle status analysis method according to claim 1, characterized in that: Before inputting the acquired cattle data into the multimodal decision model, a two-stage strategy is required to train the fully connected classifier, including: In the first stage, the backbone parameters of the multimodal decision model are frozen, and only the fully connected layers of the fully connected classifier are trained; In the second stage, all parameters are fine-tuned with a preset learning rate, and the loss function uses FocalLoss to balance the category samples.

3. A cattle status analysis method according to claim 1, characterized in that: Before inputting the acquired cattle data into the multimodal decision-making model, historical baseline modeling was performed, including: Construct behavioral baseline curves for individual cattle based on a bidirectional LSTM network; The deviation between the real-time monitoring data and the baseline is calculated. When the deviation is greater than the deviation threshold, an alarm message is issued and the multimodal decision model is activated to analyze the current monitoring data.

4. A cattle status analysis method according to claim 3, characterized in that: The deviation threshold is dynamically adjusted according to environmental data, cattle category and age level.

5. The method for analyzing cattle status according to claim 3, characterized in that: The method also includes an edge-cloud collaborative decision-making mechanism: Deploy a hybrid architecture consisting of edge computing nodes and cloud servers. Edge nodes perform real-time data collection, feature extraction, and lightweight model inference, while the cloud is responsible for multimodal fusion and global optimization. When the network is interrupted, it switches to the edge local decision-making mode, calls the cloud analysis results of the last 7 days stored in the edge node, combines the real-time output of the edge lightweight model, and generates the final warning result through weighted voting.

6. A cattle status analysis method according to claim 1, characterized in that: Before inputting the acquired cattle data into the multimodal decision-making model, the acquired data is preprocessed, including: The combination of spectral subtraction and wavelet transform is used to eliminate environmental noise from voiceprint data; Image data denoising is performed based on keyframe extraction and optical flow analysis.

7. A cattle status analysis device, characterized in that: The device comprises: Construction modules for building pre-trained voiceprint analysis models, image behavior recognition models, temperature anomaly detection models, and historical data analysis models; A fusion module is used to decompose the weight matrices of the four pre-trained models using LoRA low-rank adaptation technology to generate adapter matrices for each model, and fuse the adapter matrices through channel dimension splicing to obtain the fusion layer of the multimodal decision model; The recognition module is used to input the acquired voiceprint signals, video streams, body temperature data and breeder log text of farm cattle into a multimodal decision model, and respectively extract the voiceprint feature vector, image feature vector, temperature feature vector and text feature vector through the fusion layer; utilize the multi-head cross-modal attention mechanism to calculate the joint weight matrix of the voiceprint features, image features, temperature features and historical text feature vectors; dynamically adjust the modal weight ratio according to environmental parameters; generate fusion features based on the features and joint weight matrix of four different modalities and the modal weight ratio, and identify the output abnormal category and corresponding probability through a fully connected classifier.

8. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Bridge state evaluation method and system based on large model

    CN120995145A

  • Cow breeding named entity recognition method based on LERT multi-feature fusion

    CN122113922A