Live pig health monitoring method and system based on multi-modal fusion
Through multimodal data acquisition and fusion technology, combined with efficient model fine-tuning and prompt templates, the problems of low detection efficiency and insufficient accuracy in pig health monitoring are solved, and efficient and accurate pig disease monitoring and diagnosis are achieved, which is suitable for mobile devices.
Patent Information
- Application Number
- CN202510519659.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-29
AI Technical Summary
The existing pig health monitoring technology has problems such as low detection efficiency, high cost, susceptibility to environmental interference, insufficient detection accuracy and real-time performance. It is especially difficult to achieve real-time monitoring and rapid response of diseases in high-density breeding environments, resulting in worsening of the disease and spreading the epidemic.
The multimodal fusion method is adopted to collect pig video, body temperature and sound data through surveillance cameras, infrared thermal imaging cameras and microphone arrays, and individual segmentation and tracking are combined with Grounded-SAM and DeepSORT algorithms. The Qwen2-VL-7B model is used for fine-tuning, and the prompt template and LoRA technology optimization model are designed to achieve efficient fusion and accurate monitoring of multimodal data.
It significantly improves the accuracy and comprehensiveness of pig disease detection, reduces the interference of complex environments on detection, improves the generalization ability and computing resource efficiency of the model, provides explainable diagnostic results and treatment suggestions, and is suitable for mobile devices.
Smart Images

Figure CN120387074A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of live pig health monitoring, and in particular to a live pig health monitoring method and system based on multimodal fusion. Background Art
[0002] Health monitoring of live pigs is crucial for early warning of diseases, timely intervention, and reducing mortality. With the rapid development of China's live pig breeding industry towards large-scale and intelligent directions, at present, a large number of farms still use the traditional method of combining manual observation and laboratory testing for health monitoring, including means such as body temperature measurement, clinical symptom observation, and serological testing. However, this method has problems of low detection efficiency and high cost. Especially in a high-density breeding environment, it is difficult to achieve real-time monitoring of diseases and rapid response, easily delaying the best treatment opportunity, resulting in the deterioration of the condition and the spread of the epidemic, seriously affecting the health and production performance of the pig herd, and ultimately affecting the economic benefits and biosecurity level of the farm.
[0003] Currently, pig health monitoring technologies can be mainly divided into two categories: traditional manual judgment and intelligent detection. The traditional manual judgment method mainly relies on veterinarians to make judgments through clinical symptom observation and laboratory testing. This method not only takes time and effort, but also is difficult to achieve real-time monitoring of a large number of pig herds, easily delaying the best treatment opportunity. In the field of intelligent detection, existing technologies are mainly divided into two detection methods: contact type and non-contact type. Contact detection usually uses wearable devices or implanted sensors to monitor physiological indicators such as the body temperature, heart rate, and exercise amount of pigs, and combines machine learning algorithms for disease early warning. However, this type of method is likely to cause stress reactions to pigs, and the devices have problems such as being easily damaged and high maintenance costs. In contrast, non-contact detection technologies mainly use technologies such as computer vision, infrared thermal imaging, and sound analysis to collect the behavior, body surface temperature, and sound characteristics of pigs through devices such as cameras, thermal imagers, and microphones, and use deep learning algorithms to achieve early identification and classification of diseases. Although this type of method avoids direct contact with pigs, the detection accuracy and real-time performance in a complex breeding environment still need to be improved.
[0004] In summary, the existing technologies have the following problems:
[0005] (1) Manual detection is time-consuming and laborious, and the discovery of pig diseases is not timely enough, resulting in a lag in disease early warning, thus delaying the best treatment opportunity.
[0006] (2) Contact detection methods such as ear tags and sensors are likely to cause stress reactions to pigs, and the devices have problems such as being easily damaged and high maintenance costs, restricting their large-scale application.
[0007] (3) Existing pig health monitoring models rely on a large amount of manually labeled data and are vulnerable to interference from the breeding environment background, including pigsty layout, floor color, equipment placement, etc. This difference makes it difficult for the model to adapt to the environmental differences of different farms, and the generalization ability of the model is weak.
[0008] (4) Existing methods mostly rely on a single data source, such as body temperature, sound, or behavior, lack effective fusion of multi-dimensional data, are prone to missed detections or misjudgments, and it is difficult to comprehensively evaluate the health status of pigs.
[0009] (5) Existing pig health monitoring systems mostly rely on high-performance computing resources and are difficult to be deployed on mobile devices, which limits the feasibility of real-time monitoring and on-site applications. Summary of the Invention
[0010] In view of the above deficiencies in the prior art, a pig health monitoring method and system based on multi-modal fusion provided by the present invention solve the problems existing in the prior art.
[0011] To achieve the above invention purpose, the technical solution adopted by the present invention is: A pig health monitoring method based on multi-modal fusion, comprising the following steps:
[0012] S1: Use a multi-modal data acquisition module to obtain pig body temperature data, pig video data, and pig sound data;
[0013] S2: Use a pig feature extraction module to extract features from the pig body temperature data, pig video data, and pig sound data to obtain pig body temperature data features, pig video data features, and pig sound data features;
[0014] S3: Use a large model fine-tuning module to fine-tune a basic large model to obtain a pig health monitoring large model;
[0015] S4: Use a pig health monitoring module to input the pig body temperature data features, pig video data features, and pig sound data features into the pig health monitoring large model to obtain a pig health monitoring result and give an early warning, completing the pig health monitoring based on multi-modal fusion.
[0016] Further, the multi-modal data acquisition module in S1 includes a monitoring camera, an infrared thermal imaging camera, a microphone array, a network transmission sub-module, and a data storage sub-module. S1 includes the following sub-steps:
[0017] S11: Use the infrared thermal imaging camera, monitoring camera, and microphone array to collect pig body temperature data, pig video data, and pig sound data respectively;
[0018] S12: Upload the collected live pig body temperature data, live pig video data, and live pig sound data to the data storage sub-module through the network transmission sub-module.
[0019] Further, the S2 includes the following sub-steps:
[0020] S21: Use the instance segmentation model Grounded-SAM to perform individual segmentation on the pigs in each frame of the live pig video data, and use the tracking algorithm DeepSORT to track the pigs to obtain the video data of the i-th live pig And obtain the live pig video data segment through the division criterion Realize the extraction of the features of the live pig video data;
[0021] S22: Obtain the position information of the pigs in the infrared video based on the position information of the pigs in the video data, and use the instance segmentation model Grounded-SAM to perform individual segmentation on the pigs in each frame of the infrared video to obtain the body temperature data of the i-th live pig And obtain the live pig body temperature data segment through the division criterion Then perform temperature change detection through a sliding window to realize the extraction of the features of the live pig body temperature data;
[0022] S23: Based on the live pig sound data, calculate the time difference of arrival of the sound signal through the method based on time delay estimation, determine the approximate direction of the sound source, and combine the tracking algorithm DeepSORT to obtain the pig trajectory information to realize the matching of the sound and the pigs, and obtain the sound data of the i-th live pig And obtain the live pig sound data segment through the division criterion
[0023] S24: Filter the background noise from the live pig sound data segment And perform sound classification based on the trained CNN+LSTM model to realize the extraction of the features of the live pig sound data.
[0024] Further, the large model fine-tuning module in the S3 includes a live pig disease dataset construction sub-module, a live pig prompt engineering sub-module, and a large model fine-tuning sub-module. The S3 includes the following sub-steps:
[0025] S31: Use the live pig disease dataset construction sub-module to integrate multiple pig disease datasets, expand the dataset scale through various data augmentation strategies, and obtain the live pig disease image dataset;
[0026] S32: Use the live pig prompt engineering sub-module to pre-define the classification and multi-modal feature representations of various common pig diseases, design a prompt template, and obtain the live pig disease semantic dataset;
[0027] S33: Fine-tune the base large model using the large model fine-tuning sub-module based on the pig disease image dataset and the pig disease semantic dataset to obtain a pig health monitoring large model.
[0028] Further, the pig disease image dataset in S31 specifically includes:
[0029] Healthy pigs;
[0030] Pigs with porcine reproductive and respiratory syndrome (PRRS): The characteristics include purple skin, listlessness, and body temperature higher than 40°C;
[0031] Pigs with influenza: The characteristics include coughing, wheezing, increased nasal discharge, and body temperature higher than 39.5°C;
[0032] Pigs with porcine circovirus disease: The characteristics include emaciation, pale or purple skin patches, and decreased immunity;
[0033] Pigs with erysipelas: The characteristics include diamond-shaped erythema on the skin, high fever, and body temperature higher than 41°C;
[0034] Pigs with pasteurellosis: The characteristics include rapid breathing, coughing, and cyanosis of the skin;
[0035] Pigs with foot-and-mouth disease: The characteristics include hoof, oral ulcers, salivation, and lameness.
[0036] Further, the prompt templates in S32 include basic prompt templates, structural prompt templates, and comparison prompt templates;
[0037] The basic prompt templates include simple queries and inferential queries;
[0038] The structural prompt templates include role-based queries;
[0039] The comparison prompt templates are used to compare two images of pigs.
[0040] Further, S33 includes the following sub-steps:
[0041] S331: Read the pig disease image dataset and the pig disease semantic dataset, and construct image-text pairs;
[0042] S332: Process the input image through the Qwen2-VL-7B vision encoder to calculate the feature vector of the input image. The formula is:
[0043] F(I) = VisionEncoder(I)
[0044] where F(·) is the feature vector of the input image, I is the input image, and VisionEncoder(·) is the Qwen2-VL-7B vision encoder;
[0045] The input text is processed by the Qwen2-VL-7B text encoder to calculate the feature vector of the input text. The formula is:
[0046] G(T) = TextEncoder(T)
[0047] where G(·) is the feature vector of the input text, T is the input text, and TextEncoder(·) is the Qwen2-VL-7B text encoder;
[0048] S333: The input text is structured through a prompt template. The formula is:
[0049] T prompt = PromptTemplate(T)
[0050] where T prompt is the input text after being processed by the prompt template, and PromptTemp late(·) is the text normalization operation;
[0051] S334: Contrastive learning is used to align the multi-modalities of images and text. The Qwen2-VL-7B model is fine-tuned through the LoRA technique and optimized in combination with the InfoNCE loss function to obtain a large model for pig health monitoring.
[0052] Furthermore, the process of fine-tuning the Qwen2-VL-7B model through the LoRA technique in S334 is expressed as:
[0053] Qwen2(X v , X t ) = f θ (E v (X v ), E t (X t ))
[0054] where Qwen2(·) is the Qwen2-VL model, X v is the visual input, X t is the text input, f θ (·) is a function, and E v (·) and E t (·) respectively represent the visual embedding function and the text embedding function implemented by ViT and Tokenizer;
[0055] The LoRA technique decomposes the weight matrix into a low-rank representation. The formula is:
[0056] W = W0 + UV t
[0057] Among them, W is the weight matrix in the Transformer layer, W0 is the original weight matrix, and U and V t are low-rank matrices;
[0058] The LoRA technique is used in the attention layer and the feed-forward layer of the Transformer, and the formula is:
[0059]
[0060] Among them, Attention(·) is the attention layer of the Transformer, Q, K, and V are the matrices of query, keyword, and value respectively, softmax(·) is the activation function, ΔK represents the change made to the keyword matrix during the fine-tuning process, and d k is the key vector, and the superscript T represents the transpose of the matrix.
[0061] The technical solution adopted by the present invention is also: a system for a multi-modal fusion-based pig health monitoring method, and the system includes:
[0062] A multi-modal data acquisition module for acquiring pig body temperature data, pig video data, and pig sound data;
[0063] A pig feature extraction module for extracting features from the pig body temperature data, pig video data, and pig sound data to obtain pig body temperature data features, pig video data features, and pig sound data features;
[0064] A large model fine-tuning module for fine-tuning the basic large model to obtain a pig health monitoring large model;
[0065] A pig health monitoring module for inputting the pig body temperature data features, pig video data features, and pig sound data features into the pig health monitoring large model to obtain a pig health monitoring result and give an early warning, thereby completing the pig health monitoring based on multi-modal fusion.
[0066] The beneficial effects of the present invention are as follows: The present invention provides a method and system for pig health monitoring based on multimodal fusion. The present invention uses multimodal data acquisition devices such as video cameras, infrared thermal imaging, and microphone arrays. Through multimodal data fusion technology, the accuracy and comprehensiveness of pig disease detection are significantly improved. Then, combined with the improved Grounded-SAM model for lesion area segmentation, the interference of complex breeding environments on detection is effectively reduced. The present invention uses a variety of data augmentation strategies to expand the dataset size from 6,400 images to 25,000 images, significantly enhancing the generalization ability of the model under different lighting conditions, shooting angles, and background complexities. And through prompt engineering techniques, basic prompt templates, structural prompt templates, and comparison prompt templates are designed. Combined with the pig disease semantic dataset, the large model can not only more accurately and comprehensively identify disease types, but also provide etiology analysis and treatment suggestions, significantly enhancing the interpretability and clinical practicability of the diagnostic results. Aiming at the problem of limited computing resources of mobile devices, the present invention uses the LoRA technology to efficiently fine-tune the Qwen2-VL-7B large model, significantly reducing the computing resource requirements of the model.
[0067] In summary, the present invention has the advantages of multimodal fusion, strong interpretability, and efficient mobile deployment in pig disease detection, significantly improving the accuracy and practicability of disease detection, and reducing the epidemic risk and economic losses of farms. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 It is a flowchart of a method for pig health monitoring based on multimodal fusion according to the present invention.
[0069] Figure 2 It is an overall flowchart block diagram of a method for pig health monitoring based on multimodal fusion according to the present invention.
[0070] Figure 3 It is a schematic diagram of the multimodal data acquisition module.
[0071] Figure 4 It is a flowchart of the working process of Grounded-SAM.
[0072] Figure 5 It is an overall flowchart of the working process of PDLM.
[0073] Figure 6 It is a diagram of the fine-tuning process of PDLM. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments.
[0075] Example 1, as Figure 1 and Figure 2As shown in the figure, a pig health monitoring method based on multi-modal fusion includes the following steps:
[0076] S1: Use the multi-modal data acquisition module to obtain pig body temperature data, pig video data, and pig sound data;
[0077] S2: Use the pig feature extraction module to extract features from the pig body temperature data, pig video data, and pig sound data to obtain pig body temperature data features, pig video data features, and pig sound data features;
[0078] S3: Use the large model fine-tuning module to fine-tune the basic large model to obtain a pig health monitoring large model;
[0079] S4: Use the pig health monitoring module to input the pig body temperature data features, pig video data features, and pig sound data features into the pig health monitoring large model to obtain pig health monitoring results and issue warnings, completing pig health monitoring based on multi-modal fusion.
[0080] Multi-modal data acquisition module: The multi-modal data acquisition module mainly consists of a monitoring camera, an infrared thermal imaging camera, a microphone array, a network transmission sub-module, and a data storage sub-module. The monitoring camera is responsible for continuously shooting the pigs in the pigsty for 24 hours to collect video image data; the infrared thermal imaging camera is installed in the pigsty to monitor the body temperature of the pigs in real time and collect the body temperature data of the pigs; the microphone array collects the sound data of the pigs, including audio signals such as coughing, sneezing, and hissing, for detecting respiratory diseases or abnormal behaviors. The collected data is transmitted to the server through a switch and a router for use by the pig feature extraction module. This module realizes non-contact pig health monitoring, provides three types of multi-modal data: video, body temperature, and sound, and provides comprehensive information support for disease detection.
[0081] Pig feature extraction module: In order to effectively extract the health features of pigs and reduce the impact of environmental complexity on the model performance, the present invention designs a pig feature extraction module for processing the video images, body temperature, and sound data collected by the multi-modal data acquisition module and extracting key features to provide high-quality input for disease identification.
[0082] Large model fine-tuning module: The present invention designs a large model fine-tuning module, aiming to fine-tune the basic large model based on multi-modal data and combine the prompt engineering built into the large model, so that the large model can provide an interpretable diagnostic report and scientific treatment suggestions to improve the accuracy and applicability of pig disease detection and achieve automatic identification. This module includes a pig disease dataset construction sub-module, a pig prompt engineering sub-module, and a large model fine-tuning sub-module.
[0083] The pig disease dataset construction sub-module is used to construct a high-quality pig disease dataset, covering image data of various pig diseases. After the dataset is constructed, data augmentation is used to improve the generalization ability of the model.
[0084] The pig prompt engineering sub-module constructs a semantic dataset of pig diseases based on prompt engineering technology, and pre-defines the characteristics of various diseases. For example, the image characteristics of classical swine fever include skin erythema and conjunctival congestion; the body temperature characteristics are manifested as fever (>41°C); the sound characteristics include groaning and weak barking. The semantic dataset of pig diseases provides text data for targeted fine-tuning of the large model, aiming to improve the model's visual feature recognition ability, etiology explanation ability and treatment plan recommendation ability for pig diseases. This sub-module also designs three types of prompt templates, including basic prompt templates, structural prompt templates and comparison prompt templates, to ensure that the model can provide accurate, interpretable and scientific diagnostic results and treatment suggestions after disease detection. Among them, the basic prompt template is used to achieve disease classification and core feature recognition, the structural prompt template is used to construct causal reasoning relationships and optimize disease management strategies, and the comparison prompt template is used to strengthen the model's ability to distinguish similar diseases, thereby further improving the accuracy and practicality of diagnosis. Through an efficient prompt engineering mechanism, this sub-module significantly improves the adaptability and professionalism of the large model in pig disease diagnosis tasks, ensuring the reliability and clinical feasibility of diagnostic results.
[0085] The large model fine-tuning sub-module is used to fine-tune the large model based on multi-modal big data, enabling it to have the ability to learn disease characteristics and classify and predict. To improve training efficiency and reduce video memory occupancy, this module uses the LoRA (Low-Rank Adaptation) technology for efficient fine-tuning, optimizing the model performance and enhancing its adaptability.
[0086] Pig health monitoring module: The present invention designs a pig health monitoring module, which deeply analyzes the data extracted by the pig feature extraction module and combines with the pig health monitoring large model fine-tuned by the large model fine-tuning module to achieve accurate identification of the health status of pigs. This module supports users to upload pig video image data through mobile devices for pig disease analysis reports. This module can also conduct real-time monitoring of the pigs in the pigsty. If any abnormalities in pig disease characteristics are found, it will integrate the pig disease abnormal information and send the specific pigsty location and health warning information to the farm managers.
[0087] The overall workflow diagram of the present invention is as Figure 2 shown, and the specific description is as follows:
[0088] In S1, the multi-modal data acquisition module includes a monitoring camera, an infrared thermal imaging camera, a microphone array, a network transmission sub-module and a data storage sub-module. S1 includes the following sub-steps:
[0089] S11: Collect the live pig body temperature data, live pig video data, and live pig sound data respectively using an infrared thermal imaging camera, a surveillance camera, and a microphone array.
[0090] S12: Upload the collected live pig body temperature data, live pig video data, and live pig sound data to the data storage sub-module via the network transmission sub-module.
[0091] In this embodiment, as Figure 3 shown, install a surveillance camera with a frame rate of 25fps above the side of the live pig pen to ensure that all live pigs in the pigsty can be completely photographed. The surveillance camera continuously photographs the live pigs in the pigsty for 24 hours to collect video image data in real time for subsequent disease analysis. Deploy an infrared thermal imaging camera with a frame rate of 25fps in the pigsty. This temperature measurement device can monitor the body temperature of live pigs in real time, accurately capture the temperature change trend, and collect body temperature data for disease prediction and health management. At the same time, install a microphone array in the pigsty to collect the sound data of pigs, including audio signals such as coughing, sneezing, and hissing, in order to identify whether there are respiratory diseases or abnormal behaviors in the pigs. All the collected data is transmitted to the server via a switch and a first-level router for use by the live pig feature extraction module. This module realizes non-contact health monitoring of pigs and can synchronously provide three types of multi-modal data: video, body temperature, and sound, providing comprehensive and stable data support for accurate disease detection and health assessment.
[0092] The live pig feature extraction module will perform feature extraction operations on the data collected in S1 to obtain live pig video data, live pig body temperature data, and live pig sound data. The specific process is as follows:
[0093] First, construct an empty live pig data D, where D i represents all the information of the i-th live pig. Among them, represents the live pig video data of the i-th live pig, represents the j-th live pig video clip in and represents the k-th frame image in represents the live pig infrared data of the i-th live pig, represents the j-th live pig infrared data clip in and represents the k-th frame image in represents the live pig audio data of the i-th live pig, represents the j-th live pig audio data clip in
[0094] Step S2 includes the following sub-steps:
[0095] S21: Use the instance segmentation model Grounded-SAM to perform individual segmentation on pigs in each frame of the live pig video data, and use the tracking algorithm DeepSORT to track the pigs, obtaining the video data of the i-th live pig And obtain the live pig video data segments through the division criterion (100 frames) Realize the extraction of the features of the live pig video data;
[0096] Grounded-SAM is an advanced open-vocabulary detection and segmentation model that combines Grounding DINO and SAM, as Figure 4 shown. This model uses open-vocabulary object detection and segmentation technology to accurately identify and segment the various parts of pigs. This model ensures that the segmentation process is both accurate and applicable to different types of pig images. Grounded-SAM shows excellent performance in open-vocabulary benchmark tests and can complete more complex visual tasks compared to SAM alone. SAM requires point or box prompts to operate and cannot identify occluded objects based on arbitrary text inputs. Instead, Grounded-SAM uses Grounding DINO to generate precise boxes for objects in the image using text information, and then SAM uses these boxes to generate precise mask annotations. This integration makes Grounded-SAM more effective in scenarios that require precise and flexible object segmentation.
[0097] S22: Based on the position information of the pigs in the video data, obtain the position information of the pigs in the infrared video, and use the instance segmentation model Grounded-SAM to perform individual segmentation on the pigs in each frame of the infrared video, obtaining the body temperature data of the i-th live pig And obtain the live pig body temperature data segments through the division criterion (100 frames) Then, perform temperature change detection through a sliding window to realize the extraction of the features of the live pig body temperature data;
[0098] If the body temperature exceeds the normal body temperature threshold, convert the abnormal body temperature situation into structured text. For example: "The body temperature of pig No. 3 exceeded 41°C during the period from 10:35 to 10:50 on March 4, 2025, suspected of having a fever."
[0099] S23: Based on the live pig sound data, calculate the time difference of arrival of the sound signal by a method based on time delay estimation, determine the approximate direction of the sound source, and combine the tracking algorithm DeepSORT to obtain the pig trajectory information, realize the matching of the sound and the pigs, and obtain the sound data of the i-th live pig And obtain the live pig sound data segments through the division criterion (100 frames)
[0100] S24: For the pig sound data segment Perform noise filtering to remove background noise, and perform sound classification based on the trained CNN+LSTM model to extract the characteristics of pig sound data.
[0101] Among them, CNN is responsible for extracting time-frequency features, and LSTM is responsible for time series modeling to ensure that the model can identify different types of sound events such as coughs and calls. Finally, record the time, category, and frequency of the sound event and convert it into a text description. For example: "Pig No. 2 coughed continuously 5 times during the period from 11:10 to 11:15 on March 4, 2025, and there may be a risk of respiratory infection."
[0102] Through this process, this module can perform multi-modal fusion of the video, body temperature data, and sound data of pigs in the time and space dimensions, so as to extract the multi-modal data privileges of each pig.
[0103] In step S3, the present invention aims to fine-tune a large pig health monitoring model, hereinafter referred to as PDLM. The working process of PDLM is as Figure 5 shown. This process starts with the original images of pigs, which show various symptoms or characteristics indicating pig diseases. These images are processed using a modified version of SAM called Grounded-SAM. This model uses a masking technique to isolate the key pig features from the complex background environment, with particular attention to disease symptoms. The result is a segmented image in which the main pig features are emphasized and the irrelevant background details that may hinder accurate disease identification are removed. After that, prompt engineering for image description is carried out. A class of pig diseases is predefined for each segmented image. This classification is crucial for prompt engineering, and the designed templates can effectively describe the pig disease symptoms observed in the images. The structure of these prompt templates semantically represents disease characteristics, which helps to generate standardized and comprehensive descriptions for each image category. Subsequently, Qwen2-VL-7B is used to generate accurate and informative descriptions using the segmented images and their prompt templates. These descriptions strive to capture the basic characteristics and background of the pig diseases depicted in each image. The synthesis of these descriptions with their corresponding segmented images constitutes the Pig Disease Semantic Dataset (PDSD).
[0104] PDSD is divided into two subsets: the training dataset accounts for 80% of the data, and the test dataset accounts for 20%. The training dataset is used to fine-tune Qwen2-VL-7B. Then, this fine-tuned model is used to classify new pig disease images. The test dataset is used to evaluate the model's ability to accurately diagnose diseases through the fine-tuning scenario or using the zero-shot classification method. The specific implementation of this step includes a pig disease dataset construction sub-module, a pig prompt engineering sub-module, and a large model fine-tuning sub-module.
[0105] The large model fine-tuning module in S3 includes a pig disease dataset construction sub-module, a pig prompt engineering sub-module, and a large model fine-tuning sub-module. S3 includes the following sub-steps:
[0106] S31: Use the pig disease dataset construction sub-module to integrate multiple pig disease datasets, expand the dataset scale through various data augmentation strategies, and obtain a pig disease image dataset;
[0107] To construct the pig disease dataset, the present invention integrates multiple publicly available pig disease datasets, covering real pig disease image data. The dataset sources include a healthy and African swine fever dataset, a pig ear disease dataset, a pigsty disease pig dataset, and a self-collected dataset from a pig farm in Guangdong, China. Using these datasets, the present invention uses an image hashing algorithm to eliminate duplicate images from the original data.
[0108] The pig disease image dataset in S31 specifically includes:
[0109] Healthy pigs (1200 pieces);
[0110] Porcine reproductive and respiratory syndrome (PRRS) pigs (1100 pieces): Features include purple skin, listlessness, and body temperature higher than 40°C;
[0111] Influenza pigs (900 pieces): Features include coughing, wheezing, increased nasal discharge, and body temperature higher than 39.5°C;
[0112] Porcine circovirus disease pigs (750 pieces): Features include emaciation, pale or purple skin patches, and decreased immunity;
[0113] Erysipelas suis pigs (600 pieces): Features include diamond-shaped erythema on the skin, high fever, and body temperature higher than 41°C;
[0114] Swine pasteurellosis pigs (850 pieces): Features include rapid breathing, coughing, and cyanosis of the skin;
[0115] Foot-and-mouth disease pigs (1000 pieces): Features include hoof, oral cavity ulcers, salivation, and lameness;
[0116] The dataset consists of a total of 6,400 images, which are divided according to the ratio of 80% for training and 20% for testing.
[0117] To meet the input requirements of Qwen2-VL-7B and maintain the uniformity and high quality of the data, the present invention performs standardized preprocessing on all images. The specific steps include resizing the images so that the longer side is scaled to 490 pixels, and at the same time scaling the other side proportionally to maintain the integrity and clarity of the images.
[0118] To improve the generalization ability of the large model fine-tuned using the pig disease dataset, the present invention adopts data augmentation methods to expand the data scale, enhance the robustness of the model, reduce the risk of overfitting, and improve the adaptability to different environments, lighting, and individual variations. In terms of visible light images, first, geometric transformations are adopted, including random rotation (±15°), horizontal and vertical flipping, random scaling (0.8× to 1.2×), and random cropping (80%-100%) to simulate different camera angles, viewing perspectives, and partial occlusion situations. Secondly, color enhancement methods are applied, such as brightness adjustment (±30%), contrast adjustment (±25%), color jitter (hue and saturation ±20%), and Gaussian blur to adapt to different lighting conditions. In addition, to improve the model's recognition ability for defective areas, three methods, namely Cutout, Mixup, and CutMix, are introduced to enhance the model's learning ability of local features through random occlusion, image mixing, and region replacement respectively. To further improve the robustness of the model in complex environments, the present invention introduces adversarial data augmentation methods, uses StyleGAN2 to generate realistic disease images to expand the diversity of the dataset, and at the same time uses FGSM (Fast Gradient Sign Method) to generate slight pixel perturbations to improve the model's resistance to adversarial attacks. In addition, background replacement technology is also adopted to replace the background of pig disease images with different farm environments, such as cement floors and muddy fields, thereby enhancing the model's scene adaptability. Finally, through the above various data augmentation strategies, the scale of the dataset of this system is expanded from the original 6,400 to approximately 25,000, improving the generalization ability of the model in terms of different lighting, shooting angles, individual differences, and background complexity, and ensuring its efficient performance in complex real environments.
[0119] S32: Use the pig prompt engineering sub-module to pre-define the classification and multi-modal feature manifestations of various common pig diseases, design prompt templates, and obtain a pig disease semantic dataset;
[0120] The present invention creates a semantic dataset for swine diseases using prompt engineering for targeted fine-tuning of a base large model. The main goal is to enhance the large model's ability to accurately classify and interpret swine diseases based on visual cues, thereby providing effective treatment strategies. The semantic dataset for swine diseases predefines the classification of various common swine diseases and their multimodal characteristic manifestations according to veterinary medical knowledge and breeding practice experience, including appearance (image), body temperature, and sound characteristics. For example, the image characteristics of Porcine Reproductive and Respiratory Syndrome include purple skin and weakness; the body temperature characteristic is fever (>40°C); the sound characteristics include coughing and wheezing. These predefined multimodal data representations of swine diseases are used as text data for fine-tuning the large model. And the present invention develops three types of prompt templates - basic prompt templates, structured prompt templates, and comparison prompt templates. These templates are used to elicit specific responses from Qwen2-VL-7B, ensuring that the responses are tailored to meet the needs of swine disease classification.
[0121] The prompt templates in S32 include basic prompt templates, structured prompt templates, and comparison prompt templates;
[0122] The basic prompt templates include simple queries and reasoning queries;
[0123] The basic prompt templates are constructed to ensure the diversity of the large model for pig health monitoring by providing direct queries that focus on disease classification and the reasoning behind it. Among them, the simple query asks "What is the category of the pig image?" to obtain a direct answer, such as "Classical swine fever". The reasoning query enables the large model for pig health monitoring to evaluate the image as a swine pathology expert, identify and describe visual features, and provide a diagnosis, linking these features to the specified category, ensuring that the analysis is concise and within 150 words. For example, it might describe the purple-red patches that appear on the skin, which consist of hemorrhagic lesions characteristic of "Classical swine fever", and explain how these features lead to the high fever and systemic infection of the pig, ultimately resulting in a high mortality rate. These templates ensure that the model can accurately identify the disease and provide the reasoning behind its classification, thereby improving the accuracy and interpretability of the large model of this system.
[0124] The structured prompt templates include role-based queries;
[0125] The structured prompt template contains role-based queries that explore the reasons for assigning specific categories to images and provide detailed treatment plans. This template enhances the logical reasoning ability of the large model for pig health monitoring by integrating the Visual Chain of Thought (CoT) mechanism. For example, a specific structured prompt template focuses on classification reasoning, enabling the model to identify "classical swine fever" and describe specific features visible in the image, such as purplish-red patches and hemorrhagic lesions on the skin, and clarify the relationship between these features and the diagnosis. Another specific structured prompt template focuses on the treatment plan, prompting the model to diagnose "classical swine fever", clarify the reasons, and propose a detailed and actionable treatment plan, including vaccination, isolating infected pigs, strengthening biosecurity measures, and regularly monitoring the health status of the pig herd. These structured templates ensure that the large model for pig health monitoring can not only accurately classify diseases but also provide comprehensive insights for effective disease management, thus significantly improving the practical applicability of the model in this field.
[0126] The comparison prompt template is used to compare two images of pigs;
[0127] The comparison prompt template focuses on comparing two images of pigs to identify common or unique features that help determine the category of each image, thereby enhancing the discriminative ability of the model. This template effectively enhances the large model's ability to highlight the subtle differences between similar diseases. For example, when the user inputs two relatively similar images, the comparison prompt template instructs the model to analyze how the elements in one image correspond to the elements in the other image, pay attention to the shared features that may imply a common category, while pointing out the unique aspects that distinguish them. Then, the task of this template is to provide the most appropriate category for each image separately, ensuring that the classification accurately reflects its unique features. This method helps improve the accuracy of the model in diagnosing and differentiating similar disease symptoms, thus enhancing its practicality in this field.
[0128] S33: Use the large model fine-tuning sub-module to fine-tune the basic large model based on the pig disease image dataset and the pig disease semantic dataset to obtain the large model for pig health monitoring;
[0129] The function of this sub-module is to enable the basic model Qwen2-VL-7B to learn the mapping relationship between pig disease images and corresponding text descriptions, so that it can automatically generate complete and accurate disease diagnosis reports based on the input pig disease images during inference. This module uses three types of prompt templates (basic prompt template, structural prompt template, comparison prompt template) designed in the pig prompt engineering sub-module to construct the training data of the model from the pig disease dataset and the corresponding medical text descriptions, and these data are stored in a multi-modal (image + text) format. During the fine-tuning process, this module improves the disease recognition ability of Qwen2-VL-7B through contrastive learning, and improves the efficiency of fine-tuning the large model through the LoRA technique. Finally, it is optimized by combining the InfoNCE loss, so that the basic model can effectively align visual-language information and possess the capabilities of disease classification, diagnostic reasoning, and interpretable text generation. The specific implementation of this module is as follows:
[0130] The S33 mentioned above includes the following sub-steps:
[0131] S331: Read the pig disease image dataset and the pig disease semantic dataset, and construct image-text pairs;
[0132] In this embodiment, the data format follows the Hugging Face compatible standard and is stored in JSON format. Each sample contains {"image":"pig_disease_001.jpg","text":"What is the disease type of this pig? Please answer in combination with the image features.","target":"This pig is suspected of being infected with Porcine Reproductive and Respiratory Syndrome (PRRS), manifested as purple skin and listlessness."}. In terms of data preprocessing, relevant processing has been carried out in both the pig disease dataset construction sub-module and the pig prompt engineering sub-module.
[0133] S332: Process the input image through the Qwen2-VL-7B visual encoder, and calculate the feature vector of the input image. The formula is:
[0134] F(I) = VisionEncoder(I)
[0135] Among them, F(·) is the feature vector of the input image, I is the input image, and VisionEncoder(·) is the Qwen2-VL-7B visual encoder;
[0136] Process the input text through the Qwen2-VL-7B text encoder, and calculate the feature vector of the input text. The formula is:
[0137] G(T) = TextEncoder(T)
[0138] Among them, G(·) is the feature vector of the input text, T is the input text, and TextEncoder(·) is the Qwen2-VL-7B text encoder;
[0139] Three prompt templates designed based on prompt engineering in the pig health monitoring sub-module are introduced - the basic prompt template, the structural prompt template, and the comparison prompt template. During the training process, the text input is structured by the prompt template to ensure that the model can learn how to use visual information to answer medical questions in various scenarios.
[0140] S333: Structurally process the input text through the prompt template, and the formula is:
[0141] T prompt = PromptTemplate(T)
[0142] Among them, T prompt is the input text processed by the prompt template, and PromptTemp late(·) is the text normalization operation, which normalizes the input text according to the specific task scenario to enhance the model's adaptability to the task;
[0143] S334: Align the multi-modal of images and text by contrastive learning, fine-tune the Qwen2-VL-7B model through the LoRA technique, and optimize it in combination with the InfoNCE loss function to obtain the pig health monitoring large model.
[0144] Such as Figure 6 shown, the process of fine-tuning the Qwen2-VL-7B model through the LoRA technique in S334 is expressed as:
[0145] Qwen2(X v ,X t ) = f θ (E v (X v ),E t (X t ))
[0146] Among them, Qwen2(·) is the Qwen2-VL model, X v is the visual input, X t is the text input, f θ (·) is the function, and E v (·) and E t (·) respectively represent the visual embedding function and text embedding function implemented by ViT and Tokenizer;
[0147] LoRA (Low-Rank Adaptation) improves the efficiency of fine-tuning large language models by training low-rank matrices instead of the full set of model parameters. This approach significantly reduces the computational resources required for fine-tuning. This technique is particularly beneficial for fine-tuning vision large models, enabling task-specific optimization without substantial computational resources. LoRA enhances the parameter efficiency of Transformer layers by decomposing the weight matrix into a low-rank representation.
[0148] The LoRA technique decomposes the weight matrix into a low-rank representation, with the formula:
[0149] W = W0 + UV t
[0150] where W is the weight matrix in the Transformer layer, W0 is the original weight matrix, and U and V t are low-rank matrices;
[0151] This low-rank adaptation is applied to the attention and feed-forward layers of the Transformer, enabling the model to learn task-specific patterns while leveraging pre-trained knowledge. The adaptation in the attention layer is as follows: The low-rank adaptation technique is used in the attention and feed-forward layers of the Transformer to utilize pre-trained model insights while customizing them for specific tasks.
[0152] The LoRA technique is used in the attention layer and feed-forward layer of the Transformer, with the formula:
[0153]
[0154] where Attention(·) is the attention layer of the Transformer, Q, K, and V are the matrices of queries, keys, and values respectively, softmax(·) is the activation function, ΔK represents the changes made to the key matrix during the fine-tuning process, d k is the key vector, and the superscript T represents the transpose of the matrix.
[0155] Multi-modal alignment of images and texts is achieved through contrastive learning to optimize the Qwen2-VL model's understanding ability of image-text relationships. The core idea of contrastive learning is to construct positive and negative sample pairs so that the model can distinguish correctly matched disease image-text descriptions and suppress incorrect matches. Specifically, maximize the similarity of positive sample pairs (such as the match between a classical swine fever image and its correct symptom description), while minimizing the similarity of negative sample pairs (such as the match between a classical swine fever image and a foot-and-mouth disease symptom description). First, input a batch of live pig disease images and their annotated texts to construct positive and negative samples: Positive sample pairs: For example, a classical swine fever image and its corresponding text description "This pig is suspected of being infected with classical swine fever, showing skin erythema and conjunctival congestion", ensuring semantic consistency through manual annotation or expert knowledge; Negative sample pairs: Randomly combine images with incorrect texts (such as a classical swine fever image with a foot-and-mouth disease description of "oral ulcers and salivation"), or sample interfering texts from other disease categories. Then, extract the image feature vector F(I) through the visual encoder of Qwen2-VL, where I is the input live pig disease image, and extract the text feature vector G(T) through the text encoder, where T is the corresponding text description, and calculate the cosine similarity between the two as the matching score to measure the matching degree between the image and the text. The calculation formula of the cosine similarity sim(I,T) is as follows:
[0156]
[0157] Then, adopt the InfoNCE loss (InfoNCE: Information Noise Contrastive Estimation), and by comparing the similarity distributions of all sample pairs in the same batch of data, force the model to focus on the discrimination of positive sample pairs. The formula of the InfoNCE loss function is as follows:
[0158]
[0159] where, T + represents the positive sample text, and T j contains all negative samples and the positive sample text. The loss function is optimized through gradient descent, making the similarity of positive sample pairs significantly higher than that of negative sample pairs, enabling the model to continuously update the model parameters and iteratively optimize the alignment ability of the feature encoder. Through contrastive learning, Qwen2-VL can accurately associate the visual features of live pig diseases (such as skin erythema) with the corresponding text descriptions (such as classical swine fever symptoms), reduce cross-modal mis-matching, and thus improve the accuracy and interpretability of disease diagnosis.
[0160] In step S4, the present invention continuously monitors the disease and health status of live pigs in the farm for 24 hours and automatically triggers a health warning when an abnormality is detected. The system processes the multi-modal data monitored in real time through the pig feature extraction module and then transmits it to the large live pig health monitoring model. The large model evaluates the disease status. When the large model detects that a pig has corresponding disease characteristics, it is included in the list of abnormal live pigs. When the health abnormality of a live pig is confirmed as a disease risk, the system generates a disease warning message and sends a text message to the pig farm breeders. The disease warning message is, for example: "Time: 2025-03-02 14:35; Suspected disease: Porcine Reproductive and Respiratory Syndrome (PRRS); Affected area: Pigsty B3-2; Details: 3 live pigs were found with body temperature > 40°C and showed symptoms of listlessness and abnormal breathing; Suggested measures: Immediately isolate the live pigs, conduct virus detection, and initiate a vaccine prevention and control plan."
[0161] In addition, the system designs and develops a live pig health monitoring APP based on multi-modal fusion. Users can upload pig images or video data through this APP and call the fine-tuned Qwen2-VL-7B live pig health monitoring large model to achieve remote live pig disease detection and intelligent diagnosis report generation. The APP adopts a collaborative architecture of mobile applications and the cloud. The front end constructs a cross-platform mobile application based on the Flutter framework, supporting dual compatibility for iOS / Android. The back end uses FastAPI to build a high-concurrency microservice architecture, combined with a PostgreSQL relational database and MinIO distributed object storage to achieve data persistent management. The user side supports the shooting, uploading, and preprocessing of JPEG / PNG images and MP4 / AVI videos through the multi-modal data collection module. It uses the FFmpeg component based on WebAssembly to extract local video frames and perform H.265 encoding and compression, and transmits them to the cloud server through the HTTPS+WebSocket dual-channel encryption.
[0162] Embodiment 2, a system for a live pig health monitoring method based on multi-modal fusion, the system includes:
[0163] A multi-modal data collection module for obtaining live pig body temperature data, live pig video data, and live pig sound data;
[0164] A pig feature extraction module for extracting features from the live pig body temperature data, live pig video data, and live pig sound data to obtain live pig body temperature data features, live pig video data features, and live pig sound data features;
[0165] A large model fine-tuning module for fine-tuning the basic large model to obtain a live pig health monitoring large model;
[0166] The pig health monitoring module is used to input the pig body temperature data characteristics, pig video data characteristics, and pig sound data characteristics into the pig health monitoring large model, obtain the pig health monitoring results, and issue warnings, so as to complete the pig health monitoring based on multimodal fusion.
[0167] Those of ordinary skill in the art will realize that the embodiments described herein are to assist the reader in understanding the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the invention.
Claims
1. A method for monitoring the health of live pigs based on multimodal fusion, characterized in that, It includes the following steps: S1: Use the multi-modal data acquisition module to obtain the live pig body temperature data, live pig video data, and live pig sound data; S2: Use the live pig feature extraction module to extract features from the live pig body temperature data, live pig video data, and live pig sound data, and obtain the live pig body temperature data features, live pig video data features, and live pig sound data features; S3: Use the large model fine-tuning module to fine-tune the basic large model to obtain the live pig health monitoring large model; S4: Use the live pig health monitoring module to input the live pig body temperature data features, live pig video data features, and live pig sound data features into the live pig health monitoring large model, obtain the live pig health monitoring results, and give an alarm to complete the live pig health monitoring based on multi-modal fusion.
2. The method for monitoring the health of live pigs based on multimodal fusion according to claim 1, wherein In the S1, the multi-modal data acquisition module includes a monitoring camera, an infrared thermal imaging camera, a microphone array, a network transmission sub-module, and a data storage sub-module. The S1 includes the following sub-steps: S11: Use the infrared thermal imaging camera, monitoring camera, and microphone array to collect the live pig body temperature data, live pig video data, and live pig sound data respectively; S12: Upload the collected live pig body temperature data, live pig video data, and live pig sound data to the data storage sub-module through the network transmission sub-module.
3. The method for monitoring the health of live pigs based on multimodal fusion according to claim 1, wherein The following sub-steps are included in the S2: S21: Use the instance segmentation model Grounded-SAM to perform individual segmentation on pigs in each frame of the live pig video data, and use the tracking algorithm DeepSORT to track the pigs to obtain the video data of the i-th live pig and obtain video data segments of live pigs through the division criteria Realize the extraction of the features of the live pig video data; S22: Obtain the position information of the pigs in the infrared video based on the position information of the pigs in the video data, and use the instance segmentation model Grounded-SAM to perform individual segmentation on the pigs in each frame of the infrared video to obtain the body temperature data of the i-th live pig And obtain the body temperature data segment of the live pig through the division standard After that, perform temperature change detection through a sliding window to extract the characteristics of the live pig body temperature data; S23: Based on the sound data of live pigs, calculate the time difference of arrival of the sound signal by means of time delay estimation, determine the approximate direction of the sound source, and combine the tracking algorithm DeepSORT to obtain the trajectory information of the pigs, so as to realize the matching of the sound and the pigs and obtain the sound data of the i-th live pig And obtain the sound data segment of the live pig through the division standard S24: Perform noise filtering on the pig sound data segment to remove background noise, and perform sound classification based on the trained CNN+LSTM model to extract the characteristics of the pig sound data.
4. The method for monitoring the health of live pigs based on multimodal fusion according to claim 1, wherein, In the S3, the large model fine-tuning module includes a live pig disease dataset construction sub-module, a live pig prompt engineering sub-module, and a large model fine-tuning sub-module. The S3 includes the following sub-steps: S31: Use the live pig disease dataset construction sub-module to integrate multiple pig disease datasets, expand the dataset scale through various data augmentation strategies, and obtain the live pig disease image dataset; S32: Use the live pig prompt engineering sub-module to pre-define the classification and multi-modal feature manifestations of various common pig diseases, design a prompt template, and obtain the live pig disease semantic dataset; S33: Use the large model fine-tuning sub-module to fine-tune the basic large model based on the live pig disease image dataset and the live pig disease semantic dataset to obtain the live pig health monitoring large model.
5. The method for monitoring the health of live pigs based on multimodal fusion according to claim 4, characterized in that The live pig disease image dataset in the S31 specifically includes: Healthy pigs; Porcine reproductive and respiratory syndrome (PRRS) pigs: The characteristics include purple skin, listlessness, and body temperature higher than 40°C; Influenza pigs: The characteristics include coughing, wheezing, increased nasal secretions, and body temperature higher than 39.5°C; Porcine circovirus disease pigs: The characteristics include emaciation, pale or purple skin patches, and decreased immunity; Erysipelas suis pigs: The characteristics include diamond-shaped erythema on the skin, high fever, and body temperature higher than 41°C; Swine pasteurellosis pigs: The characteristics include rapid breathing, coughing, and cyanosis of the skin; Foot-and-mouth disease pigs: The characteristics include hoof and mouth ulcers, salivation, and lameness.
6. The method for monitoring the health of live pigs based on multi-modal fusion according to claim 4, characterized in that, The prompt template in the S32 includes a basic prompt template, a structure prompt template, and a comparison prompt template; The basic prompt template includes a simple query and an inference query; The structure prompt template includes a role-based query; The comparison prompt template is used to compare two images of pigs.
7. The method for monitoring the health of live pigs based on multimodal fusion according to claim 4, characterized in that The following sub-steps are included in the S33: S331: Read the live pig disease image dataset and the live pig disease semantic dataset, and construct image-text pairs; S332: Process the input image through the Qwen2-VL-7B vision encoder to calculate the feature vector of the input image. The formula is: F(I) = VisionEncoder(I) where F(·) is the feature vector of the input image, I is the input image, and VisionEncoder(·) is the Qwen2-VL-7B vision encoder; Process the input text through the Qwen2-VL-7B text encoder to calculate the feature vector of the input text. The formula is: G(T) = TextEncoder(T) where G(·) is the feature vector of the input text, T is the input text, and TextEncoder(·) is the Qwen2-VL-7B text encoder; S333: Structurally process the input text through a prompt template. The formula is: T prompt = PromptTemplate(T) Among them, T prompt is the input text after being processed by the prompt template, and PromptTemplate(·) is the text standardization operation; S334: Use contrastive learning to align the multi-modalities of the image and text. Fine-tune the Qwen2-VL-7B model through the LoRA technique and optimize it in combination with the InfoNCE loss function to obtain a large model for pig health monitoring.
8. The method for monitoring the health of live pigs based on multi-modal fusion according to claim 7, wherein The process of fine-tuning the Qwen2-VL-7B model through the LoRA technique in S334 is expressed as: Qwen2(X v ,X t ) = f θ (E v (X v ),E t (X t )) Among them, Qwen2(·) is the Qwen2-VL model, and X v is the visual input, and X t is the text input, and f θ (·) is a function, and E v (·) and E t (·) respectively represent the visual embedding function and the text embedding function implemented by ViT and Tokenizer; The LoRA technique decomposes the weight matrix into a low-rank representation. The formula is: W = W0 + UV t Among them, W is the weight matrix in the Transformer layer, W0 is the original weight matrix, and U and V t are low-rank matrices; Use the LoRA technique in the attention layer and feed-forward layer of Transformer. The formula is: Among them, Attention(·) is the attention layer of the Transformer, Q, K, and V are the matrices of queries, keys, and values respectively, softmax(·) is the activation function, ΔK represents the changes made to the key matrix during the fine-tuning process, and d k is the key vector, and the superscript T represents the transpose of the matrix.
9. A system for the method of monitoring the health of live pigs based on multimodal fusion as described in claims 1-8, characterized in that, The system includes: A multi-modal data acquisition module for obtaining pig body temperature data, pig video data, and pig sound data; A pig feature extraction module for extracting features from the pig body temperature data, pig video data, and pig sound data to obtain pig body temperature data features, pig video data features, and pig sound data features; A large model fine-tuning module for fine-tuning the basic large model to obtain a large model for pig health monitoring; A pig health monitoring module for inputting the pig body temperature data features, pig video data features, and pig sound data features into the large model for pig health monitoring to obtain pig health monitoring results and issue warnings, completing pig health monitoring based on multi-modal fusion.