A multi-modal fusion-based early warning system and method for pig respiratory diseases
Patent Information
- Application Number
- CN202611269441.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-20
- Publication Date
- 2026-09-25
AI Technical Summary
[0007]本申请提供一种基于多模态融合的猪呼吸道疾病早期预警系统及方法,旨在解决现有技术多模态数据缺乏有效融合、缺乏疾病渐进演变过程的时序建模以及预警决策缺乏不确定性评估的问题
[0056]本申请通过构建“多模态同步采集—时序对齐—深度特征提取—跨模态深度融合—时序状态评估—分级预警决策”的全链路智能预警架构,实现了对猪呼吸道疾病从潜伏期到发病期的全流程早期预警。
Smart Images

Figure CN122822307A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of pig farming technology, specifically relating to an early warning system and method for pig respiratory diseases based on multimodal fusion. Background Technology
[0002] With the rapid development of intensive and large-scale pig farming, health monitoring and early disease warning for pigs have become core requirements in pig farming management. Porcine respiratory diseases (such as porcine reproductive and respiratory syndrome, porcine mycoplasmal pneumonia, and porcine contagious pleuropneumonia) are among the most prevalent and serious diseases in large-scale pig farms, characterized by rapid transmission, long incubation periods, and insidious early symptoms. Once an outbreak occurs in a herd, it often causes large-scale infection and even death, resulting in huge economic losses for pig farms. Therefore, early warning before the disease develops obvious clinical symptoms is of great practical significance for timely isolation, treatment, and prevention measures, and for controlling the spread of the epidemic.
[0003] Currently, monitoring and early warning technologies for swine respiratory diseases have the following main shortcomings:
[0004] There is a lack of effective fusion mechanisms for multimodal data. Although some studies have attempted to combine data from multiple sensors (such as sound, images, and environmental parameters) for pig health assessment, existing solutions typically employ simple feature concatenation or weighted summation for multimodal fusion. Because the information structures of each modality are fundamentally different—audio signals are time-domain waveforms, infrared and visible light images are two-dimensional pixel matrices, and environmental parameters are scalar time series—direct feature concatenation or simple weighted summation fails to fully exploit the complementary information between modalities and cannot eliminate the differences in scale and semantic hierarchy among the modalities, resulting in insufficient discriminative power of the fused features.
[0005] There is a lack of temporal modeling for the gradual evolution of diseases. Swine respiratory diseases typically take several hours to several days to progress from pathogen invasion and changes in the immune response during the incubation period to the gradual appearance of clinical symptoms. Most existing technologies rely solely on sensor data at a single moment for health status assessment, making it difficult to capture subtle trends in the early stages of disease. This can easily lead to misinterpreting abnormal signals during the incubation period as normal, thus missing the optimal early warning opportunity.
[0006] Existing early warning systems lack reliable uncertainty assessment mechanisms. Due to factors such as sensor data quality and model prediction confidence, a single early warning result may carry the risk of false alarms or missed alarms. Current technical solutions typically do not provide confidence information on the prediction results when outputting early warning signals, making it impossible for aquaculture personnel to judge the reliability of the early warning results and thus difficult to make accurate on-site handling decisions. Summary of the Invention
[0007] This application provides an early warning system and method for porcine respiratory diseases based on multimodal fusion, aiming to solve the problems of existing technologies such as lack of effective fusion of multimodal data, lack of time-series modeling of the progressive evolution of diseases, and lack of uncertainty assessment in early warning decisions.
[0008] In a first aspect, there is an early warning system for swine respiratory diseases based on multimodal fusion. The system includes a multimodal data acquisition layer, a data preprocessing and alignment module, a multimodal feature extraction module, a cross-modal feature fusion module, a temporal state evaluation module, an early warning decision module, and a terminal interaction module connected in sequence.
[0009] The multimodal data acquisition layer synchronously acquires audio signals, infrared thermal imaging images, visible light images, and environmental parameters, and adds a hardware-level timestamp to each frame of data acquired.
[0010] The data preprocessing and alignment module preprocesses the original data of each modality and aligns the preprocessed modal data in the time dimension based on the timestamp to generate a multimodal time-series data packet.
[0011] The multimodal feature extraction module extracts audio global feature vectors, infrared feature vectors, visible light feature vectors, and environmental feature vectors from the multimodal time-series data packets, respectively.
[0012] The cross-modal feature fusion module performs deep fusion of the audio global feature vector, infrared feature vector, visible light feature vector, and environmental feature vector to generate cross-modal fused features;
[0013] The cross-modal feature fusion module includes a multi-scale serial fusion attention pyramid submodule and a cross-modal interaction submodule;
[0014] The multi-scale serial fusion attention pyramid submodule is used to perform multi-level cascaded attention filtering on the infrared and visible light features from coarse to fine, refine the region of interest step by step, and output infrared salient features and visible light salient features.
[0015] The cross-modal interaction submodule is used to calculate the bidirectional cross-attention output between any two modalities using the infrared salient features, visible light salient features, audio global features, and environmental features as four modal features participating in the interaction, and to fuse the cross-attention outputs of other modalities received by each modality with its own features to obtain cross-modal fusion features;
[0016] The temporal state assessment module performs temporal modeling on the cross-modal fusion features at multiple consecutive time points, captures the gradual evolution pattern of the disease, and outputs a health status score.
[0017] The early warning decision module generates a graded early warning signal based on the health status score, and sends the graded early warning signal to the terminal interaction module for visualization.
[0018] Optionally, the multimodal data acquisition layer includes:
[0019] The sound acquisition unit consists of an array of multiple microphones deployed on the top of the pigsty, used to collect the audio signals of the pigs.
[0020] The infrared thermal imaging unit, consisting of an infrared thermal imaging camera, is used to acquire infrared thermal images of the pig's face.
[0021] The visible light image acquisition unit, consisting of a high-definition visible light camera, is used to acquire visible light images of pigs;
[0022] The environmental sensing unit consists of a temperature and humidity sensor, an ammonia sensor, and a carbon dioxide sensor, and is used to collect data on the temperature, humidity, ammonia concentration, and carbon dioxide concentration of the pigsty environment.
[0023] The sound acquisition unit, infrared thermal imaging unit, visible light image acquisition unit, and environmental sensing unit acquire data at a unified sampling frequency. Each unit has a built-in unified synchronous clock source, and each frame of data acquired is accompanied by the hardware-level timestamp.
[0024] Optionally, the data preprocessing and alignment module includes:
[0025] The audio preprocessing submodule pre-emphasizes, frames, and windows the audio signal, and extracts the Mel spectrogram.
[0026] The image preprocessing submodule performs spatial registration, size normalization, and grayscale value normalization on the infrared thermal imaging image and the visible light image.
[0027] The environmental data preprocessing submodule performs missing value filling, outlier detection and correction, and normalization on the environmental parameters.
[0028] The timing alignment submodule, based on the timestamp and using the frame rate of the audio signal as the reference clock, matches the infrared image frame, visible light image frame, and environmental parameter record with the closest timestamp for each frame of audio data, forming a multimodal timing data packet.
[0029] Optionally, the multimodal feature extraction module includes:
[0030] The audio feature extraction branch employs a hybrid architecture based on convolutional neural networks and Transformers to extract global audio feature vectors from the Mel spectrogram.
[0031] The infrared feature extraction branch employs a residual network based on dilated convolution to extract infrared feature vectors from the infrared thermal imaging image. The infrared feature extraction branch performs adaptive weighted fusion of the output features of multiple different convolutional blocks.
[0032] The visible light feature extraction branch adopts a Vision Transformer-based architecture to extract visible light feature vectors from the visible light image. The output of the visible light feature extraction branch is connected to the serial fusion attention module, which consists of a compression-excitation module and a triple attention module connected in series to enhance attention to the discriminative regions of pigs.
[0033] The environmental feature extraction branch employs a multi-layer fully connected network to encode the normalized environmental parameters into environmental feature vectors.
[0034] Optionally, the multi-scale serial fusion attention pyramid submodule adopts a multi-level pyramid structure. Each layer divides the input feature map into multiple non-overlapping feature map patches in the spatial dimension. Each feature map patch is input into the serial fusion attention module for processing and then merged according to its original spatial position to obtain the attention map of that layer.
[0035] The output feature map of the previous layer is reweighted by the attention map of the current layer and then passed to the next layer, achieving a progressive refinement of the attention region from coarse to fine.
[0036] Optionally, the cross-modal interaction submodule first maps the four modal features involved in the interaction to the same embedding dimension through a linear projection layer, and then calculates all directed cross-attention outputs;
[0037] For any target modality, the cross-attention outputs received from the other three modalities are concatenated and fused with its own projection features to obtain the enhanced features of that modality;
[0038] The enhanced features of the four interactive modalities are concatenated along the channel dimension and then mapped through a fully connected layer to obtain the final cross-modal fusion feature.
[0039] Optionally, the timing state evaluation module includes:
[0040] The temporal feature aggregation submodule uses a Transformer-based temporal encoder to perform temporal modeling on the cross-modal fusion feature sequence within the sliding time window, capture long-range dependencies across time points, and output the aggregated temporal features.
[0041] The temporal feature aggregation submodule also constructs a Laplace-type template feature set, which is fused with the cross-modal fusion feature at the current moment through a cross-attention mechanism. This propagates historical temporal prior knowledge to the current moment, and finally outputs temporal features that fuse long-range temporal dependencies and historical prior knowledge.
[0042] The health status assessment submodule uses a multi-layer fully connected network to map the time-series features into the health status score.
[0043] Optionally, the early warning decision module includes:
[0044] The graded early warning submodule divides the health status of pigs into four levels based on the numerical range of the health status score: healthy status, attention status, suspected disease status, and high-risk disease status, and triggers no warning, level three warning, level two warning, and level one warning respectively.
[0045] The uncertainty estimation submodule uses the Monte Carlo dropout method to quantify and estimate the cognitive uncertainty of the model prediction. It performs multiple independent random forward propagations on the same input sample and calculates the variance of the multiple outputs as an uncertainty measure. When the variance exceeds a preset threshold, it outputs a warning signal and adds a manual review prompt.
[0046] Secondly, a method for early warning of porcine respiratory diseases based on multimodal fusion is applied to the system described above, the method comprising:
[0047] The multimodal data acquisition layer synchronously acquires audio signals, infrared thermal imaging images, visible light images, and environmental parameters, and adds a hardware-level timestamp to each frame of data acquired.
[0048] The data preprocessing and alignment module preprocesses the original data of each modality, and aligns the preprocessed modal data in the time dimension based on the timestamp to generate a multimodal time-series data packet.
[0049] The multimodal feature extraction module extracts audio global feature vectors, infrared feature vectors, visible light feature vectors, and environmental feature vectors from the multimodal time-series data packets, respectively.
[0050] The audio global feature vector, infrared feature vector, visible light feature vector, and environmental feature vector are deeply fused by the cross-modal feature fusion module to generate cross-modal fused features;
[0051] Specifically, the infrared and visible light features are subjected to multi-level cascaded attention filtering from coarse to fine through a multi-scale serial fusion attention pyramid submodule, which refines the region of interest step by step and outputs infrared and visible light salient features.
[0052] The cross-modal interaction submodule uses infrared salient features, visible light salient features, audio global features, and environmental features as four modal features participating in the interaction. It calculates the bidirectional cross-attention output between any two modalities and fuses the cross-attention outputs of other modalities received by each modality with its own features to obtain cross-modal fusion features.
[0053] The time-series state assessment module performs time-series modeling on the cross-modal fusion features at multiple consecutive time points, captures the gradual evolution pattern of the disease, and outputs a health status score.
[0054] The early warning decision module generates graded early warning signals based on the health status score, and presents them visually through the terminal interaction module.
[0055] Compared with the prior art, this application has at least the following beneficial effects:
[0056] This application constructs a full-link intelligent early warning architecture that includes "multimodal synchronous acquisition, temporal alignment, deep feature extraction, cross-modal deep fusion, temporal state assessment, and hierarchical early warning decision-making," thereby achieving early warning of swine respiratory diseases throughout the entire process from the incubation period to the onset period.
[0057] This application overcomes the shortcomings of single-modal acquisition in complex pigsty environments by setting up a multimodal data acquisition layer, which synchronously acquires four types of modal data—audio, infrared thermal imaging, visible light images, and environmental parameters—with a unified sampling frequency and hardware-level timestamps. This provides time-consistent and information-rich multidimensional raw data for subsequent multimodal fusion analysis.
[0058] This application sets up a multi-scale serial fusion attention pyramid submodule to perform multi-level cascaded attention filtering of infrared and visible light features from coarse to fine, realizing a step-by-step transition from global perception at the whole image level to fine filtering of local areas. It can simultaneously capture discriminative feature regions at different spatial scales, effectively improving the model's sensitivity to early minor symptoms (such as local small-scale temperature anomalies).
[0059] This application sets up a cross-modal interaction submodule to calculate the bidirectional cross-attention output between any two modalities, realizing full interaction and complementary enhancement of four heterogeneous modal features (audio, infrared, visible light, and environment) at the semantic level. This overcomes the shortcomings of traditional simple splicing or weighted summation fusion methods that are difficult to mine complementary information between modalities, and enables deep alignment and collaborative enhancement of discriminative cues of each modality at the feature level.
[0060] This application sets up a temporal state evaluation module and uses a Transformer-based temporal encoder to dynamically model cross-modal fusion features at multiple consecutive time points. It also combines a Laplace-type template feature set to propagate historical temporal prior knowledge to the current time point, effectively capturing the gradual evolution of swine respiratory diseases from the incubation period to the onset period. This improves the problem that single-time point determination is difficult to identify early subtle change trends. Attached Figure Description
[0061] Figure 1 A schematic diagram of module connections for an early warning system for swine respiratory diseases based on multimodal fusion, provided as an embodiment of this application;
[0062] Figure 2 This is a flowchart illustrating an early warning method for swine respiratory diseases based on multimodal fusion, provided as an embodiment of this application. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments.
[0064] This application provides an early warning system for porcine respiratory diseases based on multimodal fusion, such as... Figure 1 As shown, the system includes: a multimodal data acquisition layer, a data preprocessing and alignment module, a multimodal feature extraction module, a cross-modal feature fusion module, a time-series state evaluation module, an early warning decision module, and a terminal interaction module, which are connected in sequence.
[0065] The multimodal data acquisition layer is deployed inside the pigsty to simultaneously collect multi-dimensional raw data related to swine respiratory diseases. Pigsties are typically equipped with fans, feed lines, and other equipment. The background noise generated during operation (fan noise, feed line vibration, and a mixture of pig noises) is often higher than the energy level of pig coughs. Single-modal acquisition methods are unlikely to obtain effective disease characteristic signals in such a noisy environment.
[0066] Therefore, this layer adopts a multimodal collaborative acquisition architecture, specifically including:
[0067] The sound acquisition unit consists of an array of multiple high-sensitivity microphones deployed on the top of the pigpen, used to collect audio signals such as coughing, panting, and sneezing sounds from the pigs. The microphone array is arranged in a specific geometric configuration above the pigpen, and a beamforming algorithm is used to form a directional sound pickup beam in space, amplifying the sound signal from the target direction and suppressing background interference from other directions.
[0068] Building upon this foundation, a noise suppression algorithm is used to filter out environmental noise from fans, feed lines, and other sources. Furthermore, a sound source localization algorithm based on Time Difference of Arrival (TDOA) is employed. This algorithm calculates the spatial location of the sound source by determining the time delay difference between the arrival times of the same sound signal at different microphones, thereby enabling the pointing and localization of specific coughing sounds within a group environment and distinguishing the sound sources from different pens or different pigs. The microphone array's sampling frequency is no less than 16kHz to ensure effective acquisition of the high-frequency components of the coughing sound.
[0069] The infrared thermal imaging unit, composed of an infrared thermal imaging camera, is used to acquire infrared thermal images of the pig's face and obtain data on the pig's body surface temperature distribution. Infrared thermal imaging technology is a non-invasive and highly efficient method of body temperature measurement, enabling remote sensing of pig body temperature without contact or stress. When pigs have respiratory infections, early abnormal changes in facial temperature (especially around the eyes and ears) will appear.
[0070] Infrared thermal imaging cameras automatically trigger data acquisition when pigs are feeding or drinking. At this time, the pig's head position is relatively fixed, making it easy to obtain standardized facial thermal images. After image processing, the acquired thermal infrared images can extract temperature information from characteristic areas at the base of the pig's ears, serving as an important basis for judging the pig's physiological state.
[0071] The visible light image acquisition unit, composed of a high-definition visible light camera, is used to acquire visible light images of pigs, obtaining visual information such as their body shape, behavior, and mental state. Visible light images and infrared thermal images are acquired simultaneously by dual-light cameras, and the two are pre-registered spatially to ensure pixel-level correspondence between the two modalities at any given time. The visible light images provide detailed information such as the pig's outline, color, and posture, complementing the temperature information from the infrared thermal imaging.
[0072] Environmental sensing unit: Composed of temperature and humidity sensors, ammonia sensors, and carbon dioxide sensors, it is used to collect environmental parameters such as temperature, humidity, ammonia concentration, and carbon dioxide concentration in the pigsty. The sensor group is installed on the two side walls of the pigsty, at a height level with the pigs' breathing zone (approximately 0.5 to 1.0 meters from the ground) to ensure that the collected environmental parameters accurately reflect the gas concentration levels in the pigs' breathing area.
[0073] The real-time collected data is aggregated and uploaded through the master node. Changes in environmental parameters (such as increased ammonia concentration and abnormal temperature and humidity) are significantly correlated with the occurrence of respiratory diseases in pigs, and environmental sensor data provides important contextual information for subsequent multimodal fusion analysis.
[0074] The sound acquisition unit, infrared thermal imaging unit, visible light image acquisition unit, and environmental sensing unit all acquire data at a uniform sampling frequency (preferably 30 frames / second). Each unit has a built-in high-precision crystal oscillator to provide a unified synchronous clock source, and each frame of data acquired is accompanied by a hardware-level timestamp.
[0075] The data acquisition layer receives data from each unit simultaneously through the microcontroller's sensor interface, ensuring precise alignment of the modal data in the time dimension and providing time-consistent input data for subsequent multimodal feature fusion.
[0076] The input of the data preprocessing and alignment module is connected to the output of the multimodal data acquisition layer, and is used to preprocess and time-series align the acquired multimodal raw data.
[0077] Since the raw data output by each sensor unit in the multimodal data acquisition layer has significant differences in format, dimension, and sampling characteristics, for example, audio signals are time-domain waveforms, infrared and visible light images are two-dimensional pixel matrices, and environmental parameters are scalar time series, if these heterogeneous data are directly input into the feature extraction network, not only will the network training be difficult due to the difference in dimension, but the correlation of cross-modal features will also be weakened due to the slight temporal shift of each modal data.
[0078] Therefore, this module includes the following sub-modules for targeted processing of various types of raw data, specifically:
[0079] The audio preprocessing submodule performs preprocessing operations such as pre-emphasis, framing, and windowing on the acquired audio signal. The pre-emphasis processing uses a first-order high-pass filter with the following transfer function: α is taken as 0.95~0.97, which is used to compensate for the attenuation of high-frequency components of the sound signal during propagation, so that the high-frequency features (such as plosive components) in the cough sound can be preserved.
[0080] Frame segmentation divides the audio signal into several short frames with a frame length of 25ms (corresponding to 400 sampling points and a sampling rate of 16kHz), a frame shift of 10ms, and maintains an overlap rate of about 60% between adjacent frames to ensure the temporal continuity between frames.
[0081] Windowing processing employs a Hamming window to weight each frame of signal, reducing the spectral leakage effect introduced by frame division. After preprocessing, a Fast Fourier Transform (FFT) is performed on each frame to convert the time-domain signal into a frequency-domain representation. Then, a set of triangular bandpass filter banks (Melfilterbank) is used for Mel-frequency mapping. The number of filters is set to 40 or 64 to cover the frequency range sensitive to human hearing. Finally, the logarithmic energy is taken to obtain the Mel-spectrogram as the audio feature representation.
[0082] While extracting features, this submodule uses spectral subtraction or speech enhancement algorithms based on psychoacoustic models to suppress background noise. It performs spectral subtraction on stable background interference such as fan noise and feed line vibration noise that are present in the pig house environment. It also sets an adaptive endpoint detection threshold based on the temporal energy envelope characteristics of pig coughing sounds to effectively separate the coughing sound segment from the silent segment and the non-coughing sound segment.
[0083] The image preprocessing submodule performs preprocessing operations such as size normalization, grayscale value normalization, and data augmentation on infrared thermal imaging images and visible light images. Due to the spatial discrepancy between the installation positions of the infrared thermal imaging camera and the visible light camera above the pigpen in the multimodal data acquisition layer, a fixed geometric transformation relationship exists between the images acquired by the two cameras.
[0084] Before normalizing the image size (uniformly adjusting it to 256×128 pixels), this submodule first performs spatial registration of the infrared and visible light images using the factory calibration parameters of the dual-light imaging module: using a checkerboard calibration board or a dedicated dual-light calibration target as a reference, it pre-obtains the rotation matrix between the two cameras. Translation vector Based on this, an affine transformation model is established to map the visible light image to the coordinate system of the infrared image through affine transformation, thereby achieving spatial alignment between the infrared image and the visible light image at the pixel level.
[0085] After spatial registration, the registered infrared thermal imaging image and visible light image were normalized to a uniform size of 256×128 pixels. The maximum-minimum normalization method was used to normalize the image pixel values to the [0,1] interval to eliminate the dimensional effects caused by the difference in dynamic range between different lighting conditions and different thermal imaging devices. During the training phase, data augmentation strategies such as random horizontal flipping, random cropping, and random rotation (±5°) were used to expand the sample diversity and improve the model's generalization ability to different shooting angles and pig postures.
[0086] The environmental data preprocessing submodule performs missing value filling, outlier detection and correction, and normalization on environmental sensor data.
[0087] Since environmental sensors may experience data loss or anomalies due to communication interruptions, sensor drift, or occasional malfunctions during long-term operation, this submodule first performs quality checks on the four types of environmental parameters collected: temperature, humidity, ammonia concentration, and carbon dioxide concentration. For temporary data loss caused by communication packet loss (the loss time does not exceed 5 consecutive sampling points), linear interpolation is used to fill the gaps.
[0088] For isolated outliers caused by occasional sensor malfunctions (i.e., the reading at a certain moment deviates from the mean of the preceding and following moments by more than 3 times the standard deviation), the median of the preceding and following moments is used for replacement and correction; for continuous long-term (more than 30 sampling points) data loss or unrecoverable sensor malfunctions, the environmental data for that period is marked as invalid, and the environmental feature input for that period is automatically masked in subsequent fusion processing.
[0089] After completing missing filling and anomaly correction, the maximum and minimum value normalization method is used to map the four types of environmental parameters to the [0,1] interval respectively, eliminating the adverse effects of differences in the units and numerical ranges of different environmental parameters on network training;
[0090] The time alignment submodule, based on a unified timestamp, precisely aligns the preprocessed data of each modality in the time dimension to generate multimodal time-series data packets;
[0091] Since the sound acquisition unit, infrared thermal imaging unit, visible light image acquisition unit and environmental sensing unit have marked the acquisition time of each frame of data in the data acquisition layer through a hardware-level timestamp synchronization mechanism, this submodule uses the frame rate of the audio signal as the reference clock (30 frames / second) to find the infrared image frame, visible light image frame and environmental parameter record with the closest timestamp for each frame of audio data on the time axis, forming a homogeneous time series indexed by the audio frame.
[0092] For system-level time deviations caused by inconsistent sampling start times of each unit, this submodule corrects them by sending a synchronization pulse signal during the system initialization phase: the main controller sends a synchronization trigger signal to each acquisition unit at the same time, and each unit records its own response delay after receiving the trigger signal. The timestamp is then compensated and corrected based on the delay value during subsequent data acquisition.
[0093] Each aligned multimodal time-series data packet contains a Mel spectrogram (extracted from the audio signal) at the same time, a spatially registered infrared thermal image, a spatially registered visible light image, and a normalized four-dimensional environmental parameter vector. The data packets are arranged in chronological order to form a multimodal time-series data stream, which is then input to the multimodal feature extraction module.
[0094] The input of the multimodal feature extraction module is connected to the output of the data preprocessing and alignment module, and is used to extract the depth discrimination features of each modality from the aligned multimodal time-series data packets.
[0095] Since the modal data output by the data preprocessing and alignment modules have fundamental differences in information structure, the Mel spectrogram is a two-dimensional time-frequency representation, the infrared and visible light images are two-dimensional spatial representations, and the environmental parameters are one-dimensional vectors. If a unified feature extraction network is used to process the data of different modalities, it will be difficult to adapt to the data characteristics of each modality and fully explore its discriminative information.
[0096] To this end, this module sets up four parallel feature extraction branches based on the structural characteristics of each modality of data. After each branch independently completes feature extraction, the extracted feature vectors are output uniformly to the cross-modal feature fusion module. Specifically, this includes:
[0097] The audio feature extraction branch employs a hybrid architecture based on Cn-Transformer to extract deep features from the Mel spectrogram of the audio signal, which contain both local time-frequency details and global temporal dependencies. Specifically, this branch consists of a concatenated convolutional front end and a Transformer encoder.
[0098] The convolutional front end contains four sequentially connected two-dimensional convolutional blocks. Each convolutional block consists of a convolutional layer, a batch normalization layer, and a ReLU activation function layer. The first convolutional block uses a 7×7 convolutional kernel with a stride of 2 and 64 output channels to capture large-scale time-frequency texture patterns in the Mel spectrogram. The second convolutional block uses a 5×5 convolutional kernel with a stride of 1 and 128 output channels to extract mesoscale local features. The third and fourth convolutional blocks both use 3×3 convolutional kernels with a stride of 1 and 256 and 512 output channels, respectively, to refine the extraction of high-frequency detail features.
[0099] Each convolutional block is followed by a 2×2 max-pooling layer to progressively reduce the spatial resolution of the feature maps and expand the receptive field of subsequent convolutional layers. The output of the convolutional front end is of size [size missing]. The three-dimensional feature tensor, in which and These represent the height and width of the feature maps after convolution and pooling, respectively.
[0100] The three-dimensional feature tensor is expanded along the channel dimension into a one-dimensional sequence with a sequence length of [value missing]. Each location has a feature dimension of 512, which serves as the input to the Transformer encoder. The Transformer encoder consists of four stacked coding layers, each containing a multi-head self-attention layer and a feedforward network.
[0101] The multi-head self-attention layer uses eight attention heads, each with a feature dimension of 64, to capture long-range time-frequency dependencies in different subspaces in parallel. The feedforward network consists of two fully connected layers: the intermediate layer has a dimension of 2048, and the output layer has a dimension of 512, using the GELU activation function. Residual connections and layer normalization are applied after each sublayer to ensure gradient stability during deep network training. The output of the Transformer encoder is aggregated along the sequence dimension after global average pooling to obtain the global audio feature vector. ,in Take 512;
[0102] The infrared feature extraction branch employs a residual network based on dilated convolution, preferably a ResNet50-IBN-a network pre-trained on the ImageNet dataset that incorporates dilated convolution, to extract high-resolution depth features from infrared thermal imaging images while maintaining spatial resolution.
[0103] Temperature differences in different areas of a pig's face in infrared thermal imaging images often manifest as localized hot spots. The network needs to have a sufficient receptive field to capture these localized temperature anomalies while maintaining high spatial resolution.
[0104] Specifically, the infrared feature extraction branch contains the first to fifth convolutional blocks (Block1 to Block5) connected in sequence.
[0105] The first convolutional block contains a 7×7 convolutional layer (stride 2, output channels 64) and a 3×3 max pooling layer (stride 2).
[0106] The second convolutional block contains 3 residual units. Each residual unit consists of three convolutional layers of 1×1, 3×3, and 1×1 forming a bottleneck structure, with 256 output channels.
[0107] The third convolutional block contains 4 residual units and has 512 output channels;
[0108] The fourth convolutional block contains 6 residual units and has 1024 output channels;
[0109] The fifth convolutional block contains 3 residual units and has 2048 output channels;
[0110] To avoid the problem of low feature map resolution and loss of information in small temperature anomaly regions in deep features due to stepwise downsampling in standard ResNet50, the fourth and fifth convolutional blocks do not perform downsampling operations (stride set to 1), and dilated convolutions with dilation rates of 2 and 4 are used instead of standard 3×3 convolutions, respectively. This expands the receptive field of the fourth convolutional block to an equivalent 9×9 convolutional kernel and the receptive field of the fifth convolutional block to an equivalent 13×13 convolutional kernel, thereby capturing a wider range of contextual temperature distribution information while maintaining a high-resolution feature map of 14×14 (for a 256×128 input image).
[0111] The infrared feature extraction branch adaptively weights and fuses the output features of the second, fourth, and fifth convolutional blocks. The feature map size of the second convolutional block is 32×16×256, providing local edge and contour information of the pig; the feature map size of the fourth convolutional block is 14×14×1024, providing mid-level temperature region distribution information; and the feature map size of the fifth convolutional block is 14×14×2048, providing high-level semantic temperature pattern information. Before fusion, the feature map of the second convolutional block is first upsampled to 14×14 resolution using bilinear interpolation, and the number of channels is adjusted to 2048 using a 1×1 convolution to maintain consistency with the feature map of the fifth convolutional block in terms of spatial size and number of channels. Fusion weights. , , Obtained through network adaptive learning, satisfying The fused infrared features Represented as:
[0112]
[0113] in , , These are the output features of the second, fourth, and fifth convolutional blocks after size and channel adjustments, respectively. The fused features are then subjected to global average pooling to obtain the infrared feature vector. ,in Take 2048;
[0114] The visible light feature extraction branch adopts a Vision Transformer (ViT) based architecture to extract deep features rich in global contextual information and fine-grained appearance details from visible light images.
[0115] Visible light images contain richer color, texture, and contour information than infrared images. This information is distributed across various regions of the image, requiring the network to have global modeling capabilities to fully capture it. Specifically, this branch first divides the input visible light image of size 256×128 into non-overlapping image patches of size 16×16 pixels, resulting in a total of (256 / 16)×(128 / 16)=16×8=128 image patches;
[0116] Each image patch is flattened into a 16×16×3=768-dimensional vector, which is mapped to a 512-dimensional embedding space using a learnable linear projection matrix, resulting in 128 image patch embedding vectors. Based on this, a learnable one-dimensional positional encoding is added, with each positional encoding being 512-dimensional. This encoding is added element-wise to the corresponding image patch embedding vector to preserve the spatial arrangement information of each image patch in the original image. The 128 positionally encoded embedding vectors, along with a learnable class token, are concatenated into an input sequence of length 129, which is then fed into the Transformer encoder.
[0117] The Transformer encoder consists of six stacked coding layers. Each coding layer contains a multi-head self-attention layer and a feedforward network. The multi-head self-attention layer has 12 attention heads, each with a feature dimension of 512 / 12≈43, used to capture global dependencies between different image patches in parallel. The feedforward network consists of two fully connected layers, with the intermediate layer having a dimension of 2048 and the output layer having a dimension of 512, using the GELU activation function. Each sub-layer is followed by residual connections and layer normalization structures.
[0118] To further enhance the ability of visible light features to focus on discriminative regions of pigs (such as the eyes, ears, nose, and other areas closely related to respiratory diseases), a Serial Fusion Atention Module (SFAM) is introduced after the Transformer encoder.
[0119] The SFAM consists of a squeeze-and-excitation (SE) module and a triple attention module connected in series. The SE module first performs global average pooling on the input feature map. The spatial information of each channel is compressed into a one-dimensional descriptor. The c-th element is calculated as follows: ;
[0120] Then, the channel dependencies of z are modeled through two fully connected layers. The first fully connected layer reduces the dimension from C to C / r (r is the reduction ratio, taken as 16) and uses ReLU activation. The second fully connected layer restores the dimension to C and uses Sigmoid activation to generate the modulation weights of each channel. Finally, the modulation weights are multiplied with the original feature map channel by channel to achieve feature recalibration of the channel dimension.
[0121] The output of the SE module is then fed into the triple attention module. The triple attention module contains three parallel attention branches: the first branch establishes an attention interaction between the height dimension and the channel dimension, rotates the input features 90° counterclockwise along the height axis, performs Z-pooling (concatenates the results of average pooling and max pooling along the channel dimension) and 7×7 convolution, and then rotates them back to the original direction along the height axis;
[0122] The second branch establishes attention interaction between the width dimension and the channel dimension, and the processing flow is symmetrical to that of the first branch.
[0123] The third branch learns attention weights directly along the channel dimension. The outputs of the three branches are first summed element-wise, then divided by 3 and averaged to obtain a refined 3D attention weight tensor. This attention weight tensor is then multiplied element-wise with the output features of the SE module to obtain the enhanced visible light features. After global average pooling Take 512;
[0124] The environmental feature extraction branch employs a multi-layer fully connected network to encode the preprocessed four-dimensional environmental parameters (temperature, humidity, ammonia concentration, and carbon dioxide concentration) into an environmental feature vector. This branch consists of three fully connected layers connected in series: the first layer maps the 4-dimensional input to 64 dimensions using the ReLU activation function; the second layer maps the 64-dimensional input to 128 dimensions using the ReLU activation function; and the third layer maps the 128-dimensional input to 32 dimensions without using an activation function, directly outputting the environmental feature vector. ,in Take 32;
[0125] The temperature, humidity, ammonia concentration, and carbon dioxide concentration information encoded by the environmental feature vector serve as important contextual references for determining the health status of pigs. Studies have shown that elevated ammonia concentration and abnormal temperature and humidity are important environmental risk factors that induce respiratory diseases in pigs.
[0126] The audio feature vector Infrared feature vectors Visible light feature vectors and environmental feature vectors The features are concatenated along the channel dimension to form a multimodal feature set, which is then uniformly output to the input of the cross-modal feature fusion module. The dimensions of the four feature vectors are configured as follows: The total dimension of the spliced multimodal features is 3104.
[0127] The input of the cross-modal feature fusion module is connected to the output of the multimodal feature extraction module, and is used to process the audio feature vector output by the multimodal feature extraction module. Infrared feature vectors Visible light feature vectors and environmental feature vectors Perform deep fusion to generate cross-modal fusion features ;
[0128] Since the output features of each modality feature extraction branch have significant differences in semantic level and spatial scale, audio features are represented in the time-series frequency domain, infrared and visible light features are represented in the spatial structure, and environmental features are represented in the low-dimensional vector, it is difficult to fully explore the complementary information between modalities by directly concatenating features or simply weighting and summing them.
[0129] To this end, this module sets up a multi-scale serial fusion attention pyramid submodule and a cross-modal interaction submodule. Through a progressive attention filtering process from coarse to fine and bidirectional cross-attention interaction, it achieves deep alignment and complementary aggregation of heterogeneous modal features. Specifically, it includes:
[0130] The multi-scale serial fusion attention pyramid submodule is used to extract salient regions from the infrared and visible light features output by the multimodal feature extraction module, from coarse to fine. Since the discriminative features related to respiratory diseases on the pig's face (such as abnormal eye temperature, changes in nasal moisture, and ear congestion) are distributed across different spatial scales, early mild symptoms typically manifest as subtle local temperature changes, which are fine-grained features; while obvious symptoms manifest as abnormal body surface temperature and changes in behavior and posture over larger areas, which are coarse-grained features. A single-scale attention mechanism cannot simultaneously cover discriminative regions at different scales. Therefore, this submodule adopts a multi-layered cascaded segmentation-attention-merging architecture to progressively refine the region of interest.
[0131] Specifically, let the spatial dimensions of the infrared feature map and the visible light feature map input to this submodule both be H×W, and the number of channels be C. For the i-th layer of the pyramid (i=0,1,2,…,L−1, where L is the total number of layers in the pyramid, and L is 3 in this embodiment), the input feature map is first uniformly divided in the spatial dimension into… Number of non-overlapping feature map patches, number of segments By cardinality Decide( Take 2), satisfying That is, the 0th layer is divided into 1×1=1 blocks (no division), the 1st layer is divided into 2×2=4 blocks, and the 2nd layer is divided into 4×4=16 blocks;
[0132] As the number of layers increases, the number of segments grows exponentially, and the spatial area covered by each segment gradually shrinks, thus achieving a gradual transition from "coarse-grained" global perception at the whole image level to "fine-grained" fine filtering of local areas.
[0133] In each layer, the resulting segments will be... Each feature map block is input into the Serial Fusion Attention (SFAM) module for processing;
[0134] SFAM first recalibrates the channel dimensions of each patch through the compression-excitation module, highlighting the feature channels that contribute significantly to the discriminative power of the current patch;
[0135] Then, a triple attention module is used to establish attention interactions in the three dimensions of height-channel, width-channel, and pure channel, generating spatial-channel joint attention weights for each patch. The attention-weighted feature maps output from each patch after SFAM processing are merged according to their original spatial positions to obtain the attention map of that layer. ;
[0136] Output feature map of the previous layer Through attention maps After reweighting, the data is passed to the next level. This process can be represented as follows:
[0137]
[0138] in Indicates attention map Softmax normalization is performed along the channel dimension to make the sum of attention weights at each spatial location equal to 1; This indicates element-wise multiplication; This is the original infrared or visible light feature map input to this submodule;
[0139] Through this coarse-to-fine progressive transmission mechanism, the roughly identified regions in the previous layer are further refined into more precise local salient regions in subsequent layers. The network gradually focuses on the fine-grained spatial locations most closely associated with respiratory diseases. This submodule will process the infrared salient feature map after the three-layer pyramid. and visible light salient feature map Global average pooling is performed separately to obtain the infrared salient feature vectors. and visible light salient feature vector Along with audio feature vectors and environmental feature vectors Input the cross-modal interaction submodule together;
[0140] The cross-modal interaction submodule is used to perform bidirectional information exchange and deep fusion between modes of infrared salient features, visible light salient features, audio global features and environmental features after pyramid refinement;
[0141] Because the four modal features are complementary at the semantic level—infrared features reflect abnormal body temperature distribution, visible light features reflect changes in body shape and behavior, audio features reflect abnormal respiratory sounds, and environmental features reflect external risk factors that induce disease—they are not independent of each other, but rather have complex relationships (e.g., increased environmental ammonia concentration may cause respiratory irritation in pigs, leading to coughing, which in turn manifests as abnormal body temperature).
[0142] To model the causal relationships and collaborative change patterns across modalities, this submodule employs a fully connected cross-attention architecture to enable bidirectional information transfer between any two modalities.
[0143] Specifically, let the feature vectors of the four modalities involved in the interaction be as follows: The corresponding feature dimensions are respectively Before interaction, the modal features are first mapped to the same embedding dimension through four independent linear projection layers. (In this embodiment) (Using 256) to eliminate differences in dimensionality and semantic level among modal features:
[0144]
[0145] in Let be the learnable projection parameters of the m-th mode;
[0146] For any two modes m and n (m ≠ n), calculate the cross-attention output from mode n to mode m. Taking the transmission from mode n to mode m as an example, the projection features of mode m are... Generate query vector through linear transformation The projection features of mode n Generate key vectors through linear transformation Sum value vector :
[0147]
[0148] in It is a learnable linear transformation matrix. To query the dimensions of the vector and key vector (in this embodiment) (Take 64). Then the cross-attention output from mode n to mode m is:
[0149]
[0150] The physical meaning of this formula is: using the features of mode m as query conditions, retrieving the information fragments most relevant to the current mode m from the features of mode n and aggregating them, so as to realize the information supplementation of mode m by mode n.
[0151] For the four modalities, this submodule calculates all 4×3=12 directed cross-attention outputs. For each modality m, the cross-attention outputs received from the other three modalities are combined with its own projected features. The features are spliced together and then fused through a fully connected network to obtain the enhanced features of this modality. :
[0152]
[0153] in This indicates that the components are joined along the feature dimension. and These are learnable fusion parameters;
[0154] Enhancement features of four modes After concatenation along the channel dimension, the output dimension is... (In this embodiment) By taking the fully connected layer mapping of 512, the final cross-modal fusion features are obtained. :
[0155]
[0156] The cross-modal fusion features By integrating infrared temperature anomaly information, visible light body behavior information, audio respiratory sound information, and environmental risk factor information, the discriminative cues of each modality achieve full interaction and complementary enhancement at the feature level. The output of the cross-modal feature fusion module is connected to the input of the temporal state evaluation module, which will... These serve as input features for the time-series state evaluation module at each time step.
[0157] The input of the time-series state evaluation module is connected to the output of the cross-modal feature fusion module, and is used to evaluate the cross-modal fused features output by the cross-modal feature fusion module at each time step. Dynamic modeling is performed over time to capture the gradual evolution of swine respiratory diseases from the incubation period to the onset period;
[0158] Since the development of respiratory diseases is not a sudden event, it usually takes several hours to several days from the invasion of pathogens, changes in immune response during the incubation period to the gradual appearance of clinical symptoms. It is difficult to capture the subtle changes in the early stage of the disease by using only the fusion characteristics at a single moment to judge the health status, and it is easy to misjudge the abnormal signals during the incubation period as normal.
[0159] To this end, this module includes a temporal feature aggregation submodule and a health status assessment submodule. By fusing historical time-series information with current information, it achieves effective modeling of the gradual progression of diseases. Specifically, it includes:
[0160] The temporal feature aggregation submodule is used to perform temporal modeling of cross-modal fusion features across multiple consecutive time points, aggregating disease evolution clues from historical time-series information. This submodule adopts a Transformer-based temporal encoder architecture, combining learnable positional encoding and a Laplacian template feature set to achieve the fusion of historical and current information.
[0161] Specifically, let the current time be t, and use a sliding time window of length T (in this embodiment, T is 16, corresponding to a historical window of approximately 0.53 seconds) to extract the cross-modal fusion feature sequence. ,in ( Take 512) as the cross-modal fusion feature at time τ;
[0162] Since the Transformer encoder itself does not have the ability to perceive temporal order, this submodule first adds a learnable one-dimensional positional code to each position τ in the sequence. The positional encoding is added element-wise to the corresponding cross-modal fusion features to form an input sequence with temporal information:
[0163]
[0164] Location coding During network training, the parameters are updated together with the model parameters, enabling the network to adaptively learn the relative importance of different time steps based on the training data.
[0165] Position-encoded input sequence The input is then fed into a Transformer temporal encoder. This encoder consists of four stacked coding layers, each containing a multi-head self-attention layer and a feedforward network. The multi-head self-attention layer has eight attention heads, each with a feature dimension of 64, used to compute attention weights between any two time points within a time window in parallel. This captures long-range dependencies across time steps; for example, a feature change at the current time step may be causally related to an anomaly that occurred 5 seconds ago. The multi-head self-attention mechanism can directly establish such attention connections across time steps without being limited by time distance.
[0166] The feedforward network consists of two fully connected layers: a 2048-dimensional middle layer and a 512-dimensional output layer, using the GELU activation function. Each sub-layer is followed by residual connections and layer normalization. The Transformer encoder outputs aggregated temporal features. ;
[0167] Concurrently, this submodule also constructs a Laplacian Template Feature Set (LTFS) for explicitly modeling disease evolution priors in historical time-series information. This template set is built during the model training phase, extracting typical temporal variation patterns under healthy states from the training data as templates. Specifically, a Laplacian template set is constructed for different template frame features in the time series. Where n is the number of templates (n is 8 in this embodiment), each template Defined as a prior weight function centered at a health status baseline and following a Laplace distribution:
[0168]
[0169] Where u is the baseline value of health status (set to 1.0 in this embodiment). Let be the health assessment value of the i-th template frame, and b be a scaling parameter (0.3 in this embodiment) that controls the steepness of the Laplace distribution. The smaller b is, the more sensitive the template is to deviations from a healthy state. Compared to the Gaussian distribution, the Laplace distribution has a sharper peak and a thicker tail at the center point, making it more sensitive to subtle deviations near the health baseline and suitable for capturing weak abnormal signals in the early stages of disease.
[0170] LTFS and cross-modal fusion features at the current time step After multiplication, prior temporal knowledge is forward-propagated to the current time step through a cross-attention mechanism. Specifically, LTFS is used as the source of the key and value vectors, and the current time-step features are used as the source of the query vector. The cross-attention output from LTFS to the current time step is calculated, which enhances or suppresses the parts of the current time-step features that match or deviate from the historical healthy baseline pattern. The result of this cross-attention output is compared with the temporal features output by the Transformer encoder. By concatenating along the channel dimension and fusing them through a fully connected network layer, the final time-series features, which aggregate long-range temporal dependencies and historical prior knowledge, are obtained. (In this embodiment) Take 256);
[0171] The health status assessment submodule is used to aggregate the time-series features output by the time-series feature aggregation submodule. This is mapped to a health status score for pigs.
[0172] This submodule consists of a three-layer fully connected network: the first layer maps 256-dimensional temporal features to 128-dimensional features using the ReLU activation function and batch normalization layers; the second layer maps 128-dimensional features to 64-dimensional features using the ReLU activation function and batch normalization layers; the third layer maps 64-dimensional features to 1-dimensional features using the Sigmoid activation function, outputting a scalar value S∈[0,1], representing the pig's health status score at the current moment. A score closer to 0 indicates a poorer health condition and a higher risk of disease; a score closer to 1 indicates a good health condition.
[0173] The output of the health status assessment submodule is connected to the input of the early warning decision module, using the health status score S as input to trigger tiered early warnings. Simultaneously, this submodule employs supervised learning during the training phase, using manually labeled health status tags (healthy / sub-healthy / ill) as supervisory signals and optimizing network parameters using a binary cross-entropy loss function.
[0174] The input of the early warning decision module is connected to the output of the time-series state assessment module. It receives the health status score (SS) output by the time-series state assessment module and generates a tiered early warning signal based on this score. Since the health status score from a single inference may be affected by factors such as the randomness of model parameters and instantaneous fluctuations in input data, directly relying on a single score for early warning decisions can easily lead to false alarms or missed alarms. Therefore, this module, based on the tiered early warning mechanism, further introduces an uncertainty estimation submodule. This submodule quantifies the confidence level of the model prediction to provide a reliability assessment basis for early warning decisions. When the uncertainty of the prediction result exceeds a preset threshold, it actively prompts for manual review. Specifically, this includes:
[0175] The graded early warning submodule classifies the pigs' health status into four levels based on the health status score S∈[0,1] and triggers corresponding early warning signals. The specific judgment rules are as follows:
[0176] When S≥0.8, the pig is considered to be in a healthy state, indicating that all the multimodal characteristic indicators of the pig are within the normal fluctuation range and no warning signal is triggered. The system only records the current score value for subsequent trend analysis.
[0177] When 0.6≤S<0.8, it is judged as a state of concern, indicating that some characteristic indicators of the pigs show a slight trend of deviating from the normal range, but it is not enough to confirm a pathological change, triggering a level three warning (blue warning). At this time, the system marks the pig with a blue icon on the real-time monitoring interface of the terminal interaction module and generates a prompt message "It is recommended to strengthen daily observation and increase the frequency of inspections" to remind the farmers to pay attention to the behavior and physiological changes of the pigs.
[0178] When 0.4≤S<0.6, the pig is judged to be in a suspected disease state, indicating that multiple characteristic indicators of the pig have deviated significantly from the healthy baseline and there is a high risk of respiratory diseases, triggering a level 2 warning (yellow warning). At this time, the system marks the pig with a yellow icon on the terminal interaction module and issues a pop-up reminder. At the same time, it generates a treatment suggestion of "recommending key monitoring and arranging on-site veterinary examination", prompting the farm staff to implement isolation observation and targeted diagnosis of the pig.
[0179] When S < 0.4, the pig is identified as being in a high-risk disease state, indicating that the pig's characteristic indicators have deviated significantly from the healthy baseline and are highly consistent with the characteristic patterns of typical respiratory disease samples, triggering a Level 1 warning (red warning). At this time, the system marks the pig with a flashing red icon on the terminal interaction module and continuously issues an audible and visual alarm signal. At the same time, it generates an emergency treatment instruction of "Immediate isolation recommended, initiation of treatment procedures" to prompt the farm staff to immediately isolate the pig and take appropriate treatment measures.
[0180] The four thresholds (0.8, 0.6, and 0.4) were determined during the model training phase based on the statistical distribution of scores of healthy and diseased samples on the validation set. Specifically, the 10th percentile of the manually labeled healthy sample scores was used as the boundary threshold between healthy and concerned status; the 50th percentile of the labeled early diseased sample scores was used as the boundary threshold between concerned and suspected disease status; and the 90th percentile of the labeled confirmed diseased sample scores was used as the boundary threshold between suspected disease and high-risk disease status.
[0181] In actual deployment, farmers can fine-tune the above thresholds through the configuration interface of the terminal interaction module according to the actual management needs of the farm, so as to adapt to the individual differences of different breeds and ages of pigs.
[0182] The uncertainty estimation submodule uses the Monte Carlo dropout method to quantify and estimate the cognitive uncertainty of the model predictions.
[0183] Since the time-series state evaluation module randomly discards neurons after each fully connected layer during training with a certain probability (dropout rate is 0.2 in this embodiment) to play a regularization role and prevent overfitting, this sub-module keeps each dropout layer in the on state during the model inference stage and performs T independent random forward propagation (T is 20 times in this embodiment) on the same input sample.
[0184] Since the combination of neurons randomly dropped in each forward propagation is different, it is equivalent to performing T independent samplings from the approximate posterior distribution of the model weights, and each forward propagation outputs a health status score. (k=1,2,…,T). This submodule calculates the mean of T outputs. and variance :
[0185]
[0186] The mean The final health status score is used for graded early warning determination, variance As a measure of cognitive uncertainty, the larger the variance, the more sensitive the model parameters are to different dropout masks, the worse the stability of the prediction results, and the lower the reliability.
[0187] This submodule also includes an uncertainty threshold. (In this embodiment) Take 0.05), when When the uncertainty of the current prediction result is too high, the model lacks sufficient confidence in the discrimination of the input sample. Possible reasons include: poor quality of input data (such as blurry infrared images, low audio signal-to-noise ratio), large difference between the current feature pattern of pigs and the distribution of training data, or difficulty in making clear judgments in the boundary area.
[0188] At this time, regardless of the warning level output by the graded warning submodule, the warning decision module will add the prompt message "the result is uncertain and manual review is recommended" when outputting the warning signal, and specifically mark the confidence level of the warning as "low confidence" in the warning notification of the terminal interaction module, reminding the farmers to rely on the results of manual observation and diagnosis and not to rely solely on the system's judgment;
[0189] The early warning decision module integrates the early warning level determination results from the hierarchical early warning submodule and the confidence assessment results from the uncertainty estimation submodule to generate the final early warning decision information package. This information package includes: pig identification (pen number or ear tag number), and health status score. Warning level (blue / yellow / red / none), confidence level (high confidence / low confidence), and uncertainty variance. The warning decision module includes a timestamp and suggested handling measures. The output of the warning decision module is connected to the input of the terminal interaction module, sending the aforementioned warning decision information package to the terminal interaction module for visualization.
[0190] In one embodiment, a multimodal fusion-based early warning system and method for porcine respiratory diseases are also provided, such as... Figure 2 As shown, the method includes:
[0191] S1: Synchronous multimodal data acquisition. A multimodal data acquisition layer deployed within the pigsty synchronously acquires multi-dimensional raw data at a uniform sampling frequency. Specifically, this includes:
[0192] Audio signals from pigs were collected using a microphone array;
[0193] Infrared thermal imaging images of pig faces were captured using an infrared thermal imaging camera.
[0194] Visible light images of pigs are captured using a high-definition visible light camera;
[0195] Environmental parameters such as temperature, humidity, ammonia concentration, and carbon dioxide concentration in the pigsty are collected using environmental sensors.
[0196] All collected data is accompanied by a uniform timestamp to ensure time sequence alignment in subsequent processing;
[0197] S2: Data preprocessing and time-series alignment. The data preprocessing and alignment module performs preprocessing and time-series alignment on the acquired multimodal raw data.
[0198] Audio preprocessing includes pre-emphasis, framing, windowing of the audio signal, and extraction of Mel spectrograms.
[0199] Image preprocessing includes size normalization (256×128 pixels), grayscale normalization, and data augmentation for both infrared thermal imaging and visible light images.
[0200] Environmental data preprocessing includes filling missing values, detecting and correcting outliers, and normalizing environmental sensor data.
[0201] Time alignment, based on a unified timestamp, precisely aligns the preprocessed data of each modality in the time dimension to generate multimodal time-series data packets;
[0202] S3: Multimodal depth feature extraction. The multimodal feature extraction module extracts depth features from the aligned modal data, specifically including:
[0203] Audio feature extraction involves inputting the Mel spectrogram into the audio feature extraction branch based on a CNN-Transformer hybrid architecture. After processing by a convolutional front end and a Transformer encoder, the output is a global audio feature vector. ;
[0204] Infrared feature extraction involves inputting the infrared thermal imaging image into the infrared feature extraction branch based on a dilated convolutional residual network. The outputs of the second, fourth, and fifth convolutional blocks are adaptively weighted and fused to output an infrared feature vector. ;
[0205] Visible light feature extraction involves inputting a visible light image into the visible light feature extraction branch based on the Transformer architecture. After processing by the Transformer encoder and the Serial Fusion Attention (SFAM) module, the output is a visible light feature vector. ;
[0206] Environmental feature extraction involves inputting environmental parameters into a multi-layer fully connected network and outputting an environmental feature vector. ;
[0207] S4: Cross-modal feature deep fusion. The cross-modal feature fusion module performs deep fusion of features from various modalities. The specific steps are as follows:
[0208] Multi-scale attention pyramid processing inputs infrared and visible light features into the multi-scale serial fusion attention pyramid submodule for salient region extraction, resulting in infrared salient features and visible light salient features. Infrared salient features, visible light salient features, audio global features, and environmental features are then input into the cross-modal interaction submodule for deep fusion. A pyramid structure from coarse to fine is used to progressively refine the feature maps at different scales, capturing salient discriminative regions at different scales.
[0209] Cross-modal interaction, through a cross-attention mechanism, enables full information exchange between four modalities: audio, infrared, visible light, and environment, generating cross-modal fusion features. ;
[0210] S5: Dynamic evaluation of temporal state. The temporal state evaluation module performs temporal modeling and health status evaluation on cross-modal fusion features, as detailed below:
[0211] Temporal feature aggregation inputs cross-modal fusion features from multiple consecutive time points into a Transformer-based temporal encoder to capture long-range dependencies between time points. Simultaneously, a Laplacian template feature set is constructed, and historical temporal prior knowledge is propagated forward through a cross-attention mechanism to achieve the fusion of historical and current information.
[0212] The health status score is obtained by inputting the aggregated temporal features into a multi-layer fully connected network and mapping them to a health status score S∈[0,1].
[0213] S6: Tiered early warning decision-making, the early warning decision-making module generates tiered early warning signals based on the health status score:
[0214] If S≥0.8, the system is considered healthy and no warning is triggered.
[0215] If 0.6≤S<0.8, a Level 3 warning (blue warning) is triggered.
[0216] If 0.4≤S<0.6, a Level II warning (yellow warning) is triggered.
[0217] If S < 0.4, a Level 1 warning (red warning) is triggered.
[0218] Meanwhile, the Monte Carlo dropout method is used to perform T random forward propagations to calculate the cognitive uncertainty variance of the prediction results;
[0219] If the variance exceeds the preset threshold, output the message "Result is uncertain, manual review is recommended";
[0220] S7: Visual presentation of early warning results. The terminal interaction module presents the early warning results to aquaculture managers in a visual manner, including a real-time monitoring interface, an early warning notification interface, a historical trend interface, and a data analysis report interface.
[0221] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
Claims
1. An early warning system for porcine respiratory diseases based on multimodal fusion, characterized in that, The system includes a multimodal data acquisition layer, a data preprocessing and alignment module, a multimodal feature extraction module, a cross-modal feature fusion module, a time-series state evaluation module, an early warning decision module, and a terminal interaction module, which are connected in sequence. The multimodal data acquisition layer synchronously acquires audio signals, infrared thermal imaging images, visible light images, and environmental parameters, and adds a hardware-level timestamp to each frame of data acquired. The data preprocessing and alignment module preprocesses the original data of each modality and aligns the preprocessed modal data in the time dimension based on the timestamp to generate a multimodal time-series data packet. The multimodal feature extraction module extracts audio global feature vectors, infrared feature vectors, visible light feature vectors, and environmental feature vectors from the multimodal time-series data packets, respectively. The cross-modal feature fusion module performs deep fusion of the audio global feature vector, infrared feature vector, visible light feature vector, and environmental feature vector to generate cross-modal fused features; The cross-modal feature fusion module includes a multi-scale serial fusion attention pyramid submodule and a cross-modal interaction submodule; The multi-scale serial fusion attention pyramid submodule is used to perform multi-level cascaded attention filtering on the infrared and visible light features from coarse to fine, refine the region of interest step by step, and output infrared salient features and visible light salient features. The cross-modal interaction submodule is used to calculate the bidirectional cross-attention output between any two modalities using the infrared salient features, visible light salient features, audio global features, and environmental features as four modal features participating in the interaction, and to fuse the cross-attention outputs of other modalities received by each modality with its own features to obtain cross-modal fusion features; The temporal state assessment module performs temporal modeling on the cross-modal fusion features at multiple consecutive time points, captures the gradual evolution pattern of the disease, and outputs a health status score. The early warning decision module generates a graded early warning signal based on the health status score, and sends the graded early warning signal to the terminal interaction module for visualization.
2. The early warning system for porcine respiratory diseases based on multimodal fusion according to claim 1, characterized in that, The multimodal data acquisition layer includes: The sound acquisition unit consists of an array of multiple microphones deployed on the top of the pigsty, used to collect the audio signals of the pigs. The infrared thermal imaging unit, consisting of an infrared thermal imaging camera, is used to acquire infrared thermal images of the pig's face. The visible light image acquisition unit, consisting of a high-definition visible light camera, is used to acquire visible light images of pigs; The environmental sensing unit consists of a temperature and humidity sensor, an ammonia sensor, and a carbon dioxide sensor, and is used to collect data on the temperature, humidity, ammonia concentration, and carbon dioxide concentration of the pigsty environment. The sound acquisition unit, infrared thermal imaging unit, visible light image acquisition unit, and environmental sensing unit acquire data at a unified sampling frequency. Each unit has a built-in unified synchronous clock source, and each frame of data acquired is accompanied by the hardware-level timestamp.
3. The early warning system for porcine respiratory diseases based on multimodal fusion according to claim 1, characterized in that, The data preprocessing and alignment module includes: The audio preprocessing submodule pre-emphasizes, frames, and windows the audio signal, and extracts the Mel spectrogram. The image preprocessing submodule performs spatial registration, size normalization, and grayscale value normalization on the infrared thermal imaging image and the visible light image. The environmental data preprocessing submodule performs missing value filling, outlier detection and correction, and normalization on the environmental parameters. The timing alignment submodule, based on the timestamp and using the frame rate of the audio signal as the reference clock, matches the infrared image frame, visible light image frame, and environmental parameter record with the closest timestamp for each frame of audio data, forming a multimodal timing data packet.
4. The early warning system for porcine respiratory diseases based on multimodal fusion according to claim 3, characterized in that, The multimodal feature extraction module includes: The audio feature extraction branch employs a hybrid architecture based on convolutional neural networks and Transformers to extract global audio feature vectors from the Mel spectrogram. The infrared feature extraction branch employs a residual network based on dilated convolution to extract infrared feature vectors from the infrared thermal imaging image. The infrared feature extraction branch performs adaptive weighted fusion of the output features of multiple different convolutional blocks. The visible light feature extraction branch adopts a Vision Transformer-based architecture to extract visible light feature vectors from the visible light image. The output of the visible light feature extraction branch is connected to the serial fusion attention module, which consists of a compression-excitation module and a triple attention module connected in series to enhance attention to the discriminative regions of pigs. The environmental feature extraction branch employs a multi-layer fully connected network to encode the normalized environmental parameters into environmental feature vectors.
5. The early warning system for porcine respiratory diseases based on multimodal fusion according to claim 1, characterized in that, The multi-scale serial fusion attention pyramid submodule adopts a multi-level pyramid structure. Each layer divides the input feature map into multiple non-overlapping feature map blocks in the spatial dimension. Each feature map block is input into the serial fusion attention module for processing and then merged according to its original spatial position to obtain the attention map of that layer. The output feature map of the previous layer is reweighted by the attention map of the current layer and then passed to the next layer, achieving a progressive refinement of the attention region from coarse to fine.
6. The early warning system for porcine respiratory diseases based on multimodal fusion according to claim 1, characterized in that, The cross-modal interaction submodule first maps the four modal features involved in the interaction to the same embedding dimension through a linear projection layer, and then calculates all directed cross-attention outputs. For any target modality, the cross-attention outputs received from the other three modalities are concatenated and fused with its own projection features to obtain the enhanced features of that modality; The enhanced features of the four interactive modalities are concatenated along the channel dimension and then mapped through a fully connected layer to obtain the final cross-modal fusion feature.
7. The early warning system for porcine respiratory diseases based on multimodal fusion according to claim 1, characterized in that, The timing state evaluation module includes: The temporal feature aggregation submodule uses a Transformer-based temporal encoder to perform temporal modeling on the cross-modal fusion feature sequence within the sliding time window, capture long-range dependencies across time points, and output the aggregated temporal features. The temporal feature aggregation submodule also constructs a Laplace-type template feature set, which is fused with the cross-modal fusion feature at the current moment through a cross-attention mechanism. This propagates historical temporal prior knowledge to the current moment, and finally outputs temporal features that fuse long-range temporal dependencies and historical prior knowledge. The health status assessment submodule uses a multi-layer fully connected network to map the time-series features into the health status score.
8. The early warning system for porcine respiratory diseases based on multimodal fusion according to claim 1, characterized in that, The early warning decision module includes: The graded early warning submodule divides the health status of pigs into four levels based on the numerical range of the health status score: healthy status, attention status, suspected disease status, and high-risk disease status, and triggers no warning, level three warning, level two warning, and level one warning respectively. The uncertainty estimation submodule uses the Monte Carlo dropout method to quantify and estimate the cognitive uncertainty of the model prediction. It performs multiple independent random forward propagations on the same input sample and calculates the variance of the multiple outputs as an uncertainty measure. When the variance exceeds a preset threshold, it outputs a warning signal and adds a manual review prompt.
9. A method for early warning of porcine respiratory diseases based on multimodal fusion, applied to the system described in any one of claims 1 to 8, characterized in that, The method includes: The multimodal data acquisition layer synchronously acquires audio signals, infrared thermal imaging images, visible light images, and environmental parameters, and adds a hardware-level timestamp to each frame of data acquired. The data preprocessing and alignment module preprocesses the original data of each modality, and aligns the preprocessed modal data in the time dimension based on the timestamp to generate a multimodal time-series data packet. The multimodal feature extraction module extracts audio global feature vectors, infrared feature vectors, visible light feature vectors, and environmental feature vectors from the multimodal time-series data packets, respectively. The audio global feature vector, infrared feature vector, visible light feature vector, and environmental feature vector are deeply fused by the cross-modal feature fusion module to generate cross-modal fused features; Specifically, the infrared and visible light features are subjected to multi-level cascaded attention filtering from coarse to fine through a multi-scale serial fusion attention pyramid submodule, which refines the region of interest step by step and outputs infrared and visible light salient features. The cross-modal interaction submodule uses infrared salient features, visible light salient features, audio global features, and environmental features as four modal features participating in the interaction. It calculates the bidirectional cross-attention output between any two modalities and fuses the cross-attention outputs of other modalities received by each modality with its own features to obtain cross-modal fusion features. The time-series state assessment module performs time-series modeling on the cross-modal fusion features at multiple consecutive time points, captures the gradual evolution pattern of the disease, and outputs a health status score. The early warning decision module generates graded early warning signals based on the health status score, and presents them visually through the terminal interaction module.