Poultry health condition detection method, system and equipment and storage medium
Through spectrum diagram analysis and MAE self-supervised learning mechanism, a poultry health detection system based on Vision Transformer was built, which solved the problems of poor real-time performance, low degree of automation, and insufficient anti-interference ability in the existing technology, and achieved high-precision poultry health status detection.
Patent Information
- Application Number
- CN202510592858.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-22
Smart Images

Figure CN120526809A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of poultry health status detection, and in particular to a poultry health status detection method, system, equipment and storage medium. Background Art
[0002] Chickens are a vital poultry species in the global livestock industry, particularly in meat and egg production. Timely and effective monitoring of chicken health is crucial for ensuring flock productivity and reducing the spread of disease. However, existing chicken health monitoring methods have numerous shortcomings, particularly in areas such as automation, real-time performance, and data analysis.
[0003] Currently, chicken health monitoring relies primarily on the following methods: Physiological indicators: These methods rely on monitoring physiological parameters such as body temperature, heart rate, and respiratory rate to determine health. While these methods can provide some health information, they present numerous challenges in practical application. For example, the testing process requires significant labor and equipment costs, making them unsuitable for large-scale flocks. Furthermore, these physiological indicators often only change significantly when a chicken exhibits serious health issues, lacking early warning capabilities. Behavioral observation: This traditional method uses behavioral changes such as activity and diet to infer health. However, this method is highly subjective and relies on staff experience, making standardization and automation difficult. Furthermore, behavioral changes are often late-stage symptoms of disease, making timely intervention difficult and increasing the risk of disease transmission within the flock. Traditional audio analysis: Sound can reflect a chicken's health. For example, when a chicken is sick, its vocalizations may change. Traditional audio analysis methods analyze specific frequency bands in the chicken's vocalizations to assess health. However, these methods often rely on simple spectrum analysis techniques, which cannot accurately capture subtle changes in complex acoustic signals. Furthermore, traditional audio analysis methods are sensitive to background noise and environmental factors, making them susceptible to interference and difficult to apply on a large scale in real-world production environments. Image analysis-based health monitoring methods: Some methods analyze images of chickens' appearance (such as feathers and body shape) to determine their health. However, these methods are susceptible to factors such as external lighting and camera angle. Furthermore, some potential health issues (such as respiratory diseases) are difficult to detect through images, resulting in a limited detection range. Machine learning-based detection methods: With the advancement of machine learning technology, some studies have attempted to utilize machine learning algorithms to monitor chicken health. These methods, by training models with large-scale data, can improve detection accuracy and automation. However, most current machine learning-based methods rely on large amounts of labeled data, resulting in complex training processes and high computational resource requirements. Furthermore, the black-box nature of the models makes the results difficult to interpret, hindering their widespread application in real-world production. Existing chicken health monitoring methods have significant shortcomings in terms of real-time performance, automation, and interference resistance, making them unable to meet the demands of modern farms for efficient large-scale chicken health management. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by the present invention is: the existing poultry health status detection technology has the problems of poor real-time performance, low degree of automation, insufficient anti-interference ability, and weak model interpretability, and how to achieve high-precision automatic detection of poultry health status by introducing spectrum analysis and MAE self-supervised learning mechanism.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: a method for detecting the health status of poultry, comprising collecting poultry audio data and preprocessing the poultry audio data; constructing a MAE model based on the time-frequency characteristics of the audio data; pre-training and fine-tuning the MAE model, and evaluating the pre-trained MAE model.
[0007] As a preferred embodiment of the poultry health status detection method of the present invention, the collecting of poultry audio data includes collecting tag audio data at a preset frequency, and the recording devices are evenly arranged in the aisles of the chicken house.
[0008] As a preferred embodiment of the poultry health status detection method described in the present invention, the labeled audio data includes the natural sounds of healthy poultry and environmental background noise as positive labels, and the abnormal sounds of poultry with confirmed diseases as negative labels. The labeled audio data is divided into a training set and a test set, and a preset number of unlabeled audio data are collected as a pre-training set for the MAE model.
[0009] As a preferred embodiment of the poultry health status detection method described in the present invention, the poultry audio data is preprocessed, including processing high-frequency components through pre-emphasis filtering, framing the audio data using a Hamming window, extracting time-frequency features that conform to auditory characteristics through a 40-dimensional Mel filter bank, and converting them into a standardized Mel spectrum.
[0010] As a preferred embodiment of the poultry health status detection method described in the present invention, the construction of the MAE model includes selecting Vision Transformer as the basic framework based on the time-frequency characteristics of the sound signal. The MAE model adopts an encoder-decoder structure and uses the encoder to extract high-level feature representations from partially visible spectrograms.
[0011] The decoder reconstructs the masked spectral region using the features extracted by the encoder.
[0012] As a preferred embodiment of the poultry health status detection method described in the present invention, the pre-training and fine-tuning of the MAE model include: pre-training includes using a self-supervised learning method to train the MAE model with large-scale unlabeled audio data to learn a general spectral feature representation, randomly masking 75% of the blocks in the input Mel-spectrogram, setting the MAE model to reconstruct the complete spectrum based on the remaining 25% of visible blocks, selecting the MAE model architecture, setting the training batch size, obtaining a stable gradient estimate under GPU memory limitations, performing learning rate scheduling through a cosine annealing strategy, controlling the weight attenuation strength by setting key regularization parameters, and normalizing the spectrum using norm_pix_loss.
[0013] Fine-tuning includes adapting the general acoustic features learned in pre-training to health detection scenarios, using a progressive unfreezing strategy for fine-tuning, unfreezing the last two layers of Transformer modules and classification heads in the initial stage, setting a basic learning rate, and gradually unfreezing more underlying modules while pre-training is in progress. The learning rate is attenuated layer by layer, data enhancement is performed through linear interpolation and area occlusion, and multiple regularization techniques are used to control overfitting.
[0014] As a preferred embodiment of the poultry health status detection method described in the present invention, the evaluation of the pre-trained MAE model includes performing end-to-end testing on the MAE model using a test set, evaluating the performance of the MAE model at different thresholds by drawing curves, monitoring the real-time inference speed during the deployment phase, determining whether the timeliness requirements of online detection of the farm are met, and outputting a visual report.
[0015] Another object of the present invention is to provide a poultry health status detection system that can perform pre-emphasis filtering, frame segmentation and Mel-spectrogram conversion on the collected poultry audio data, combine the MAE model based on self-supervised learning to perform mask reconstruction learning of audio time-frequency features, and fine-tune the adaptation classification task after training. This solves the problems of current technologies such as reliance on manual labor, weak anti-interference ability, lack of early warning and model generalization.
[0016] As a preferred solution of the poultry health status detection system described in the present invention, it includes: an audio acquisition and preprocessing module, an MAE modeling module, and a model training and evaluation module; the audio acquisition and preprocessing module includes an audio acquisition unit and an audio preprocessing unit, the audio acquisition unit is used to collect the call data of poultry in different states and synchronously record the sampling environment information, the audio preprocessing unit is used to pre-emphasize, frame, Mel filtering and other processing on the collected audio signal, and convert it into a standardized Mel spectrum graph; the MAE modeling module includes a spectrum graph division unit and a coding structure construction unit, the spectrum graph division unit is used to divide the Mel spectrum graph into image blocks of equal size as model input, and the coding structure construction unit is used to build a Transformer-based MAE model structure, including an encoder and decoder architecture; the model training and evaluation module includes a pretraining unit and an evaluation and verification unit, the pretraining unit adopts a self-supervised learning method to train the MAE model through a masked reconstruction task, and the evaluation and verification unit is used to fine-tune and verify the accuracy of the trained model on a labeled data set.
[0017] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of a poultry health status detection method.
[0018] A computer-readable storage medium stores a computer program, which implements the steps of a poultry health status detection method when executed by a processor.
[0019] Beneficial effects of the present invention: This method adopts spectrum analysis technology, which can accurately identify the health status of poultry, which is significantly improved compared with traditional methods, and has the characteristics of high precision and high reliability. Through pre-emphasis filtering and Mel spectrum conversion, it can effectively overcome the interference of environmental noise on the detection results, and can still maintain stable detection performance under low signal-to-noise ratio conditions, fully meeting the application requirements of complex farm environments, and realizing a fully automated detection process. It only needs to collect poultry sound data and upload it to the system to output a health status analysis report in real time, greatly reducing the cost of manual detection. By combining self-supervised pre-training with the Transformer architecture, the designed MAE-ViT model can automatically learn the time-frequency feature representation of poultry sounds, and has higher accuracy than traditional CNN models. This method uses attention visualization technology to intuitively display the key frequency bands based on which the model makes decisions, which is highly consistent with veterinary clinical experience and provides an explainable scientific basis for disease diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 This is an overall flow chart of a poultry health status detection method provided by the first embodiment of the present invention.
[0022] Figure 2 A diagram showing the distribution of recording equipment in a poultry farm for a method for detecting the health status of poultry provided in accordance with the second embodiment of the present invention.
[0023] Figure 3 A diagram of data collection equipment for a poultry health status detection method provided by the second embodiment of the present invention.
[0024] Figure 4 A chicken sound spectrum diagram of a poultry health status detection method provided by the second embodiment of the present invention.
[0025] Figure 5 This is a diagram showing the classification training results of Group A of a poultry health status detection method provided by the second embodiment of the present invention.
[0026] Figure 6 This is a diagram of the pre-training results of Group B of a poultry health status detection method provided by the second embodiment of the present invention.
[0027] Figure 7 This is a diagram showing the classification training results of Group B of a poultry health status detection method provided by the second embodiment of the present invention.
[0028] Figure 8 This is an overall flow chart of a poultry health status detection system provided by the third embodiment of the present invention. DETAILED DESCRIPTION
[0029] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0030] Example 1, reference Figure 1 , as one embodiment of the present invention, provides a poultry health status detection method, comprising:
[0031] S1: Collect poultry audio data and pre-process the poultry audio data.
[0032] Furthermore, collecting poultry audio data includes collecting tag audio data at a preset frequency, and recording devices are evenly arranged in the aisles of the chicken house.
[0033] It should also be noted that the present invention collects sound data of different varieties of poultry in healthy and unhealthy states. To ensure data quality, the collection process should be carried out in a standardized breeding environment to avoid strong noise interference, and key parameters such as ambient temperature and humidity should be recorded at the same time. The data must cover different time periods (morning, noon, evening) and multiple behavioral states (eating, resting, activity, etc.). All samples must be marked with health status by veterinary professionals, and metadata such as original audio and environmental parameters must be retained. Professional recording equipment should be used for collection equipment. The data collection process must follow animal welfare principles to avoid causing additional stress to poultry.
[0034] A preferred solution for the preset frequency is that the sampling frequency is 22.05kHz, the depth is 16 bits, and the duration of each sound segment is 5 seconds.
[0035] Furthermore, the labeled audio data includes the natural sounds of healthy poultry and environmental background noise as positive labels, and the abnormal sounds of poultry with confirmed diseases as negative labels. The labeled audio data is divided into training sets and test sets, and a preset number of unlabeled audio data are collected as the MAE model pre-training set.
[0036] It should also be noted that the labeled audio data includes samples of natural vocalizations of healthy poultry, samples of typical environmental background noise, and samples of abnormal vocalizations of diseased poultry confirmed by veterinarians. Based on the annotations, healthy poultry and environmental background sound samples are labeled as positive, while diseased poultry vocalizations are labeled as negative, forming the binary classification dataset. This dataset is divided into training and test sets in a predetermined ratio for evaluating model performance during the fine-tuning phase.
[0037] Furthermore, the poultry audio data were preprocessed, including processing the high-frequency components through pre-emphasis filtering, framing the audio data using a Hamming window, extracting time-frequency features that conform to auditory characteristics through a 40-dimensional Mel filter bank, and converting them into a standardized Mel spectrum.
[0038] It should also be noted that the pre-emphasis filtering process for high-frequency components is expressed as:
[0039] H(Z)=1-αz -1
[0040] Where H(Z) is the system transfer function of the chicken call signal after passing through the pre-emphasis filter, and α is the high-frequency enhancement coefficient of the chicken call. Abnormal chicken calls (such as wheezing sounds caused by respiratory diseases) are mostly concentrated in the low frequency band (<500Hz), while the high-frequency harmonics (1-5kHz) of healthy calls are the key features. -1 is a delay operator, which represents the sound signal of the previous sampling point (the chicken's call has short-term correlation, such as continuous crowing or coughing). The processed signal satisfies the difference equation expressed as:
[0041] y(n)=x(n)-ax(n-1)
[0042] Here, x(n) is the audio value of the chicken's call at the nth moment of the original sampling, and y(n) is the output signal after differential processing. The differential equation enhances the high-frequency component (by 6-12dB) and suppresses common low-frequency noise in farms (such as fan noise and footsteps). This ensures that subsequent analysis focuses on the chicken's biological characteristic frequency band, effectively suppressing low-frequency interference (<200Hz) such as breathing sounds and equipment background noise, while retaining the main energy band (1-5kHz) characteristic of poultry calls.
[0043] It should also be noted that the use of a Hamming window to frame the audio data includes the following: a single rooster's call lasts approximately 30-50ms, and a 25ms Hamming window can completely cover a call cycle (satisfying the short-term stationarity assumption). A 10ms frame shift ensures a 15ms overlap between adjacent frames to avoid missing short-term abnormal sounds (such as a cough, which lasts approximately 20ms). To meet the short-term stationarity assumption, the continuous audio signal is segmented into 25ms analysis frames with a 10ms frame shift. The rooster's calls are framed using a Hamming window, and the window function is expressed as:
[0044]
[0045] Where w(n) is the window function weight for the nth sampling point of the crowing signal, and N is the total number of sampling points in a frame, representing the length of the window function. In this invention, N = 400, ensuring that each frame covers the typical short, stationary period of crowing (e.g., a single crowing duration of approximately 30-50ms). This parameter setting ensures an optimal balance between time domain resolution (per frame) and frequency domain resolution while avoiding spectral leakage. The window function coefficients (0.54, -0.46) optimize spectral leakage and reduce the interference of sudden noise in the chicken house (such as collisions) on the spectral analysis.
[0046] The power spectrum is calculated based on the Mel-scale filter bank. Based on the chicken's auditory sensitive frequency band, a 40-dimensional Mel filter bank is designed and expressed as:
[0047] M(m)=∑|X(k)| 2 H chicken (m,k), m=1,...,40
[0048] Among them, M(m) is the weighted power spectrum value output by the mth Mel filter, which represents the energy of the chicken sound signal in the mth perceptual frequency band, forming one dimension of the 40-dimensional Mel spectrum graph, H chicken (m, k) is the passband weight of the mth Mel filter designed for the poultry sound characteristics at the kth frequency point. In the present invention, the low-frequency part (200-500 Hz) is linearly distributed, and the high-frequency part (3-5 kHz) is logarithmically densely distributed.
[0049] The Mel spectrum cepstrum obtained after logarithmic compression is expressed as:
[0050] MFCC = DCT(log(M(m)))
[0051] Among them, the filter bank design follows the auditory critical bandwidth theory, maintaining a linear distribution below 1kHz and expanding on a logarithmic scale above 1kHz, which is consistent with the auditory perception characteristics of poultry.
[0052] The dynamic range compression technology is used to process the Mel spectrum representation as follows:
[0053]
[0054] in, Where is the normalized mel-spectrogram feature value, S is the original mel-spectrogram energy value, and μ and σ are calculated from the training set. The call intensity of different chickens can vary by up to 20 dB. Normalization eliminates individual differences in call intensity, allowing the model to focus on frequency features rather than volume (e.g., the hoarseness of a sick chicken versus the high-pitched call of a healthy chicken). The final feature matrix has a dimension of 40 × T (where T is the number of time frames), retaining the effective bioacoustic frequency band of 0-8 kHz.
[0055] S2: Construct a MAE model based on the time-frequency characteristics of audio data.
[0056] Furthermore, the construction of the MAE model includes selecting VisionTransformer as the basic framework based on the time-frequency characteristics of the sound signal. The MAE model adopts an encoder-decoder structure and uses the encoder to extract high-level feature representations from partially visible spectrograms.
[0057] The decoder reconstructs the masked spectral region using the features extracted by the encoder.
[0058] It should be noted that the poultry health status detection model adopts the Masked Autoencoder (MAE) architecture based on Transformer. In the input processing stage of the model, the spectrogram is first divided into 16×16 non-overlapping patch units. Let the input spectrogram be X∈R H×W , where H and W represent the height and width of the time-frequency dimension respectively. Each patch is converted into a 768-dimensional feature vector by linear projection. The projection process can be expressed as:
[0059] z i =W e Vec(X i )+p i
[0060] Among them, X i is a 16×16 spectrogram segment corresponding to the time-frequency unit of the chicken’s call, W e ∈R 768×256 is a learnable projection matrix that maps acoustic features to a high-dimensional space (768 dimensions) to separate healthy or pathological features. Vec(.) represents the operation of flattening the patch into a vector. i It is a position encoding vector that marks the position of the segment in the time-frequency diagram and is used to model the temporal regularity of chicken calls.
[0061] The position encoding system uses two-dimensional sine and cosine encoding to encode the time and frequency dimensions respectively. Chicken calls have temporal regularity, and two-dimensional position encoding enables the model to understand the time-frequency relationship. The position encoding of the time-frequency coordinate (t, f) is calculated as follows:
[0062]
[0063] Among them, t and f represent time and frequency coordinates respectively, and i,j∈[0,383] is the dimension index.
[0064] During the feature extraction phase, the model uses a high-ratio random masking strategy of 75%. This strategy simulates common sound occlusion scenarios in farms (such as overlapping chicken calls and equipment noise), randomly masks 75% of the sound spectrum, and forces the model to reconstruct the complete chicken call features from the remaining 25% of the fragments. Let the mask matrix M∈{0,1} N , where N is the total number of patches, M i ~Bernoulli(0.75), only 25% of the visible patches are retained and fed into the 12-layer Transformer encoder. Each layer of the encoder contains the following calculation process:
[0065] The multi-head self-attention mechanism is expressed as:
[0066]
[0067] Where Q, K, V represent query, key, and value matrices respectively, and d=64 is the dimension of each attention head. is the learnable parameter matrix.
[0068] The feedforward network adopts a two-layer structure as follows:
[0069] FFN(x)=W2GELU(W1x+b1)·+b2
[0070] Where W1∈R 3072×768 , W2∈R 768×2072 is the weight of the fully connected layer, b1 and b2 are bias terms, and the GELU activation function is defined as: GELU(x) = xφ(x), where φ is the standard normal distribution CDF.
[0071] The decoder part adopts a lightweight 4-layer Transformer structure, and its reconstruction loss function is designed as a joint loss in the time domain and frequency domain, expressed as:
[0072]
[0073] Where X′ is the reconstructed output, STFT(.) represents the short-time Fourier transform, which is optimized for chicken calls, ||.|| F represents the Frobenius norm.
[0074] The preferred approach for constructing the MAE model in this paper is an asymmetric encoder-decoder design. The encoder uses a deep 12-layer Transformer architecture (each with 12 attention heads) to focus on feature extraction, while the decoder uses a lightweight 4-layer architecture to perform reconstruction. The latent space features output by the model can be directly used for downstream health status classification tasks or can be further adapted to specific detection scenarios through fine-tuning.
[0075] S3: Pre-train and fine-tune the MAE model, and evaluate the MAE model after pre-training.
[0076] Furthermore, the MAE model is pre-trained and fine-tuned. Pre-training includes using self-supervised learning to train the MAE model on large-scale unlabeled audio data to learn general spectral feature representations, randomly masking 75% of the blocks in the input Mel-spectrogram, and setting the MAE model to reconstruct the complete spectrum based on the remaining 25% of visible blocks. By selecting the MAE model architecture and setting the training batch size, a stable gradient estimate is obtained under the GPU memory limit. The learning rate is scheduled through the cosine annealing strategy, the weight attenuation strength is controlled by setting key regularization parameters, and the spectrum is normalized using norm_pix_loss.
[0077] It should also be noted that this invention uses a self-supervised learning paradigm to pre-train the model, with the core goal of learning a universal time-frequency feature representation from large-scale unlabeled spectrogram data. The pre-training process is based on a masked reconstruction task, which randomly masks 75% of the input spectrogram, forcing the model to reconstruct the complete time-frequency features from the remaining visible 25%. The effectiveness of this method is based on the theoretical assumption that a model that can accurately reconstruct the masked time-frequency regions must have already grasped the essential characteristics of the sound signal.
[0078] The pre-training data processing process first normalizes the input spectrogram, which is expressed as:
[0079]
[0080] in, is the normalized input spectrogram feature matrix, which is the direct input of the pre-trained MAE model. X is the original input Mel-spectrogram feature matrix, μ and σ are the mean and standard deviation of the training dataset. Then, a random masking strategy is used to generate a binary mask matrix M∈{0,1} N , where N is the total number of patches, and each element M i Independent Bernoulli distribution is expressed as:
[0081] M i ~Bernoulli (p=0.75)
[0082] The optimization objective of the model consists of two key components: time domain reconstruction loss and frequency domain consistency loss. The time domain reconstruction loss is calculated using the mean square error (MSE) and is expressed as:
[0083]
[0084] Among them, L time is the time domain loss, which represents the mean square error of the reconstructed signal in the time domain, reflecting the accuracy of the reconstruction of the time-frequency features of the chicken call. |M| represents the number of masked patches, and X′ is the model reconstruction output.
[0085] The frequency domain consistency loss is calculated by short-time Fourier transform and is expressed as:
[0086] L freq =||STFT(X′)-STFT(X)|| F
[0087] The final composite loss function is the weighted sum of the two and is expressed as:
[0088] L total =αL time +(1-α)L freq
[0089] Among them, α=0.7 is the empirical weight coefficient.
[0090] Jointly optimize the time domain and frequency domain errors. The time domain loss (70%) ensures that the waveform of the reconstructed call is consistent with the real chicken's crowing (such as the intermittent sound of a sick chicken). The frequency domain loss (30%) focuses on retaining the spectral characteristics of the healthy diagnostic frequency band (1-5kHz).
[0091] The training process uses the AdamW optimizer, and its parameter update rule is:
[0092]
[0093] Among them, θ t+1 The updated value of the model parameters at the next moment t+1; θ t is the model parameter value at the current time t, representing the weight in the Transformer neural network used to encode the chicken call features; η t is the dynamic learning rate at the current moment, m t is the first-order moment estimate, v t is the second-order moment estimate, and ε=1e-8 is the numerical stability term.
[0094] Chicken call data has high diversity (breed / age / disease type). The learning rate scheduling adopts the cosine annealing strategy to prevent the model from falling into local optimality (such as overfitting the call of a certain chicken species). It is expressed as:
[0095]
[0096] Among them, T is the total number of training steps, η max =1e -3 , the initial high learning rate quickly fits the common characteristics of healthy chickens, η min =1e -5 In the later stage, a low learning rate is used to fine-tune the sparse abnormal pattern of sick chickens; a mixed precision training (AMP) strategy is adopted in the process, and cosine annealing is used to prevent the model from falling into local optimality (such as overfitting a certain disease). FP16 precision is used for forward calculation, and FP32 precision is used for gradient calculation and parameter update to take into account both training speed and numerical stability.
[0097] To prevent overfitting, DropPath regularization is used, and its drop probability increases linearly with the network depth as follows:
[0098]
[0099] Among them, p l is the probability of the path being randomly discarded (skipped) in the l-th layer Transformer encoder, l is the current layer index, and L is the total number of layers. After pre-training is completed, the encoder part of the model can be migrated to the downstream health status classification task, and its learned time-frequency feature representation shows strong generalization ability.
[0100] Fine-tuning includes adapting the general acoustic features learned in pre-training to health detection scenarios, using a progressive unfreezing strategy for fine-tuning, unfreezing the last two layers of Transformer modules and classification heads in the initial stage, setting a basic learning rate, and gradually unfreezing more underlying modules while pre-training is in progress. The learning rate is attenuated layer by layer, data enhancement is performed through linear interpolation and area occlusion, and multiple regularization techniques are used to control overfitting.
[0101] It should be noted that the present invention employs a transfer learning strategy for supervised fine-tuning of a pre-trained model. Its core objective is to adapt the general time-frequency feature representations obtained through pre-training to the specific task of poultry health classification. The fine-tuning process is based on a labeled spectrogram dataset and optimizes model parameters by minimizing classification loss. Theoretical evidence suggests that the first few layers of the pre-trained model have already learned the fundamental time-frequency pattern features, requiring only fine-tuning of the higher-level networks to achieve efficient task adaptation.
[0102] The data processing flow before fine-tuning first normalizes the input spectrogram, maintaining the same normalization parameters as the pre-training stage. The data augmentation strategy includes time-frequency masking: randomly masking 20% of the time-frequency region; MixUp enhancement: linearly mixing two samples to simulate a mixed vocalization scenario of a flock of chickens and enhance the model's ability to distinguish overlapping sounds (such as the simultaneous presence of panting sounds of sick chickens and crowing of healthy chickens).
[0103] The model structure is adjusted to a two-stage fine-tuning strategy:
[0104] The first stage freezes the front layers of the encoder and only optimizes the subsequent layers and the classification head, which is expressed as:
[0105] θ ft ={θ enc [k:],θ head}
[0106] Among them, θ ft is the parameter set used in the fine-tuning stage, i.e., the network weights used for training optimization in the poultry health status classification task, θ enc [k:] represents all parameters from the kth layer to the last layer of the encoder; freezing the first k layers of the encoder and only training the subsequent layers is beneficial to retain the spectral structure features extracted from the supervised pre-training, θ head Represents all parameters of the classification head, which is used to judge the health status of chickens based on Mel spectrum features.
[0107] In the second stage, all parameters are unfrozen for end-to-end fine-tuning. The classification loss function is expressed as label-smoothed cross entropy:
[0108]
[0109] Among them, L cls is the classification loss value, C is the number of categories, and α is the smoothing coefficient.
[0110] By introducing a regularization term to prevent overfitting, we can include feature distribution consistency constraints, which can be expressed as:
[0111] L fd =KL(p(z ft )||p(z pretrain ))
[0112] Among them, L fd is the feature distillation loss value, which is used to measure the difference between the audio time-frequency feature distribution extracted after fine-tuning the model and the extraction result in the pre-training stage; KL(…||…) is used to measure the distance between two probability distributions, p(z ft ) is the distribution of poultry mel-spectrogram features extracted during the model fine-tuning phase, p(z pretrain) is the feature distribution learned by the model from a large amount of unlabeled chicken call audio through the MAE structure during the pre-training stage.
[0113] The attention focus penalty is expressed as:
[0114]
[0115] Among them, L att is the attention focus loss value, A l is the attention weight matrix of the lth layer, and I is the identity matrix, which serves as the target reference matrix.
[0116] The final compound loss is expressed as:
[0117] L total =L cls +0.1L fd +0.01L att
[0118] Among them, L total is the final compound loss value.
[0119] The optimizer uses AdamW with layer-wise learning rate as:
[0120] η l =η base γ l
[0121] Among them, η base is the basic learning rate, defining the initial learning rate of the shallowest layer, γ l is the learning rate decay factor.
[0122] The learning rate schedule uses cosine annealing with warm restart, which is expressed as:
[0123]
[0124] Among them, T cur is the number of training rounds currently conducted, indicating the current training progress, T i The total length of the current cycle.
[0125] Fine-tuning evaluation uses five-fold cross validation, and the main indicators include:
[0126] The classification accuracy is expressed as:
[0127]
[0128] Among them, Acc is the classification accuracy, N is the total number of test samples, which means the total number of chicken samples for health status prediction. The health status label of the i-th sample predicted by the model, y iis the true label of the i-th sample.
[0129] In order to address the class imbalance (healthy samples are more than sick chickens), we calculated the F1 score for each class and averaged it to ensure the model's detection sensitivity for the minority class (sick chickens). The formula for the macro-average F1 score is:
[0130]
[0131] Among them, F1 macro The macro-average F1 score is used to measure the overall detection performance of the model for each category (especially the minority class, such as "sick chicken"). c Recall is the precision of category C, which indicates how many of the samples predicted by the model as this category are actually chicken calls of this category. c is the recall rate of class C.
[0132] Furthermore, the evaluation of the pre-trained MAE model includes end-to-end testing of the MAE model using a test set, evaluating the performance of the MAE model at different thresholds by drawing curves, monitoring the real-time inference speed during the deployment phase, determining whether it meets the timeliness requirements of online detection in farms, and outputting a visual report.
[0133] It should also be noted that this invention uses a standard machine learning evaluation process to systematically verify model performance. The evaluation process strictly adheres to the functional modules provided in the code implementation and mainly includes the following three aspects of testing: Basic metrics such as accuracy and recall are calculated using evaluation functions. The code uses the classification metric calculation function provided by the torchmetrics library to output top-1 and top-5 accuracy on the validation set. The evaluation process adopts a distributed computing mode to ensure the accuracy of metric calculation in a multi-GPU environment. The code framework supports k-fold cross-validation and can enable distributed evaluation mode. During the evaluation process, the model independently calculates performance metrics on each validation fold, and ultimately outputs the average result and standard deviation, reflecting the stability of model performance. The timing function is used to measure the single-sample inference time of the model under the specified hardware configuration. The test includes the entire process of preprocessing, model forward calculation, and post-processing to ensure that the evaluation results reflect the performance in the actual deployment scenario. The evaluation function implemented in the code is developed entirely based on actual needs and does not introduce unimplemented evaluation methods or metrics. All test results can be reproduced using the log files in output_dir to ensure the repeatability of the evaluation process.
[0134] Example 2, reference Figure 2-Figure 7 , which is an embodiment of the present invention, provides a poultry health status detection method. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0135] The present invention provides an example of a method for detecting the health status of chickens. The method can be divided into six main parts according to the process: data collection, data preprocessing, model construction, model pretraining, model fine-tuning, and model evaluation. The following are the main steps included in the specific implementation example:
[0136] Data Collection: In a standardized farm environment, a professional recording system was used to collect three key audio samples: natural vocalizations of healthy poultry, abnormal sounds of poultry with confirmed illnesses, and environmental background noise. The recording equipment was evenly distributed throughout the aisles of the chicken coop, sampling at a frequency of 22.05kHz and a depth of 16 bits. Each sound segment lasted 5 seconds and was saved as a WAV file. In this example, a total of 5,000 valid samples were collected, of which 1,000 were classified as positive for healthy poultry sounds and 1,000 as negative for sick poultry sounds. An additional 3,000 unlabeled data items were used for pre-training. Figure 2 The distribution of recording equipment in the farm is shown. Figure 3 To collect device images.
[0137] Data preprocessing: The collected audio data undergoes a multi-stage processing flow: first, pre-emphasis filtering is performed to enhance high-frequency components, followed by framing using a 25ms Hamming window and a 10ms frame shift to ensure time-frequency continuity. A 40-dimensional Mel filter bank (0-8kHz) is used to extract time-frequency features that conform to auditory characteristics, and finally converted into a standardized Mel spectrum. The preprocessed spectrum is shown below: Figure 4 As shown, the difference in the spectrograms between healthy and sick poultry can be clearly observed.
[0138] Model construction: Based on the time-frequency characteristics of sound signals, we chose Vision Transformer (ViT) as the basic framework, mainly considering that its self-attention mechanism can well capture long-range dependencies in the spectrogram.
[0139] The model employs an encoder-decoder architecture, where the encoder extracts high-level feature representations from the partially visible spectrogram. Specifically, the input spectrogram is divided into 16×16 non-overlapping blocks, each of which is converted to a feature vector via linear projection. The encoder consists of 12 Transformer layers, each equipped with a multi-head self-attention mechanism and a feedforward neural network, gradually building a hierarchical representation of sound features.
[0140] The decoder is designed as a lightweight structure, consisting of only four Transformer layers. Its primary task is to reconstruct the masked spectral region based on the features extracted by the encoder. This asymmetric design (12 encoder layers + 4 decoder layers) ensures sufficient feature extraction while limiting the model's computational complexity. The model ultimately outputs a reconstructed complete spectrum, and the pre-training loss is calculated by comparing the reconstruction with the original spectrum.
[0141] In terms of the dataset, healthy poultry sounds and environmental sounds were classified as positive, while sick poultry sounds were classified as negative. The dataset was divided into two groups, A and B. Each group had 800 healthy poultry / environmental sounds (400 healthy and 400 environmental) and 800 sick poultry sounds as training sets; 200 healthy poultry / environmental sounds and 200 sick poultry sounds as testing sets. Group B also had 3,000 unlabeled poultry sounds as pre-training data. Specifically:
[0142] Group A:
[0143] Training set: 800 healthy poultry sounds / environmental sounds, 800 sick poultry sounds.
[0144] Test set: 200 healthy poultry sounds / environmental sounds, 200 sick poultry sounds.
[0145] Group B:
[0146] Pre-training set: 3000 unlabeled poultry sounds.
[0147] Training set: 800 healthy poultry sounds / environmental sounds, 800 sick poultry sounds.
[0148] Test set: 200 healthy poultry sounds / environmental sounds, 200 sick poultry sounds.
[0149] Model pre-training: This stage uses self-supervised learning to train the model on large-scale unlabeled audio data to learn general spectral feature representations. The parameter settings in the example are as follows:
[0150] batch_size=64,Mmodel=mea_vit_large_patch16,mask_ratio=0.75
[0151] epochs=5000,blr=1.5e-4,patience=100,norm_pix_loss,
[0152] warmup_epochs=40,weight_decay=0.05
[0153] The pre-training task is designed to be masked spectrogram reconstruction. This randomly masks 75% of the input Mel-spectrogram (mask_ratio = 0.75), requiring the model to reconstruct the complete spectrum based on the remaining 25% of visible blocks. This ratio has been experimentally verified to achieve the optimal balance between feature learning difficulty and reconstruction feasibility.
[0154] The model architecture and training batch size were selected to ensure stable gradient estimates within GPU memory limitations. A base learning rate was used for 5000 training epochs (epochs = 5000), with a cosine annealing strategy for learning rate scheduling, including 40 for a smooth training start. This long-term, low learning rate configuration helped the model gradually converge to a better local minimum.
[0155] Key regularization parameters were set to: weight_decay = 0.05 to control weight decay and prevent overfitting; norm_pix_loss was enabled to normalize the spectrogram pixels to improve reconstruction quality. During training, patience = 100 was set to implement an early stopping mechanism, terminating training when the validation loss has not decreased for 100 consecutive rounds to avoid wasted computation. These parameter combinations were verified through grid search to maximize the model's feature extraction capabilities.
[0156] After pre-training is completed, the encoder part will learn discriminative time-frequency feature representations, providing high-quality initialization parameters for downstream classification tasks.
[0157] Model fine-tuning: After pre-training, this stage performs supervised fine-tuning on the model to adapt it to the specific health status classification task. The fine-tuning process mainly achieves three key goals: first, adapt the general acoustic features learned in pre-training to the health detection scenario, second, optimize the classification decision boundary, and finally, maintain the model's generalization ability to avoid overfitting. The parameter settings in this example are as follows:
[0158] batch_size=32,model=vit_large_patch16,epochs=2000,blr=1E-3,
[0159] patience=100,layer_decay=0.75,weight_drop=0.05,drop_path=0.1,
[0160] mixup=0.8,cutmix=1.0,reprob=0.25,dist_eval
[0161] This paper uses a progressive unfreezing strategy for fine-tuning. Initially, only the last two layers of Transformer modules and the classification head are unfrozen, and the base learning rate is set to 1e-3. This setting can both leverage pre-trained features and quickly adapt to new tasks. As training progresses, more underlying modules are gradually unfrozen, and the learning rate is decayed layer by layer (layer_decay = 0.75), reducing the bottom-level learning rate to 1e-5. This layered optimization strategy effectively protects the common acoustic features extracted at the bottom level.
[0162] For data augmentation, we configured MixUp (α = 0.8) and CutMix (α = 1.0) with a random erasure probability of 0.25 (reprob = 0.25). These enhancements significantly improve the model's robustness to sound distortion through linear interpolation and regional occlusion. We set the batch size to 32, halving the pre-training phase, to increase the randomness of parameter updates and help escape local optima.
[0163] To prevent overfitting, several regularization techniques were employed: the DropPath rate was set to 0.1, randomly dropping some connections during attention calculations; the label smoothing coefficient was set to 0.1 to mitigate the problem of overconfidence in classifications; and weight decay was maintained at 0.05 to control parameter size. The optimizer continued to use AdamW, but with β = (0.9, 0.95) adjusted to balance the gradient moment estimates.
[0164] The training process was set to 2000 epochs, with an early stopping mechanism (patience = 100). Training was terminated when the validation loss did not decrease for 100 consecutive epochs to avoid inefficient computation. This parameter combination has been thoroughly validated and achieves an optimal balance between computational efficiency and model performance.
[0165] Model Evaluation
[0166] In this phase, the trained model is systematically verified for performance, focusing on evaluating its classification accuracy and robustness in real-world farming scenarios. End-to-end testing is performed using an independent test set (200 healthy / environmental sounds and 200 sick poultry sounds), and core indicators such as accuracy, recall, and F1-score are calculated. The performance of the model at different thresholds is evaluated by drawing curves, and the real-time inference speed is monitored during the deployment phase to ensure that the timeliness requirements of online detection in farms are met. Finally, a visual report is output to provide a basis for model optimization. The evaluation results show that Figure 5 This is the classification training accuracy curve for group A. The accuracy of group A is 67.25%; Figure 6 is the pre-training loss curve for group B, Figure 7 This is the classification training accuracy curve for group B. The accuracy of group B is 75.00%. The model accuracy before and after pre-training is shown in Table 1.
[0167] Table 1 Comparison of Group A and Group B
[0168]
[0169]
[0170] From the comparison in Table 1, we can see that the model with pre-training (Group B) has an accuracy improvement of 7.75 percentage points compared to the baseline (Group A), and shows stronger adaptability to device differences and environmental noise.
[0171] Example 3, reference Figure 8 , which is an embodiment of the present invention, provides a poultry health status detection system, including an audio collection and preprocessing module 100, a MAE modeling module 200, and a model training and evaluation module 300.
[0172] Among them, S4: audio acquisition and preprocessing module 100 includes audio acquisition unit 101 and audio preprocessing unit 102. The audio acquisition unit 101 is used to collect the call data of poultry in different states and synchronously record the sampling environment information. The audio preprocessing unit 102 is used to perform pre-emphasis, framing, Mel filtering and other processing on the collected audio signal and convert it into a standardized Mel spectrum.
[0173] It should also be noted that the audio acquisition unit 101 collects the call data of poultry in different states and synchronously records the sampling environment information to generate raw audio data. The audio preprocessing unit 102 receives the raw audio data, completes pre-emphasis, framing and Mel filtering, and outputs a standardized Mel spectrum.
[0174] S5: The MAE modeling module 200 includes a spectrum graph division unit 201 and a coding structure construction unit 202. The spectrum graph division unit 201 is used to divide the Mel spectrum graph into image blocks of equal size as model input. The coding structure construction unit 202 is used to build a Transformer-based MAE model structure, including an encoder and decoder architecture.
[0175] It should also be noted that the spectrum graph division unit 201 receives the Mel spectrum graph, divides it into image blocks of equal size, and generates a patch sequence. The encoding structure construction unit 202 receives the patch sequence, constructs a Transformer-based MAE model structure, and performs mask encoding and reconstruction operations.
[0176] S6: The model training and evaluation module 300 includes a pre-training unit 301 and an evaluation and verification unit 302. The pre-training unit 301 adopts a self-supervised learning method to train the MAE model through the masked reconstruction task. The evaluation and verification unit 302 is used to fine-tune and verify the accuracy of the trained model on a labeled dataset.
[0177] It should also be noted that the pre-training unit 301 receives the MAE model structure output by the encoding structure construction unit, performs training based on the self-supervised task, generates a pre-trained model, and the evaluation and verification unit 302 receives the pre-trained model, performs fine-tuning and verification based on the label data set, and outputs the final detection model.
[0178] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0179] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0180] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0181] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc. It should be noted that the above embodiments are merely illustrative of the technical solutions of the present invention and are not intended to be limiting. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced with equivalents without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications should be encompassed by the claims of the present invention.
Claims
1. A method for detecting the health status of poultry, characterized in that: include: collecting poultry audio data and preprocessing the poultry audio data; Construct a MAE model based on the time-frequency characteristics of audio data; Pre-train and fine-tune the MAE model, and evaluate the MAE model after pre-training.
2. The poultry health status detection method according to claim 1, wherein: The poultry audio data collection includes collecting tag audio data at a preset frequency, and the recording devices are evenly arranged in the aisles of the chicken house.
3. The poultry health status detection method according to claim 1 or 2, characterized in that: The labeled audio data includes the natural sounds of healthy poultry and environmental background noise as positive labels, and the abnormal sounds of poultry with confirmed diseases as negative labels. The labeled audio data is divided into a training set and a test set, and a preset number of unlabeled audio data are collected as a pre-training set for the MAE model.
4. The poultry health status detection method according to claim 3, wherein: The poultry audio data is preprocessed by processing high-frequency components through pre-emphasis filtering, framing the audio data using a Hamming window, extracting time-frequency features that conform to auditory characteristics through a 40-dimensional Mel filter bank, and converting the features into a standardized Mel spectrum.
5. The method for detecting the health status of poultry according to any one of claims 1, 2 or 4, wherein: The MAE model is constructed by selecting Vision Transformer as a basic framework based on the time-frequency characteristics of the sound signal. The MAE model adopts an encoder-decoder structure and uses the encoder to extract high-level feature representations from partially visible spectrograms. The decoder reconstructs the masked spectral region using the features extracted by the encoder.
6. The poultry health status detection method according to claim 5, wherein: The pre-training and fine-tuning of the MAE model include: pre-training includes using a self-supervised learning method to train the MAE model on large-scale unlabeled audio data to learn a general spectral feature representation, randomly masking 75% of the blocks in the input Mel-spectrogram, setting the MAE model to reconstruct the complete spectrum based on the remaining 25% of visible blocks, selecting the MAE model architecture, setting the training batch size, obtaining a stable gradient estimate under GPU memory limitations, scheduling the learning rate through a cosine annealing strategy, controlling the weight decay strength by setting a key regularization parameter, and normalizing the spectrum using norm_pix_loss; Fine-tuning includes adapting the general acoustic features learned in pre-training to health detection scenarios, using a progressive unfreezing strategy for fine-tuning, unfreezing the last two layers of Transformer modules and classification heads in the initial stage, setting a basic learning rate, and gradually unfreezing more underlying modules while pre-training is in progress. The learning rate is attenuated layer by layer, data enhancement is performed through linear interpolation and area occlusion, and multiple regularization techniques are used to control overfitting.
7. The method for detecting the health status of poultry according to any one of claims 1, 2, 4 or 6, wherein: The evaluation of the pre-trained MAE model includes performing end-to-end testing on the MAE model using a test set, evaluating the performance of the MAE model at different thresholds by drawing curves, monitoring the real-time inference speed during the deployment phase, determining whether the timeliness requirements of online detection of farms are met, and outputting a visual report.
8. A system using the poultry health status detection method according to any one of claims 1 to 7, characterized in that: It includes an audio acquisition and preprocessing module (100), a MAE modeling module (200), and a model training and evaluation module (300); The audio collection and preprocessing module (100) comprises an audio collection unit (101) and an audio preprocessing unit (102), wherein the audio collection unit (101) is used to collect the call data of poultry in different states and synchronously record the sampling environment information, and the audio preprocessing unit (102) is used to perform pre-emphasis, framing, Mel filtering and other processing on the collected audio signal and convert it into a standardized Mel spectrum. The MAE modeling module (200) includes a spectrum graph division unit (201) and a coding structure construction unit (202), wherein the spectrum graph division unit (201) is used to divide the Mel spectrum graph into image blocks of equal size as model input, and the coding structure construction unit (202) is used to build a Transformer-based MAE model structure, including an encoder and a decoder architecture; The model training and evaluation module (300) includes a pre-training unit (301) and an evaluation and verification unit (302). The pre-training unit (301) uses a self-supervised learning method to train the MAE model through a masked reconstruction task. The evaluation and verification unit (302) is used to fine-tune and verify the accuracy of the trained model on a labeled data set.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the poultry health status detection method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the poultry health status detection method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Large model fine tuning method and legal affair information analysis method based on task awareness
CN120973957A
Animal respiratory tract health detection method and system
CN121606281A
Livestock abnormal sound field identification method, device, equipment and medium
CN121811924A
Cage-rearing laying hen abnormal sound monitoring method and system
CN122050398A
A duck early disease monitoring system and method based on sound spectrum features
CN122474084A