Pig online monitoring and health state intelligent evaluation method and system based on acoustics

CN122531674APending Publication Date: 2026-08-07XIANGTAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIANGTAN UNIV
Filing Date
2026-07-01
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

当前,业内针对生猪健康状态的监测方式主要依赖人工巡视与视觉检测,两种传统监测方式存在固有技术缺陷,难以满足现代化智能养殖的精准监测与早期预警需求

Benefits of technology

本发明摒弃人工主观判断模式,采用全天候动静双模式协同音频采集方式,可实现圈舍无间断常态化监测,不受人工工作时长与精力限制;依托精细化声学信号处理与智能模型识别能力,能够精准捕捉人工肉眼无法识别的早期轻微咳嗽、喘鸣、隐性应激等隐蔽性病理发声信号,从根源上避免人工经验判断带来的漏判、误判问题,实现生猪健康异常的早发现、早预警,大幅前移疫病防控干预时机。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531674A_ABST
    Figure CN122531674A_ABST
Patent Text Reader

Abstract

The application discloses a pig online monitoring and health state intelligent evaluation method and system based on acoustics, which adopts dynamic and static dual-mode cooperation to complete global audio collection and preprocessing of the pig house, relies on an underdetermined blind source separation algorithm of breeding behavior prior constraint to strip the aliasing sound source, combines voiceprint matching and spatial positioning verification, and realizes individual accurate binding of the pig sound signal. Through frequency band layering, time-frequency domain static characteristics and dynamic differential characteristics are extracted, a multi-scale acoustic texture and time sequence dual-mode fusion model is constructed, and a special loss function, optimization scheduling and multiple regularization strategies are matched, so that the noise resistance and generalization performance of the model are effectively improved, and the pig sound type and four-level health state results are accurately output. Meanwhile, according to the abnormal grade, three-level grading early warning and equipment linkage intervention are triggered, the pig health state is all-weather and fine intelligent monitoring and closed-loop prevention and control are realized, and the application is suitable for large-scale intelligent breeding scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent monitoring technology for pig farming, and in particular to an acoustic-based method and system for online monitoring and intelligent assessment of the health status of pigs. Background Technology

[0002] In large-scale pig farming, real-time monitoring of pig health status is a core element of disease prevention and control and refined management. Currently, the industry mainly relies on manual inspection and visual detection to monitor pig health status. These two traditional monitoring methods have inherent technical limitations and cannot meet the precise monitoring and early warning requirements of modern intelligent farming.

[0003] The manual inspection method relies heavily on on-site observation and the personal experience of the staff, making it impossible to achieve continuous 24-hour monitoring in the pigpens. The timeliness and continuity of monitoring are extremely poor. Furthermore, subtle abnormalities such as early pathological vocalizations and mild stress-induced vocalizations in pigs cannot be identified by the naked eye, easily leading to missed or incorrect diagnoses. This results in delayed detection of health abnormalities in pigs and hinders early disease intervention.

[0004] Visual detection monitoring methods have significant limitations in terms of application scenarios. They can only collect information on the appearance of pigs and their behavior within the visible range. Factors such as obstruction by pen space, stacking of individual pigs, and changes in lighting can all create blind spots, making it impossible to achieve full-area, comprehensive monitoring. Furthermore, in the early stages of many respiratory diseases and internal inflammation in pigs, there are no obvious abnormal changes in their appearance or behavior. The pathological characteristics are only revealed through vocalizations and breathing patterns. Visual detection cannot capture these latent health risks, resulting in serious monitoring lag and limitations.

[0005] Meanwhile, existing swine health monitoring technologies have not fully explored the strong correlation between swine acoustic signals and their own health status, have not established a comprehensive acoustic identity management system for individual swine, and lack standardized technical solutions for sound source separation, precise individual matching, and extraction of specific acoustic features tailored to farming scenarios. This makes it impossible to achieve accurate acoustic monitoring of individual swine and to complete differentiated assessments of individual health status. The overall technical system lacks sufficient intelligence and refinement, failing to meet the all-weather, high-precision health monitoring needs of large-scale swine farming. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides an acoustic-based online monitoring and intelligent health status assessment method and system for pigs. By constructing a multi-scale acoustic texture and time-series dual-modal fusion model, it resists interference from pen environmental noise, accurately captures early hidden pathological vocalization characteristics, and combines an acoustic identity database to complete individual tracing, thereby achieving intelligent health status assessment and closed-loop prevention and control.

[0007] In a first aspect, the present invention provides an acoustic-based method for online monitoring and intelligent assessment of the health status of pigs, specifically including: The system performs coordinated full-area audio signal acquisition in both dynamic and static modes in pig pens, and preprocesses the acquired audio signals. The effective sound segments are detected and located. A sound source separation algorithm based on prior constraints of breeding behavior is used to extract a single pig sound source from the aliased signal. The sound signal is then uniquely matched with the individual pig by combining the pig acoustic identity database. The matched single pig sound source signal is subjected to pig pathological vocalization-specific frequency band layered feature extraction to obtain time-frequency domain static features and dynamic differential features; and these features are fused to form initial acoustic features, which are then optimized to obtain complete acoustic features. The complete acoustic features are input into a multi-scale acoustic texture and temporal dual-modal fusion model, which outputs the probability distribution of pig vocalization types and health status assessment results, and triggers graded early warnings based on the health status assessment results.

[0008] This method overcomes the shortcomings of traditional monitoring methods, which cannot distinguish individuals and are difficult to capture subtle pathological sounds. It can separate target sound sources from complex aliased audio and complete individual matching. Combined with hierarchical feature extraction and dual-modal fusion model, it enhances the ability to identify anomalies, can promptly detect hidden health problems in pigs and trigger early warnings, and significantly reduce the probability of missed or false alarms, thus meeting the intelligent monitoring needs of large-scale farming.

[0009] In this embodiment, the method of extracting a single pig sound source from the aliased signal using a sound source separation algorithm based on prior constraints of farming behavior specifically includes: An underdetermined blind source separation algorithm based on prior constraints of breeding behavior is adopted, which integrates spatial positioning information and the spatiotemporal distribution pattern of pig behavior to construct a non-uniform hybrid matrix estimation prior. The clustering estimation process of the mixture matrix is ​​guided by prior information on breeding behavior. The mixture matrix is ​​estimated by a prior-constrained clustering method, taking into account the time-frequency sparsity of the acoustic signal of pigs. The time-frequency sparsity means that there is usually only one dominant sound source at the same time-frequency point. Based on the estimated mixing matrix, the L1 norm minimization method is used to reconstruct the components of each independent sound source, and the single pig sound source is adaptively calculated from the multi-source aliased signal.

[0010] It should be noted that this sound source separation method relies on the prior knowledge of breeding behavior and the time-frequency domain sparsity optimization algorithm logic to accurately complete the solution of the mixing matrix and the reconstruction of independent sound sources. It can efficiently extract the single pig vocalization signal from the aliased audio, has strong anti-interference ability, and ensures the effectiveness and stability of the sound source separation results.

[0011] In this embodiment, the step of combining the pig acoustic identity database to uniquely match the vocal signal with an individual pig specifically includes: Acoustic features were extracted from the single pig sound source; Calculate the Euclidean distance and cosine similarity between the acoustic features and each standard acoustic feature template of pigs in the pig acoustic identity database; The feature matching degree of all pig individuals is ranked, and the pig individual with the highest matching degree is selected as the preliminary matching result; Spatial positioning data is introduced to verify the preliminary matching results, thus completing the unique binding of the vocal signal to the individual pig.

[0012] It should be noted that by combining multi-dimensional acoustic feature comparison with spatial location joint verification, the limitations of a single matching algorithm are overcome, effectively reducing matching errors in aliased scenarios, reliably establishing a unique association between the vocal signal and the individual pig, and ensuring the correspondence of individual-level monitoring data.

[0013] In this embodiment, the matched single swine sound source signal undergoes layered feature extraction based on the specific frequency band of swine pathological vocalization to obtain time-frequency domain static features and dynamic differential features; these features are then fused to form initial acoustic features, specifically including: The sound signal from a single pig source is divided into low-frequency band, mid-frequency band, and high-frequency band. The first correlation feature is extracted in the low-frequency band, the second correlation feature is extracted in the mid-frequency band, and the third correlation feature is extracted in the high-frequency band. By merging the first correlation feature, the second correlation feature, and the third correlation feature, the time-frequency domain static feature is obtained; The first-order and second-order differences of the static time-frequency domain features are calculated to obtain the dynamic difference features; the first-order difference reflects the rate of change of the features, and the second-order difference reflects the acceleration of the change of the features. The time-frequency domain static features are fused with the corresponding dynamic difference features to form the initial acoustic features.

[0014] It is evident that this extraction method closely aligns with the frequency band distribution patterns of pathological vocalizations in pigs. It not only preserves the basic acoustic features of the signal but also supplements dynamic information such as the rate of feature change and fluctuation amplitude. This solves the problem that single static features are insufficient to identify latent and intermittent abnormal vocalizations, significantly improving the feature's ability to identify abnormal health states.

[0015] In this embodiment, the construction process of the multi-scale acoustic texture and temporal bimodal fusion model includes: A network architecture for a multi-scale acoustic texture and temporal dual-modal fusion model was constructed, including an adaptive time-frequency enhancement module, a multi-scale wavelet acoustic texture encoder, a lightweight temporal correlation gating module, a temporal-frequency domain hetero-weighted dual-modal fusion module, and a global average pooling layer, a Dropout layer, and a fully connected classification layer, which were cascaded in sequence according to their functional logic. The adaptive time-frequency enhancement module divides the complete acoustic features of the input into sub-bands according to the frequency dimension, distinguishes the effective sound segment from the noise segment by calculating the energy entropy of each sub-band, generates an adaptive noise mask, performs feature gain on the effective sound area and feature attenuation on the periodic noise area, and obtains the noise-reduced and enhanced feature tensor. The multi-scale wavelet acoustic texture encoder uses a multi-scale discrete wavelet convolution group to extract corresponding acoustic texture features from the feature tensor according to the low frequency band, mid frequency band and high frequency band and perform weighted calculation. The lightweight temporal correlation gating module introduces a temporal decay factor to adaptively and iteratively update the historical temporal state, mines the temporal dependency of continuous abnormal vocalizations of pigs based on the multi-scale acoustic texture features output by the encoder, and filters out the temporal interference of periodic environmental noise in the pigsty. The time-frequency domain hetero-weighted dual-modal fusion module performs independent attention weight learning on the features after time-series modeling from the time-series dimension and the frequency domain dimension, respectively, and performs hetero-weighted deep weighted fusion of time-series dynamic features and frequency domain acoustic texture features. The fused features are sequentially processed through a global average pooling layer for feature dimensionality reduction and aggregation, and a Dropout layer to suppress model overfitting. Finally, the output is mapped through a fully connected classification layer to obtain the probability distribution of various vocalization types in pigs and the health status assessment results of normal, mildly abnormal, moderately abnormal, and severely abnormal.

[0016] It should be noted that this model integrates time-frequency enhancement, multi-scale feature extraction, temporal modeling and dual-modal weighted fusion technologies, which can effectively filter environmental noise, fully explore the frequency domain texture and temporal variation law of acoustic signals, and optimize model performance by combining regularization methods. It can accurately identify the vocalization type of pigs and classify the level of health abnormality, and adapt to the online recognition needs of complex acoustic environment in pig farms.

[0017] In this embodiment, the multi-scale acoustic texture and temporal bimodal fusion model further includes: The training process of the multi-scale acoustic texture and temporal dual-modal fusion model adopts FocalLoss combined with label smoothing as the loss function, uses AdamW adaptive optimizer and sets weight decay coefficient, and is matched with cosine annealing learning rate scheduling strategy. At the same time, the EMA exponential moving average method is used to smoothly update the model weights. The training process incorporates batch normalization, Dropout mechanism and gradient clipping strategy for regularization constraints, and adopts early stopping strategy. When the performance of the validation set does not improve for a preset number of consecutive rounds, training is stopped to retain the optimal weights of the model. After the model training is completed, a 5-fold cross-validation method is used to evaluate its performance.

[0018] It should be noted that, compared to conventional model training methods, this training strategy specifically addresses issues such as imbalanced acoustic samples in aquaculture, easy overfitting of the model, poor training convergence, and biased evaluation results. Multiple optimization and regularization techniques work synergistically to fully explore sample features and standardize the training process, thereby improving the model's ability to recognize various vocalization states and ensuring its generalization performance in complex real-world scenarios.

[0019] In this embodiment, triggering a tiered early warning based on the health status assessment result includes: Based on the health status assessment results, corresponding Level 3 early warning levels are triggered for abnormal intervention, including: Mild abnormalities trigger a Level 1 warning, increasing ventilation in the pigpens and highlighting the target pigs. A moderate abnormality triggers a Level 2 alert, prompting the isolation of pigs and adjustment of the feeding program; In case of severe abnormality, an emergency alarm will be triggered, the unit fan will be shut down, and the disinfection equipment will be turned on.

[0020] Secondly, the present invention also provides an acoustic-based online monitoring and health status intelligent assessment system for pigs, the intelligent assessment system comprising: an audio acquisition and preprocessing module, a sound source separation and individual matching module, a feature extraction and optimization module, and an intelligent assessment and early warning linkage module; The audio acquisition and preprocessing module is used to perform dynamic and static dual-mode coordinated full-domain audio signal acquisition in the pigsty and to complete the preprocessing operation of the acquired signals. The sound source separation and individual matching module is used to detect and locate the effective sound segments in the audio. It uses a sound source separation algorithm based on the prior constraints of breeding behavior to extract a single pig sound source from the aliased signal, and combines the pig acoustic identity database to achieve a unique match between the sound signal and the individual pig. The feature extraction and optimization module is used to extract pathological vocalization-specific frequency band layered features from the matched single pig sound source signal, obtain time-frequency domain static features and dynamic differential features, and fuse them into initial acoustic features. Then, the initial acoustic features are optimized to output complete acoustic features. The intelligent assessment and early warning linkage module is used to input complete acoustic features into a multi-scale acoustic texture and temporal dual-modal fusion model, output the probability distribution of pig vocalization types and health status assessment results, and trigger graded early warnings and corresponding linkage operations based on the assessment results.

[0021] The intelligent assessment and early warning linkage module also includes: a model building unit and a model optimization unit; The model building unit is used to build a network architecture for a multi-scale acoustic texture and temporal dual-modal fusion model, including an adaptive time-frequency enhancement module, a multi-scale wavelet acoustic texture encoder, a lightweight temporal correlation gating module, a temporal-frequency domain hetero-weighted dual-modal fusion module, and a global average pooling layer, a Dropout layer, and a fully connected classification layer cascaded in sequence according to functional logic. The model optimization unit is used to perform model training, regularization constraints, performance verification, generalization testing, and parameter fine-tuning.

[0022] The intelligent evaluation system also includes: an acoustic identity database module; The acoustic identity database module uses the pig entry stage as a unified initialization node. It constructs a unique acoustic identity for each pig through audio collection. The unique acoustic identity is uniquely associated with the unique individual number bound to the pig's ear tag, the pen and stall location information, and the pig's age and growth cycle basic information.

[0023] Compared with the prior art, the present invention has the following advantages and beneficial effects: This invention abandons the subjective judgment mode and adopts a 24 / 7 dynamic and static dual-mode collaborative audio acquisition method, which can realize uninterrupted and normalized monitoring of pig pens, without being limited by the working hours and energy of humans. Relying on refined acoustic signal processing and intelligent model recognition capabilities, it can accurately capture early mild cough, wheezing, and latent stress and other hidden pathological vocal signals that are not visible to the naked eye. It avoids the problems of missed and misjudgment caused by human experience judgment from the root, realizes early detection and early warning of abnormal pig health, and significantly advances the timing of disease prevention and control intervention.

[0024] Furthermore, this invention uses the acoustic signals of pigs as the monitoring carrier, which is not affected by factors such as obstruction by pen space, stacking of pigs, and changes in lighting, and has no blind spots in monitoring. At the same time, it makes full use of the strong correlation between the vocalization state of pigs and internal pathology, respiratory lesions, and stress state, and can accurately detect potential health risks in the early stages when there are no obvious abnormalities in the appearance and behavior of pigs. This makes up for the technical limitations of visual monitoring, which can only identify visible abnormalities on the body surface, and greatly improves the depth and predictive ability of pig health monitoring.

[0025] A fully intelligent closed-loop system has been constructed, which realizes standardized graded early warning and intelligent linkage intervention of breeding equipment through quantitative health scoring. At the same time, the system relies on the feedback of on-site labeled data to realize incremental iterative evolution of the model, so that the monitoring accuracy of the system can be continuously optimized with long-term use. This completely solves the problems of poor standardization, lack of closed-loop management and inability to upgrade performance in traditional monitoring methods, and effectively meets the needs of refined, intelligent and normalized health monitoring and management in large-scale pig farming. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of an acoustic-based online monitoring and intelligent health status assessment method for pigs.

[0027] Figure 2 This is a schematic diagram of an acoustic-based online monitoring and intelligent health status assessment system for pigs.

[0028] Figure 3 This is a schematic diagram of the intelligent assessment and early warning linkage module. Detailed Implementation

[0029] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention. It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0030] Example 1: This invention provides an acoustic-based method for online monitoring and intelligent health status assessment of pigs, such as... Figure 1 As shown, it specifically includes: Step 1: Collect dynamic and static dual-mode coordinated full-area audio signals from the pigsty and preprocess the collected audio signals; Step 2: Detect and locate the effective sound segment, and use a sound source separation algorithm based on the prior constraints of breeding behavior to extract the single pig sound source from the aliased signal. Combine the pig acoustic identity database to complete the unique matching of the sound signal with the individual pig. Step 3: Perform layered feature extraction of the specific frequency band for pathological vocalization of pigs on the matched single pig sound source signal to obtain static features and dynamic differential features in the time and frequency domains; and fuse them to form initial acoustic features, and optimize the initial acoustic features to obtain complete acoustic features; Step 4: Input the complete acoustic features into the multi-scale acoustic texture and temporal dual-modal fusion model, output the probability distribution of pig vocalization types and health status assessment results, and trigger graded early warning based on the health status assessment results.

[0031] This method overcomes the shortcomings of traditional monitoring methods, which cannot distinguish individuals and are difficult to capture subtle pathological sounds. It can separate target sound sources from complex aliased audio and complete individual matching. Combined with hierarchical feature extraction and dual-modal fusion model, it enhances the ability to identify anomalies, can promptly detect hidden health problems in pigs and trigger early warnings, and significantly reduce the probability of missed or false alarms, thus meeting the intelligent monitoring needs of large-scale farming.

[0032] In this embodiment, in step 1, considering the characteristics of large-scale pig pens with large spatial spans, dispersed pig activity ranges, complex environmental noise, and strong randomness of sound signals, this invention adopts a dual-mode (static and dynamic) collaborative acquisition method to achieve full coverage acquisition of audio signals across the entire pen area, completely avoiding the problems of signal omission, local monitoring blind spots, and incomplete signal acquisition that exist in traditional single-point audio acquisition. Specifically, the static acquisition mode is a fixed-point, routine acquisition method. By deploying fixed audio acquisition devices at key locations such as above the pens, on both sides of the passageways, and at the four corners of the pen, and setting fixed sampling frequencies and accuracy, this achieves 24-hour uninterrupted, routine audio signal acquisition across the entire pen area, covering the fixed sound field monitoring of the entire area where pigs are feeding, resting, and active. The dynamic acquisition mode is a mobile auxiliary acquisition method. Using a mobile audio acquisition terminal, it dynamically supplements acquisition by following the pig activity trajectory and pen inspection path. For special scenarios such as pigs clustering together, resting in corners, or exhibiting localized activity, it supplements audio signal acquisition in signal blind spots and weak signal areas of the fixed acquisition points, compensating for the limitations of static fixed acquisition.

[0033] In this embodiment, the dynamic and static dual-mode collaborative acquisition method mainly uses a microphone array, which is deployed according to the standard pen layout of a large-scale pig pen. Typically, a set of microphones is placed every 5-8 meters on the top of the pen to ensure that there is an effective pickup point above each pen. The microphones use industrial-grade audio equipment with high signal-to-noise ratio, anti-interference, and wide frequency response, which can stably capture various acoustic signals such as pig vocalization, coughing, panting, and stress screaming, achieving 24 / 7 uninterrupted fixed-point acquisition, without being limited by time, personnel, or environmental factors, and making up for the shortcomings of manual inspection that cannot be on duty around the clock. The mobile data acquisition unit utilizes an automated inspection robot. Equipped with a directional microphone, positioning module, and environmental perception module, the robot circulates within the pigpen along a pre-set route. This route covers blind spots inaccessible to fixed microphones, such as passageways between pens, feeding areas, drinking areas, and rest areas. The robot's acquisition frequency can be flexibly adjusted according to the stocking density. In high-density areas, a full-area inspection is completed every 15-30 minutes, while in low-density areas, it is completed every 30-60 minutes. This mobile acquisition further enhances the completeness and coverage of acoustic signal acquisition. All acquisition devices adhere to standardized acquisition parameters: a sampling rate of 16000Hz to meet the complete acquisition requirements of the high-frequency band of pig vocalization; a monophonic acquisition mode to reduce data redundancy and transmission pressure; and 16-bit quantization precision, using 16-bit binary digits to digitally represent the acoustic signal amplitude, ensuring complete preservation of signal details while balancing data storage and transmission efficiency. The raw audio signals acquired are transmitted in real time to the backend data server via the internal WiFi network or 5G mobile communication network of the pigpens. Simultaneously, local caching is performed at the edge computing unit at the acquisition end to prevent data loss due to network fluctuations. Data storage employs a hierarchical and categorized management mechanism. The server constructs a multi-level directory structure based on pen number, acquisition date, acquisition time, device number, and acquisition type. Audio files are stored in standard WAV format. Auxiliary information such as the coordinates of the acquisition point, pen environmental temperature and humidity, stocking density, and pig growth cycle are also recorded, forming a complete and traceable acoustic data resource pool. This provides a unified data entry point for subsequent signal preprocessing, feature extraction, model training, and online evaluation.

[0034] After the audio signal acquisition is completed, the raw audio signal undergoes standardized preprocessing to filter out invalid interference noise in the complex environment of the pig farm, normalize the signal format, enhance the effective sound signal, and improve the signal quality for subsequent sound source separation and feature extraction. The specific preprocessing process is strictly executed in a fixed order, namely fixed-length framing, amplitude normalization, dereverberation, adaptive noise reduction, and frame-level standardization.

[0035] First, fixed-length framing is performed, dividing the continuous audio stream into fixed segments of 1 second each. This length matches the typical duration of a single vocalization, cough, or wheezing in a pig, preserving acoustic characteristics while avoiding multi-source noise caused by excessively long segments. Segments shorter than 1 second are padded with zeros at the end, while segments longer than 1 second are truncated at the front end to ensure uniform length for all signal segments.

[0036] Next, amplitude normalization is performed to convert the acquired int16 integer audio data into float32 floating-point data and normalize the amplitude range to the [-1,1] interval to eliminate amplitude differences caused by different microphones and different acquisition distances, ensuring uniform signal amplitude scale and facilitating subsequent feature calculation.

[0037] Subsequently, weighted spectral subtraction was used to remove reverberation interference from the pigsty. Pig pens are mostly enclosed spaces, and reflections from walls, floors, and roofs can cause sound superposition and reverberation, making the signal unclear. Weighted spectral subtraction effectively suppresses reflection interference and improves signal clarity by estimating the reverberation component and subtracting it from the original signal.

[0038] Adaptive noise reduction is achieved through the minimum statistical noise estimation algorithm. This algorithm can track the environmental noise level in real time, automatically distinguish between effective sound and background noise, and filter out stable environmental noise such as fans, water curtains, and mechanical operation while retaining the sound of the pig target. This improves the audio signal-to-noise ratio by more than 10dB, meeting the requirements for high-precision feature extraction.

[0039] Finally, zero-mean, unit-variance standardization is performed on each frame of signal. The calculation formula is as follows: ; In the formula: The average amplitude of a single frame of audio signal. This represents the standard deviation of the amplitude of a single frame of audio signal. Standardization eliminates differences in signal amplitude distribution, ensuring consistent distribution characteristics for signals across different segments and environments, thereby improving model training stability and inference consistency.

[0040] In this embodiment, the effective sound segments of the preprocessed audio signal are detected and accurately located to distinguish between invalid noise segments from the pigsty and actual sound segments from the pigs. This invention employs a joint detection algorithm of short-time energy and short-time zero-crossing rate to process the entire audio signal in frames, calculating the corresponding short-time energy and zero-crossing rate features frame by frame. Through a preset dual threshold screening mechanism, audio frames whose energy and zero-crossing rate both meet the characteristics of pig sounding are selected. Continuous effective audio frames are spliced ​​together to form a complete effective sound segment. At the same time, the start time, end time, and corresponding spatial acquisition position of each sound segment are marked, completing the accurate time and spatial positioning of the effective sound segments from the pigs. Invalid signal segments such as residual environmental noise and equipment noise are accurately eliminated, thus locking in the effective analysis range for subsequent sound source separation.

[0041] In this embodiment, step 2 employs a sound source separation algorithm based on prior constraints of farming behavior to extract a single pig sound source from the aliased signal, specifically including: Step 2.1 adopts an underdetermined blind source separation algorithm based on prior constraints of breeding behavior, integrates spatial positioning information with the spatiotemporal distribution pattern of pig behavior, and constructs a non-uniform hybrid matrix estimation prior; Step 2.2 Utilize prior information on breeding behavior to guide the clustering estimation process of the mixture matrix. Combine the time-frequency sparsity of the pig acoustic signal and estimate the mixture matrix through a prior-constrained clustering method. Wherein, the time-frequency sparsity means that at the same time-frequency point, usually only one sound source is dominant. Step 2.3 Based on the estimated mixing matrix, the L1 norm minimization method is used to reconstruct the independent sound source components, and the single pig sound source is adaptively calculated from the multi-source aliased signal.

[0042] In this embodiment, step 2.1, in its specific implementation, firstly relies on the layout of the audio acquisition equipment in the pigpen to input the precise spatial coordinates of each acquisition terminal, the pen zoning range, and the regional sound field coverage range, etc., to establish a spatial sound field coordinate system for the pigpen and clarify the spatial region affiliation relationship corresponding to different audio signals. Simultaneously, combining the standardized feeding, resting, and activity rhythms of large-scale pig farming, the spatiotemporal distribution patterns of pig behavior are summarized as core prior constraints. In terms of time, pigs exhibit fixed behavioral characteristics at different times. During resting periods, pigs vocalize less frequently and in concentrated areas, while during feeding and activity periods, pigs vocalize more frequently, interact more individually, and have a higher probability of mixing. In terms of spatial dimension, the activity range of pigs is strictly limited by the pen zoning. The location of individual vocalizations is only distributed within their own pen and a small adjacent area. There is no situation of random vocalizations over long distances across pens.

[0043] Based on the aforementioned spatial positioning information and the prior knowledge of the spatiotemporal distribution of pig behavior, a non-uniform mixing matrix estimation prior adapted to pig farming scenarios is constructed. Specifically, according to the pig distribution density in different spatial regions and the vocal activity at different times, the mixing weight constraints of sound sources in each space and time period are set differently. Higher sound source mixing iteration weights are assigned to areas with high-frequency vocalization by pigs and peak activity periods, while noise suppression weight constraints are set for areas without pig distribution and quiet resting periods. This ensures that the estimation process of the mixing matrix fully conforms to the actual sound source distribution and pig behavior patterns in the pigpen, effectively narrowing the solution range of the mixing matrix estimation, avoiding invalid parameter iterations, and significantly improving the accuracy and convergence efficiency of subsequent mixing matrix estimation. This provides reliable prior support for the accurate separation of single pig sound sources in complex aliased audio scenarios.

[0044] In this embodiment, step 2.2 completes the accurate solution of the hybrid matrix based on priors. In specific implementation, the spatiotemporal priors of breeding behavior and spatial positioning constructed in step 2.1 are embedded in the entire clustering iteration process to forcibly constrain the update range and iteration direction of the cluster centers, avoiding abnormal solutions in the clustering results that do not conform to the spatial and behavioral patterns of pig breeding. During the clustering iteration process, effective clustering samples are continuously screened based on the spatial range of the pen and the characteristics of pig activity periods, and invalid noise clustering points that cross regions or are not active during periods are eliminated. At the same time, the inherent time-frequency domain sparsity characteristics of pig acoustic signals are fully utilized to assist in matrix estimation. The time-frequency domain sparsity is specifically defined as follows: after time-frequency transformation of the mixed audio signals in the pig pen, the signal energy corresponding to each independent time-frequency point is usually dominated by a single pig sound source, and there is no situation where the energy of multiple sound sources is superimposed and dominant. This characteristic provides the core basis for multi-source separation under underdetermined conditions. Based on this sparsity characteristic, the algorithm selects pure time-frequency points with high signal-to-noise ratio and single sound source dominance as effective clustering samples, and discards inferior time-frequency samples with weak aliasing from multiple sources and severe noise interference, which greatly improves the purity and effectiveness of clustering samples.

[0045] This invention employs a priori-constrained clustering method to perform iterative estimation of the hybrid matrix. Unlike traditional unconstrained global clustering, this method uses farming behavior as a priori to limit the cluster solution boundary and uses time-frequency sparsity to screen high-quality feature samples. When iteratively updating the cluster centers, it always adheres to the actual sound source distribution patterns in the pigsty, effectively avoiding problems such as cluster divergence, local optima, and center shift. Finally, iterative convergence yields a high-precision hybrid matrix adapted to the complex and overlapping scenarios of pig farming, providing a reliable matrix parameter foundation for subsequent accurate reconstruction of independent pig sound sources.

[0046] In this embodiment, step 2.3 involves simultaneously modeling the preprocessed aliased audio time-frequency signal from the pigsty with the estimated mixing matrix to construct an underdetermined sound source reconstruction equation set adapted to the pig farming scenario. Considering the inherent time-frequency sparsity of pig acoustic signals, this invention uses the L1 norm minimization method as the core constraint to solve for the sound sources. Unlike the conventional L2 norm, which tends to smooth and weaken subtle pathological vocal details and cause the loss of low-level abnormal vocalizations, the L1 norm minimization constraint can impose extreme sparsity constraints on the time-frequency redundancy coefficients, forcing the time-frequency coefficients of non-dominant sound sources to approach zero. This maximizes the preservation of subtle and effective acoustic features such as slight wheezing, low-frequency coughing, and weak stress sounds in pigs. By iteratively solving the L1 norm minimization objective function, the time-frequency component coefficients of each sound source are continuously optimized and corrected, gradually decoupling multiple superimposed pig sound sources and environmental noise components in the aliased audio, and batch reconstructing multiple independent sound source components. Simultaneously, it incorporates an adaptive verification and filtering mechanism, combining the specific vocal frequency band range, vocal energy threshold, and vocal temporal continuity characteristics of pigs to intelligently identify each sound source component obtained from the reconstruction. It automatically filters out non-pig noise components such as fans, water curtains, and equipment vibrations, eliminating residual noise interference from the reconstruction. Ultimately, it adaptively calculates a single pig sound source with no crosstalk, low distortion, and complete features from the complex multi-source aliased audio, completely achieving accurate separation of overlapping vocal signals from multiple pigs.

[0047] In this embodiment, step 2 further includes combining the pig acoustic identity database to complete a unique match between the vocal signal and the individual pig, specifically including: 2.4 Extracting acoustic features from the single pig sound source; In the specific implementation process, for the separated pure single pig sound source signal, specific acoustic features suitable for individual pig identification are extracted, abandoning general audio feature extraction methods and conforming to the acoustic characteristics of pigs' daily vocalizations and pathological vocalizations. This step mainly extracts multi-dimensional stable acoustic features including Mel frequency cepstral coefficients, spectral centroid, spectral bandwidth, short-time energy, zero-crossing rate, etc., to form a specific voiceprint feature set for each pig. During the extraction process, a feature screening threshold is set simultaneously to remove weak noise features and invalid fluctuation features remaining after sound source reconstruction, retaining core voiceprint features with high recognizability and strong stability, forming a standardized and dimensionally unified acoustic feature vector for the individual to be tested, providing regular feature data for subsequent accurate matching and comparison.

[0048] 2.5 Calculate the Euclidean distance and cosine similarity between the acoustic features and the standard acoustic feature templates of each pig in the pig acoustic identity database. This invention uses a dual-algorithm joint comparison mechanism to replace a single matching algorithm, overcoming the shortcomings of low matching accuracy and poor fault tolerance of a single algorithm. The pig acoustic identity database pre-stores a unique standard acoustic feature template for each pig. The template data consists of the average values ​​of massive voiceprint features collected and calibrated under different behavioral scenarios and in a healthy state, possessing strong individual uniqueness. In specific implementation, the acoustic feature vector to be tested extracted in step 2.4 is compared one by one with the standard feature templates of all pigs in the database: the amplitude difference between the feature to be tested and the standard template features is calculated using Euclidean distance to accurately quantify the magnitude of the deviation at the feature numerical level; the spatial angle between the two sets of feature vectors is calculated using cosine similarity to characterize the overall trend and morphological similarity of the acoustic features, achieving feature matching verification from both numerical differences and morphological trends, comprehensively covering the differentiated features of individual voiceprints, and significantly improving the comprehensiveness and accuracy of the matching.

[0049] 2.6 The feature matching degree of all pig individuals is ranked, and the pig individual with the highest matching degree is selected as the preliminary matching result. In the specific implementation process, the two indicators are normalized and fused by combining the evaluation rules that the smaller the Euclidean distance and the larger the cosine similarity, the higher the matching degree, to calculate the comprehensive matching score of each pig individual. According to the comprehensive matching score from high to low, all pig individuals in the pen that have entered the database are uniformly sorted, and the pig individual with the highest comprehensive matching score is selected as the preliminary matching individual corresponding to the current voice source. This completes the preliminary identity tracing at the voiceprint feature level, locks the candidate matching target, and achieves the effect of quickly selecting the optimal matching object from a massive number of individuals.

[0050] 2.7 Introducing spatial positioning data to verify the preliminary matching results and complete the unique binding of the vocal signal to the individual pig. This invention adds a spatial positioning cross-verification mechanism to achieve dual fault tolerance correction. In specific implementation, the spatial coordinates of the device corresponding to this audio acquisition, the sound source positioning area, and the real-time distribution data of pig pens in the pen are retrieved to obtain the true spatial location of the current sound source. The permanent pen and real-time activity space range of the pig individual obtained from the preliminary matching are compared and verified with the spatial location of the sound source to determine whether the pig individual is within the effective spatial area of ​​the sound source, following the breeding behavior rule that pigs do not randomly vocalize across areas. If the activity space of the initially matched individual matches the sound source positioning area, the matching result is deemed valid; if there are situations such as spatial location mismatch or cross-pen matching that do not conform to the breeding rules, the preliminary result is discarded, and the second-best matched individual is selected for spatial verification again. Through the dual verification mechanism of voiceprint feature matching + spatial location verification, the problems of voiceprint approximate mismatch and algorithm random mismatch are completely avoided, and finally, the unique and accurate binding of the vocal signal to the individual pig is achieved, completing the individual-level sound source tracing.

[0051] In this embodiment, the present invention establishes a pig acoustic identity database in the background, and creates a unique acoustic identity identifier for each pig. This identifier serves as the pig's acoustic ID card, realizing the accurate and automatic binding of sound signals with individual pigs.

[0052] (1) Establish a pig acoustic identity database in the background; The database assigns a unique identifier to each pig and stores information including: ① A unique identification number for each pig, which is consistent with the pig's ear tag and individual identification; ② Information on pen number, stall number, pig age, and growth cycle; ③Basic acoustic characteristics of pigs, specifically including fundamental frequency, formants, energy distribution and vocal tract characteristics; ④ Templates for normal vocalization in pigs and templates for typical coughing and wheezing; ⑤ Historical health status records of pigs.

[0053] The database is initialized when the pigs enter the pen. Through the preliminary audio collection work, a standard acoustic feature template for each pig is established, which serves as the core basis for subsequent individual matching of pigs.

[0054] (2) VAD endpoint detection; First, VAD (Voice Activity Detection) is performed. The energy of a single frame signal is calculated to distinguish between silent segments and effective vocal segments. The energy calculation formula is as follows: ; In the formula: This refers to the number of sampling points per frame. The amplitude at the nth sampling point. Set the noise energy threshold. When the frame energy is greater than the threshold, it is determined to be a valid sound frame, the target sound range is located, and silent segments and weak interference segments are excluded.

[0055] (3) Separation of underdetermined blind sources; Subsequently, an underdetermined blind source separation algorithm based on prior constraints of pig farming behavior was employed to separate the aliased signals within the effective sound emission interval. This algorithm, building upon traditional underdetermined blind source separation based on sparse component analysis, introduces the daily behavioral patterns of pigs as prior constraints. By fusing the spatial positioning information of the microphone array with the spatiotemporal distribution patterns of pigs' feeding, drinking, and resting behaviors, a non-uniform mixing matrix estimation prior is constructed. For example, during feeding periods, sound sources are more likely to be concentrated in the feeding trough area; during rest periods, sound sources are more likely to be distributed in the lying area. Utilizing this prior information on farming behavior to guide the clustering estimation process of the mixing matrix can significantly improve the separation accuracy of target pig sound sources in multi-pig aliasing scenarios, especially suitable for extracting weak signals such as occasional coughing. This algorithm takes advantage of the sparsity of the acoustic signal of pigs in the time-frequency domain, that is, only one sound source usually dominates at the same time-frequency point. It estimates the mixing matrix through a priori-constrained clustering method, and then uses L1 norm minimization to reconstruct each independent sound source component. It can adaptively solve each independent sound source component from the mixed signal and extract the pure vocal signal of a single pig.

[0056] (4) Sound is precisely linked to individual pigs; After extracting acoustic features from the separated single-source signals, the extracted real-time features are matched with the standard acoustic feature templates of each pig stored in the background acoustic identity database for feature similarity matching.

[0057] ① Calculate the Euclidean distance and cosine similarity between the acoustic features of the single-source signal extracted in real time and the standard acoustic feature template of each pig in the database, so as to quantify the degree of feature matching between the two. ② Sort the feature matching degree of all pig individuals and select the pig individual with the highest matching degree as the preliminary matching result of the sound signal; ③ Introduce spatial positioning data, combine the preset placement of the microphone with the specific location information of the pig's pen, verify the above preliminary matching results, eliminate matching errors, and finally complete the unique binding of the sound signal with the individual pig.

[0058] This enables full traceability and location tracking of individual pig vocalizations, vocalization types, and health status, providing accurate data support for subsequent pig health monitoring.

[0059] In this embodiment, the present invention employs an acoustic identity database matching and calibration method, including: (1) Feature extraction; This invention employs a layered feature extraction scheme for pathological vocalizations in pigs, targeting the frequency band characteristics of different pathological vocalizations. For the bound single-source signal, time-frequency domain features are extracted in parallel using Short-Time Fourier Transform (STFT) and Constant Q Transform (CQT). Simultaneously, the fundamental frequency and formants are extracted using cepstral analysis and autocorrelation. Specifically, the following features are extracted: ① Mel spectrum: By mapping the STFT spectrum to a Mel scale filter bank, the auditory perception characteristics of pig ears are simulated, highlighting the frequency band energy distribution of abnormal sounds such as coughing and wheezing.

[0060] ② Constant Q Transform (CQT): It adopts the center frequency of geometric interval, with high frequency resolution in the low frequency band and high time resolution in the high frequency band, and is adapted to the wide frequency characteristics of pig vocalization from low-frequency snorting to high-frequency screaming.

[0061] ③ Fundamental frequency: Reflects the period of vocal cord vibration. The fundamental frequency is stable during normal vocalization, but changes occur when vocalization is stimulated or breathing is abnormal.

[0062] ④ First and second formants: These reflect the shape and degree of expansion and contraction of the vocal tract. The formant frequency and bandwidth shift during coughing and wheezing.

[0063] ⑤ Short-term energy: measures the intensity of vocalization; occasional coughs have lower energy, while continuous coughs accumulate energy.

[0064] ⑥ Zero crossing rate: Reflects the proportion of high-frequency components in the signal; the zero crossing rate of wheezing is significantly increased.

[0065] Building upon the aforementioned general features, this invention emphasizes tiered processing of three specific frequency bands: the low-frequency band (<500Hz) focuses on extracting features reflecting basal metabolism and comfort, such as breathing sounds and snoring; the mid-frequency band (500-2000Hz) focuses on capturing features with sudden and transient characteristics, such as coughing sounds and eating sounds; and the high-frequency band (>2000Hz) focuses on extracting sharp, high-pitched features, such as wheezing sounds and stress sounds. This tiered extraction strategy can more effectively enhance acoustic cues related to specific health problems.

[0066] Simultaneously, the first-order and second-order difference dynamic characteristics of the above static features are calculated. The first-order difference reflects the rate of change, and the second-order difference reflects the acceleration of change. The formula for calculating the difference characteristics is: ; In the formula: This is the original static feature sequence. The difference calculation window is typically set to 2. A linear weighted summation is used, where sampling points further from the current time point within the window have a larger weight (i.e., a larger absolute value of n), which highlights the trend of feature changes over time. The denominator is the sum of squared weights to ensure gain normalization. Compared to simple frame difference calculations, this formula is more sensitive to short-term abrupt changes and outputs near zero for stable noise segments, thus enhancing the distinguishability between pathological and normal vocalizations.

[0067] (2) Missing feature imputation; After feature extraction, feature values ​​for some frames may be missing due to brief signal loss or environmental interference. This invention employs a hierarchical missing feature imputation strategy to ensure the completeness of the feature matrix and its acoustic-physical consistency.

[0068] Random missing values: loss of individual frequency components, local missing values ​​of features in a single frame. The K Nearest Neighbor (KNN) algorithm is used to fill in the missing values. The K nearest neighbors (KNN) algorithm is used to search for the K complete samples that are closest to the missing sample in the feature space in the training sample set. The missing values ​​are filled in by the mean of their features, thus maintaining the continuity of the local structure and the consistency of the distribution in the feature space.

[0069] Spatial field-related deficiencies: Regional feature loss due to microphone placement, sound field attenuation, and spatial obstruction. These are filled using the Poisson equation as a physical constraint, the formula being: ; In the formula: The acoustic field function to be reconstructed represents spatial field quantities such as acoustic energy density and amplitude distribution. The acoustic guiding vector field is obtained by estimating the characteristic gradients of effective measurement points. The formula follows the steady-state physical field evolution law and has a unique smooth solution under given boundary conditions, which can guarantee that the reconstructed field conforms to the acoustic propagation mechanism.

[0070] This invention uses the measured characteristic values ​​of the microphone array measurement points as Dirichlet boundary conditions and numerically solves the full-field acoustic field distribution using the finite difference method to achieve interpolation reconstruction of missing location features. This method strictly adheres to physical constraints such as sound wave propagation energy attenuation and boundary reflection, resulting in more reasonable and reliable filling results compared to purely numerical interpolation methods.

[0071] (3) Data labeling based on acoustic identity database; ① Individual identification: The acoustic features extracted in real time are compared with the feature information in the acoustic identity database to determine the unique identifier of the vocalizing pig; ② Attribute identification: Based on the determined unique identifier of the pig, automatically read the corresponding pen information, stall information, age information and growth cycle information of the pig in the database; ③ Vocalization labeling: Automatically label the vocalization type of the current pig. Vocalization types include normal vocalization, occasional coughing, continuous coughing, wheezing, and stress-induced vocalization. ④ Health labeling: Based on the historical health status records of the pig in the database, its current health status is automatically labeled, and the health status is divided into normal, mildly abnormal, moderately abnormal and severely abnormal. ⑤ Data entry and labeling: Write all the labeled data into the standard dataset and associate it with the unique identifier of the pig to form a structured sample that can be used for model training, data query and full traceability.

[0072] All calibration processes are completed automatically by the system without human intervention, effectively ensuring the consistency and accuracy of calibration data.

[0073] (4) Dataset construction; After data calibration, all calibrated data is stored according to a unified structure, including: ① unique identifier for each pig; ② acoustic feature matrix; ③ individual pig attribute information; ④ vocalization type label; and ⑤ health status label. This structured data is then divided into training and validation sets in an 8:2 ratio for subsequent model training and performance evaluation.

[0074] In this embodiment, step 3 involves extracting the specific frequency band features of pig pathological vocalization from the matched single pig sound source signal to obtain time-frequency domain static features and dynamic differential features; these are then fused to form initial acoustic features, specifically including: 3.1 The single pig sound source signal is divided into low-frequency, mid-frequency, and high-frequency bands. In practice, based on the physiological vocalization mechanism of pigs and the frequency distribution characteristics of various pathological vocalizations, the matched pure single pig sound source signal is finely divided into frequency bands to achieve partitioned analysis of different types of vocalization signals. Combining the statistical patterns of massive acoustic samples from pig farming, the audio spectrum is precisely segmented: the low-frequency band corresponds to the low-frequency pathological vocalizations such as panting, low-pitched breathing, and muffled sounds from mild inflammation; the mid-frequency band corresponds to the core frequency band of coughing, grunting, and normal communication vocalizations, and is the key band for distinguishing between healthy pigs and those with mild respiratory abnormalities; the high-frequency band corresponds to the sudden abnormal vocalizations such as screaming, stress-induced roaring, and severe coughing and wheezing. Through precise three-band division, different pathological types and degrees of abnormality in vocalization signals are partitioned and classified.

[0075] 3.2 First relevant features are extracted in the low-frequency band, second relevant features in the mid-frequency band, and third relevant features in the high-frequency band. The first relevant features extracted in the low-frequency band include core features such as low-frequency energy proportion, low-frequency spectrum stability, and low-frequency waveform duration, which mainly characterize low-frequency weak pathological states such as chronic respiratory distress, mild wheezing, and latent inflammation in pigs. The second relevant features extracted in the mid-frequency band include features such as mid-frequency spectrum peak value, mid-frequency energy concentration, waveform regularity, and fundamental frequency stability, which accurately reflect the differences between pigs' routine breathing and normal vocalization and mild coughing and mild respiratory discomfort, and are the core features for identifying normal mild lesions. The third relevant features extracted in the high-frequency band include features such as high-frequency pulse intensity, high-frequency energy mutation rate, and high-frequency duration frame count, which specifically correspond to sudden abnormal behaviors and severe pathological states in pigs such as sudden stress, severe coughing and wheezing, and painful screaming.

[0076] 3.3 The first, second, and third relevant features are merged to obtain the time-frequency domain static features. Specifically, the three types of exclusive features extracted from the low-frequency, mid-frequency, and high-frequency bands are dimensionally aligned, feature-stitched, and normalized to retain the independent and identifiable core features of each frequency band, constructing a complete time-frequency domain static feature set. This static feature set fully preserves the inherent attributes of a single pig sound source in each frequency band, such as spectral distribution, energy level, waveform morphology, and frequency composition. It can comprehensively characterize the static inherent characteristics of a pig's vocalization state at a given moment, intuitively reflecting the pig's current basic vocal health status.

[0077] 3.4 Calculate the first-order and second-order differences of the time-frequency domain static features to obtain dynamic difference features; the first-order difference reflects the rate of change of the features, and the second-order difference reflects the acceleration of the change of the features; in specific implementation, according to the audio time sequence, the static features at consecutive moments are differentially calculated frame by frame: the numerical change of the static features in adjacent frames is calculated by the first-order difference to accurately characterize the real-time rate of change of the vocalization features of pigs, reflecting the fast and slow fluctuations of breathing and vocalization states; the second-order difference is used to subtract the first-order difference result again to obtain the acceleration of the feature change, which is used to capture the sudden trend, sudden fluctuations and abnormal oscillation characteristics of the vocalization state. Compared with static features, dynamic difference features can effectively identify intermittent, sudden, and short-term pathological abnormal vocalizations.

[0078] 3.5 The time-frequency domain static features are fused with the corresponding dynamic difference features to form initial acoustic features. Specifically, a feature channel splicing and weight balancing fusion method is used to deeply fuse the time-frequency domain static features representing the inherent properties of sound production with the first-order and second-order dynamic difference features representing the temporal evolution of sound production. By setting balanced weights to weaken redundant features and strengthen core differentiated features, the final initial acoustic features possess both static spatial representation capabilities and dynamic temporal characterization capabilities, preserving the detailed static features of pathological sound production in each frequency band while also encompassing dynamic information such as the rate of change and abrupt acceleration of sound production states.

[0079] In this embodiment, the present invention constructs a multi-scale acoustic texture-temporal fusion model, where "acoustic texture" refers to the structured energy distribution pattern related to vocalization type in the time-frequency domain. Analogous to the statistical characteristics of image texture at different scales and directions, the local energy distribution patterns of acoustic signals at different frequency bands and temporal resolutions are extracted through multi-scale wavelet transform. The model takes a fused acoustic feature tensor as input and outputs the probability distribution of vocalization type and health status of pigs. The overall architecture consists of an adaptive time-frequency enhancement module, a multi-scale wavelet acoustic texture encoder, a lightweight temporal correlation gating module, a temporal-frequency domain hetero-weighted dual-modal fusion module, a global average pooling layer, a Dropout layer, and a fully connected classification layer, cascaded sequentially. The adaptive time-frequency enhancement module of this model achieves effective signal enhancement and noise suppression at the feature domain level, forming a progressive signal optimization link. The construction process of the multi-scale acoustic texture-temporal fusion model includes: S1 constructs a multi-scale acoustic texture and temporal dual-modal fusion model network architecture, including an adaptive time-frequency enhancement module, a multi-scale wavelet acoustic texture encoder, a lightweight temporal correlation gating module, a temporal-frequency domain hetero-weighted dual-modal fusion module, and a global average pooling layer, Dropout layer, and fully connected classification layer cascaded sequentially according to functional logic. In specific implementation, the front end uses the adaptive time-frequency enhancement module to achieve noise reduction and detail enhancement of the input acoustic features. The middle layer sequentially connects the multi-scale wavelet acoustic texture encoder, the lightweight temporal correlation gating module, and the temporal-frequency domain hetero-weighted dual-modal fusion module to perform frequency domain texture feature mining and temporal correlation feature modeling in parallel and achieve deep fusion. The back end connects the global average pooling layer, Dropout layer, and fully connected classification layer to complete feature optimization, regularization constraints, and classification judgment. Each module is functionally independent, sequentially linked, and complementary in gain. The overall network structure is lightweight and adaptable to the engineering deployment requirements of online real-time monitoring of pigs.

[0080] The adaptive time-frequency enhancement module described in S2 divides the input complete acoustic features into sub-bands according to the frequency dimension. It distinguishes between effective sound segments and noise segments by calculating the energy entropy of each sub-band, generates an adaptive noise mask, and applies feature gain to the effective sound area and feature attenuation to the periodic noise area to obtain the noise-enhanced feature tensor. Specifically, after the complete acoustic features are input into the module, uniform sub-band segmentation is first performed along the global frequency dimension, fully covering the entire frequency range of low-frequency breathing, mid-frequency coughing, and high-frequency stress in pigs, achieving refined frequency domain partitioning. For each sub-band corresponding to each frame of audio, the sub-band energy entropy index is calculated one by one. Utilizing the high sensitivity of energy entropy to signal disorder, it accurately distinguishes between effective sound and environmental noise: the effective sound signal from pigs has concentrated energy and regular waveform, resulting in a lower energy entropy value; periodic noise from pen fans, water curtains, and equipment operation has dispersed energy and strong disorder, resulting in a higher energy entropy value. By preset an adaptive energy entropy threshold, signal attributes are determined frame by frame and sub-band by sub-band, dynamically generating a dedicated adaptive noise mask. Based on this mask, differentiated feature modulation processing is performed. Feature gain amplification is applied to sub-band regions identified as valid pig vocalizations to enhance the detailed features of early, weak pathological vocalizations and prevent subtle abnormalities from being masked. Feature attenuation suppression is applied to sub-band regions identified as periodic steady-state noise to filter out fixed noise interference. The final output is a high-precision feature tensor with high signal-to-noise ratio, strong feature discriminative power, and sufficient noise reduction, providing clean input for subsequent multi-scale feature extraction.

[0081] The multi-scale wavelet acoustic texture encoder described in S3 employs a multi-scale discrete wavelet convolution group to extract corresponding acoustic texture features from the feature tensor according to the low-frequency band, mid-frequency band, and high-frequency band, and then performs weighted calculations. In specific implementation, the encoder is equipped with a multi-scale discrete wavelet convolution group, configured with wavelet convolution kernels of different granularities to adapt to the pathological vocal texture features of pigs of different types and severity. For the feature tensor after pre-enhancement, segmented feature extraction is carried out strictly according to the partitioning rules of low-frequency band, mid-frequency band, and high-frequency band: for low-frequency band signals, stable low-frequency texture features such as chronic respiratory distress, latent inflammation, and low-pitched panting are extracted; for mid-frequency band signals, core pathological texture features such as common coughing, mild respiratory inflammation, and abnormal snoring in pigs are extracted; for high-frequency band signals, pulsed and abrupt high-frequency texture features such as severe coughing, painful screaming, and sudden stress are accurately captured. After completing the extraction of multi-scale texture features in three frequency bands, the contribution weight of pathological vocalization in each frequency band and the sample recognition degree are combined to carry out multi-dimensional feature weighted fusion calculation, weaken the background texture and noise residue features with no difference, strengthen the core acoustic texture features with pathological discrimination value, and obtain a multi-granular, full-coverage, and highly discriminative multi-scale acoustic texture spatial feature set.

[0082] The lightweight temporal correlation gating module described in S4 introduces a temporal attenuation factor to adaptively and iteratively update historical temporal states. It mines the temporal dependencies of continuous abnormal vocalizations in pigs based on the multi-scale acoustic texture features output by the encoder, while filtering out temporal interference from periodic environmental noise in the pigpen. Specifically, considering the continuous, intermittent, and temporally correlated characteristics of abnormal vocalizations in pigs, an adaptively adjustable temporal attenuation factor is introduced to construct a lightweight temporal gating iterative mechanism. After receiving the multi-scale acoustic texture feature sequence output by the encoder, the module iteratively updates the historical temporal feature states frame by frame according to the audio timeline. The temporal attenuation factor dynamically adjusts the weights of historical features based on time proximity, weakening interference from ineffective features in the distant past and strengthening effective vocalization features in the recent past, thus conforming to the evolution of real-time vocalization states in pigs. Simultaneously, relying on the gating temporal modeling logic, it deeply mines the feature correlations between continuous audio frames, accurately capturing the temporal changes in continuous pathological behaviors such as intermittent coughing and wheezing, persistent abnormal breathing, and prolonged stress in pigs. Furthermore, by combining the characteristics of fixed periodic and fixed temporal fluctuations in pen environmental noise, the differential patterns between noise temporal characteristics and the temporal characteristics of pathological vocalization in pigs are distinguished, and the temporal interference of periodic environmental noise is adaptively filtered out, outputting high-quality temporal modeling features with strong temporal correlation and no noise interference.

[0083] The temporal-frequency domain hetero-weighted dual-modal fusion module described in S5 performs independent attention weight learning on the features after temporal modeling from both the temporal and frequency domain dimensions, and performs hetero-weighted deep fusion of temporal dynamic features and frequency domain acoustic texture features. In specific implementation, this module is the core fusion unit of the model, abandoning the traditional simple splicing fusion method with the same weights, and adopting a dual-branch independent attention learning mechanism to achieve hetero-weighted deep fusion. Temporal attention branches and frequency domain attention branches are constructed separately, and weights are trained and adaptively allocated independently for the dynamic temporal features obtained from temporal modeling and the multi-scale frequency domain acoustic texture features output by the encoder. The fusion weights are configured differently according to the vocal characteristics of pigs in different health states: for early latent and intermittent mild pathological abnormalities, the weight of temporal dynamic features is increased to prioritize capturing subtle temporal fluctuation differences; for overt and severe pathological abnormalities and stress sounds, the weight of frequency domain texture features is increased to strengthen significant texture features such as high-frequency mutations and spectral distortions. By combining the advantages of fine-grained static representation of frequency domain texture features with the advantages of dynamic evolution characterization of temporal features through dual-dimensional heterogeneous weighted fusion, complementary gains of the two types of modal features are achieved, and a dual-modal fusion feature with complete dimensions and strong discriminative ability is constructed.

[0084] The S6 fused features are sequentially processed through a global average pooling layer for feature dimensionality reduction and aggregation, and a Dropout layer to suppress model overfitting. Finally, the output is mapped through a fully connected classification layer to obtain the probability distribution of various vocalization types in pigs and the health status assessment results of normal, mildly abnormal, moderately abnormal, and severely abnormal. In specific implementation, the high-dimensional features after deep fusion of the two modalities are input into the backend normalization and classification network. First, the redundant features are reduced and aggregated through a global average pooling layer to simplify the model parameter size, retain the core discriminative features, reduce the model's computational complexity, and ensure the real-time performance of online monitoring. Then, a Dropout layer randomly deactivates some network neurons to break the fixed correlation of parameters, effectively suppressing the overfitting problem in the model training and inference process, and improving the model's generalization ability in complex and variable farming scenarios. Finally, the optimized and normalized features are input into a fully connected classification layer. Through nonlinear mapping and Softmax probability normalization, the probability distribution of various vocalization types of pigs is accurately output. At the same time, the health level is quantitatively judged based on feature differences. Finally, four categories of refined health status assessment results—normal, mildly abnormal, moderately abnormal, and severely abnormal—are stably output, completing the intelligent, automated, and quantitative assessment of the health status of individual pigs.

[0085] In this embodiment, the adaptive time-frequency enhancement module specifically includes: This module is used to automatically enhance effective sound generation and suppress steady-state mechanical noise in noisy pig farm environments, providing high-purity time-frequency features for subsequent feature extraction.

[0086] First, a short-time Fourier transform is performed on the input single-source audio signal to obtain a two-dimensional time-frequency diagram; then, the frame energy entropy is calculated, and the effective sound segment and noise segment are dynamically distinguished by the entropy value to generate an adaptive noise mask.

[0087] Frame energy entropy calculation formula: ; In the formula: For the single-frame time-frequency diagram Energy percentage of body size This represents the number of sub-bands.

[0088] The lower the energy entropy, the more concentrated the signal, and it is judged as effective sound generation; the higher the energy entropy, the more chaotic the signal, and it is judged as noise. The module applies feature gain to the effective sound generation area and attenuates periodic noise areas such as fans and water curtains, significantly improving the identifiability of weak pathological signals (such as occasional coughs and mild wheezing).

[0089] In this embodiment, the scale wavelet acoustic texture encoder specifically includes: The core feature extraction unit of this model employs a one-dimensional multi-scale discrete wavelet convolution group to extract texture features of pig sounds in different frequency bands in parallel, specifically adapted to the frequency distribution characteristics of different vocalizations such as coughing, breathing, wheezing, and stress screaming. ①Scale 1 (low frequency): Extract steady-state vocal features such as breathing sounds and snoring sounds; ②Scale 2 (Mid-frequency): Extract abrupt vocal features such as coughing sounds and eating sounds; ③Scale 3 (high frequency): Extract sharp vocal features such as wheezing and stress sounds.

[0090] After each scale output, channel energy weighting is performed, and weights are automatically assigned based on the signal-to-noise ratio of each channel, strengthening the pathology-related channels and suppressing the noise channels.

[0091] Channel energy weighting formula: ; In the formula: For the first Channel feature map, For channel weights, These are weighted output features.

[0092] This module is designed to adapt to acoustic monitoring scenarios for pigs and can suppress on-site noise and extract effective acoustic features.

[0093] In this embodiment, the Lightweight Timing Association Gating Module (LTAG) specifically includes: Abnormal vocalizations in pigs (persistent coughing, continuous wheezing) have a clear temporal dependence. This module is specifically designed to capture temporal patterns while suppressing temporal interference from environmental noise.

[0094] The module constructs a lightweight timing correlation gate and introduces a timing decay factor. It adaptively forgets historical states and strengthens the temporal correlation of effective vocal segments.

[0095] Timing state update formula: ; In the formula: This represents the current time sequence state. The texture encoder outputs features. For time-series decay factor, This is the Sigmoid activation function.

[0096] Through the temporal decay mechanism, the module can effectively suppress the temporal interference of periodic noises such as fans and water curtains, and significantly improve the modeling accuracy of abnormal temporal patterns such as continuous coughing and persistent wheezing.

[0097] In this embodiment, the dual-modal cross-attention fusion module specifically includes: Acoustic texture features and temporal dynamic features are adaptively fused across modalities to enhance and complement each other.

[0098] Calculate the bidirectional attention weights for texture→temporal and temporal→texture respectively: ; ; Final fusion output: ; This fusion mechanism is adapted to dual-modal acoustic data of pigs and can effectively highlight key features related to health status while suppressing noise and irrelevant interference.

[0099] In this embodiment, the model input and output specifically include: The model input is a three-dimensional tensor that fuses static acoustic features and dynamic difference features: ; Where: C is the number of feature channels, F is the frequency dimension, and T is the time dimension.

[0100] The model outputs two types of results: ① Probability distribution of vocalization types: normal vocalization, occasional cough, persistent cough, wheezing, and stress-induced vocalization; ② Health status assessment: normal, mildly abnormal, moderately abnormal, severely abnormal.

[0101] In this embodiment, the multi-scale acoustic texture and temporal bimodal fusion model further includes: The training process of the multi-scale acoustic texture and temporal bimodal fusion model uses FocalLoss combined with label smoothing as the loss function to specifically address the sample imbalance problem. The formula is as follows: ; In the formula: γ is the model's predicted probability of the true category, and γ is the focusing coefficient, which is usually set to 2. This can reduce the loss weight of easily classified samples and improve the learning power of difficult-to-classify pathological samples.

[0102] The optimizer uses the AdamW adaptive optimizer with a weight decay factor set to 1e. -4 This effectively alleviates model overfitting; the learning rate scheduling adopts a cosine annealing strategy, with the initial learning rate set to 1e. -3 The model is dynamically adjusted with each training round to ensure more stable convergence; weight smoothing uses the EMA (Exponential Moving Average) method, with the following formula: ; In the formula: This is a smoothing coefficient, usually set to 0.999, which makes the model weight updates smoother and improves generalization performance.

[0103] The FocalLoss loss function automatically reduces the loss weight of a large number of easily classifiable normal vocal samples, focusing on mining the loss features of a small number of difficult-to-classify early mild pathological vocal samples and intermittent abnormal vocal samples. It forces the model to focus on learning the differentiated details of abnormal vocal sounds, effectively solving the problems of insufficient learning of niche abnormal samples and model bias towards normal sample predictions caused by traditional cross-entropy loss. It should be noted that the optimizer and learning rate adopt a refined combined scheduling scheme during training. The AdamW adaptive optimizer is used instead of the traditional Adam optimizer, and a reasonable weight decay coefficient is set. The weight decay mechanism regularizes the model parameters, effectively suppressing parameter redundancy and overfitting, and improving the stability of model parameter updates. Simultaneously, a cosine annealing learning rate scheduling strategy is used, abandoning the fixed learning rate training mode. The learning rate is dynamically adjusted according to the cosine curve pattern during the training cycle. In the early stage of training, the learning rate steadily decreases, quickly converging to the optimal parameter range. In the later stage of training, it undergoes small oscillations and fine-tuning to avoid the model getting trapped in local optima, significantly improving the model's convergence accuracy and training stability. Meanwhile, the EMA exponential moving average method is used to smoothly update the model weights. The model weights of different iterations during training are weighted and averaged to filter out the weight oscillation noise of a single iteration, retain the optimal weight features, and effectively improve the stability of the model's actual inference.

[0104] The training process incorporates batch normalization, Dropout, and gradient pruning for regularization constraints, and employs an early stopping strategy: training stops when the validation set performance fails to improve for a preset number of consecutive rounds, preserving the model's optimal weights. Notably, batch normalization normalizes the mean and variance of each batch of training features, unifying the feature distribution scale, accelerating model convergence, and weakening the coupling between internal network parameters, thus improving training efficiency and stability. The Dropout mechanism randomly deactivates some neurons, breaking fixed dependencies between neurons and suppressing overfitting. Furthermore, a gradient pruning strategy with a preset gradient threshold truncates and scales gradients exceeding the threshold during training, effectively resolving the gradient explosion and vanishing problems that commonly occur during deep bimodal network training. In addition, the present invention adopts an early stopping strategy to control the training termination node, presets the monitoring rounds of the validation set evaluation index, continuously monitors the recognition accuracy and loss value changes of the validation set, and automatically terminates the training process when the validation set performance does not improve for a preset number of rounds, thus avoiding overfitting caused by overtraining. At the same time, the optimal weight parameters of the model during the training process are saved in real time, inferior iterative weights are eliminated, and the optimal model file is retained.

[0105] After model training, performance is evaluated using 5-fold cross-validation. The accuracy calculation formula is as follows: ; The formula for calculating the macro average F1 score is: ; In the formula: For the number of categories, Let be the precision of the i-th class. Let be the recall rate for class i. The macro-average F1 score can fairly reflect the overall performance of the model across all classes, and is especially suitable for imbalanced datasets.

[0106] In practice, the preprocessed full-volume pig acoustic dataset is randomly shuffled and divided into five subsets. One subset is selected as the validation set, and the remaining four subsets are used as the training set. This training and validation process is repeated five times. Key evaluation metrics such as accuracy, precision, recall, and F1 score are recorded for each round. The average of the five validation results is then used as the final evaluation result for the model's overall performance. This approach effectively avoids the randomness and bias inherent in traditional single-dataset partitioning, eliminates the impact of dataset partitioning bias on performance evaluation, and accurately reflects the model's generalization performance and recognition stability under different data distributions. This ensures that the trained model is adaptable to various complex farming scenarios.

[0107] In this embodiment, triggering a tiered early warning based on the health status assessment results specifically includes: triggering a corresponding three-level early warning level for abnormal intervention based on the health status assessment results, including: Mild anomalies trigger a Level 1 alert, increasing ventilation in the pigpens and highlighting the target pigs. Mild anomalies typically correspond to early-stage sub-health conditions in pigs, such as mild respiratory discomfort, intermittent mild stress, or latent mild inflammation, which do not pose an acute risk of transmission. While no emergency treatment is needed, continuous monitoring and environmental optimization are required. Upon triggering the alert, the system generates a Level 1 anomaly record, highlighting the ear tag number and pen location of the target pig. This facilitates quick location of the abnormal individual and tracking of monitoring data by farm staff. Simultaneously, the system activates the pigpens' environmental control equipment, adaptively increasing ventilation in the corresponding pen unit to improve air quality, reduce harmful gas concentrations, and alleviate mild respiratory discomfort in pigs. This environmental optimization intervention prevents further deterioration of the early sub-health condition. The system also continuously retains monitoring data for routine tracking and monitoring.

[0108] A moderate abnormality triggers a Level 2 alert, prompting the isolation of pigs and adjustments to the feeding plan. This means that moderate abnormalities often correspond to persistent coughing, frequent panting, and moderate stress-induced agitation in pigs—symptoms indicating a certain risk of disease and potential for small-scale transmission—requiring manual intervention and adjustments to the farming plan. Upon triggering the alert, the system sends pop-up notifications, log entries, and mobile message alerts, accurately indicating the pen number, individual information, type of abnormal vocalizations, and duration of the abnormality to farmers. This guides staff to promptly isolate and observe the target pigs, preventing the spread of risk due to cross-contamination. Simultaneously, the system integrates with the feeding management module, adaptively adjusting the feed amount, frequency, and water ratio for the corresponding pen based on the pig's abnormal condition. This reduces the burden on the gastrointestinal and respiratory systems, and, combined with manual verification and diagnosis, enables targeted care intervention to prevent moderate abnormalities from worsening into severe pathological symptoms.

[0109] Severe anomalies trigger an emergency alarm, shutting down unit fans and activating disinfection equipment. Severe anomalies often correspond to high-risk, highly infectious emergency pathological states in pigs, such as violent coughing and wheezing, persistent respiratory distress, severe pain stress, and acute respiratory lesions. Immediate emergency response is required to prevent widespread disease transmission and pig mortality. Upon triggering the alarm, the system issues an audible and visual emergency alarm signal, simultaneously pushing high-risk anomaly warning information to alert farm personnel for emergency action. At the same time, it automatically activates the pen environmental equipment, shutting down the corresponding unit's fans to prevent airflow from accelerating the spread of pathogen aerosols and expanding the disease's reach. It immediately activates the pen's spray disinfection equipment and air purification equipment to disinfect the entire unit, blocking pathogen transmission paths, minimizing the risk of group infection, and buying valuable time for emergency diagnosis and treatment of sick pigs, achieving a rapid, closed-loop emergency response to high-risk health risks.

[0110] In one embodiment, online monitoring processes the real-time audio stream frame by frame with a fixed window of 1 second, outputting probability distributions for five vocal types. A health score is calculated based on the vocal type probability; a typical calculation method is as follows: ; In the formula: Assess the health of pigs. , , , Weighting coefficients for various abnormal vocalizations. This represents the probability of occasional coughing. To determine the probability of persistent coughing, The probability of wheezing. The probability of the excited sound.

[0111] The weights can be set based on veterinary experience. When the health score falls below a preset threshold, the system automatically triggers a tiered early warning mechanism. ① Mild abnormality (health score 0.6-0.85): Triggers a level one warning. The system automatically increases the ventilation of the corresponding pen and highlights the individual pig on the monitoring board, prompting the feeder to strengthen observation. ② Moderate abnormality (health score 0.4-0.6): Triggers a level 2 warning. The system recommends isolating the pig and automatically adjusting its feeding plan (such as reducing feed intake and adding electrolytes), while also sending a message to the farm management personnel. ③ Severe abnormality (health score <0.4): Triggers a level 3 warning, the system activates an emergency alarm, shuts down the unit's fan and starts the disinfection spray, and notifies a veterinarian to take emergency measures.

[0112] This tiered early warning and intervention mechanism is based on a gradient matching strategy for the severity of abnormalities in pigs. It achieves a refined, tiered prevention and control model that adjusts the environment for early abnormalities, manages and controls for moderate abnormalities, and responds urgently to severe abnormalities. This completely changes the problems of indiscriminate treatment, excessive intervention, or delayed treatment in traditional farming, and effectively improves the ability to prevent and control diseases in pigs and the level of intelligent management in large-scale farming.

[0113] Example 2: This invention also provides an acoustic-based online monitoring and intelligent health status assessment system for pigs, such as... Figure 2-3 As shown, the intelligent assessment system includes: an audio acquisition and preprocessing module 100, a sound source separation and individual matching module 200, a feature extraction and optimization module 300, and an intelligent assessment and early warning linkage module 400; The audio acquisition and preprocessing module is used to perform dynamic and static dual-mode coordinated full-domain audio signal acquisition in the pigsty and to complete the preprocessing operation of the acquired signals. The sound source separation and individual matching module is used to detect and locate the effective sound segments in the audio. It uses a sound source separation algorithm based on the prior constraints of breeding behavior to extract a single pig sound source from the aliased signal, and combines the pig acoustic identity database to achieve a unique match between the sound signal and the individual pig. The feature extraction and optimization module is used to extract pathological vocalization-specific frequency band layered features from the matched single pig sound source signal, obtain time-frequency domain static features and dynamic differential features, and fuse them into initial acoustic features. Then, the initial acoustic features are optimized to output complete acoustic features. The intelligent assessment and early warning linkage module is used to input complete acoustic features into a multi-scale acoustic texture and temporal dual-modal fusion model, output the probability distribution of pig vocalization types and health status assessment results, and trigger graded early warnings and corresponding linkage operations based on the assessment results.

[0114] The intelligent assessment and early warning linkage module 400 further includes: a model building unit 401 and a model optimization unit 402; The model building unit is used to build a network architecture for a multi-scale acoustic texture and temporal dual-modal fusion model, including an adaptive time-frequency enhancement module, a multi-scale wavelet acoustic texture encoder, a lightweight temporal correlation gating module, a temporal-frequency domain hetero-weighted dual-modal fusion module, and a global average pooling layer, a Dropout layer, and a fully connected classification layer cascaded in sequence according to functional logic. The model optimization unit is used to perform model training, regularization constraints, performance verification, generalization testing, and parameter fine-tuning.

[0115] The intelligent evaluation system also includes: an acoustic identity database module 500; The acoustic identity database module uses the pig entry stage as a unified initialization node. It constructs a unique acoustic identity for each pig through audio collection. The unique acoustic identity is uniquely associated with the unique individual number bound to the pig's ear tag, the pen and stall location information, and the pig's age and growth cycle basic information.

[0116] All content not described in detail in this specification is prior art known to those skilled in the art, and the model parameters of each electrical appliance are not specifically limited; conventional equipment can be used. Electrical control components not mentioned in this technical solution are not shown in the figures because they are prior art, and will not be described further here.

[0117] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An acoustic-based online monitoring and intelligent health status assessment method for pigs, characterized in that, include: The system performs coordinated full-area audio signal acquisition in both dynamic and static modes in pig pens, and preprocesses the acquired audio signals. The effective sound segments are detected and located. A sound source separation algorithm based on prior constraints of breeding behavior is used to extract a single pig sound source from the aliased signal. The sound signal is then uniquely matched with the individual pig by combining the pig acoustic identity database. The matched single pig sound source signal is subjected to pig pathological vocalization-specific frequency band layered feature extraction to obtain time-frequency domain static features and dynamic difference features; The initial acoustic features are then integrated and optimized to obtain complete acoustic features. The complete acoustic features are input into a multi-scale acoustic texture and temporal dual-modal fusion model, which outputs the probability distribution of pig vocalization types and health status assessment results, and triggers graded early warnings based on the health status assessment results.

2. The method for online monitoring and intelligent health status assessment of pigs based on acoustics according to claim 1, characterized in that, The method of extracting a single pig sound source from an aliased signal using a sound source separation algorithm based on prior constraints of farming behavior specifically includes: An underdetermined blind source separation algorithm based on prior constraints of breeding behavior is adopted, which integrates spatial positioning information and the spatiotemporal distribution pattern of pig behavior to construct a non-uniform hybrid matrix estimation prior. The clustering estimation process of the mixture matrix is ​​guided by prior information on breeding behavior. The mixture matrix is ​​estimated by a prior-constrained clustering method, taking into account the time-frequency sparsity of the acoustic signal of pigs. The time-frequency sparsity means that there is usually only one dominant sound source at the same time-frequency point. Based on the estimated mixing matrix, the L1 norm minimization method is used to reconstruct the components of each independent sound source, and the single pig sound source is adaptively calculated from the multi-source aliased signal.

3. The acoustic-based online monitoring and intelligent health status assessment method for pigs according to claim 2, characterized in that, The process of combining the acoustic identity database of pigs to achieve a unique match between vocal signals and individual pigs specifically includes: Acoustic features were extracted from the single pig sound source; Calculate the Euclidean distance and cosine similarity between the acoustic features and each standard acoustic feature template of pigs in the pig acoustic identity database; The feature matching degree of all pig individuals is ranked, and the pig individual with the highest matching degree is selected as the preliminary matching result; Spatial positioning data is introduced to verify the preliminary matching results, thus completing the unique binding of the vocal signal to the individual pig.

4. The acoustic-based online monitoring and intelligent health status assessment method for pigs according to claim 3, characterized in that, The matched single pig sound source signal is subjected to pig pathological vocalization-specific frequency band layered feature extraction to obtain time-frequency domain static features and dynamic difference features; And they are integrated to form the initial acoustic features, specifically including: The sound signal from a single pig source is divided into low-frequency band, mid-frequency band, and high-frequency band. The first correlation feature is extracted in the low-frequency band, the second correlation feature is extracted in the mid-frequency band, and the third correlation feature is extracted in the high-frequency band. By merging the first correlation feature, the second correlation feature, and the third correlation feature, the time-frequency domain static feature is obtained; The first-order and second-order differences of the static time-frequency domain features are calculated to obtain the dynamic difference features; the first-order difference reflects the rate of change of the features, and the second-order difference reflects the acceleration of the change of the features. The time-frequency domain static features are fused with the corresponding dynamic difference features to form the initial acoustic features.

5. The acoustic-based online monitoring and intelligent health status assessment method for pigs according to claim 4, characterized in that, The construction process of the multi-scale acoustic texture and temporal bimodal fusion model includes: A network architecture for a multi-scale acoustic texture and temporal dual-modal fusion model was constructed, including an adaptive time-frequency enhancement module, a multi-scale wavelet acoustic texture encoder, a lightweight temporal correlation gating module, a temporal-frequency domain hetero-weighted dual-modal fusion module, and a global average pooling layer, a Dropout layer, and a fully connected classification layer, which were cascaded in sequence according to their functional logic. The adaptive time-frequency enhancement module divides the complete acoustic features of the input into sub-bands according to the frequency dimension, distinguishes the effective sound segment from the noise segment by calculating the energy entropy of each sub-band, generates an adaptive noise mask, performs feature gain on the effective sound area and feature attenuation on the periodic noise area, and obtains the noise-reduced and enhanced feature tensor. The multi-scale wavelet acoustic texture encoder uses a multi-scale discrete wavelet convolution group to extract corresponding acoustic texture features from the feature tensor according to the low frequency band, mid frequency band and high frequency band and perform weighted calculation. The lightweight temporal correlation gating module introduces a temporal decay factor to adaptively and iteratively update the historical temporal state, mines the temporal dependency of continuous abnormal vocalizations of pigs based on the multi-scale acoustic texture features output by the encoder, and filters out the temporal interference of periodic environmental noise in the pigsty. The time-frequency domain hetero-weighted dual-modal fusion module performs independent attention weight learning on the features after time-series modeling from the time-series dimension and the frequency domain dimension, respectively, and performs hetero-weighted deep weighted fusion of time-series dynamic features and frequency domain acoustic texture features. The fused features are sequentially processed through a global average pooling layer for feature dimensionality reduction and aggregation, and a Dropout layer to suppress model overfitting. Finally, the output is mapped through a fully connected classification layer to obtain the probability distribution of various vocalization types in pigs and the health status assessment results of normal, mildly abnormal, moderately abnormal, and severely abnormal.

6. The acoustic-based online monitoring and intelligent health status assessment method for pigs according to claim 5, characterized in that, The multi-scale acoustic texture and temporal dual-modal fusion model also includes: The training process of the multi-scale acoustic texture and temporal dual-modal fusion model adopts FocalLoss combined with label smoothing as the loss function, uses AdamW adaptive optimizer and sets weight decay coefficient, and is matched with cosine annealing learning rate scheduling strategy. At the same time, the EMA exponential moving average method is used to smoothly update the model weights. The training process incorporates batch normalization, Dropout mechanism and gradient clipping strategy for regularization constraints, and adopts early stopping strategy. When the performance of the validation set does not improve for a preset number of consecutive rounds, training is stopped to retain the optimal weights of the model. After the model training is completed, a 5-fold cross-validation method is used to evaluate its performance.

7. The acoustic-based online monitoring and intelligent health status assessment method for pigs according to claim 6, characterized in that, The triggering of tiered early warnings based on health status assessment results includes: Based on the health status assessment results, corresponding Level 3 early warning levels are triggered for abnormal intervention, including: Mild abnormalities trigger a Level 1 warning, increasing ventilation in the pigpens and highlighting the target pigs. A moderate abnormality triggers a Level 2 warning, prompting the isolation of pigs and adjustment of the feeding program; In case of severe abnormality, an emergency alarm will be triggered, the unit fan will be shut down, and the disinfection equipment will be turned on.

8. An intelligent assessment system for an acoustic-based online monitoring and health status assessment method for pigs according to any one of claims 1-7, characterized in that, The intelligent assessment system includes: an audio acquisition and preprocessing module, a sound source separation and individual matching module, a feature extraction and optimization module, and an intelligent assessment and early warning linkage module. The audio acquisition and preprocessing module is used to perform dynamic and static dual-mode coordinated full-domain audio signal acquisition in the pigsty and to complete the preprocessing operation of the acquired signals. The sound source separation and individual matching module is used to detect and locate the effective sound segments in the audio. It uses a sound source separation algorithm based on the prior constraints of breeding behavior to extract a single pig sound source from the aliased signal, and combines the pig acoustic identity database to achieve a unique match between the sound signal and the individual pig. The feature extraction and optimization module is used to extract pathological vocalization-specific frequency band layered features from the matched single pig sound source signal, obtain time-frequency domain static features and dynamic differential features, and fuse them into initial acoustic features. Then, the initial acoustic features are optimized to output complete acoustic features. The intelligent assessment and early warning linkage module is used to input complete acoustic features into a multi-scale acoustic texture and temporal dual-modal fusion model, output the probability distribution of pig vocalization types and health status assessment results, and trigger graded early warnings and corresponding linkage operations based on the assessment results.

9. The intelligent evaluation system according to claim 8, characterized in that, The intelligent assessment and early warning linkage module also includes: a model building unit and a model optimization unit; The model building unit is used to build a network architecture for a multi-scale acoustic texture and temporal dual-modal fusion model, including an adaptive time-frequency enhancement module, a multi-scale wavelet acoustic texture encoder, a lightweight temporal correlation gating module, a temporal-frequency domain hetero-weighted dual-modal fusion module, and a global average pooling layer, a Dropout layer, and a fully connected classification layer cascaded in sequence according to functional logic. The model optimization unit is used to perform model training, regularization constraints, performance verification, generalization testing, and parameter fine-tuning.

10. The intelligent evaluation system according to claim 9, characterized in that, The intelligent evaluation system also includes: an acoustic identity database module; The acoustic identity database module uses the pig entry stage as a unified initialization node. It constructs a unique acoustic identity for each pig through audio collection. The unique acoustic identity is uniquely associated with the unique individual number bound to the pig's ear tag, the pen and stall location information, and the pig's age and growth cycle basic information.