Dairy cow voiceprint feature real-time monitoring system and method based on anomaly identification

By acquiring acoustic nodes and enhancing physiological characteristics, combined with the inversion of vocal tract physical parameters and comparison of optimal physiological state parameters, the problem of inaccurate identification in dairy cow voiceprint monitoring systems under complex environments has been solved, achieving high interpretability and traceability in early pathological monitoring.

CN122392579APending Publication Date: 2026-07-14河南省种业发展中心

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
河南省种业发展中心
Filing Date
2026-04-29
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing dairy cow voiceprint monitoring systems struggle to distinguish between mechanical noise interference and pathological evolution of the shape of the cow's vocal tract cavity and the physical state of its inner wall in complex environments, leading to inaccurate identification and false or missed reports.

Method used

The system synchronously collects ambient sound pressure signals through acoustic nodes, generates raw audio stream data, performs spatial orientation of sound sources and time-frequency domain feature matching, extracts physiological characteristics to enhance audio, synthesizes simulated waveforms, identifies voiceprint coordination states, retrieves optimal physiological state parameters, compares them with individual historical health benchmarks, and outputs real-time monitoring reports.

Benefits of technology

The system can keenly detect extremely weak acoustic shifts in dairy cows under complex environments, enabling early pathological monitoring with high interpretability and traceability, thus improving the accuracy and reliability of dairy cow health monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122392579A_ABST
    Figure CN122392579A_ABST
Patent Text Reader

Abstract

The application discloses a dairy cow voiceprint feature real-time monitoring system and method based on abnormality identification, relates to the technical field of voiceprint identification, and comprises the following steps: performing time-frequency characteristic enhancement and signal fine reduction processing on an effective dairy cow sound audio segment, and outputting physiological characteristic enhanced audio as an observation waveform; extracting sound channel physical parameters of the physiological characteristic enhanced audio under physiological structure constraints, and synthesizing an analog waveform; identifying cooperative change information between dairy cow sound rhythm and sound channel resonance envelopes, determining a voiceprint cooperative state, and outputting a composite voiceprint feature vector in combination with the sound channel physical parameters; based on waveform distortion components between the observation waveform and the analog waveform, the optimal physiological state parameters corresponding to the composite voiceprint feature vector are inverted, and are compared with individual historical health benchmarks, and a dairy cow voiceprint feature real-time monitoring report is outputted; and the application realizes the conversion of dairy cow intangible calling into high-explainability and traceable early pathological monitoring reports.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voiceprint recognition technology, and in particular to a real-time monitoring system and method for dairy cow voiceprint features based on anomaly recognition. Background Technology

[0002] In recent years, assessing dairy cow health using non-contact acoustic monitoring methods has become an important direction for smart farms. Existing related technologies are mainly based on classical statistical features or deep neural network architectures. They extract Mel-frequency cepstral coefficients (MFCCs) or perform spectral mapping on collected cow mooing sounds to construct classification models from acoustic signals to behavioral modes (such as rumination, estrus, and stress). These methods, relying on automated screening techniques, have improved the efficiency of livestock monitoring to some extent.

[0003] Existing monitoring methods struggle to detect changes in vocal resonance position caused by tracheal mucosal congestion or weakened vocal cord tension in dairy cows during specific phases such as the peripartum period. In real-world dairy farm settings, the low-frequency pulsations of milking equipment and the resonant noise of cooling fans often highly overlap with the fundamental sound wave frequency of cow mooing, severely masking subtle timbre fluctuations carrying early pathological information. Due to the lack of effective tracking of the physical evolution of vocal organs, the system cannot distinguish, even under strong noise interference, whether changes in the vocal waveform originate from mechanical disturbances or from pathological changes in the shape of the cow's vocal tract cavity and the physical state of its inner walls. This lack of identification logic makes the system prone to false alarms or missed alarms in complex dynamic environments, hindering accurate tracing of the causes and early warning of early lesions in dairy cows. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a real-time monitoring method for cow voiceprint features based on anomaly recognition to solve the problem of inaccurate and unstable anomaly recognition in cows under complex environments due to the lack of coupling of physiological mechanisms.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a real-time monitoring method for cow voiceprint features based on anomaly recognition, comprising:

[0008] The system synchronously acquires ambient sound pressure signals based on acoustic nodes to generate raw audio stream data.

[0009] The sound source spatial orientation is performed on the original audio stream data by utilizing the time difference of sound wave arrival, and the vocalization interval is extracted by combining time-frequency domain feature matching to generate effective cow vocalization audio segments.

[0010] For valid cow vocalization audio segments, time-frequency characteristic enhancement and signal refinement restoration processing are performed, and the output is physiological characteristic enhanced audio as the observed waveform;

[0011] Extract physiological characteristics to enhance the vocal tract physical parameters under physiological structural constraints and synthesize simulated waveforms; identify the coordinated change information between the vocal rhythm and vocal tract resonance envelope of dairy cows, determine the voiceprint coordination state, and output a composite voiceprint feature vector in combination with vocal tract physical parameters.

[0012] Based on the waveform distortion components between the observed and simulated waveforms, the optimal physiological state parameters corresponding to the composite voiceprint feature vector are inverted and compared with the individual's historical health benchmark, outputting a real-time monitoring report of dairy cow voiceprint features.

[0013] Preferably, the method for generating raw audio stream data includes:

[0014] By using multiple acoustic acquisition nodes arranged in a grid, the sound pressure analog signal in the environment is sensed synchronously, and the sound pressure analog signal is modulated and mapped into a digital acoustic sample;

[0015] Based on the timestamps, multiple digital acoustic samples are clock-aligned and combined with spatial coordinate information for channel mapping and encapsulation to generate raw audio stream data containing multi-channel spatial orientation information.

[0016] Preferably, the method for spatial orientation of the sound source includes:

[0017] The raw audio stream data is matched between channels to determine the phase delay state of the sound waves arriving at different acoustic acquisition nodes.

[0018] By combining the spatial locations of each acoustic acquisition node, the spatial geometric location association of the sound-emitting target is established using the phase delay state; spatial orientation matching is performed on the spatial geometric location association to obtain the sound-emitting spatial coordinates of the sound-emitting target within the monitoring area.

[0019] Preferably, the method for generating effective cow vocalization audio segments includes:

[0020] Identify the short-term energy evolution trend of the raw audio stream data, and perform signal energy jump logic verification based on the ambient noise floor benchmark to lock the start and end positions of amplitude instability in the raw audio stream data and delineate the active sound range.

[0021] The spectral energy distribution of the active sound range is obtained, and the spectral energy distribution is compared with the inherent formant frequency range of cow vocalization. From the active sound range, the vocalization range of cows that conforms to the acoustic characteristics of cows is selected.

[0022] The vocalization spatial coordinates are used as spatial constraints to map onto the vocalization range of cows. Based on the mapping results, the corresponding signal segments are cut from the original audio stream data to obtain valid cow vocalization audio segments.

[0023] Preferably, the method for enhancing audio as a physiological characteristic of the observed waveform includes:

[0024] High-pass filtering and frame windowing are performed on valid cow vocalization audio segments to obtain short-time audio frames; frequency feature stripping is performed on short-time waveform frames to obtain signal component carriers;

[0025] Based on the ambient noise floor standard, the environmental noise in the signal component carrier is filtered out; the filtered signal component carrier is then waveform synthesized to obtain a clean audio waveform.

[0026] The amplitude of a pure audio waveform is adaptively adjusted by utilizing the peak intensity of the pure audio waveform, and the output is an enhanced audio signal that reflects the physiological characteristics of the observed waveform.

[0027] Preferably, the method for synthesizing analog waveforms includes:

[0028] Physiological characteristics enhanced audio is deconstructed into glottal excitation components and vocal tract modulation components to obtain glottal characteristic parameters of the vocal cord vibration state of dairy cows;

[0029] The spectral envelope of physiological characteristics is extracted to enhance the audio, and multiple resonance parameters of the vocal tract cavity physical structure are determined based on the energy peak distribution in the spectral envelope, thus obtaining the vocal tract physical parameters.

[0030] The basic excitation sequence is reconstructed based on the vibration period determined by the glottal characteristic parameters, and the glottal characteristic parameters are used to modulate the instantaneous excitation intensity of the basic excitation sequence to generate a simulated pulse signal. The driving pulse signal is driven through the resonant path constructed by the resonant parameters to restore the waveform and synthesize a simulated waveform.

[0031] Preferably, the method for outputting the composite voiceprint feature vector includes:

[0032] Extract the state evolution trajectory of vocal tract physical parameters during the vocalization process to serve as the vocal rhythm of cows; analyze the resonance intensity distribution of physiological characteristics at different frequency positions to obtain the vocal tract resonance envelope;

[0033] A cross-dimensional correlation mapping was performed on the vocal prosody of dairy cows and the vocal tract resonance envelope to extract the cooperative variation information between vocal duration and resonance frequency intensity as the voiceprint cooperative state.

[0034] The physical parameters of the vocal tract and the co-state of the voiceprint are fused in vector space and their dimensions are reduced to output a composite voiceprint feature vector.

[0035] Preferably, the method for retrieving the optimal physiological state parameters corresponding to the composite voiceprint feature vector includes:

[0036] Physiologically enhanced audio is phase-aligned and feature-stripped with analog waveforms to obtain waveform distortion components;

[0037] Identify the trend evolution of waveform distortion components relative to the composite voiceprint feature vector, and automatically iteratively optimize the physiological state values ​​in the composite voiceprint feature vector;

[0038] When the waveform distortion component is reduced to the preset waveform matching range, the corrected physiological state value is extracted as the optimal physiological state parameter.

[0039] Preferably, the method for outputting a real-time monitoring report of cow voiceprint features includes:

[0040] The identity of a cow is determined based on valid audio clips of cow vocalizations, and the individual's historical health benchmark corresponding to the cow's identity is retrieved; the optimal physiological state parameters and the multidimensional attribute offset of the composite voiceprint feature vector relative to the individual's historical health benchmark are identified.

[0041] The multidimensional attribute offset is compared with the preset abnormal threshold to determine the abnormality level. Based on the multidimensional attribute offset and the optimal physiological state parameters, the corresponding pathological cause label and treatment suggestions are matched. The abnormality level, pathological cause label and treatment suggestions are integrated to output a real-time monitoring report of dairy cow voiceprint features.

[0042] Secondly, the present invention provides a real-time monitoring system for cow voiceprint features based on anomaly recognition, comprising:

[0043] The acquisition module is used to synchronously acquire ambient sound pressure signals based on acoustic nodes and generate raw audio stream data containing multi-channel spatial orientation information.

[0044] The extraction module is used to spatially orient the sound source of the original audio stream data by utilizing the time difference of sound wave arrival, and to extract the vocalization interval by combining time-frequency domain feature matching to generate effective cow vocalization audio segments.

[0045] The enhancement module is used to perform time-frequency characteristic enhancement and signal refinement on effective cow vocal audio segments, and output physiological characteristic enhanced audio as the observed waveform;

[0046] The modeling module is used to extract the vocal tract physical parameters of enhanced audio under physiological structural constraints and synthesize simulated waveforms; identify the coordinated change information between the vocal rhythm of dairy cows and the vocal tract resonance envelope, determine the voiceprint coordination state, and output a composite voiceprint feature vector in combination with the vocal tract physical parameters.

[0047] The judgment module is used to invert the optimal physiological state parameters corresponding to the composite voiceprint feature vector based on the waveform distortion components between the observed waveform and the simulated waveform, compare them with the individual's historical health benchmark, and output a real-time monitoring report of the cow's voiceprint features.

[0048] The beneficial effects of this invention are as follows: By deconstructing features under physiological structural constraints, the vocal patterns of dairy cows are extracted from chaotic waveform statistics and extracted into a physical semantic space, deeply capturing the spatiotemporal collaborative mapping between the power source and the resonance cavity, thereby constructing a composite feature base with mechanistic robustness. On this basis, a dynamic inversion mechanism based on waveform distortion components is introduced, utilizing the sampling-level deviation between the observed waveform and the simulated waveform to guide feedback gain, and achieving precise reverse tracing from "sound afterimage" to "physiological essence" through iterative correction. The coupling and synergy of these two aspects greatly mitigates the intertwined interference of pasture environmental background and individual heterogeneity, enabling keen insight into extremely weak acoustic shifts during key stages such as the peripartum period of dairy cows, and realizing the transformation of intangible calls into highly interpretable and traceable early pathological monitoring reports. Attached Figure Description

[0049] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a flowchart of the real-time monitoring method for cow voiceprint features based on anomaly recognition in this invention.

[0051] Figure 2 This is a schematic diagram of the real-time monitoring system for cow voiceprint features based on anomaly recognition in this invention.

[0052] Figure 3 This is a flowchart of the synthesized analog waveform in this invention;

[0053] Figure 4 This is a flowchart of the process for outputting composite voiceprint feature vectors in this invention. Detailed Implementation

[0054] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0055] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0056] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0057] Reference Figure 1 , Figure 2 , Figure 3 and Figure 4 This is one embodiment of the present invention, which provides a real-time monitoring method for cow voiceprint features based on anomaly recognition, including the following steps:

[0058] Methods for generating raw audio stream data include:

[0059] By utilizing multiple acoustic acquisition nodes arranged in a grid, the system synchronously senses analog sound pressure signals in the environment and modulates and maps these analog sound pressure signals into digital acoustic samples.

[0060] It should be noted that above the dairy cow activity area or around the fence, multiple acoustic acquisition nodes (such as MEMS digital microphone arrays or piezoelectric pickup devices) are arranged in a grid pattern in the form of an equidistant matrix (e.g., 10-meter spacing). After the grid pattern is completed, the absolute geographical location of each acoustic acquisition node in the three-dimensional space of the monitoring area is recorded as spatial coordinate information.

[0061] Each acoustic acquisition node uses a built-in piezoelectric sensing diaphragm to sense the simulated sound pressure signal in the environment in real time. The sound wave pressure drives the piezoelectric sensing diaphragm to produce mechanical deformation, and the piezoelectric effect, which induces charges on the surface by the relative displacement of the internal charge centers of the material after being compressed, converts mechanical energy into electrical energy. This generates a weak potential difference that fluctuates with the strength of the simulated sound pressure signal and amplifies it (e.g., by 100 times). Thus, the change in the simulated sound pressure signal is converted into a continuous voltage signal that fluctuates with time. A high-precision clock circuit is used to mark a uniform start acquisition time for the continuous voltage signal.

[0062] The instantaneous amplitude of the continuous voltage signal is truncated at a fixed frequency according to an exemplary sampling period of 62.5 microseconds to obtain discrete time point amplitude values;

[0063] Map discrete time point amplitude values ​​to an example On a linearly distributed discrete step level, the voltage threshold range of the amplitude value at each time point is compared, that is, it is determined which quantization level the discrete time point amplitude value falls within, divided by the 0.0001 volt level interval of the example value. The corresponding quantization value is output using binary encoding, and the continuous voltage signal is converted into a binary value sequence to generate a digital acoustic sample carrying timestamp and spatial coordinate information.

[0064] Based on the timestamps, multiple digital acoustic samples are clock-aligned and combined with spatial coordinate information for channel mapping and encapsulation to generate raw audio stream data containing multi-channel spatial orientation information.

[0065] It should be noted that the timestamp of one acoustic acquisition node in the digital acoustic sample is selected as the global reference time axis. The time deviation distance of the remaining acoustic acquisition node sampling points relative to the global reference time axis is calculated. Based on the time deviation distance, a normalized weight coefficient is generated. The amplitude values ​​of adjacent sampling points are proportionally allocated and weighted and summed to reconstruct the amplitude values ​​corresponding to the reference time and complete the clock alignment.

[0066] Spatial coordinate information is extracted from multiple digital acoustic samples, and the three-dimensional geographical location values ​​in the spatial coordinate information are mapped to the physical channel index of the corresponding digital acoustic sample to establish a topological association between the digital acoustic sample and the physical space of the monitoring area.

[0067] According to the arrangement order of the physical channel index, multiple digital acoustic samples carrying spatial coordinate information are spliced ​​in parallel and encapsulated at the frame level, integrating independent and scattered data streams into a unified data body under the same time slice, generating raw audio stream data containing multi-channel spatial orientation information.

[0068] Methods for spatial localization of sound sources include:

[0069] Signal matching between channels is performed on the raw audio stream data to determine the phase delay state of the sound waves arriving at different acoustic acquisition nodes.

[0070] It should be noted that by performing relative displacement sliding on the time axis on the audio signal sequences corresponding to different physical channel indices in the original audio stream data, and multiplying and accumulating the overlapping sampling points of the two audio signal sequences point by point at each movement step, the sum of the inner products is obtained;

[0071] Arrange the sums of the inner products corresponding to different movement step sizes in order to obtain the correlation curve. Find the horizontal axis offset value corresponding to the position with the largest value in the correlation curve and determine the phase delay state of the sound wave reaching different acoustic acquisition nodes.

[0072] By combining the spatial locations of each acoustic acquisition node, the spatial geometric location association of the sound-emitting target is established using the phase delay state; spatial orientation matching is performed on the spatial geometric location association to obtain the sound-emitting spatial coordinates of the sound-emitting target within the monitoring area.

[0073] It should be noted that the spatial coordinate information corresponding to different physical channel indices is extracted, and the phase delay state is multiplied by a preset sound speed value to obtain the sound path difference value. For example, the sound speed value is 340 meters per second. Using the sound path difference value and the corresponding spatial coordinate information, multiple sets of hyperboloid equations are constructed with each acoustic acquisition node as the origin and the distance difference equal to the sound path difference value. This establishes the spatial geometric position association of the sound-emitting target. An example of the expression of multiple sets of hyperboloid equations is as follows:

[0074] ;

[0075] in, , , For the spatial coordinates of the sound source, , , It is the first The spatial coordinates of the acoustic acquisition nodes corresponding to each physical channel index within the three-dimensional space of the monitoring area. , , It is the first The spatial coordinates of the acoustic acquisition nodes corresponding to each physical channel index within the three-dimensional space of the monitoring area. This is the numerical value for the speed of sound. For sound waves to reach the acoustic acquisition node Acoustic acquisition node The phase delay state between;

[0076] Find the spatial coordinate solution that minimizes the sum of squared residuals of all hyperboloid equations. An example of the expression is:

[0077] ;

[0078] in, It is the coordinate value corresponding to the minimum value of the sum of squared residuals. , It is the distance from the sound source to the acoustic acquisition node. Acoustic acquisition node The difference function of geometric distance between them; by solving the optimal estimated value of the intersection point of multiple hyperboloid equations, the spatial coordinates of the sound-emitting target in the monitoring area are obtained.

[0079] Methods for generating valid audio clips of cow vocalizations include:

[0080] Identify the short-term energy evolution trend of the raw audio stream data, and perform signal energy jump logic verification based on the ambient noise baseline to lock the start and end positions of amplitude instability in the raw audio stream data and delineate the active sound range.

[0081] It should be noted that the original audio stream data containing multi-channel spatial orientation information is processed by frame segmentation. The continuous audio is divided into discrete time segments with an overlap rate of 50% according to the example time step of 20 milliseconds. The sum of squares of the values ​​of all audio sampling points in each discrete time segment is calculated to obtain the short-time energy evolution trend of the original audio stream data.

[0082] By continuously sampling the raw audio stream data containing multi-channel spatial orientation information in an empty fenced environment (such as without cows), the arithmetic mean of the short-time energy evolution trend of all frames within an exemplary 10 seconds is calculated to obtain the average environmental noise baseline energy; the energy difference is obtained by subtracting the short-time energy evolution trend of each frame from the average environmental noise baseline energy.

[0083] The signal energy jump logic verification is as follows: by statistically analyzing the energy difference, the mean of the energy difference plus three times the standard deviation of the example value is set as the start threshold; when the energy difference exceeds the start threshold for 5 consecutive example values, it is determined as the starting point; when the energy difference is continuously (10 example values) below the preset stop threshold, it is determined as the ending point, thereby defining the active sound range.

[0084] The spectral energy distribution of the active sound range is obtained, and the spectral energy distribution is compared with the inherent formant frequency range of cow vocalization. From the active sound range, the vocalization range of cows that conforms to the acoustic characteristics of cows is selected.

[0085] Specifically, a Fast Fourier Transform is performed on each discrete time segment within the active sound region to obtain the corresponding complex spectrum sequence; the real and imaginary part values ​​corresponding to each frequency point in the complex spectrum sequence are extracted, and the sum of the squares of the real and imaginary part values ​​is calculated to obtain the energy amplitude corresponding to each frequency point, thus obtaining the spectral energy distribution of the active sound region; the inherent formant frequency range of cow vocalization includes the first formant concentration band with values ​​from 80Hz to 800Hz (example) and the second formant concentration band with values ​​from 800Hz to 3000Hz (example);

[0086] The overlap between the energy extreme points in the spectral energy distribution and the inherent resonant frequency range of cow vocalization is checked. If the proportion of energy extreme points in the spectral energy distribution falling within the inherent resonant frequency range of cow vocalization exceeds 80% of the example value, and the duration of the sound active interval is between 0.5 seconds and 5 seconds of the example value, then the corresponding sound active interval is determined to be a signal that conforms to the acoustic characteristics of cows. From the sound active intervals, the cow vocalization intervals that conform to the acoustic characteristics of cows are selected.

[0087] The vocalization spatial coordinates are used as spatial constraints to map onto the vocalization range of cows. Based on the mapping results, the corresponding signal segments are cut from the original audio stream data to obtain valid cow vocalization audio segments.

[0088] It should be noted that the spatial coordinates of the sound emission are matched with the vocalization intervals of the cows on the timestamp axis; the spatial coordinates of the sound emission are mapped to the set of spatial coordinates of the preset logical area of ​​cow activity in the three-dimensional space of the monitoring area (obtained by manually surveying the boundary values ​​of the spatial coordinates of the cowshed fence, milking parlor, and activity area during the initial deployment of acoustic acquisition nodes). By verifying whether the spatial coordinates of the sound emission are within the spatial range covered by the set of spatial coordinates of the preset logical area of ​​cow activity, for example, identifying that the sound source is located in the area of ​​the No. 3 bedding in the No. 1 cowshed, rather than in mechanical operation areas such as corridors, interference noise caused by background mechanical collisions is eliminated, thus forming the mapping result;

[0089] If the mapping results verify that the vocalization spatial coordinates are within the preset set of spatial coordinates of the cow's activity logical area, then based on the sample index values ​​corresponding to the start and end positions of the cow's vocalization interval, digital acoustic samples of the corresponding time length are accurately extracted from the original audio stream data, and the sound signals under the specific physical channel index corresponding to the vocalization spatial coordinates are retained to obtain valid cow vocalization audio segments.

[0090] The method of enhancing audio by using the physiological characteristics of the observed waveform as the output includes:

[0091] High-pass filtering and frame-by-frame windowing are performed on valid cow vocalization audio segments to obtain short-time audio frames; frequency feature stripping is performed on short-time waveform frames to obtain signal component carriers.

[0092] It should be noted that a high-pass filter is performed on the effective cow vocalization audio segments. For example, a cutoff frequency of 80Hz is set in the example. The effective cow vocalization audio segments are recursively calculated using the difference equation to filter out low-frequency mechanical vibration interference caused by the ventilation equipment in the barn. The filtered effective cow vocalization audio segments are then divided into segments with a frame length of 25 milliseconds, and a Hamming window is applied to each segment to obtain short audio frames with smooth edges.

[0093] Performing a fast Fourier transform on short audio frames yields a complex spectrum sequence containing different frequency components, thereby achieving frequency feature stripping and obtaining signal component carriers that reflect the composition of sound fluctuations.

[0094] The ambient noise in the signal component carrier is filtered based on the ambient noise floor standard; the filtered signal component carrier is then used for waveform synthesis to obtain a clean audio waveform.

[0095] It should be noted that the average energy value of the ambient noise floor reference is retrieved, and the energy amplitude of each frequency component in the signal component carrier is extracted; a subtraction operation is performed, subtracting the average energy value of the ambient noise floor reference at the corresponding frequency position from the energy amplitude of each frequency component. If the subtraction result is less than zero, the amplitude of that frequency component is reset to 0.01 times the original amplitude of the example value, thereby filtering the ambient noise in the signal component carrier according to the ambient noise floor reference; the filtered signal component carrier is then used to perform an inverse fast Fourier transform, converting the frequency domain components back into a discrete amplitude sequence that varies with time. By overlapping and adding adjacent discrete amplitude sequences, the discontinuity caused by the frame segmentation operation is eliminated, waveform synthesis is completed, and a clean audio waveform is obtained.

[0096] The amplitude of a pure audio waveform is adaptively adjusted by utilizing the peak intensity of the pure audio waveform, and the output is an enhanced audio signal that reflects the physiological characteristics of the observed waveform.

[0097] Specifically, the process involves iterating through all sample point values ​​in the pure audio waveform, finding the largest absolute value among the sample point values, and determining it as the peak intensity of the pure audio waveform. The ratio of the maximum boundary value of the target normalized interval (e.g., the example value [-1, 1]) to the peak intensity is calculated to obtain the amplitude gain coefficient. Each sample point value in the pure audio waveform is multiplied by the amplitude gain coefficient to eliminate volume fluctuations caused by differences in the distance between the cow and the acoustic acquisition node, achieving adaptive amplitude adjustment. The output is then used as physiological characteristic enhanced audio of the observed waveform, providing a standardized observation benchmark for subsequent extraction of physiological parameters.

[0098] Existing technologies for voiceprint recognition or health monitoring in dairy cows typically extract only superficial statistical features such as Mel-frequency cepstral coefficients, failing to isolate complex environmental noise interference from barn conditions and lacking the ability to deeply characterize the pathophysiological changes in the vocal organs. This results in insufficient sensitivity for the early diagnosis of respiratory diseases. Especially when dealing with subtle acoustic feature drifts caused by changes in physiological state, traditional methods lack the ability to reconstruct the vocalization mechanism, making it difficult to accurately trace the source of physiological parameters. Therefore, this invention deconstructs the coupling relationship between the vocal source dynamics and the physical structure of the vocal tract, and uses bioacoustic modeling techniques to reconstruct the mechanism of the dairy cow vocalization process. The specific steps are as follows:

[0099] Methods for synthesizing analog waveforms include:

[0100] Physiological characteristics enhanced audio is deconstructed into glottal excitation components and vocal tract modulation components to obtain glottal characteristic parameters of the vocal cord vibration state of dairy cows.

[0101] It should be noted that the glottal excitation component refers to the original vibration generated when the airflow impacts the edge of the vocal cords during a cow's exhalation, which determines the initial power and fundamental frequency of the sound production; the vocal tract modulation component refers to the frequency filtering and resonance effect caused by the change in the shape of the cavities such as the cow's throat, mouth, and nasal cavity when the original vibration passes through them, which determines whether the final sound produced is a deep "moo" or a short panting sound.

[0102] Utilizing the continuity of sound signals within a short period, physiologically enhanced audio is subjected to multiple small, equidistant displacements along the time axis. The example displacement step size is one sampling point interval. The current sampling point value is multiplied point-to-point with the sampling point values ​​from the past (12 points in the example). The results of these multiplications are then summed and averaged to obtain the vocal tract resonance attenuation coefficient, which reflects the effect of sound propagation within the cow's vocal tract cavity on the vibration of the soft tissue of the vocal tract wall and the reflection from the cavity wall. The vocal tract resonance attenuation coefficient represents the physical filtering law of sound waves produced by the cow's vocal tract cavity; that is, to what extent the sound emitted from the cow's mouth and nose at the current moment is accumulated from the aftershocks of sound reverberating in the throat and mouth at previous moments.

[0103] Take the actual sampling point value of each of the 12 past example times, multiply it by its corresponding duct resonance attenuation coefficient, and then sum the 12 multiplication results to output the duct cavity feedback amplitude that should be at the current time.

[0104] Subtract the vocal tract cavity feedback amplitude from the actual sampling point value at the current moment in the physiological characteristic enhanced audio. Through this difference operation, the repetitive oscillation fluctuations caused by the physical structure of the cow's vocal tract in the physiological characteristic enhanced audio are removed from the original signal. The remaining difference part is the residual value sequence representing the direct impact of the cow's exhaled airflow on the edge of the vocal cords, i.e., the glottic excitation component.

[0105] By retrieving the number of sampling points between two adjacent energy maxima in the residual numerical sequence corresponding to the glottal excitation component, multiplying it by the sampling period (e.g., 62.5 microseconds) used when generating the original audio stream data, the fundamental frequency value is obtained. The repetition frequency of the residual pulse per unit time is then counted to obtain the glottal characteristic parameters of the cow's vocal cord vibration state. The glottal characteristic parameters include the fundamental frequency, pulse width, and excitation slope.

[0106] The spectral envelope of physiological characteristics is extracted to enhance the audio, and multiple resonance parameters of the vocal tract cavity physical structure are determined based on the energy peak distribution in the spectral envelope, thus obtaining the vocal tract physical parameters.

[0107] It should be noted that the physiological characteristic enhanced audio is framed with a frame length of 25 milliseconds and a frame shift of 10 milliseconds. After adding a Hamming window, a Discrete Fourier Transform is performed to obtain a complex spectrum sequence. The amplitude squared of the complex spectrum sequence is used to obtain the short-time energy spectrum. After taking the natural logarithm, an Inverse Discrete Fourier Transform is performed to obtain the real cepstral sequence. The values ​​in the range of the lowest cepstral order from 20th to 30th are retained, and a Discrete Fourier Transform is performed again with exponential operation to obtain the spectral envelope.

[0108] Local maxima are searched for frequency-by-frequency on the spectral envelope. The frequency position corresponding to the local maxima is determined as the center frequency of the formant. The frequency interval where the amplitude on both sides of the maximum drops to half the square root of the peak amplitude is determined as the formant bandwidth. The amplitude of the maximum is determined as the formant energy intensity. The three together constitute the resonance parameters. The center frequency ranges of the first four formants are approximately 80 Hz to 800 Hz, 800 Hz to 3000 Hz, 2000 Hz to 4000 Hz, and 3000 Hz to 5000 Hz, respectively. The center frequency of the formant reflects the contraction and expansion of the cross-sectional area of ​​the vocal tract along the length of the vocal tract, and the formant bandwidth reflects the acoustic damping characteristics of the soft tissue of the vocal tract wall. Based on this, the physical parameters of the vocal tract are obtained.

[0109] The basic excitation sequence is reconstructed based on the vibration period determined by the glottal characteristic parameters, and the glottal characteristic parameters are used to modulate the instantaneous excitation intensity of the basic excitation sequence to generate a simulated pulse signal. The driving pulse signal is driven through the resonant path constructed by the resonant parameters to restore the waveform and synthesize a simulated waveform.

[0110] It should be noted that the pitch period value in the glottal characteristic parameters is retrieved. The pitch period is approximately 5 to 20 milliseconds, corresponding to a fundamental frequency of 50 Hz to 200 Hz. An empty array of the same length as the physiological characteristic enhancement audio is created, and discrete impact points with an amplitude of 1 are filled in at pitch period intervals, with zeros filled in the gaps to form the basic excitation sequence.

[0111] The pulse width and excitation slope are adjusted. The pulse width is approximately 30% to 60% of the fundamental period, and the excitation slope reflects the asymmetric ratio of the rising and falling segments. The waveform shape of each discrete impact point is shaped according to the Liljencrants Fant glottal wave segmentation expression to complete the instantaneous excitation intensity modulation and generate an analog pulse signal.

[0112] The analog pulse signal is used as the driving pulse signal and input into the resonant path constructed by the resonant parameters. The resonant path consists of four sets of cascaded second-order resonant equations. The center frequency of the first set of second-order resonant equations is set to the center frequency of the first resonant peak and the bandwidth is set to the bandwidth of the first resonant peak. The center frequency of the second set of second-order resonant equations is set to the center frequency of the second resonant peak and the bandwidth is set to the bandwidth of the second resonant peak. The center frequency of the third set of second-order resonant equations is set to the center frequency of the third resonant peak and the bandwidth is set to the bandwidth of the third resonant peak. The center frequency of the fourth set of second-order resonant equations is set to the center frequency of the fourth resonant peak and the bandwidth is set to the bandwidth of the fourth resonant peak.

[0113] The operational coefficients of each set of second-order resonance equations are determined by calculating the corresponding center frequency and bandwidth values. The calculation method is to substitute the center frequency and bandwidth into the second-order resonance transfer function expression and then solve the difference equation to derive the recursive coefficients. The driving pulse signal is sequentially processed by the cascaded recursive operation of the four sets of second-order resonance equations. Within each set of second-order resonance equations, the current output value and the historical input and output values ​​are weighted and summed according to the difference equation, so that the driving pulse signal generates frequency-selective oscillation and energy gain at the center frequency of each resonance peak. The sampling point sequence output from the end of the four sets of cascaded second-order resonance equations is the analog waveform synthesized after waveform evolution and restoration.

[0114] By employing a dual deconstruction design targeting both the glottis and vocal tract, a technological leap from pure digital signal analysis to physiological and physical simulation has been achieved. Utilizing a cascaded second-order resonant path to morphologically reconstruct the simulated pulse signal not only enables high-fidelity reproduction of the spectral characteristics of real cow vocalization, but more importantly, by decoupling the resonance parameters from the excitation parameters, the system can independently monitor vocal cord fatigue and the degree of vocal tract lesions. This mechanism-based synthesis method significantly improves the robustness of voiceprint features, providing a standardized physical benchmark for subsequent acquisition of optimal physiological state parameters through waveform residual inversion, ensuring the accuracy and scientific rigor of the monitoring report in identifying early health risks in dairy cows.

[0115] Methods for outputting composite voiceprint feature vectors include:

[0116] The evolution trajectory of vocal tract physical parameters during vocalization is extracted as the vocal rhythm of cows; the resonance intensity distribution of physiological characteristics at different frequency positions is analyzed to obtain the vocal tract resonance envelope.

[0117] It should be noted that the physical parameters of the vocal tract, including values ​​reflecting the cross-sectional area distribution and length of the vocal tract, are retrieved and arranged on the timeline according to the order of the timestamps. By observing the fluctuations, jumps, and smooth transitions of these physical structure values ​​over time, the evolution trajectory of the state reflecting the rhythm of muscle contraction and relaxation during vocalization is obtained and used as the vocal rhythm of cows.

[0118] By extracting physiological characteristics to enhance the energy intensity of multiple resonance peaks defined by resonance parameters in the audio, calculating the energy contribution ratio at each frequency position, mapping these energy peaks and their distribution patterns onto the global frequency axis, and outlining the overall contour reflecting the restricted oscillation law of sound waves in a specific physical cavity, the resonance envelope of the vocal tract is obtained.

[0119] A cross-dimensional correlation mapping was performed on the vocal prosody of dairy cows and the vocal tract resonance envelope to extract the cooperative variation information between vocal duration and resonance frequency intensity as the voiceprint cooperative state.

[0120] It should be noted that by placing the vocal rhythm of the cow and the vocal tract resonance envelope in a unified time-frequency coordinate system for overlapping matching, the linear correlation between the changes in vocal duration and the increase or decrease in resonance frequency intensity can be calculated to identify the synchronous fluctuation pattern of the two on the time axis. For example, when the vocal duration reaches the example value of 1 second, it can be analyzed whether the intensity ratio of the first resonance peak and the second resonance peak changes by a specific proportion, and the information on the coordinated change reflecting the mutual constraint relationship between vocal duration and energy distribution can be extracted as the vocal pattern coordination state that can characterize the vocal habits of a specific cow.

[0121] The physical parameters of the vocal tract and the co-state of the voiceprint are fused in vector space and their dimensions are reduced to output a composite voiceprint feature vector.

[0122] It should be noted that the vocal tract physical parameter values ​​and the voiceprint co-state values ​​are concatenated and spliced ​​(for example, 12-dimensional vocal tract physical parameters and 128-dimensional voiceprint co-state values ​​are concatenated sequentially) to form a high-dimensional multi-dimensional numerical set. The values ​​of this multi-dimensional numerical set are obtained over 10 consecutive sampling batches. The average of the sum of squared deviations of each dimension from its corresponding batch average is calculated to obtain the variance of each dimension's values ​​over the 10 sample batches of the example values. Dimensions with variances higher than the mean variance of all dimensions are identified as critical dimensions, while those with variances lower than the mean variance of all dimensions are identified as critical dimensions. Dimensions with 0% are considered redundant. Key dimensions are assigned a high retention weight of 0.9, and redundant dimensions are assigned a low retention weight of 0.1. The high and low retention weights are used to multiply and accumulate the high-dimensional multidimensional numerical set, that is, the numerical value of each dimension is multiplied by its corresponding weight and then summed in groups to merge into a single numerical value with stronger representativeness. This reduces the compression of data with more than 100 dimensions to a low-dimensional numerical sequence with 12 or 24 dimensions, removes background interference and retains the core identity difference components, and outputs a composite voiceprint feature vector.

[0123] Existing technologies for assessing the physiological health of dairy cows often rely on direct threshold comparisons or pattern classifications of extracted acoustic features. This approach frequently overlooks the mapping loss between sound signals and the physical structure of vocal organs, resulting in extracted feature parameters that remain only on the signal surface and fail to accurately reconstruct the pathological evolution of the cow's respiratory tract and vocal cords. Especially in the complex acoustic environment of livestock sheds, single feature extraction is highly susceptible to nonlinear interference, making it difficult to guarantee the accuracy and reliability of physiological assessments. Therefore, this invention constructs a closed-loop feedback mechanism between simulated waveforms and real observed signals, utilizing residual-driven numerical optimization to achieve deep source tracing of physiological parameters. The specific steps are as follows:

[0124] Methods for retrieving the optimal physiological state parameters corresponding to the composite voiceprint feature vector include:

[0125] Physiologically enhanced audio is phase-aligned and feature-stripped with analog waveforms to obtain waveform distortion components.

[0126] Specifically, the process involves retrieving the physiologically enhanced audio and the synthesized analog waveform. By sliding the analog waveform along the time axis and calculating the cross-correlation coefficient between the analog waveform and the physiologically enhanced audio, the offset position corresponding to the maximum value of the cross-correlation coefficient is found, thus achieving phase alignment between the physiologically enhanced audio and the analog waveform. Subsequently, the amplitude value of the physiologically enhanced audio at each moment is subtracted from the amplitude value of the corresponding moment in the aligned analog waveform. This point-by-point subtraction operation removes the simulated feature components from the original signal, thereby obtaining the waveform distortion components composed of unreconstructed noise or nonlinear fluctuations.

[0127] The system identifies the trend evolution of waveform distortion components relative to the composite voiceprint feature vector and automatically iteratively optimizes the physiological state values ​​in the composite voiceprint feature vector.

[0128] Specifically, a correlation matrix is ​​established between the energy magnitude of the waveform distortion component and each numerical dimension in the composite voiceprint feature vector. The physical state values ​​representing the vocal tract's physical structure are fine-tuned one by one in the composite voiceprint feature vector (the vocal tract length value is increased or decreased in steps of 0.1 mm, as shown in the example), and a simulated waveform is regenerated to observe the energy change trend of the waveform distortion component. Using a numerical gradient descent method, the direction of descent of the waveform distortion component as the values ​​in the composite voiceprint feature vector are identified. This allows for automatic iterative optimization of the physiological state values ​​in the composite voiceprint feature vector, enabling the adjusted combination of physiological state values ​​to produce a waveform output that more closely resembles physiologically enhanced audio.

[0129] When the waveform distortion component is reduced to the preset waveform matching range, the corrected physiological state value is extracted as the optimal physiological state parameter.

[0130] Specifically, during the automatic iterative optimization process, the root mean square error (RMSE) of the waveform distortion component is calculated in real time. When the RMS error value continues to decrease and enters the preset waveform matching range (for example, the value is less than 5% of the total energy of the physiological characteristic enhanced audio), it is determined that the simulated waveform and the physiological characteristic enhanced audio have reached physical consistency. At this point, the iterative calculation is stopped, and the values ​​of vocal tract length, cross-sectional area distribution, and glottal feature parameters after multiple corrections are extracted from the composite vocal feature vector as the optimal physiological state parameters that can accurately reflect the current physical state of the cow's vocal organs.

[0131] This iterative optimization process, based on minimizing waveform distortion components, effectively eliminates parameter mismatch between the simulation model and the actual physiological structure. This ensures that the optimal physiological parameters obtained are no longer abstract mathematical indicators, but rather physical values ​​with clear anatomical significance. This design significantly enhances the system's ability to detect subtle pathological changes in dairy cows, such as abnormal glottal closure, vocal tract swelling, or secretion accumulation, providing crucial technical support for achieving non-contact, non-destructive, and highly accurate individual health monitoring of dairy cows.

[0132] Methods for generating real-time monitoring reports of cow voiceprint features include:

[0133] The identity of a cow is determined based on valid audio clips of cow vocalizations, and the individual's historical health benchmark corresponding to the cow's identity is retrieved; the optimal physiological state parameters and the multidimensional attribute offset of the composite voiceprint feature vector relative to the individual's historical health benchmark are identified.

[0134] It should be noted that the composite voiceprint feature vector corresponding to the valid cow vocalization audio segment is compared with the known individual voiceprint templates in the pre-existing database, and the target index with the highest similarity is identified as the cow's identity.

[0135] Retrieve the individual's historical health baseline values ​​(including historical average fundamental frequency, historical formant center frequency, etc.) recorded in a healthy state and associated with the cow's identity; calculate the difference of each dimension attribute by subtracting the currently acquired optimal physiological state parameters and composite voiceprint feature vector from the individual's historical health baseline, thereby identifying the multi-dimensional attribute offset of the optimal physiological state parameters and composite voiceprint feature vector relative to the individual's historical health baseline.

[0136] The multidimensional attribute offset is compared with the preset abnormal threshold to determine the abnormality level. Based on the multidimensional attribute offset and the optimal physiological state parameters, the corresponding pathological cause label and treatment suggestions are matched. The abnormality level, pathological cause label and treatment suggestions are integrated to output a real-time monitoring report of dairy cow voiceprint features.

[0137] It should be noted that the multi-dimensional attribute offset is input into the judgment logic, and by checking the threshold range in which the offset value falls, the abnormality levels of "normal", "warning", and "severe" are divided:

[0138] Normal level: When all values ​​in the multidimensional attribute offset (such as fundamental frequency offset rate, formant frequency drift, etc.) are in the low range of the preset abnormal threshold, for example, the offset ratio is less than or equal to 10%, it is determined that the cow's current physiological state is stable, the voiceprint features are highly consistent with its individual historical health benchmark, and the abnormal level is determined to be "normal".

[0139] Warning level: When at least one key value in the multidimensional attribute offset exceeds the low range but does not reach the high warning line, for example, the offset ratio is between 10% and 20%, it is determined that the cow's physiological state has potential fluctuations, and it may be in a state of fatigue, environmental stress or disease incubation period, and the abnormal level is determined to be "warning".

[0140] Severity level: When the core value of the multidimensional attribute offset exceeds the high warning line, or when multiple offset values ​​show significant jumps at the same time, such as when the offset ratio example value exceeds 20%, it is determined that the vocal organs of the cow have obvious pathological features or functional disorders, and the abnormality level is determined to be "severe".

[0141] Based on the specific numerical change characteristics in the multidimensional attribute offset (e.g., abnormal increase in fundamental frequency and specific shift in vocal tract resonance parameters), the corresponding pathological cause tags (e.g., matching "suspected upper respiratory tract inflammation") and corresponding treatment suggestions (e.g., "veterinary intervention recommended" or "isolation and observation") are retrieved from the preset pathological cause tag library; the determined abnormality level, pathological cause tags and treatment suggestions are automatically filled and summarized according to the preset text format, and a real-time monitoring report of dairy cow voiceprint characteristics is output.

[0142] It should be noted that the preset process for the pathological cause database is as follows: collect a large number of clinically diagnosed dairy cow pathological voiceprint samples, extract the feature deviation pattern of the composite voiceprint feature vector corresponding to each type of disease relative to the health benchmark, and establish a mapping index between the feature deviation pattern and the pathological cause label and expert treatment suggestions.

[0143] This embodiment also provides a real-time monitoring system for cow voiceprint features based on anomaly recognition, including:

[0144] The acquisition module is used to synchronously acquire ambient sound pressure signals based on acoustic nodes and generate raw audio stream data containing multi-channel spatial orientation information.

[0145] The extraction module is used to spatially orient the sound source of the original audio stream data by utilizing the time difference of sound wave arrival, and to extract the vocalization interval by combining time-frequency domain feature matching to generate effective cow vocalization audio segments.

[0146] The enhancement module is used to perform time-frequency characteristic enhancement and signal refinement on effective cow vocal audio segments, and output physiological characteristic enhanced audio as the observed waveform;

[0147] The modeling module is used to extract the vocal tract physical parameters of enhanced audio under physiological structural constraints and synthesize simulated waveforms; identify the coordinated change information between the vocal rhythm of dairy cows and the vocal tract resonance envelope, determine the voiceprint coordination state, and output a composite voiceprint feature vector in combination with the vocal tract physical parameters.

[0148] The judgment module is used to invert the optimal physiological state parameters corresponding to the composite voiceprint feature vector based on the waveform distortion components between the observed waveform and the simulated waveform, compare them with the individual's historical health benchmark, and output a real-time monitoring report of the cow's voiceprint features.

[0149] This embodiment also provides a computer device applicable to the real-time monitoring method of cow voiceprint features based on anomaly recognition, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the real-time monitoring method of cow voiceprint features based on anomaly recognition as proposed in the above embodiment.

[0150] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0151] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the real-time monitoring method for cow voiceprint features based on anomaly recognition as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0152] In summary, this invention, through feature deconstruction under physiological constraints, extracts and degrades cow vocalizations from chaotic waveform statistics into a physical semantic space, deeply capturing the spatiotemporal synergistic mapping between the power source and the resonant cavity, thereby constructing a composite feature base with mechanistic robustness. Based on this, a dynamic inversion mechanism based on waveform distortion components is introduced, utilizing the sampling-level deviation between the observed and simulated waveforms to guide feedback gain, achieving precise reverse tracing from "sound afterimages" to "physiological essence" through iterative correction. The coupled and synergistic effect of these two mechanisms greatly mitigates the intertwined interference of pasture environmental background and individual heterogeneity, enabling keen detection of extremely subtle acoustic shifts during critical stages such as the peripartum period of dairy cows, and transforming intangible vocalizations into highly interpretable and traceable early pathological monitoring reports.

[0153] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A real-time monitoring method for cow voiceprint features based on anomaly recognition, characterized in that, include: The system synchronously acquires ambient sound pressure signals based on acoustic nodes to generate raw audio stream data. The sound source spatial orientation is performed on the original audio stream data by utilizing the time difference of sound wave arrival, and the vocalization interval is extracted by combining time-frequency domain feature matching to generate effective cow vocalization audio segments. For valid cow vocalization audio segments, time-frequency characteristic enhancement and signal refinement restoration processing are performed, and the output is physiological characteristic enhanced audio as the observed waveform; Extract physiological characteristics to enhance the vocal tract physical parameters of audio under physiological structural constraints, and synthesize simulated waveforms; identify the coordinated change information between the vocal rhythm of dairy cows and the vocal tract resonance envelope, determine the voiceprint coordination state, and output a composite voiceprint feature vector in combination with the vocal tract physical parameters. Based on the waveform distortion components between the observed and simulated waveforms, the optimal physiological state parameters corresponding to the composite voiceprint feature vector are inverted and compared with the individual's historical health benchmark, outputting a real-time monitoring report of dairy cow voiceprint features.

2. The real-time monitoring method for cow voiceprint features based on anomaly recognition as described in claim 1, characterized in that, The method for generating raw audio stream data includes: By using multiple acoustic acquisition nodes arranged in a grid, the sound pressure analog signal in the environment is sensed synchronously, and the sound pressure analog signal is modulated and mapped into a digital acoustic sample; Based on the timestamps, multiple digital acoustic samples are clock-aligned and combined with spatial coordinate information for channel mapping and encapsulation to generate raw audio stream data containing multi-channel spatial orientation information.

3. The real-time monitoring method for cow voiceprint features based on anomaly recognition as described in claim 2, characterized in that, The method for spatial orientation of a sound source includes: The raw audio stream data is matched between channels to determine the phase delay state of the sound waves arriving at different acoustic acquisition nodes. By combining the spatial locations of each acoustic acquisition node, the spatial geometric location association of the sound-emitting target is established using the phase delay state; spatial orientation matching is performed on the spatial geometric location association to obtain the sound-emitting spatial coordinates of the sound-emitting target within the monitoring area.

4. The real-time monitoring method for cow voiceprint features based on anomaly recognition as described in claim 3, characterized in that, The method for generating valid cow vocalization audio clips includes: Identify the short-term energy evolution trend of the raw audio stream data, and perform signal energy jump logic verification based on the ambient noise floor benchmark to lock the start and end positions of amplitude instability in the raw audio stream data and delineate the active sound range. The spectral energy distribution of the active sound range is obtained, and the spectral energy distribution is compared with the inherent formant frequency range of cow vocalization. From the active sound range, the vocalization range of cows that conforms to the acoustic characteristics of cows is selected. The vocalization spatial coordinates are used as spatial constraints to map onto the vocalization range of cows. Based on the mapping results, the corresponding signal segments are cut from the original audio stream data to obtain valid cow vocalization audio segments.

5. The real-time monitoring method for cow voiceprint features based on anomaly recognition as described in claim 4, characterized in that, The method of enhancing audio by using the physiological characteristics of the observed waveform as the output includes: High-pass filtering and frame windowing are performed on valid cow vocalization audio segments to obtain short-time audio frames; frequency feature stripping is performed on short-time waveform frames to obtain signal component carriers; Based on the ambient noise floor standard, the environmental noise in the signal component carrier is filtered out; the filtered signal component carrier is then waveform synthesized to obtain a clean audio waveform. The amplitude of a pure audio waveform is adaptively adjusted by utilizing the peak intensity of the pure audio waveform, and the output is an enhanced audio signal that reflects the physiological characteristics of the observed waveform.

6. The real-time monitoring method for cow voiceprint features based on anomaly recognition as described in claim 5, characterized in that, The method for synthesizing analog waveforms includes: Physiological characteristics enhanced audio is deconstructed into glottal excitation components and vocal tract modulation components to obtain glottal characteristic parameters of the vocal cord vibration state of dairy cows; The spectral envelope of physiological characteristics is extracted to enhance the audio, and multiple resonance parameters of the vocal tract cavity physical structure are determined based on the energy peak distribution in the spectral envelope, thus obtaining the vocal tract physical parameters. The basic excitation sequence is reconstructed based on the vibration period determined by the glottal characteristic parameters, and the glottal characteristic parameters are used to modulate the instantaneous excitation intensity of the basic excitation sequence to generate a simulated pulse signal. The driving pulse signal is driven through the resonant path constructed by the resonant parameters to restore the waveform and synthesize a simulated waveform.

7. The real-time monitoring method for cow voiceprint features based on anomaly recognition as described in claim 6, characterized in that, The method for outputting the composite voiceprint feature vector includes: Extract the state evolution trajectory of vocal tract physical parameters during the vocalization process to serve as the vocal rhythm of cows; analyze the resonance intensity distribution of physiological characteristics at different frequency positions to obtain the vocal tract resonance envelope; A cross-dimensional correlation mapping was performed on the vocal prosody of dairy cows and the vocal tract resonance envelope to extract the cooperative variation information between vocal duration and resonance frequency intensity as the voiceprint cooperative state. The physical parameters of the vocal tract and the co-state of the voiceprint are fused in vector space and their dimensions are reduced to output a composite voiceprint feature vector.

8. The real-time monitoring method for cow voiceprint features based on anomaly recognition as described in claim 7, characterized in that, The method for retrieving the optimal physiological state parameters corresponding to the composite voiceprint feature vector includes: Physiologically enhanced audio is phase-aligned and feature-stripped with analog waveforms to obtain waveform distortion components; Identify the trend evolution of waveform distortion components relative to the composite voiceprint feature vector, and automatically iteratively optimize the physiological state values ​​in the composite voiceprint feature vector; When the waveform distortion component is reduced to the preset waveform matching range, the corrected physiological state value is extracted as the optimal physiological state parameter.

9. The real-time monitoring method for cow voiceprint features based on anomaly recognition as described in claim 8, characterized in that, The method for outputting a real-time monitoring report of cow voiceprint features includes: The identity of a cow is determined based on valid audio clips of cow vocalizations, and the individual's historical health benchmark corresponding to the cow's identity is retrieved; the optimal physiological state parameters and the multidimensional attribute offset of the composite voiceprint feature vector relative to the individual's historical health benchmark are identified. The multidimensional attribute offset is compared with the preset abnormal threshold to determine the abnormality level. Based on the multidimensional attribute offset and the optimal physiological state parameters, the corresponding pathological cause label and treatment suggestions are matched. The abnormality level, pathological cause label and treatment suggestions are integrated to output a real-time monitoring report of dairy cow voiceprint features.

10. A real-time monitoring system for cow voiceprint features based on anomaly recognition, based on the real-time monitoring method for cow voiceprint features based on anomaly recognition according to any one of claims 1-9, characterized in that, include: The acquisition module is used to synchronously acquire ambient sound pressure signals based on acoustic nodes and generate raw audio stream data containing multi-channel spatial orientation information. The extraction module is used to spatially orient the sound source of the original audio stream data by utilizing the time difference of sound wave arrival, and to extract the vocalization interval by combining time-frequency domain feature matching to generate effective cow vocalization audio segments. The enhancement module is used to perform time-frequency characteristic enhancement and signal refinement on effective cow vocal audio segments, and output physiological characteristic enhanced audio as the observed waveform; The modeling module is used to extract the vocal tract physical parameters of enhanced audio under physiological structural constraints and synthesize simulated waveforms; identify the coordinated change information between the vocal rhythm of dairy cows and the vocal tract resonance envelope, determine the voiceprint coordination state, and output a composite voiceprint feature vector in combination with the vocal tract physical parameters. The judgment module is used to invert the optimal physiological state parameters corresponding to the composite voiceprint feature vector based on the waveform distortion components between the observed waveform and the simulated waveform, compare them with the individual's historical health benchmark, and output a real-time monitoring report of the cow's voiceprint features.