Ewe oestrus detection method and device based on multi-modal feature fusion

Through the multimodal feature fusion method, wearable devices are used to collect ewe behavior and audio data, and combined with artificial intelligence models, the problems of low accuracy and efficiency in ewe estrus detection are solved, real-time and accurate estrus detection is achieved, and the stress response of traditional methods is avoided.

CN120823644APending Publication Date: 2025-10-21BEIJING RES CENT FOR INFORMATION TECH & AGRI

Patent Information

Application Number
CN202510775666.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing ewe estrus detection methods have problems with low accuracy and efficiency, especially in group-rearing conditions, where target tracking is difficult to be accurate and continuous. In addition, serum progesterone detection is labor-intensive and causes stress reactions in ewes.

Method used

A multimodal feature fusion method is used to collect behavioral data and audio data of ewes, and wearable devices are used to obtain inertial speed and audio signals. Behavior recognition and time series change feature extraction are performed, and combined with artificial intelligence models to achieve real-time and accurate estrus detection.

Benefits of technology

The accuracy and efficiency of ewe estrus detection are improved, the stress response of traditional methods is avoided, and a convenient and safe data collection method is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823644A_ABST
    Figure CN120823644A_ABST
Patent Text Reader

Abstract

The invention provides an ewe oestrus detection method and device based on multi-modal feature fusion, and relates to the technical field of livestock behavior monitoring, the method comprises the following steps: acquiring behavior data and audio data of an ewe to obtain a first behavior data set and a first audio data set; inputting the first behavior data set into an ewe behavior recognition model to obtain a second behavior data set, and inputting the first audio data set into an ewe sound detection model to obtain a second audio data set; performing feature extraction on the second behavior data set and the second audio data set to obtain a multi-modal feature fusion data set; and inputting the multi-modal feature fusion data set into an ewe oestrus discrimination model to obtain an ewe oestrus detection result. According to the ewe oestrus detection method based on multi-modal feature fusion, on the basis of fully learning the multi-modal features of the ewe, whether the ewe is oestrus or not is intelligently judged through the ewe oestrus judgment model, and the accuracy and efficiency of ewe oestrus detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of livestock behavior monitoring, and in particular to a method and device for detecting ewe estrus based on multimodal feature fusion. Background Art

[0002] Ewe estrus monitoring plays a key role in ensuring efficient reproduction of ewes and improving flock production performance.

[0003] In the prior art, there are two main methods for monitoring estrus in ewes: one is a technology based on machine vision tracking of motion trajectories, and the other is a method based on electrochemiluminescence detection of progesterone levels in ewes to determine estrus.

[0004] However, in practice, methods based on machine vision tracking require continuous monitoring of walking distance and trajectory to detect estrus. Due to the high similarity in appearance of ewes, target detection and tracking in group housing can easily lead to loss or confusion, making accurate and continuous tracking difficult. Methods based on electrochemiluminescence detection of progesterone levels in ewes require blood sampling, which is labor-intensive, time-consuming, and can cause significant stress for the ewes. Consequently, existing technologies for detecting estrus in ewes are inaccurate and inefficient. Summary of the Invention

[0005] The present invention provides a method and device for detecting ewe estrus based on multimodal feature fusion, which are used to solve the technical problem of low accuracy and efficiency of ewe estrus detection in the prior art.

[0006] The present invention provides a method for detecting ewe estrus based on multimodal feature fusion, comprising the following steps:

[0007] Collecting behavioral data and audio data of the ewe within a preset time period; wherein the behavioral data includes the ewe's movement speed, movement time, and magnetic field strength at the ewe's location;

[0008] Based on the behavioral data, obtaining a first behavioral data set;

[0009] Based on the audio data, obtaining a first audio data set;

[0010] Inputting the first behavior data set into a ewe behavior recognition model to obtain a second behavior data set output by the ewe behavior recognition model;

[0011] Inputting the first audio data set into a ewe sound detection model to obtain a second audio data set output by the ewe sound detection model;

[0012] Performing feature extraction on the second behavior dataset and the second audio dataset respectively to obtain a multimodal feature fusion dataset;

[0013] The multimodal feature fusion data set is input into the ewe estrus discrimination model to obtain the ewe estrus detection result output by the ewe estrus discrimination model.

[0014] According to a method for detecting estrus in ewe based on multimodal feature fusion provided by the present invention, obtaining a first behavioral data set based on the behavioral data includes:

[0015] Calculating pre-selected multi-dimensional characteristic parameters based on the behavioral data; the multi-dimensional characteristic parameters include the activity intensity and activity range of the ewe;

[0016] Based on the multidimensional feature parameters, the first behavior data set is obtained.

[0017] According to a method for detecting estrus in ewe based on multimodal feature fusion provided by the present invention, obtaining a first audio data set based on the audio data includes:

[0018] Calculating the short-time energy and zero-crossing rate of the audio signal in the audio data;

[0019] The first audio data set is obtained based on the short-time energy and the zero-crossing rate.

[0020] According to a method for detecting ewe estrus based on multimodal feature fusion provided by the present invention, the training step of the ewe behavior recognition model includes:

[0021] Collect historical behavioral data of ewes;

[0022] Calculating pre-selected multi-dimensional feature parameters based on the historical behavior data;

[0023] The ewe behavior recognition model is trained using the multidimensional feature parameters as feature data and the corresponding historical behavior categories as label data.

[0024] According to a method for detecting ewe estrus based on multimodal feature fusion provided by the present invention, the training step of the ewe call detection model includes:

[0025] Collect historical audio data of ewes;

[0026] Calculating the short-time energy and zero-crossing rate of the audio signal in the historical audio data;

[0027] The ewe call detection model is trained using the short-time energy and zero-crossing rate of the audio signal in the historical audio data as feature data and the category of the corresponding historical audio as label data.

[0028] According to a method for detecting estrus in ewe based on multimodal feature fusion provided by the present invention, feature extraction is performed on the second behavior dataset and the second audio dataset respectively to obtain a multimodal feature fusion dataset, including:

[0029] Calculate the proportion of the duration corresponding to each behavior of the ewe in the second behavior data set to the duration of all behaviors;

[0030] Calculating the short-time energy and zero-crossing rate of the ewe sounds in the second audio data set;

[0031] Determining a temporal variation characteristic of the behavior based on a proportion of the duration corresponding to each behavior to the duration of all behaviors;

[0032] Determining the temporal variation characteristics of the call based on the short-time energy and the zero-crossing rate;

[0033] Performing feature fusion on the behavior temporal change feature and the call temporal change feature according to the current time sequence to obtain a multimodal fusion feature;

[0034] Based on the multimodal fusion features, the multimodal feature fusion dataset is obtained.

[0035] According to a method for detecting ewe estrus based on multimodal feature fusion provided by the present invention, the training step of the ewe estrus discrimination model includes:

[0036] Collecting the duration of various behaviors of the ewe within a historical time length and corresponding audio data; the historical time length includes the estrus period and the non-estrus period;

[0037] Determining the temporal variation characteristics of the behavior of the ewe over a historical time period based on the proportion of the duration of each behavior of the ewe in the duration of all behaviors;

[0038] Determining a temporal variation characteristic of the ewe's calls within a historical time length based on the short-time energy and zero-crossing rate of the ewe's calls in the audio data;

[0039] Performing feature fusion on the behavior temporal change features and the call temporal change features according to the time sequence within the historical time length to obtain historical multimodal feature fusion data;

[0040] The ewe estrus discrimination model is trained using the historical multimodal feature fusion data as feature data and the corresponding estrus and non-estrus categories as label data.

[0041] The present invention also provides a device for detecting ewe estrus based on multimodal feature fusion, comprising the following modules:

[0042] A collection module is used to collect behavioral data and audio data of the ewe within a preset time length; wherein the behavioral data includes the ewe's movement speed, movement time, and magnetic field strength at the ewe's location;

[0043] A first obtaining module, configured to obtain a first behavior data set based on the behavior data;

[0044] A second obtaining module, configured to obtain a first audio data set based on the audio data;

[0045] a first recognition module, configured to input the first behavior data set into a ewe behavior recognition model to obtain a second behavior data set output by the ewe behavior recognition model;

[0046] a second recognition module, configured to input the first audio data set into a ewe sound detection model to obtain a second audio data set output by the ewe sound detection model;

[0047] a feature fusion module, configured to extract features from the second behavior dataset and the second audio dataset to obtain a multimodal feature fusion dataset;

[0048] The detection module is used to input the multimodal feature fusion data set into the ewe estrus discrimination model to obtain the ewe estrus detection result output by the ewe estrus discrimination model.

[0049] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements any of the above-mentioned methods for detecting ewe estrus based on multimodal feature fusion.

[0050] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for detecting estrus in ewe based on multimodal feature fusion as described above is implemented.

[0051] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-described methods for detecting ewe estrus based on multimodal feature fusion.

[0052] The present invention provides a method for detecting estrus of an ewe based on multimodal feature fusion. The method collects behavioral data and audio data of an ewe within a preset time length to obtain a first behavioral dataset and a first audio dataset, thereby achieving accurate, efficient and continuous tracking and collection of the ewe and obtaining sufficient data. The first behavioral dataset is then input into an ewe behavior recognition model to obtain a second behavioral dataset output by the ewe behavior recognition model, and the first audio dataset is input into an ewe call detection model to obtain a second audio dataset output by the ewe call detection model. Feature extraction is performed on the second behavioral dataset and the second audio dataset respectively to obtain a multimodal feature fusion dataset, thereby fully characterizing the changing patterns of the ewe's behavior and calls through the multimodal feature fusion dataset. Finally, the multimodal feature fusion dataset is input into an ewe estrus discrimination model to obtain an ewe estrus detection result output by the ewe estrus discrimination model. Based on the full learning of the ewe's multimodal features, the ewe estrus discrimination model can be used to intelligently discriminate whether the ewe is in estrus, thereby improving the accuracy and efficiency of ewe estrus detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0054] Figure 1 It is a flow chart of a method for detecting ewe estrus based on multimodal feature fusion provided by the present invention.

[0055] Figure 2 It is a structural schematic diagram of a device for detecting ewe estrus based on multimodal feature fusion provided by the present invention.

[0056] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0057] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0058] Ewe estrus monitoring plays a key role in ensuring efficient breeding and improving flock performance. Real-time, accurate, and automated monitoring of ewe estrus not only helps farmers schedule breedings, improve conception rates, and shorten the reproductive cycle, but also increases the number of lambs born each year, ultimately improving overall flock productivity.

[0059] Traditional methods often rely on manual observation and ram testing, which is not only time-consuming and labor-intensive, but also easily affected by subjective factors, leading to inaccurate judgments.

[0060] With the rapid development of information technology, the livestock industry is gradually transforming from traditional manual management to intelligent and precise management. The introduction of modern technologies such as sensors, the Internet of Things, and big data analysis has greatly improved production efficiency and management levels in the livestock industry.

[0061] In the prior art, there are two main methods for automatically detecting estrus in ewes: one is a technology based on machine vision tracking of motion trajectories, and the other is a method based on electrochemiluminescence detection of progesterone levels in ewes to determine estrus.

[0062] The drawbacks of machine vision-based tracking of motion trajectories are as follows: First, due to the limited field of view of a single camera, it is only suitable for small herds in confined conditions and is not suitable for outdoor grazing or large-scale free-range farming. Second, in actual production, estrus detection requires continuous monitoring of walking distance and trajectory. However, ewes are highly similar in appearance, making target detection and tracking in group housing prone to loss or confusion, making accurate and continuous tracking difficult. Methods based on serum progesterone testing require blood sampling, which is labor-intensive and can cause severe stress for the ewes.

[0063] In response to the problems of low efficiency and strong subjectivity in traditional manual monitoring methods of ewe estrus, as well as the defects of existing image analysis methods and serum progesterone detection methods, the present invention proposes a method for detecting ewe estrus based on multimodal feature fusion. A wearable device worn on the neck is used to collect the ewe's inertial velocity and audio signals. Behavior recognition and temporal change feature derivation are performed based on the inertial information, and the ewe's audio signal features are simultaneously extracted, thereby fully extracting the behavioral change characteristics and sound change characteristics of the ewe during estrus. Through multimodal data fusion, the changes in the ewe's physiological state are more comprehensively reflected. Combined with an artificial intelligence model, real-time and accurate ewe estrus detection is achieved.

[0064] The following combination Figures 1 to 3 The present invention describes a method and device for detecting ewe estrus based on multimodal feature fusion.

[0065] Figure 1This is a flow chart of a method for detecting ewe estrus based on multimodal feature fusion provided by the present invention. Figure 1 As shown, the method includes the following steps:

[0066] Step 101: Collect behavioral data and audio data of the ewe within a preset time length; wherein the behavioral data includes the ewe's movement speed, movement time, and magnetic field strength at the ewe's location.

[0067] Specifically, the wearable device designed in the embodiment of the present invention is used to collect ewe information. The wearable device integrates a nine-axis sensor (monitoring accuracy: accelerometer 0.01g, magnetometer 1mg, gyroscope 0.1°), a sound sensor (signal-to-noise ratio 61dBA), a controller, a power supply, a data storage module, and a data transmission module. It is worn on the neck of the ewe through a freely adjustable strap. Among them, the nine-axis sensor is used to collect the acceleration, angular velocity, movement time of the ewe's movement, and the magnetic field strength at the ewe's location. The sound sensor is used to collect audio data with a timestamp. The corresponding ewe individual identification code can be set through the supporting software system to distinguish the data of different ewes.

[0068] According to actual needs, the preset time length can be different time lengths such as 24 hours, 48 ​​hours and 72 hours. After collecting behavioral data and audio data within the preset time length, LoRa, Bluetooth and WiFi hybrid communication can be used to receive data through the relay gateway and wearable device self-organizing network and upload it to the system, or the wearable device can directly upload it to the system.

[0069] When a relay gateway is used, it consists of a solar module, power supply box, base station, mobile tripod, and Lora antenna. All components are mounted on a movable tripod and can be recharged by solar energy in the field. The base station integrates a computing module, controller, power supply, data storage module, and data transmission module. The processed data is transmitted to a database, and the computer calls the data for analysis and display. The gateway receives acceleration, angular velocity, movement time, magnetic field strength, and processed audio data collected by multiple wearable devices, identifies the behavior at the gateway, and uploads the fused data features to the system database. When a relay gateway is not used, the wearable devices directly upload the collected acceleration, angular velocity, movement time, magnetic field strength, and processed audio data to the system for further processing.

[0070] The embodiment of the present invention collects behavioral data and audio data of ewes by designing a wearable device, providing a data basis for further estrus monitoring. At the same time, no blood sample needs to be collected, avoiding the stress response of ewes caused by traditional serum progesterone detection methods, and improving the convenience and safety of data collection.

[0071] Step 102: Obtain a first behavior data set based on the behavior data.

[0072] Optionally, obtaining a first behavior data set based on the behavior data includes:

[0073] Calculating pre-selected multi-dimensional characteristic parameters based on the behavioral data; the multi-dimensional characteristic parameters include the activity intensity and activity range of the ewe;

[0074] Based on the multidimensional feature parameters, the first behavior data set is obtained.

[0075] Specifically, in some embodiments, in order to further improve monitoring accuracy, the collected behavioral data may be subjected to denoising processing, including steps such as range conversion, wavelet denoising, and Z-score normalization after denoising (this step may be performed in the wearable device or at the relay gateway):

[0076] First, the signal (including the acceleration, angular velocity, and magnetic field strength signals collected by the nine-axis sensor) can be decomposed into sub-signals X(t) of different scales (frequency bands) through wavelet transform:

[0077]

[0078] Where c k are the wavelet coefficients, is the wavelet basis function, which represents the local characteristics of the signal;

[0079] After performing wavelet transform on the signal, multiple sub-signals S in different frequency ranges are obtained:

[0080] S = wavedec(x,wavelet,level)

[0081] Wherein, x is the original signal, wavedec is the selected wavelet basis, wavelet represents wavelet transform, and level is the number of decomposition levels. In the embodiment of the present invention, Daubechies wavelet db8 is selected as the wavelet basis, and the number of decomposition levels is 5.

[0082] Then, threshold processing is performed on the wavelet coefficients to remove noise. The embodiment of the present invention uses a soft threshold, which is expressed as:

[0083]

[0084] Where c k are the wavelet coefficients, is the threshold coefficient, and λ is the threshold used to control the strength of denoising. In the embodiment of the present invention, the threshold is 0.4;

[0085] Finally, the thresholded coefficients are restored to the denoised signal x by inverse wavelet transform. denoised :

[0086] x denoised =waverec(c thresholded ,wavelet)

[0087] Where waverec represents inverse wavelet transform, c thresholded is the coefficient after threshold processing.

[0088] In the embodiment of the present invention, since the fluctuation amplitude of accelerometer data is obvious and easy to observe, the accelerometer signal in the behavioral data collection process is selected, the number of sampling points is 100, the sampling frequency is 1 Hz, and the above-mentioned denoising method is used to perform denoising processing on it.

[0089] Based on the denoised data, pre-selected multi-dimensional feature parameters can be calculated. Among them, the multi-dimensional feature parameters include the activity intensity and activity range of the ewes, and the activity intensity and activity range of the ewes can be specifically characterized by thirty-dimensional feature parameters, namely: the interquartile range of the angular velocity y axis (Ggy(° / s)_iqr), the interquartile range of the angular velocity z axis (Ggz(° / s)_iqr), the standard deviation of the angular velocity y axis (Ggy(° / s)_std), the range of the angular velocity y axis (Ggy(° / s)_range), the interquartile range of the angular velocity x axis (Ggy(° / s)_iqr), the standard deviation of the angular velocity y axis (Ggy(° / s)_std), the range of the angular velocity y axis (Ggy(° / s)_range), the interquartile range of the angular velocity x axis (Ggy(° / s)_range), the interquartile range of the angular velocity Interdigit spacing (Ggx(° / s)_iqr), standard deviation of angular velocity on the z-axis (Ggz(° / s)_std), range of angular velocity on the x-axis (Ggx(° / s)_range), standard deviation of angular velocity on the x-axis (Ggx(° / s)_std), range of angular velocity on the z-axis (Ggz(° / s)_range), first quartile of acceleration on the z-axis (Aaz(g)_q1), minimum value of angular velocity on the y-axis (Ggy(° / s)_min), acceleration The minimum value of the z-axis (Aaz(g)_min), the maximum value of the angular velocity on the x-axis (Ggx(° / s)_max), the minimum value of the angular velocity on the x-axis (Ggx(° / s)_min), the standard deviation of the acceleration on the z-axis (Aaz(g)_std), the range of the acceleration on the z-axis (Aaz(g)_range), the maximum value of the angular velocity on the z-axis (Ggz(° / s)_max), the maximum value of the angular velocity on the y-axis (Ggy(° / s)_max), the standard deviation of the acceleration on the y ... standard deviation of the acceleration on the z-axis (Aaz(g)_std), the standard deviation of the acceleration on the z-axis (Aaz(g)_std), the standard deviation of the acceleration on the z-axis (Aaz(g)_std), the standard deviation of the acceleration on the z-axis (Aaz(g)_std), the standard deviation of the acceleration on the Standard deviation (Aay(g)_std), third quartile of angular velocity on the y-axis (Ggy(° / s)_q3), median of acceleration on the z-axis (Aaz(g)_median), mean of acceleration on the z-axis (Aaz(g)_mean), first quartile of acceleration on the y-axis (Aay(g)_q1), interquartile range of acceleration on the z-axis (Aaz(g)_iqr), minimum value of angular velocity on the z-axis (Ggz(° / s)_min), range of magnetic field on the y-axis The first quartile of the angular velocity on the y-axis (Ggz(° / s)_q1), the median of the acceleration on the y-axis (Aay(g)_median), and the mean of the magnetic field on the x-axis and the minimum value of the magnetic field on the y-axis The 30-dimensional feature parameters, including acceleration and angular velocity, along with the corresponding movement duration, can fully characterize the ewe's activity intensity. For example, prolonged high-speed movement indicates high activity intensity, while extended periods of low or even zero movement speed (such as lying down) indicate low activity intensity.

[0090] By analyzing the magnetic field strength at the location of the ewe in the thirty-dimensional characteristic parameters, the range of the ewe's position change can be characterized, thereby obtaining the ewe's activity range.

[0091] Based on the multidimensional feature parameters, the first behavioral data set is constructed to provide a data basis for subsequent ewe behavior identification.

[0092] Step 103: Obtain a first audio data set based on the audio data.

[0093] Optionally, obtaining a first audio data set based on the audio data includes:

[0094] Calculating the short-time energy and zero-crossing rate of the audio signal in the audio data;

[0095] The first audio data set is obtained based on the short-time energy and the zero-crossing rate.

[0096] Specifically, in some embodiments, in order to further improve the monitoring accuracy, the audio data collected by the sound sensor may be subjected to denoising processing, which mainly includes Fourier transform, filtering processing, and inverse Fourier transform.

[0097] First, perform Fourier transform on the ewe's audio data to extract frequency domain features from the audio signal:

[0098]

[0099] Where X(f) is the frequency component of the signal at frequency f; x(t) is the original time domain signal (i.e., audio signal), t is time, f is frequency, and e -j2πft is a complex exponential function used to calculate the weight (amplitude and phase) of each frequency component, and dt is a small increment of time, representing an integration operation.

[0100] Then, in order to improve the quality of the data and reduce the influence of unnecessary noise, the audio signal needs to be filtered. In the embodiment of the present invention, a bandpass filter is used to process the spectrum, retaining only the audio signal within the range of 0-2700 Hz, thereby removing interference noise outside the frequency range.

[0101] In the case of discrete signals, the Fourier transform is the discrete Fourier transform:

[0102]

[0103] Where X[k] is the amplitude and phase of the kth frequency component; x[n] is the sampling value of the signal at time point n; N is the total number of sampling points of the signal; k is the index of the frequency component; is a complex exponential, used to represent the basis function in the frequency domain.

[0104] In discrete Fourier transform, the expression corresponding to frequency is:

[0105]

[0106] Where, f k is the actual frequency value corresponding to the kth frequency, k frequency index, ranging from 0 to N-1; N is the total number of sampling points of the signal, Δt is the sampling time interval, equal to sampling_rate is the sampling frequency, the number of samples per second (unit: Hz).

[0107] Use mask to set the frequency range to 0≤f≤2700Hz:

[0108]

[0109] The frequency components of 0-2700 Hz are retained by masking, and other frequency components are set to 0, so that the frequencies below 2700 Hz (usually an important frequency band that can be heard by the human ear) are retained.

[0110] Finally, the frequency domain signal is converted back to the time domain through inverse Fourier transform.

[0111] Inverse Fourier transform formula: Convert the frequency domain signal back to the time domain:

[0112]

[0113] Where x[n] is the time domain signal after inverse transformation, X[k] is the kth frequency component, N is the number of sampling points of the signal, and k is the frequency index. is a complex exponential, which is used to reconstruct the time domain signal.

[0114] After denoising, a short time window is set based on the sampling points, and characteristic parameters are calculated for the denoised audio data. Specifically, the short-time energy E(n) and zero-crossing rate ZCR(n) of the audio signal are calculated. Short-time energy describes the energy of the signal within a short time window. In audio processing, short-time energy is used to detect the active state of audio (such as the fluctuation of speech or music).

[0115] By summing the squares of the audio signal in a short time window, the energy E(n) of each frame signal x(t) is obtained:

[0116]

[0117] Where x(t) is the sample value of the audio signal in the short time window, and N is the size of the short time window.

[0118] The zero-crossing rate is a measure of the number of sign changes in an audio signal and is used to describe the smoothness of an audio signal.

[0119] By counting the number of sign changes in the audio signal (i.e., the number of sign change intervals), the number of zero crossings in each frame of the signal, i.e., the zero crossing rate ZCR(n), is obtained:

[0120]

[0121] Where x(t) is the value of the audio signal at time t, N is the size of the short time window, and the sign() function is a sign function that returns 1 if x(t) is positive, otherwise it returns -1.

[0122] After calculating the short-time energy E(n) and zero-crossing rate ZCR(n) of the audio signal in the audio data, E(n) and ZCR(n) can be used to construct the first audio data set, thereby providing a data basis for subsequent ewe call detection.

[0123] Step 104: Input the first behavior data set into a ewe behavior recognition model to obtain a second behavior data set output by the ewe behavior recognition model.

[0124] Specifically, ewe behaviors are categorized into six main categories: standing, walking, feeding, ruminating, lying down, and exploring. By constructing a first behavioral dataset and then inputting it into an ewe behavior recognition model, we can obtain a second behavioral dataset, which is output by the model and provides a data foundation for subsequent multimodal feature fusion.

[0125] The ewe behavior recognition model may be an XgBoost model, a support vector machine model, or a random forest model, and the second behavior dataset includes six types of ewe behaviors with timestamps, that is, each type of ewe behavior and the corresponding behavior time.

[0126] Step 105: Input the first audio data set into a ewe sound detection model to obtain a second audio data set output by the ewe sound detection model.

[0127] Specifically, after using E(n) and ZCR(n) to construct the first audio dataset, the first audio dataset is input into the ewe call detection model. The ewe call detection model can screen out audio signals identified as ewe calls, that is, the short-time energy E'(n) and zero-crossing rate ZCR'(n) of ewe calls, and output them as the second audio dataset, thereby providing a data basis for subsequent further multimodal feature fusion.

[0128] Among them, the ewe call detection model can be a random forest model, a decision tree model or a support vector machine model.

[0129] Step 106: Perform feature extraction on the second behavior dataset and the second audio dataset respectively to obtain a multimodal feature fusion dataset.

[0130] Optionally, performing feature extraction on the second behavior dataset and the second audio dataset respectively to obtain a multimodal feature fusion dataset includes:

[0131] Calculate the proportion of the duration corresponding to each behavior of the ewe in the second behavior data set to the duration of all behaviors;

[0132] Calculating the short-time energy and zero-crossing rate of the ewe sounds in the second audio data set;

[0133] Determining a temporal variation characteristic of the behavior based on a proportion of the duration corresponding to each behavior to the duration of all behaviors;

[0134] Determining the temporal variation characteristics of the call based on the short-time energy and the zero-crossing rate;

[0135] Performing feature fusion on the behavior temporal change feature and the call temporal change feature according to the current time sequence to obtain a multimodal fusion feature;

[0136] Based on the multimodal fusion features, the multimodal feature fusion dataset is obtained.

[0137] Specifically, for the second behavior dataset, a time window and step size of a specific size are used to slide and extract the proportion of the duration of the behavior within the corresponding time window to the duration of all behaviors to determine the temporal change characteristics of the behavior.

[0138] For each time window, calculate the proportion of identified behaviors within the time window (the duration of all behaviors). Set each behavior as a class label laber_count(i), the expression is:

[0139]

[0140] For the current time window, after calculating the proportions of the six types of behaviors, we further extract the behavior categories with the largest and second largest proportions. We define the behavior label with the largest proportion as the dominant behavior dominant_label, and the behavior label with the second largest proportion as the second dominant behavior second_dominant_label. The expression is:

[0141] dominant_label=arg max(label_count)

[0142] second_dominant_label=arg second(label_count)

[0143] The difference between the proportion of behavior labels identified in the current time window and the average proportion of corresponding behaviors processed in all time windows within the preset time length (taking the first 24 hours as an example) is label_diff(i):

[0144]

[0145] Where, the cumulative proportion of label i in all windows is the cumulative proportion of label i in all time windows in the previous 24 hours, and the number of processed windows is the total number of time windows processed in the previous 24 hours.

[0146] If the previous time window exists, the change in the proportion of the behavior label between the current time window and the previous time window, label_change(i), is further calculated:

[0147] label_change(i) = label_count(i) - the proportion of label i in the previous window

[0148] That is, the proportion of label i in the current window is subtracted from the proportion of label i in the previous window, thereby determining the temporal change characteristics of the ewe's behavior within a preset time length.

[0149] For the second audio data set, the short-time energy E'(n) and zero-crossing rate ZCR'(n) of the ewe's calls in the second audio data set are matched with the corresponding timing information according to the timestamp of the collected data, so as to determine the timing change characteristics of the calls.

[0150] Furthermore, the behavior temporal change features and the call temporal change features are fused through feature-level fusion to obtain multimodal fusion features, and a multimodal feature fusion dataset is constructed.

[0151] In an embodiment of the present invention, multimodal feature integration in the spatiotemporal dimension is achieved by extracting alignment features within the timestamp intersection interval of two modal feature parameters and merging them, and extending the time axis to the union range of the original timestamps.

[0152] Direct concatenation and fusion are used for multimodal features within overlapping time windows, while interpolation or zero-filling is used to ensure the integrity of the feature matrix within non-overlapping time windows. Unlike the mid-term fusion (model hidden layer interaction) and late fusion (decision layer weighting) in traditional methods, the feature-level fusion (early fusion) proposed in this paper completes the information coupling between modalities during the feature extraction stage, which has the advantages of preserving the original feature correlation and enhancing cross-modal representation capabilities. The specific steps are as follows:

[0153] The behavior temporal change feature D1 and the call temporal change feature D2 are expressed as:

[0154]

[0155] Where n and m are the feature dimensions of mode 1 and mode 2 respectively. Through feature-level fusion, the feature vectors of the two modes are concatenated to obtain the fused multimodal fusion feature D f :

[0156]

[0157] The timestamp sets of the two modes are T1 and T2 respectively:

[0158]

[0159] Where p and q are the timestamp numbers of mode 1 and mode 2 respectively. In order to merge features, it is necessary to first determine the intersection of the two modal timestamps T common :

[0160] T common =T1∩T2

[0161] For the intersection part, the features under the timestamp can be extracted and spliced. The timestamp union part of the two modes T union :

[0162] T union =T1∪T2

[0163] For the timestamps in the union, if any mode has no data at that timestamp, the missing values ​​can be processed by interpolation or filling methods.

[0164] Based on the multimodal fusion feature D f , the obtained multimodal feature fusion dataset D′ f , which can be expressed as:

[0165] D′ f =concat(D1,D2)where D1,D2∈T common ∪T union

[0166] Where concat(·) represents the concatenation operation of feature vectors.

[0167] Through the above steps, feature-level fusion can merge multimodal features based on timestamp alignment and ensure the integrity and consistency of the data, thereby providing accurate data for subsequent models to detect estrus in ewes and improve detection accuracy.

[0168] Step 107: input the multimodal feature fusion data set into the ewe estrus discrimination model to obtain the ewe estrus detection result output by the ewe estrus discrimination model.

[0169] Specifically, in an embodiment of the present invention, the multimodal feature fusion data set has a total of 103 feature parameters. Before being input into the ewe estrus discrimination model, feature dimension reduction can also be performed using the model's own feature screening method, retaining only the top thirty most important feature parameters.

[0170] Among them, the feature screening method is to determine the importance of each feature by calculating the number of splits of each feature.

[0171] For each feature f, its importance is expressed by the sum of the number of splits in which feature f appears in all decision trees. i ) for characterization:

[0172] (Feature f i Whether it is used in the split of a node)

[0173] Where T is the number of trees, K t is the number of nodes in tree t, II is the indicator function, if the feature f i If it is used in the splitting of a certain node, then II=1, otherwise II=0.

[0174] After the above feature screening, the multimodal feature fusion dataset can be reduced in dimension to obtain a dataset after feature dimension reduction. The dataset after feature dimension reduction is then input into an estrus discrimination model for ewe estrus discrimination, and the estrus detection result is obtained as either non-estrus or estrus. The ewe estrus discrimination model can be a LightGBM model or a CatBoost model, for example.

[0175] The embodiment of the present invention performs feature screening on a multimodal feature fusion data set through an ewe estrus discrimination model, further reducing the amount of data processed during model detection on the basis of obtaining effective feature data, thereby improving the accuracy and efficiency of the model in detecting ewe estrus.

[0176] The present invention provides a method for detecting estrus of an ewe based on multimodal feature fusion. The method collects behavioral data and audio data of an ewe within a preset time length to obtain a first behavioral dataset and a first audio dataset, thereby achieving accurate, efficient and continuous tracking and collection of the ewe and obtaining sufficient data. The first behavioral dataset is then input into an ewe behavior recognition model to obtain a second behavioral dataset output by the ewe behavior recognition model, and the first audio dataset is input into an ewe call detection model to obtain a second audio dataset output by the ewe call detection model. Feature extraction is performed on the second behavioral dataset and the second audio dataset respectively to obtain a multimodal feature fusion dataset, thereby fully characterizing the changing patterns of the ewe's behavior and calls through the multimodal feature fusion dataset. Finally, the multimodal feature fusion dataset is input into an ewe estrus discrimination model to obtain an ewe estrus detection result output by the ewe estrus discrimination model. Based on the full learning of the ewe's multimodal features, the ewe estrus discrimination model can be used to intelligently discriminate whether the ewe is in estrus, thereby improving the accuracy and efficiency of ewe estrus detection.

[0177] Optionally, the training step of the ewe behavior recognition model includes:

[0178] Collect historical behavioral data of ewes;

[0179] Calculating pre-selected multi-dimensional feature parameters based on the historical behavior data;

[0180] The ewe behavior recognition model is trained using the multidimensional feature parameters as feature data and the corresponding historical behavior categories as label data.

[0181] Specifically, historical behavior data of ewes is first collected. Pre-selected multidimensional feature parameters are then calculated based on this historical behavior data. Finally, supervised training of an ewe behavior recognition model is performed using the multidimensional feature parameters as feature data and the six corresponding historical behavior categories (standing, walking, feeding, ruminating, lying down, and exploring) as label data. The ewe behavior recognition model can be an XgBoost model, a support vector machine model, or a random forest model.

[0182] The embodiment of the present invention uses a pre-experimental training model to enable the model to fully learn various behavioral characteristics of ewes and achieve accurate and rapid identification of the categories of various behaviors of ewes.

[0183] Optionally, the training step of the ewe sound detection model includes:

[0184] Collect historical audio data of ewes;

[0185] Calculating the short-time energy and zero-crossing rate of the audio signal in the historical audio data;

[0186] The ewe call detection model is trained using the short-time energy and zero-crossing rate of the audio signal in the historical audio data as feature data and the category of the corresponding historical audio as label data.

[0187] Specifically, historical audio data of ewes is first collected. The short-time energy and zero-crossing rate of the audio signals in the historical audio data are then calculated. Finally, supervised training of a model for detecting ewe calls is performed using the short-time energy and zero-crossing rate of the audio signals in the historical audio data as feature data and the corresponding historical audio categories as label data. The ewe call detection model can be a random forest model, a decision tree model, or a support vector machine model.

[0188] The embodiment of the present invention fully characterizes the call characteristics of ewes through the short-time energy and zero-crossing rate of the audio signal, so that the trained model can fully learn the various call characteristics of ewes and accurately and quickly detect whether the sound signal in the audio data is the category of ewe calls.

[0189] Optionally, the training step of the ewe estrus discrimination model includes:

[0190] Collecting the duration of various behaviors of the ewe within a historical time length and corresponding audio data; the historical time length includes the estrus period and the non-estrus period;

[0191] Determining the temporal variation characteristics of the behavior of the ewe over a historical time period based on the proportion of the duration of each behavior of the ewe in the duration of all behaviors;

[0192] Determining a temporal variation characteristic of the ewe's calls within a historical time length based on the short-time energy and zero-crossing rate of the ewe's calls in the audio data;

[0193] Performing feature fusion on the behavior temporal change features and the call temporal change features according to the time sequence within the historical time length to obtain historical multimodal feature fusion data;

[0194] The ewe estrus discrimination model is trained using the historical multimodal feature fusion data as feature data and the corresponding estrus and non-estrus categories as label data.

[0195] Specifically, in an embodiment of the present invention, the historical time length is taken as the time before the embolism removal and injection, 0-24 hours after the embolism removal and injection, 24-48 hours after the embolism removal and injection, and 48-72 hours after the embolism removal and injection in the synchronized estrus experiment, and the duration of various behaviors of the ewes within the historical time length and the corresponding audio data are collected, so that the collected data covers the estrus period and the non-estrus period, thereby improving the reliability and accuracy of the data.

[0196] Then, the duration of each behavior of the ewe (wherein the ewe behavior category is detected and identified by the ewe behavior recognition model) is calculated as a proportion of the duration of all behaviors to determine the temporal variation characteristics of the ewe's behavior over the historical time length; and the short-term energy and zero-crossing rate of the ewe's calls in the audio data (wherein the ewe's calls are detected and identified by the ewe call detection model) are calculated to determine the temporal variation characteristics of the ewe's calls over the historical time length;

[0197] Then, the behavior temporal change features and the call temporal change features are fused according to the time sequence within the historical time length to obtain historical multimodal feature fusion data;

[0198] Finally, supervised training of the ewe estrus discrimination model is performed using the historical multimodal feature fusion data as feature data and the corresponding estrus and non-estrus categories as label data. The ewe estrus discrimination model can be a LightGBM model or a CatBoost model.

[0199] The embodiment of the present invention fully characterizes the estrus characteristics of ewes by fusing historical multimodal feature data, so that the trained model can fully learn the estrus characteristics of ewes and accurately and quickly detect whether the ewes are in estrus.

[0200] Based on any of the above embodiments, the present invention comprehensively integrates technologies such as sensors, the Internet of Things, signal processing, and artificial intelligence to build a multimodal fusion wearable ewe estrus detection system that applies the above-mentioned ewe estrus detection method based on multimodal feature fusion. In actual applications, the multimodal fusion wearable ewe estrus detection system has an automatic recognition accuracy of 97% for six types of ewe behaviors, an estrus detection accuracy of up to 96%, and an F1 score of 0.93. Compared with the traditional method of using only behavioral features for estrus discrimination, the accuracy and F1 score are improved by 2% and 3% respectively. The embodiment of the present invention realizes automated ewe behavior monitoring and intelligent estrus recognition, and can perform continuous and uninterrupted monitoring, thereby reducing breeding failures caused by human misjudgment, improving the success rate of reproduction, providing technical support and auxiliary tools for animal husbandry production management, improving the level of informationization and intelligence of animal husbandry production, and thus ensuring the quality and effect of animal husbandry production.

[0201] The following describes a device for detecting estrus of ewe based on multimodal feature fusion provided by the present invention. The device for detecting estrus of ewe based on multimodal feature fusion described below and the method for detecting estrus of ewe based on multimodal feature fusion described above can refer to each other.

[0202] Based on any of the above embodiments, Figure 2 : is a structural diagram of a device for detecting ewe estrus based on multimodal feature fusion provided by the present invention, such as Figure 2 The embodiment of the present invention provides a device for detecting ewe estrus based on multimodal feature fusion, comprising an acquisition module 201, a first acquisition module 202, a second acquisition module 203, a first recognition module 204, a second recognition module 205, a feature fusion module 206, and a detection module 207, wherein:

[0203] The acquisition module 201 is used to collect behavioral data and audio data of the ewe within a preset time length; wherein the behavioral data includes the ewe's movement speed, movement time and magnetic field strength at the ewe's location; the first acquisition module 202 is used to obtain a first behavioral data set based on the behavioral data; the second acquisition module 203 is used to obtain a first audio data set based on the audio data; the first recognition module 204 is used to input the first behavioral data set into the ewe behavior recognition model to obtain a second behavioral data set output by the ewe behavior recognition model; the second recognition module 205 is used to input the first audio data set into the ewe call detection model to obtain a second audio data set output by the ewe call detection model; the feature fusion module 206 is used to extract features from the second behavioral data set and the second audio data set respectively to obtain a multimodal feature fusion data set; the detection module 207 is used to input the multimodal feature fusion data set into the ewe estrus discrimination model to obtain the ewe estrus detection result output by the ewe estrus discrimination model.

[0204] The present invention provides an estrus detection device for ewe based on multimodal feature fusion. The device collects behavioral data and audio data of ewes within a preset time length to obtain a first behavioral data set and a first audio data set, thereby achieving accurate, efficient and continuous tracking and collection of ewes and obtaining sufficient data. The first behavioral data set is then input into an ewe behavior recognition model to obtain a second behavioral data set output by the ewe behavior recognition model, and the first audio data set is input into an ewe call detection model to obtain a second audio data set output by the ewe call detection model. Feature extraction is performed on the second behavioral data set and the second audio data set respectively to obtain a multimodal feature fusion data set, thereby fully characterizing the changing patterns of ewe behavior and calls through the multimodal feature fusion data set. Finally, the multimodal feature fusion data set is input into an ewe estrus discrimination model to obtain an ewe estrus detection result output by the ewe estrus discrimination model. Based on the full learning of the multimodal features of the ewe, the ewe estrus discrimination model can be used to intelligently discriminate whether the ewe is in estrus, thereby improving the accuracy and efficiency of ewe estrus detection.

[0205] Figure 3 An example of a physical structure diagram of an electronic device is shown below. Figure 3As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330 and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the ewe estrus detection method based on multimodal feature fusion, which includes:

[0206] Collecting behavioral data and audio data of the ewe within a preset time period; wherein the behavioral data includes the ewe's movement speed, movement time, and magnetic field strength at the ewe's location;

[0207] Based on the behavioral data, obtaining a first behavioral data set;

[0208] Based on the audio data, obtaining a first audio data set;

[0209] Inputting the first behavior data set into a ewe behavior recognition model to obtain a second behavior data set output by the ewe behavior recognition model;

[0210] Inputting the first audio data set into a ewe sound detection model to obtain a second audio data set output by the ewe sound detection model;

[0211] Performing feature extraction on the second behavior dataset and the second audio dataset respectively to obtain a multimodal feature fusion dataset;

[0212] The multimodal feature fusion data set is input into the ewe estrus discrimination model to obtain the ewe estrus detection result output by the ewe estrus discrimination model.

[0213] In addition, the logic instructions in the above-mentioned memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0214] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the ewe estrus detection method based on multimodal feature fusion provided by the above methods, and the method includes:

[0215] Collecting behavioral data and audio data of the ewe within a preset time period; wherein the behavioral data includes the ewe's movement speed, movement time, and magnetic field strength at the ewe's location;

[0216] Based on the behavioral data, obtaining a first behavioral data set;

[0217] Based on the audio data, obtaining a first audio data set;

[0218] Inputting the first behavior data set into a ewe behavior recognition model to obtain a second behavior data set output by the ewe behavior recognition model;

[0219] Inputting the first audio data set into a ewe sound detection model to obtain a second audio data set output by the ewe sound detection model;

[0220] Performing feature extraction on the second behavior dataset and the second audio dataset respectively to obtain a multimodal feature fusion dataset;

[0221] The multimodal feature fusion data set is input into the ewe estrus discrimination model to obtain the ewe estrus detection result output by the ewe estrus discrimination model.

[0222] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the ewe estrus detection method based on multimodal feature fusion provided by the above methods, the method comprising:

[0223] Collecting behavioral data and audio data of the ewe within a preset time period; wherein the behavioral data includes the ewe's movement speed, movement time, and magnetic field strength at the ewe's location;

[0224] Based on the behavioral data, obtaining a first behavioral data set;

[0225] Based on the audio data, obtaining a first audio data set;

[0226] Inputting the first behavior data set into a ewe behavior recognition model to obtain a second behavior data set output by the ewe behavior recognition model;

[0227] Inputting the first audio data set into a ewe sound detection model to obtain a second audio data set output by the ewe sound detection model;

[0228] Performing feature extraction on the second behavior dataset and the second audio dataset respectively to obtain a multimodal feature fusion dataset;

[0229] The multimodal feature fusion data set is input into the ewe estrus discrimination model to obtain the ewe estrus detection result output by the ewe estrus discrimination model.

[0230] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0231] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0232] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0233] It should also be noted that the terms "first," "second," and the like are used in the present invention to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein. Furthermore, the terms "first" and "second" generally distinguish objects of the same type, and do not limit the number of objects. For example, the first object can be one or more.

[0234] In the embodiments of the present application, "determine B based on A" means that the factor A must be considered when determining B. It is not limited to "B can be determined based on A alone", and should also include: "determine B based on A and C", "determine B based on A, C and E", "determine C based on A, and further determine B based on C", etc. It can also include taking A as a condition for determining B, for example, "when A meets the first condition, use the first method to determine B"; for example, "when A meets the second condition, determine B", etc.; for example, "when A meets the third condition, determine B based on the first parameter", etc. Of course, it can also be a condition that takes A as a factor in determining B, for example, "when A meets the first condition, use the first method to determine C, and further determine B based on C", etc.

[0235] In the present invention, the term "plurality" refers to two or more than two, and other quantifiers are similar to it.

[0236] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for detecting estrus in ewe based on multimodal feature fusion, characterized in that: include: Collecting behavioral data and audio data of the ewe within a preset time period; wherein the behavioral data includes the ewe's movement speed, movement time, and magnetic field strength at the ewe's location; Based on the behavioral data, obtaining a first behavioral data set; Based on the audio data, obtaining a first audio data set; Inputting the first behavior data set into a ewe behavior recognition model to obtain a second behavior data set output by the ewe behavior recognition model; Inputting the first audio data set into a ewe sound detection model to obtain a second audio data set output by the ewe sound detection model; Performing feature extraction on the second behavior dataset and the second audio dataset respectively to obtain a multimodal feature fusion dataset; The multimodal feature fusion data set is input into the ewe estrus discrimination model to obtain the ewe estrus detection result output by the ewe estrus discrimination model.

2. The ewe estrus detection method based on multimodal feature fusion according to claim 1, characterized in that: The step of obtaining a first behavior data set based on the behavior data includes: Calculating pre-selected multi-dimensional characteristic parameters based on the behavioral data; the multi-dimensional characteristic parameters include the activity intensity and activity range of the ewe; Based on the multidimensional feature parameters, the first behavior data set is obtained.

3. The ewe estrus detection method based on multimodal feature fusion according to claim 1, characterized in that: The step of obtaining a first audio data set based on the audio data includes: Calculating the short-time energy and zero-crossing rate of the audio signal in the audio data; The first audio data set is obtained based on the short-time energy and the zero-crossing rate.

4. The method for detecting ewe estrus based on multimodal feature fusion according to claim 1, characterized in that: The training steps of the ewe behavior recognition model include: Collect historical behavioral data of ewes; Calculating pre-selected multi-dimensional feature parameters based on the historical behavior data; The ewe behavior recognition model is trained using the multidimensional feature parameters as feature data and the corresponding historical behavior categories as label data.

5. The ewe estrus detection method based on multimodal feature fusion according to claim 1, characterized in that: The training steps of the ewe sound detection model include: Collect historical audio data of ewes; Calculating the short-time energy and zero-crossing rate of the audio signal in the historical audio data; The ewe call detection model is trained using the short-time energy and zero-crossing rate of the audio signal in the historical audio data as feature data and the category of the corresponding historical audio as label data.

6. The ewe estrus detection method based on multimodal feature fusion according to claim 1, characterized in that: The performing feature extraction on the second behavior dataset and the second audio dataset respectively to obtain a multimodal feature fusion dataset includes: Calculate the proportion of the duration corresponding to each behavior of the ewe in the second behavior data set to the duration of all behaviors; Calculating the short-time energy and zero-crossing rate of the ewe sounds in the second audio data set; Determining a temporal variation characteristic of the behavior based on a proportion of the duration corresponding to each behavior to the duration of all behaviors; Determining the temporal variation characteristics of the call based on the short-time energy and the zero-crossing rate; Performing feature fusion on the behavior temporal change feature and the call temporal change feature according to the current time sequence to obtain a multimodal fusion feature; Based on the multimodal fusion features, the multimodal feature fusion dataset is obtained.

7. The method for detecting ewe estrus based on multimodal feature fusion according to claim 1, characterized in that: The training steps of the ewe estrus discrimination model include: Collecting the duration of various behaviors of the ewe within a historical time length and corresponding audio data; the historical time length includes the estrus period and the non-estrus period; Determining the temporal variation characteristics of the behavior of the ewe over a historical time period based on the proportion of the duration of each behavior of the ewe in the duration of all behaviors; Determining a temporal variation characteristic of the ewe's calls within a historical time length based on the short-time energy and zero-crossing rate of the ewe's calls in the audio data; Performing feature fusion on the behavior temporal change features and the call temporal change features according to the time sequence within the historical time length to obtain historical multimodal feature fusion data; The ewe estrus discrimination model is trained using the historical multimodal feature fusion data as feature data and the corresponding estrus and non-estrus categories as label data.

8. A device for detecting ewe estrus based on multimodal feature fusion, characterized in that: include: A collection module is used to collect behavioral data and audio data of the ewe within a preset time length; wherein the behavioral data includes the ewe's movement speed, movement time, and magnetic field strength at the ewe's location; A first obtaining module, configured to obtain a first behavior data set based on the behavior data; A second obtaining module, configured to obtain a first audio data set based on the audio data; a first recognition module, configured to input the first behavior data set into a ewe behavior recognition model to obtain a second behavior data set output by the ewe behavior recognition model; a second recognition module, configured to input the first audio data set into a ewe sound detection model to obtain a second audio data set output by the ewe sound detection model; a feature fusion module, configured to extract features from the second behavior dataset and the second audio dataset respectively to obtain a multimodal feature fusion dataset; The detection module is used to input the multimodal feature fusion data set into the ewe estrus discrimination model to obtain the ewe estrus detection result output by the ewe estrus discrimination model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for detecting estrus of ewe based on multimodal feature fusion as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for detecting estrus in ewe based on multimodal feature fusion as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Dairy cow behavior identification digital twin system based on ultra wide band and inertial measurement unit and digital twin method of dairy cow behavior identification digital twin system

    CN117322358A

  • Behavior recognition monitoring system based on ewe oestrus

    CN117561997A

  • Panda mating period prediction method based on multi-modal behavior information

    CN118397498A

  • Sow estrus behavior identification method and system based on multi-modal feature fusion

    CN119513697A

  • System for detecting cow estrus using recognition of behavior pattern

    KR102117092B1

Cited By

  • Ewe oestrus detection device and method

    CN121694895A

  • Ewe mating opportunity identification system and method

    CN122207610A