A personalized pet multi-modal emotion recognition method
By collecting sound and visual signals from pets in a calm state, a personalized emotional baseline is established. Combined with multimodal features, emotional state points are generated, solving the accuracy and stability problems of pet emotion recognition in existing technologies and achieving accurate perception and prediction of pet emotions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN SONGSHI INTELLIGENT TECH CO LTD
- Filing Date
- 2025-10-27
- Publication Date
- 2026-07-03
AI Technical Summary
Existing technologies rely on timing misalignment in speech recognition, which affects accuracy. Furthermore, they ignore individual behavioral habits and differences in characteristic responses, making it difficult to identify potential anxiety trends in pets and unable to create dynamic emotional profiles.
The system collects sound signals from pets in a calm state, extracts features such as Mel frequency cepstral coefficients, pitch envelope, and formants, establishes a personalized emotional homeostasis baseline, and generates multimodal emotional state points by combining visual signals of tail wagging frequency and ear posture angle. Continuous emotional fluctuations are detected through an emotional homeostasis imbalance assessment score.
It improves the accuracy and specificity of emotion recognition, enhances data robustness, and can detect and predict the stability and trend changes of pets' emotions, avoiding misjudgments caused by short-term fluctuations.
Smart Images

Figure CN121331171B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of emotion recognition technology, and more particularly to a method for recognizing multimodal emotions in pets based on individual perception. Background Technology
[0002] Emotion recognition is a technical field that intersects artificial intelligence and human behavior analysis. It mainly identifies an individual's current emotional state by perceiving, analyzing, and modeling various external expressive information.
[0003] Current technologies relying solely on voice are prone to temporal misalignment, impacting recognition accuracy. Furthermore, their generalized modeling approach ignores individual behavioral habits and characteristic response differences, leading to conflicts between model generalization and personalized expression. In emotion state recognition, current technologies often only assess the immediate state, failing to support stability analysis and trend assessment. For instance, when a pet continuously exhibits low-frequency vocalizations and slow tail movements, current technologies struggle to identify underlying anxiety trends, merely recognizing it as a temporary anomaly and failing to create a dynamic emotion profile. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a personalized perception-based multimodal emotion recognition method for pets.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for recognizing multimodal emotions in pets based on individual perception, comprising the following steps:
[0006] Sound signals and all vocalization events over a long period of time were collected when the pet was in a calm state. Mel frequency cepstral coefficients, pitch envelope, formants and short-term energy of the sound signals were extracted to establish a personalized emotional homeostasis baseline.
[0007] Based on the personalized emotional steady-state baseline, the acoustic features of newly acquired pet vocalizations are extracted, a real-time acoustic deviation metric is generated, and the real-time acoustic deviation metric is mapped to a preset multi-dimensional space of arousal and pleasure to obtain an acoustic emotion offset vector.
[0008] Based on the acoustic emotion offset vector, visual signals synchronized with the pet's vocalization are collected, and the tail wagging frequency and ear posture angle are extracted as multimodal visual behavior parameters. Multimodal visual behavior parameters are obtained, and multimodal fused emotion state points are obtained based on the multimodal visual behavior parameters.
[0009] Based on the personalized emotional stability baseline and the multimodal fused emotional state points, the multimodal fused emotional state points are collected within a preset time window to form an emotional state sequence. The vocal frequency and type of the emotional state sequence are compared to generate an emotional stability imbalance assessment score.
[0010] Preferably, the step of obtaining the personalized emotional steady-state baseline is as follows:
[0011] The acquisition is initiated when the pet is calm. The original waveform of the sound signal is recorded, the reference time information is saved, the start and end timestamps of all vocal events within a long time period are marked, the duration of each vocal event is calculated, and the interval between adjacent vocal events is counted to obtain the sound signal and vocal event records in the calm state.
[0012] Based on the sound signal in the calm state and the recording of the sound event, resampling, DC removal, framing and windowing are performed. The Mel frequency cepstral coefficient values are calculated frame by frame, the pitch envelope curve is extracted frame by frame, the formant center frequency and bandwidth are extracted frame by frame, the short-time energy value is calculated frame by frame, and the frame-by-frame results are aligned with the corresponding sound event start timestamp to obtain the acoustic feature sequence and rhythm data.
[0013] Based on the acoustic feature sequence and rhythm data, the mean, variance, skewness, kurtosis and quantile of Mel frequency cepstral coefficients, pitch envelope, formants and short-term energy are calculated respectively. An ordered sequence is established by combining the timestamp of the vocal event, and the periodic statistics of duration distribution and adjacent interval distribution are calculated. The acoustic feature statistics and rhythm distribution results are fused to generate a personalized emotional steady state baseline.
[0014] Preferably, the step of obtaining the real-time acoustic deviation measurement is as follows:
[0015] Based on the personalized emotional steady-state baseline, the waveform segments of newly acquired pet vocalizations are read, and the frames are divided and windowed according to a uniform frame length and frame shift. The Mel frequency cepstral coefficients, pitch envelope, formant center frequency and bandwidth, and short-time energy are extracted frame by frame. The histogram interval count and quantile of each feature are counted in chronological order to obtain the current acoustic feature distribution.
[0016] Based on the current acoustic feature distribution, align with the corresponding distribution in the personalized emotional steady-state baseline, calculate the quantile difference, skewness difference and kurtosis difference for each component, calculate the energy interval proportion difference and the frequency band interval proportion difference, and perform weighted summation according to a preset weight table to generate a real-time acoustic deviation metric.
[0017] Preferably, the step of obtaining the acoustic emotion offset vector is as follows:
[0018] Based on the real-time acoustic deviation metric, the components of the real-time acoustic deviation metric are transformed according to the arousal axis coefficient and the pleasure axis coefficient, the range is clipped and normalized, and the positioning in the multi-dimensional space of arousal and pleasure is completed, generating an acoustic emotion offset vector.
[0019] Preferably, the steps for obtaining the multimodal visual behavior parameters are as follows:
[0020] Based on the acoustic emotion offset vector, visual signal segments are extracted synchronously during the start and end time of vocalization, the trajectory of key points at the tail is detected, and the number of tail wagging cycles per unit time is calculated to obtain the tail wagging frequency. At the same time, the position of key points of the auricle relative to the head is detected and the auricle offset angle is calculated to form multimodal visual behavior parameters.
[0021] Preferably, the step of obtaining the multimodal fusion emotion state points is as follows:
[0022] Based on the multimodal visual behavior parameters, a visual emotion modulation vector is calculated and fused with an acoustic emotion offset vector to obtain a multimodal fused emotion state point.
[0023] Preferably, the steps for obtaining the emotional state sequence are as follows:
[0024] Based on the personalized emotional steady-state baseline and the multimodal fused emotional state points, the multimodal fused emotional state points are collected in fixed time slices within a preset time window. The number of vocalizations in each time slice is counted to obtain the instantaneous vocalization frequency. The proportion of vocalization types in each time slice is also counted to form an emotional state sequence.
[0025] Preferably, the steps for obtaining the emotional homeostasis imbalance assessment score are as follows:
[0026] Based on the emotional state sequence, the expected acoustic behaviors in the personalized emotional steady-state baseline are aligned by time slice, the expected vocal frequencies and the proportion of expected vocal types are read, and a quadruple is assembled for each time slice to form a time-series comparison data set.
[0027] Based on the time-series comparison data set, an assessment score for emotional homeostasis imbalance is calculated.
[0028] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0029] This invention effectively captures the acoustic behavior patterns of different pets under stable emotional states by collecting pet sound signals in a calm state and extracting multidimensional acoustic features. Based on a baseline, it measures the feature deviation of new vocal samples to form an acoustic emotion offset vector with directionality and individual differences, improving the accuracy and targeting of emotion perception. While processing the acoustic signals, it simultaneously collects visual signals and extracts behavioral parameters that are highly correlated with emotions, such as tail wagging frequency and ear posture angle. This establishes a logical pairing relationship between the temporal features of visual behavior and acoustic features. By fusing with the acoustic emotion offset vector, it generates the current emotional state point, enhancing the perception dimension and data robustness. At the same time, it uses statistical distance to measure periodic deviations and calculates an emotional steady-state imbalance assessment score, thereby enabling the detection and intervention judgment of continuous emotional fluctuations. By introducing the distribution comparison of behavioral frequency and type and statistical smoothing processing, it can effectively avoid misjudgments caused by short-term fluctuations. Attached Figure Description
[0030] Figure 1 This is a schematic diagram of the steps of the present invention. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0032] Please see Figure 1 This invention provides a technical solution: a method for recognizing multimodal emotions in pets based on individual perception, comprising the following steps:
[0033] Sound signals and all vocalization events over a long period of time were collected when the pet was in a calm state. Mel frequency cepstral coefficients, pitch envelope, formants and short-term energy of the sound signals were extracted to establish a personalized emotional homeostasis baseline.
[0034] Based on the personalized emotional steady-state baseline, the acoustic features of newly acquired pet vocalizations are extracted to generate a real-time acoustic deviation metric. The real-time acoustic deviation metric is then mapped to a preset multi-dimensional space of arousal and pleasure to obtain an acoustic emotion offset vector.
[0035] Based on the acoustic emotion offset vector, visual signals synchronized with the pet's vocalization are collected, and the tail wagging frequency and ear posture angle are extracted as multimodal visual behavior parameters. Multimodal visual behavior parameters are obtained, and multimodal fused emotion state points are obtained based on the multimodal visual behavior parameters.
[0036] Based on the personalized emotional stability baseline and the multimodal fusion emotional state points, the multimodal fusion emotional state points are collected within a preset time window to form an emotional state sequence. The vocal frequency and type of the emotional state sequence are compared to generate an emotional stability imbalance assessment score.
[0037] The steps for obtaining a personalized emotional homeostasis baseline are as follows:
[0038] The acquisition is initiated when the pet is calm. The original waveform of the sound signal is recorded, the reference time information is saved, the start and end timestamps of all vocal events within a long time period are marked, the duration of each vocal event is calculated, and the interval between adjacent vocal events is counted to obtain the sound signal and vocal event records in the calm state.
[0039] Based on the sound signal in the calm state and the recording of the sound event, resampling, DC removal, framing and windowing are performed. The Mel frequency cepstral coefficient values are calculated frame by frame, the pitch envelope curve is extracted frame by frame, the formant center frequency and bandwidth are extracted frame by frame, and the short-time energy value is calculated frame by frame. The frame-by-frame results are aligned with the corresponding sound event start timestamp to obtain the acoustic feature sequence and rhythm data.
[0040] Based on acoustic feature sequences and rhythm data, the mean, variance, skewness, kurtosis, and quantile of Mel frequency cepstral coefficients, pitch envelope, formants, and short-term energy are calculated. An ordered sequence is established by combining the timestamps of vocal events, and periodic statistics of duration distribution and adjacent interval distribution are calculated. The acoustic feature statistics and rhythm distribution results are integrated to generate a personalized emotional steady-state baseline.
[0041] Specifically, the acquisition program is initiated when the pet is in a calm state defined by low activity (e.g., sleeping, lying down, sitting quietly) for more than 30 minutes, as confirmed by motion sensors and video analysis. The raw waveform of the sound signal is continuously recorded using an omnidirectional microphone configured with a 44100 Hz sampling rate and 16-bit quantization precision. Simultaneously, a network time protocol service is invoked to obtain and save the reference time information, adding absolute timestamps to all subsequent events. Then, all vocalization events within the long time period are labeled. First, the mean and standard deviation of the background noise energy within a 5-second sliding window are calculated, and the initial energy detection threshold is set to the mean background noise energy plus three times the standard deviation. The zero cross-rate threshold is empirically set to 0.15. When audio frames continuously exceed 50 milliseconds... When both the short-term energy and short-term zero-crossing rate of an audio frame are higher than their respective thresholds, the start of a sound event is determined, and the start timestamp is recorded. Conversely, when the energy or zero-crossing rate of an audio frame that lasts for more than 100 milliseconds is lower than the threshold, the end of a sound event is determined, and the end timestamp is recorded. After marking a sound event, the duration of the sound event is immediately calculated, which is the difference between the end timestamp and the start timestamp. The time difference between the start timestamp of the current sound event and the end timestamp of the previous sound event is calculated as the interval between adjacent sound events. All the recorded raw waveform data are combined with a list containing the start timestamp, end timestamp, event duration, and event interval of each sound event to finally obtain the calm state sound signal and the sound event record.
[0042] Based on the sound signal in a calm state and the recording of the sound events, the original waveform in the recording is first preprocessed by resampling it to 16000 Hz to reduce computational complexity while retaining key sound information. Next, a DC removal operation is performed on each labeled sound event segment, i.e., the average amplitude of all sample points in that segment is subtracted to eliminate the influence of DC bias on subsequent Fourier transforms. Then, the processed sound segments are framed, with a frame length of 25 milliseconds and a frame shift of 10 milliseconds. A Hamming window function is applied to each frame, calculated by multiplying the nth point in a frame of length N by 0.54 and subtracting 0.46 multiplied by... As a result, to smooth frame edges and reduce spectral leakage, 13-dimensional Mel frequency cepstral coefficients are calculated frame-by-frame for each windowed frame. This process includes Fast Fourier Transform, passing through a Mel filter bank, taking the logarithm, and Discrete Cosine Transform. The pitch envelope curve is extracted frame-by-frame using the autocorrelation function method. The fundamental frequency is determined by finding the first non-zero peak of the autocorrelation function of each frame, forming pitch time-series data. Linear predictive coding analysis is used frame-by-frame to extract the center frequencies and bandwidths of the first four formants. The formants are obtained by solving the roots of the linear predictive coding polynomial. The parameters are calculated frame by frame, and the short-time energy value is calculated as the sum of the squares of the amplitudes of all sampling points within the frame. Finally, the Mel frequency cepstral coefficient value, the point on the pitch envelope curve, the formant center frequency and bandwidth, and the short-time energy value calculated for each frame are correlated and aligned with the relative time offset of the frame in the entire sound event and the start timestamp of the sound event itself to generate a feature vector sequence containing timestamps. This sequence is merged with the duration and interval data extracted from the sound signal in the calm state and the sound event record to obtain the acoustic feature sequence and rhythm data.
[0043] Based on acoustic feature sequences and rhythmic data, statistical analysis was performed on each dimension of the acoustic feature sequences. Specifically, for twenty acoustic features—13 Mel-frequency cepstral coefficients, pitch, center frequencies and bandwidths of the four formants, and short-time energy—the mean, variance, skewness, and kurtosis were calculated for all frames of all quiet-state sound events. The 25th, 50th (median), and 75th percentiles were also calculated to quantify the central tendency, dispersion, and distribution pattern of each feature. Simultaneously, using the duration of all sound events and the intervals between adjacent sound events recorded in the rhythmic data, empirical cumulative distribution functions were constructed and calculated. The mean and standard deviation of these distributions are used as periodic statistics to describe the rhythmic patterns of a pet's vocalizations in a calm state. Subsequently, the calculated statistics for all acoustic features (each feature includes mean, variance, skewness, kurtosis, and three quantiles, totaling 20 features multiplied by 7 statistics, resulting in 140 values) are fused with statistics describing vocal rhythm (mean and standard deviation of duration and interval, totaling 4 values). Specifically, this is achieved by constructing a structured data object containing all these statistical values, indexed by feature names and statistic names as keys. For example, the object's structure is {"MFCC_1":{"mean": The code snippet `val, "variance": val, ...}, "Pitch":{"mean": val, ...}, "Rhythm":{"duration_mean":val, "interval_mean": val, ...}}` combines the detailed distribution of acoustic features with macroscopic rhythmic patterns to generate a quantitative, personalized emotional homeostasis baseline that comprehensively describes the acoustic behavior characteristics of an individual pet in a calm state.
[0044] The steps for obtaining real-time acoustic deviation measurement are as follows:
[0045] Based on the personalized emotional steady-state baseline, newly acquired waveform segments of pet vocalizations are read, and the segments are divided into frames and windowed according to a uniform frame length and frame shift. Mel frequency cepstral coefficients, pitch envelope, formant center frequency and bandwidth, and short-time energy are extracted frame by frame. The histogram interval count and quantile of each feature are counted in chronological order to obtain the current acoustic feature distribution.
[0046] Based on the current acoustic feature distribution, the corresponding distribution in the personalized emotional steady-state baseline is aligned, and the quantile difference, skewness difference and kurtosis difference are calculated for each component. The energy interval proportion difference and the frequency band interval proportion difference are calculated. The weighted summaries are performed according to the preset weight table to generate a real-time acoustic deviation metric.
[0047] Specifically, based on the personalized emotional steady-state baseline, when a new pet vocalization is detected, the original waveform segment corresponding to that vocalization event is read in real time and preprocessed using parameters identical to those used when establishing the baseline. Specifically, the waveform segment is divided into frames using a frame length of 25 milliseconds and a frame shift of 10 milliseconds, and a Hamming window function is applied to each frame. Then, on each processed frame, a set of acoustic features is extracted in parallel, including 13-dimensional Mel frequency cepstral coefficients calculated using Fast Fourier Transform, Mel filter bank, and Discrete Cosine Transform, pitch envelope points determined using the autocorrelation function method, and center frequencies of the four formants obtained by solving the roots of the linear predictive coding polynomial. After extracting features frame by frame, including bandwidth and short-time energy obtained by calculating the sum of squares of signal samples within the frame, the numerical sequences of each acoustic feature are collected in chronological order for all frames in the current vocal segment. Then, statistical analysis is performed on these sequences. First, based on the numerical range set for each feature in the personalized emotional steady-state baseline, the range is divided into 20 intervals. The number of feature frames in the current vocal segment that fall into each interval is counted to form a histogram interval count. Second, the 25%, 50%, and 75% quantiles of the current feature sequence are calculated. The histogram interval counts and quantile statistics of all acoustic features are integrated to obtain the current acoustic feature distribution.
[0048] Based on the current acoustic feature distribution, it is precisely aligned with the corresponding feature distribution stored in the personalized emotional steady-state baseline. For example, the pitch feature distribution of the current vocalization is matched with the pitch feature distribution in the baseline, and a series of difference indices are calculated component by component. First, the quantile difference is calculated, which is the sum of the absolute values of the differences between the 25th, 50th, and 75th quantiles of the current feature distribution and the corresponding quantiles of the baseline. Second, the skewness and kurtosis of the current feature sequence are calculated, and the skewness and kurtosis values recorded in the baseline are subtracted respectively to obtain the skewness difference and kurtosis difference. Third, the energy interval proportion difference is calculated using histogram interval counting. This process involves converting the count values of the 20 intervals of the short-time energy feature into percentages, and then calculating the sum of the absolute values of the percentage differences between the current distribution and the baseline distribution in each interval. Similarly, for the first and second Mel frequency cepstral coefficients representing the shape of the spectral envelope, their frequency band proportion difference is calculated. Subsequently, the calculated quantile difference, skewness difference, kurtosis difference, and energy interval proportion are compared. Five categories of difference indicators—pitch quantile difference, energy interval proportion difference, and frequency band interval proportion difference—were weighted and aggregated according to a pre-defined weight table. This weight table was constructed based on principal component analysis and regression analysis of vocalization data from 50 pets of different breeds under various emotional states (jointly labeled by two senior veterinarians). The analysis results showed that pitch-related features were most sensitive to changes in emotional arousal, while energy and spectral envelope features were more sensitive to changes in pleasure. Weights were assigned accordingly; for example, pitch quantile difference was weighted at 0.3, energy interval proportion difference at 0.25, frequency band interval proportion difference at 0.2, skewness difference at 0.15, and kurtosis difference at 0.1, with a total weight of 1.0. The weighted aggregation was calculated by multiplying the value of each difference indicator by its corresponding weight and then summing all the products. For example, if the calculated pitch quantile difference was 1.2, energy interval proportion difference was 0.8, frequency band interval proportion difference was 0.5, skewness difference was -0.3, and kurtosis difference was 0.6, then the weighted aggregation value would be... The final weighted sum is the real-time acoustic deviation metric that includes multiple acoustic feature deviation information.
[0049] The steps for obtaining the acoustic emotion offset vector are as follows:
[0050] Based on the real-time acoustic deviation metric, the components of the real-time acoustic deviation metric are transformed according to the arousal axis coefficient and the pleasure axis coefficient, the range is clipped and normalized, and the positioning in the multi-dimensional space of arousal and pleasure is completed, generating an acoustic emotion offset vector.
[0051] Specifically, based on the real-time acoustic deviation measurement, this single deviation metric needs to be decomposed and mapped to a two-dimensional emotion space. To this end, before weighting and summing, the five types of difference indicators calculated in the previous step (quantile difference, skewness difference, kurtosis difference, energy range proportion difference, and frequency band proportion difference) are multiplied by their respective arousal axis coefficients and pleasure axis coefficients. These coefficients are also determined through multiple linear regression analysis on the aforementioned labeled dataset of 50 pets, aiming to quantify the independent contribution of each acoustic deviation to the arousal and pleasure dimensions. For example, the regression model might yield an arousal axis coefficient of 0.7 for pitch quantile difference and a pleasure axis coefficient of -0.2, while the energy range... The arousal axis coefficient for the difference in acoustic arousal ratio is 0.5, and the pleasure axis coefficient is 0.4. Each difference index is multiplied by its corresponding arousal axis coefficient and then summed to obtain the original acoustic arousal component. Similarly, multiplying by the pleasure axis coefficient and summing yields the original acoustic pleasure component. Next, these two original components are cropped. Based on large-scale baseline data statistics, the theoretical range for arousal and pleasure components is set between -2.0 and +2.0. Any calculation result outside this range is forcibly set as a boundary value. For example, if the calculated acoustic arousal component is 2.5, it is cropped to 2.0. Then, normalization is performed, and the cropped values are linearly mapped... Transforming to the standard range of -1.0 to +1.0, this process achieves precise localization in the multidimensional space of arousal and pleasure. Finally, the normalized acoustic arousal and pleasure components are combined to form a two-dimensional vector. ,in It is the arousal component. It is the pleasure component, and this vector is the acoustic emotion offset vector pointing from the origin (0,0) of the emotion space to the current emotion state point.
[0052] The steps for obtaining multimodal visual behavior parameters are as follows:
[0053] Based on the acoustic emotion offset vector, visual signal segments are extracted synchronously during the start and end time of vocalization. The trajectory of key points at the tail is detected and the number of tail wagging cycles per unit time is calculated to obtain the tail wagging frequency. At the same time, the position of key points of the auricle relative to the head is detected and the auricle offset angle is calculated to form multimodal visual behavior parameters.
[0054] Specifically, based on the acoustic emotion offset vector and its associated vocalization start and end timestamps, corresponding visual signal segments are precisely extracted from a synchronously recorded high-definition video stream at 30 frames per second. Then, a deep learning-based animal pose estimation algorithm is applied to each frame of this visual signal segment. This algorithm is a convolutional neural network model pre-trained on a dataset containing over 100,000 labeled dog and cat images, specifically using the DeepLabCut framework. This model identifies and outputs the two-dimensional coordinates of 20 body key points, including the base of the tail, the midpoint of the tail, the tip of the tail, the base of both ears, the tips of both ears, the center of both eyes, and the nose. Next, focusing on the horizontal coordinates of the tail tip key point, its coordinate sequence throughout the entire visual signal segment is extracted to form a time-series signal. A bandpass filter is then applied to this signal to filter out… Slow body movement noise below 0.5 Hz and high-frequency jitter noise above 5 Hz are filtered, and then a Fast Fourier Transform is performed on the filtered signal. The peak frequency in the resulting spectrum is found in the range of 0.5 to 5 Hz, and the frequency corresponding to the peak frequency is determined as the tail wagging frequency. At the same time, the ear posture angle is calculated using the coordinates of key points on the ears and head. First, a head center line is determined by the center points of the eyes and the nose key point, and its perpendicular line is defined as the vertical reference axis of the head. Then, the ear vectors of the left and right ears are calculated separately, and this vector points from the ear root key point to the ear tip key point. Finally, the angle between each ear vector and the vertical reference axis of the head is calculated, and the average of the two angles is used as the final auricle offset angle. Finally, the calculated tail wagging frequency and auricle offset angle are combined into a two-dimensional vector to form multimodal visual behavior parameters.
[0055] The steps for obtaining emotional state points through multimodal fusion are as follows:
[0056] Based on the multimodal visual behavior parameters, a visual emotion modulation vector is calculated and fused with an acoustic emotion offset vector to obtain a multimodal fused emotion state point. The calculation formula is as follows:
[0057] ;
[0058] in:
[0059] ;
[0060] ;
[0061] ;
[0062] For multimodal fusion of emotional state points, For acoustic arousal component, For acoustic pleasure component, For the arousal component of visual accommodation, For the pleasure component of visual accommodation, The frequency of tail wagging. For the ear posture angle, The average frequency of tail wagging is the baseline. The standard deviation of the tail wagging frequency is the baseline. The baseline mean of ear posture angle, The baseline standard deviation of ear posture angle, This is the standardized value for the tail wagging frequency. Standardized values for ear posture angles. The coefficient representing the influence of tail wagging frequency on arousal level. The coefficient representing the influence of ear posture angle on arousal level. The coefficient representing the influence of the interaction term between tail wagging frequency and ear posture angle on arousal level. The coefficient representing the influence of tail wagging frequency on pleasure level. The coefficient representing the influence of ear posture angle on pleasure level. The coefficient representing the influence of the interaction term between tail wagging frequency and ear posture angle on pleasure level. Adjust the intensity coefficient to increase arousal level. The intensity coefficient is adjusted to reflect the level of pleasure.
[0063] Specifically, in the multimodal fusion emotional state point calculation formula, the emotional state initially obtained from acoustic signals (represented by a two-dimensional vector of arousal and pleasure) is used as a baseline. Then, synchronized visual behavioral signals (tail wagging and ear posture) are used to dynamically and non-linearly adjust it, through an independent visual emotion adjustment vector. To correct the acoustic results, visual information plays a calibrating and supplementary role. Simultaneously, the hyperbolic tangent function (tanh) is used to process the visual accommodation, limiting the accommodation effect to a bounded range. This prevents drastic deviations in the final emotion judgment due to extreme or erroneous visual signal detection, increasing the model's robustness. Furthermore, a standardized value for the tail wagging frequency is introduced. Standardized value of ear posture angle Interactive items This design can capture the complex emotional expression when two behaviors are combined, such as rapid tail wagging (high... The presence of a feeling of excitement itself may indicate excitement, but if it is accompanied by high pressure behind the ear (high pressure)... Interaction items generate significant moderating signals, shifting the interpretation of emotions from simple "happiness" to "nervous excitement" or "anxiety".
[0064] The acoustic arousal component represents the level of emotional arousal inferred solely from the acoustic features of the pet's vocalizations. Its value originates from the first component of the acoustic emotion offset vector generated in the previous step. This vector is defined in a standardized two-dimensional emotion space, where the arousal axis represents the degree from calm to excitement, with values normalized to the interval [-1, 1]. Here, -1 represents extreme calmness or drowsiness, 0 represents a neutral or calm state, and 1 represents extreme excitement or agitation. This value is calculated based on the deviation between the newly acquired pet vocalizations and the personalized emotional steady-state baseline, incorporating changes in various acoustic features such as pitch, energy, and spectrum. In this example, the acoustic arousal component of the current vocalization event is obtained through the aforementioned real-time acoustic deviation measurement and mapping process. A value of 0.4 indicates a moderately high level of wakefulness.
[0065] The acoustic pleasure component represents the level of emotional pleasure inferred solely from the acoustic features of the pet's vocalizations. Its value originates from the second component of the acoustic emotion offset vector generated in the previous step. Within the same standardized two-dimensional emotion space, the pleasure axis represents the degree from displeasure to pleasure, with values normalized to the [-1, 1] interval. Here, -1 represents extreme displeasure, pain, or fear, 0 represents neutral emotion, and 1 represents extreme pleasure or satisfaction. The generation logic of this value is similar to... Parallelism, obtained by comparing the acoustic differences between the new vocalization and the personalized emotional steady-state baseline, and projecting them onto the pleasure dimension, reflects the positive or negative emotional coloring contained in the vocalization. In this example, based on the acoustic analysis results, the acoustic pleasure component is obtained. A value of 0.2 indicates a slightly positive emotional state.
[0066] Tail wagging frequency is a key behavioral indicator extracted in real-time from visual signal segments synchronized with vocalization. It quantifies the speed of the pet's tail wagging, measured in Hertz (Hz). The acquisition process involves using a pose estimation algorithm to track the horizontal displacement of key points at the tip of the tail across consecutive video frames, forming a time-series data. This time-series data is then analyzed using signal processing techniques (such as Fourier transform) to determine its dominant frequency. In canine behavior, tail wagging frequency is typically positively correlated with arousal levels; a higher frequency (e.g., greater than 3 Hz) is usually associated with excitement or agitation. In this example, the current tail wagging frequency of the pet is calculated through analysis of synchronized video segments. It is 3.5 Hz.
[0067] The ear posture angle is another key behavioral indicator extracted from visual signals. It quantifies the change in a pet's ear posture relative to its head, measured in degrees (°). Its calculation relies on accurately identifying the positions of the ear root, tip, and a head reference point (such as between the eyes). Geometric relationships are used to calculate the angle at which the ear deflects backward or to the side. This angle reflects the pet's alertness, compliance, or tension. For example, a large backward tilt may be associated with anxiety or fear, while a naturally erect or forward tilt may indicate relaxation or focus. In this example, the average ear offset angle of the pet is calculated by analyzing the coordinates of key points in the video frames. It is 30 degrees Celsius.
[0068] The baseline mean of tail wagging frequency represents the average tail wagging frequency of a specific pet in a calm, relaxed state. It is part of the individualized emotional baseline and is obtained by simultaneously recording the pet's visual signals over a long period (e.g., recording continuously for 7 days, collecting 2 hours of data each day during the pet's resting period) during the establishment of the individualized emotional homeostasis baseline. All video clips identified as calm (judged by activity level sensors and the absence of vigorous vocalizations) are analyzed to extract the tail wagging frequency. The arithmetic mean of these frequency values is then calculated. This baseline mean is individualized and reflects the pet's basic behavioral patterns. For example, a Golden Retriever's tail may still wag slightly and slowly when calm, and its baseline mean might be 0.8 Hz. In this example, the baseline mean of tail wagging frequency is obtained by retrieving the pet's individualized baseline data. It is 1.0 Hz.
[0069] The standard deviation of tail wagging frequency is the baseline. This parameter quantifies the dispersion or stability of tail wagging frequency in a specific pet under resting conditions, compared to the baseline mean. Together they form the baseline of visual behavior, and its calculation and Simultaneously, after collecting a large number of tail wagging frequency samples in a calm state, these sample values and the mean were calculated. The standard deviation is the square root of the mean of the squares of the differences. A smaller standard deviation indicates that the pet's tail wagging frequency is very stable when at rest, while a larger standard deviation means that the tail wagging behavior itself fluctuates more when at rest. This parameter is crucial for standardizing subsequent behavioral changes. In this example, based on its personalized baseline data, the pet's tail wagging frequency baseline standard deviation is... It is 0.5 Hz.
[0070] The baseline mean of ear pose angle is defined by this parameter, which indicates the typical pose angle of a pet's ears in a calm state. It is a core component of the personalized visual baseline, and its acquisition process is similar to... Similarly, during prolonged observation in a calm state, the ear offset angle is continuously calculated, and the arithmetic mean of all obtained valid angle values is taken to obtain the "standard" ear position when the pet is relaxed. The natural shape and position of the ears vary greatly among different breeds and individuals of dogs and cats; therefore, individualization of this parameter is crucial. For example, the baseline angle for a German Shepherd may be close to 0 degrees (ears erect), while the baseline angle for a Cocker Spaniel will have different definitions and values due to its long, drooping ears. In this example, the baseline mean of the pet's ear posture angle is... It is 15 degrees.
[0071] The baseline standard deviation of ear posture angle describes the range of variation in ear posture angle in a pet at rest. A baseline for ear posture was jointly defined by calculating the mean of all ear posture angle samples collected in a calm state and the baseline. The root mean square value of the deviation is obtained, which reflects the natural range of ear posture changes in a pet without external stimuli. This value helps determine whether the current ear posture change exceeds the normal range of random fluctuations, thereby determining whether it has emotional indicative significance. In this example, the baseline standard deviation of the ear posture angle is obtained by referring to personalized baseline data. It is 5 degrees.
[0072] , , These are the influence coefficients of tail wagging frequency, ear posture angle, and their interaction terms on arousal. These coefficients are key parameters of the model, determining how visual signals modulate the arousal component. They were obtained through machine learning training on a dataset containing hundreds of pet behavior video clips labeled with arousal and pleasure levels. Specifically, a multivariate nonlinear regression model was constructed using standardized visual parameters. , and interactive items As input features, the difference between the arousal values labeled by human experts and the arousal values predicted by the acoustic model is used as the target variable for fitting. The model weights obtained after training are these influence coefficients. According to the training results, the tail-wagging frequency usually has a strong positive influence on arousal, while the influence of ear posture and interaction terms is more complex. In this example, we set... , , .
[0073] , , These are the influence coefficients of tail wagging frequency, ear posture angle, and their interaction terms on pleasure, respectively. These parameters are obtained in the same way as those for arousal, except that when training the regression model, the target variable becomes the difference between the pleasure value labeled by human experts and the pleasure value predicted by the acoustic model. In this way, the model can learn the specific moderating effect of different combinations of visual behaviors on the pleasure dimension. For example, the model will learn high-frequency tail wagging (high... It has a positive effect on pleasure, but if there is also high pressure behind the ear (…), it can also negatively impact pleasure. The negative impact coefficient of interaction items can offset some of the positive sentiment, pulling the result towards neutrality or even negativity. In this example, based on the model training results, the following parameters are set: , , .
[0074] The arousal modulation intensity coefficient controls the overall intensity or maximum amplitude of the visual signal's modulation of arousal. It is a scaling factor that determines the visual accommodation vector. The range of values for this coefficient is determined by data-driven principles. After training the regression model for the influence coefficient, the distribution of the model's output (i.e., visual modulation) is evaluated on the validation set. Set it to 1.5 times the standard deviation of the target variable (awakeness prediction error) on the validation set. For example, if the standard deviation of the awakening prediction error on the validation set is 0.2, then you can set it to... This setting allows visual accommodation to effectively correct acoustic predictions in most cases while avoiding over-accommodation. In this example, the arousal intensity coefficient is set. It is 0.3.
[0075] The intensity coefficient is adjusted to improve pleasure levels; this parameter functions similarly to... Similarly, but acting on the pleasure dimension, it controls the overall strength of the visual signal's modulation of the pleasure component. The method for determining its value is also the same: it's set based on the distribution of the model's pleasure prediction errors on the validation set. The standard deviation of the pleasure prediction errors on the validation set is calculated and multiplied by a fixed scaling factor (e.g., 1.5) to determine its value. The value of this parameter ensures that the adjustment range of pleasure level matches the variability of the data itself, guaranteeing the rationality and effectiveness of the adjustment. In this example, based on the analysis of the model validation performance, a pleasure level adjustment intensity coefficient is set. It is 0.4.
[0076] Calculations based on parameters:
[0077] First, based on the frequency of tail wagging. Benchmark Mean and the benchmark standard deviation Calculate the standardized value of the tail wagging frequency. :
[0078] ;
[0079] Next, based on the ear posture and angle Benchmark Mean and the benchmark standard deviation Calculate the standardized value of ear posture angle :
[0080] ;
[0081] Then, the arousal component of visual accommodation is calculated using standardized values and influence coefficients. ,in , , , :
[0082] ;
[0083] ;
[0084] Subsequently, the pleasure component of visual accommodation was calculated. ,in , , , :
[0085] ;
[0086] ;
[0087] Finally, the acoustic emotion component and The calculated visual accommodation components are fused to obtain the multimodal fused emotion state points. :
[0088] ;
[0089] ;
[0090] The results show that the initial emotional state judged by sound was (0.4, 0.2), representing moderate to high arousal and slight pleasure. After incorporating visual information (very rapid tail wagging and significantly tucked-back ears), the final multimodal fused emotional state point was (0.4874, 0.0480). This new coordinate point slightly improved the arousal dimension (from 0.4 to 0.4874) but significantly decreased the pleasure dimension (from 0.2 to 0.0480), almost returning to a neutral level. This indicates that the analysis of visual behavior corrected the conclusion drawn solely from sound. Although high-frequency tail wagging increased arousal, the tense ear posture and the interaction between the two weakened the judgment of pleasure, correcting the emotion from "relatively happy excitement" to "more neutral or slightly tense excitement." This final coordinate point is the comprehensive assessment result of the pet's emotional state at the current moment.
[0091] The steps for obtaining the emotional state sequence are as follows:
[0092] Based on the personalized emotional steady-state baseline and the multimodal fusion emotional state points, the multimodal fusion emotional state points are collected in fixed time slices within a preset time window. The number of vocalizations in each time slice is counted to obtain the instantaneous vocalization frequency. The proportion of vocalization types in each time slice is also counted to form an emotional state sequence.
[0093] Specifically, based on the personalized emotional steady-state baseline and multimodal fusion emotional state points, a preset time window of 30 minutes is set, and this window is divided into 10 fixed time slices of 3 minutes each. This time setting is based on observations of the activity and rest cycles of common canine and feline pets. The 30-minute window is sufficient to capture a complete emotional fluctuation event, while the 3-minute time slices provide sufficient temporal resolution while ensuring data smoothness. As multimodal fusion emotional state points with timestamps are continuously generated, each state point is assigned to its corresponding time slice according to its timestamp. After a time slice ends, all multimodal fusion emotional state points collected within it are immediately statistically analyzed. First, the total number of state points within the time slice is calculated. Since each state point corresponds to an independent vocalization event, this total number is the instantaneous vocalization frequency. The unit is "times / 3 minutes". Next, the vocalizations within each time slice are classified into types. This classification is based on the position of the multimodal fusion emotional state point in the arousal-pleasure two-dimensional space. This space is pre-divided into three regions. For example, the region with arousal greater than 0.3 and pleasure greater than 0.3 is defined as the "positive excitement" type, the region with arousal greater than 0.3 and pleasure less than -0.3 is defined as the "negative vigilance" type, and the remaining regions are defined as the "neutral normal" type. The number of vocalizations falling into these three types within each time slice is counted, and the proportion of each type to the total number of vocalizations in that time slice is calculated to obtain the vocalization type proportion vector for each time slice, such as [positive excitement proportion, negative vigilance proportion, neutral normal proportion]. The instantaneous vocalization frequency and vocalization type proportion vector of these 10 time slices are arranged in chronological order to form an emotional state sequence.
[0094] The steps to obtain the emotional homeostasis imbalance assessment score are as follows:
[0095] Based on the emotional state sequence, the expected acoustic behaviors in the personalized emotional steady-state baseline are aligned by time slice, the expected vocal frequencies and the proportion of expected vocal types are read, and a quadruple is assembled for each time slice to form a time-series comparison data set.
[0096] Based on the time-series comparative dataset, the emotional homeostasis imbalance assessment score is calculated using the following formula:
[0097] ;
[0098] Where Q is the emotional homeostasis imbalance assessment score, U is the number of time slices, and u is the time slice index with a value from 0 to U-1. The time slice weight is defined as , The balance coefficient takes values in the range [0,1]. Let u be the frequency of sound emitted in the u-th time slice. Let u be the expected frequency of the sound in the u-th time slice. The proportion of sound type k within the u-th time slice. Let K be the expected proportion of voice type k, K be the total number of voice types, and k be the voice type index with values from 1 to K.
[0099] Specifically, based on the emotional state sequence formed in the previous step, to assess the deviation of the current emotional fluctuation from the pet's normal state, it is necessary to extract the corresponding expected acoustic behavior as a comparison benchmark from the personalized emotional homeostasis baseline. This baseline was established by statistically analyzing the pet's acoustic behavior in a confirmed calm state over several consecutive days. The expected vocalization frequency is obtained by calculating the average number of vocalizations per 3-minute time slice in all calm states. This value is usually very low, for example, an average of 0.5 vocalizations per 3 minutes. The expected vocalization type proportion is obtained by classifying all vocalizations in calm states (using the same spatial region classification standard as mentioned above) and calculating the long-term average proportion of each type of vocalization. For example, a typical calm baseline might show 95% of vocalizations as "neutral and normal," 3% as "actively excited," and 2% as "negatively alert." Next, each time slice in the emotional state sequence is aligned with these benchmark values. For the u-th time slice in the sequence, the measured instantaneous vocalization frequency is extracted. and vocal type ratio vector Simultaneously, the desired sound frequency is read from the baseline. and the proportional vector of the desired vocal type These four parts Assemble into a quadruple, repeat this operation for all 10 time slices within the entire time window, and arrange the generated 10 quadruples in chronological order to form a structured time-series comparison data set.
[0100] In the formula for calculating the assessment score of emotional homeostasis imbalance, the difference in the proportion of vocal types is measured using the total variation distance. This is a classic method for measuring the difference between two probability distributions, and its squared value is also in the interval [0, 1]. Furthermore, it introduces time slice weights based on the Hanning window. This approach allows the evaluation process to focus more on behavior within the time window, while assigning lower weight to instantaneous behavior at the window's edge. This aligns with the characteristic of emotional states having a certain degree of persistence, filtering out accidental noise at the beginning and end, resulting in more robust evaluation results. The balance coefficient... The introduction of this approach provides flexibility in adjusting the contribution of frequency bias and type bias to the total score, allowing for customized adjustments based on the behavioral characteristics of different species or individuals.
[0101] The number of time slices, this parameter defines the total number of discrete time units contained within the time window for emotion steady-state assessment. It is determined by the total duration of the preset time window and the duration of each fixed time slice, and is calculated using the following formula: The setting of this parameter affects the time scale of the evaluation. A larger number of time slices can provide more refined time dynamic analysis, but it also requires more computational resources. In the implementation of this method, the preset time window is set to 30 minutes, and the fixed time slice duration is 3 minutes. Therefore, the number of time slices... The acquisition process is as follows That is, within one evaluation period, 10 consecutive time slices will be analyzed.
[0102] Time slice weights are used to assign different importance to different time slices when calculating the total bias. The formula for their calculation is: ,in It is the index of the time slice, from 0 to... This formula is the squared form of the Hanning window function. The resulting weight sequence is bell-shaped, with low values at both ends and high values in the middle. This design ensures that time slices located in the center of the time window contribute the most to the final evaluation score, while time slices at the beginning and end of the window contribute less. In this example, ,index The weights of each time slice, from 0 to 9, are calculated as follows:
[0103] ,
[0104] ,
[0105] ,
[0106] ,
[0107] ,
[0108] to weights and to Symmetry, that is , , , , .
[0109] The balance coefficient is a dimensionless parameter ranging from [0, 1], used to adjust the relative weights of the vocal frequency deviation and vocal type ratio deviation terms in the emotional homeostasis imbalance assessment score. At that time, the evaluation score was determined entirely by the change in the frequency of the sound. At other times, it is entirely determined by changes in vocalization type. For example, by collecting a large amount of pet behavior data labeled with levels of emotional imbalance by veterinary experts, methods such as grid search can be used to find an optimal [emotional pattern]. The value is chosen to maximize the correlation between the calculated assessment score Q and the expert rating. For general applications, it can be set to 0.5, indicating that frequency and type changes are treated equally. In this example, it is set to... .
[0110] Let be the vocalization frequency of the u-th time slice. This parameter represents the actual number of times the pet vocalizes within the time slice with index u. It comes from the "emotional state sequence" generated in the previous step and is a core observation in the time-series comparison dataset. The unit is "times / time slice," which reflects the activity level of the pet's vocalization behavior within a specific time period. In this example, an observation sequence containing 10 time slices (U=10) is set, and its vocalization frequency... The values (from u=0 to 9) are as follows: .
[0111] Let be the expected vocalization frequency for the u-th time slice. This parameter represents the average number of vocalizations the pet makes within a time slice of the same length in a calm state as defined in the personalized emotional homeostasis baseline. Read from the personalized emotional homeostasis baseline, it represents the vocalization level of the pet when it is "normal" or "quiet." In this example, the expected vocalization frequency obtained from the pet's personalized emotional homeostasis baseline is 0.5 times per 3-minute time slice. Therefore, for all time slices u, .
[0112] The total number of vocalization types defines the total number of discrete categories into which pet vocalization behavior can be divided. This value was determined during the method design phase, based on an understanding of the complexity of pet emotional expression and a trade-off between the practicality of the classification system. In this method, as mentioned before, vocalizations are divided into three types (positive excitement, negative alertness, and neutral normality). .
[0113] Let be the proportion of vocalization type k within the u-th time slice. This parameter is a vector representing the proportion of each of the K different vocalization types within the u-th time slice, with a sum of 1. It also originates from the "emotional state sequence" and reflects the "qualitative" characteristics of the pet's vocalization behavior within that time slice. In this example, The three types are neutral (k=1), positive excitement (k=2), and negative vigilance (k=3). For time slices where the vocal frequency is not zero, the type ratios are set as follows:
[0114] u=0, f=1: ;
[0115] u=1, f=2: ;
[0116] u=2, f=4: ;
[0117] u=3, f=6: ;
[0118] u=4, f=5: ;
[0119] u=5, f=3: ;
[0120] u=6, f=2: ;
[0121] u=7, f=1: ;
[0122] For u=8 and u=9, since the sound frequency is 0, the type ratio is meaningless, and the contribution of this term in the calculation is 0.
[0123] Let be the expected proportion of vocalization type k. This parameter is a vector extracted from the personalized emotional homeostasis baseline, representing the long-term average distribution proportion of various vocalization types in a pet in a calm state. Its sum is also 1, providing a "normal" reference for comparing the current vocalization type distribution. In this example, based on the baseline data, the expected proportion vectors of the three types (neutral / normal, positive / excited, negative / vigilant) are: .
[0124] Calculations based on parameters:
[0125] First, calculate the two dimensionless deviation terms for each time slice u: the frequency deviation term. and type deviation item :
[0126] u=0: , , ; ; .but The total contribution is 0.
[0127] u=1: , , ; ; ......
[0128] Calculate the weighted sum:
[0129] ;
[0130] Calculate the weighted sum:
[0131] ;
[0132] Calculate Q:
[0133] ;
[0134] The results indicate that within this 30-minute time window, the calculated emotional homeostasis imbalance assessment score Q is approximately 0.604. Since the Q value is calculated using a normalized bias term, its theoretical range is close to [0, 1]. Based on this, assessment criteria can be established, for example: [0, 0.2) represents "homeostasis," [0.2, 0.4) represents "mild imbalance," and [0.4, 0.7) represents "moderate imbalance." A score of 0.7 indicates "severe imbalance," while a score of 0.604 falls into the "moderate imbalance" range. This indicates that the pet's emotional state has deviated significantly and persistently during this period, possibly indicating excitement, anxiety, or neediness. This warrants attention and concern from pet owners or caregivers.
[0135] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for recognizing multimodal emotions in pets based on individual perception, characterized in that, Includes the following steps: Sound signals and all vocalization events over a long period of time were collected when the pet was in a calm state. Mel frequency cepstral coefficients, pitch envelope, formants and short-term energy of the sound signals were extracted to establish a personalized emotional homeostasis baseline. Based on the personalized emotional steady-state baseline, the acoustic features of newly acquired pet vocalizations are extracted, a real-time acoustic deviation metric is generated, and the real-time acoustic deviation metric is mapped to a preset multi-dimensional space of arousal and pleasure to obtain an acoustic emotion offset vector. Based on the acoustic emotion offset vector, visual signals synchronized with the pet's vocalization are collected, and the tail wagging frequency and ear posture angle are extracted as multimodal visual behavior parameters. Multimodal visual behavior parameters are obtained, and multimodal fused emotion state points are obtained based on the multimodal visual behavior parameters. Based on the personalized emotional stability baseline and the multimodal fused emotional state points, the multimodal fused emotional state points are collected within a preset time window to form an emotional state sequence. The vocal frequency and type of the emotional state sequence are compared to generate an emotional stability imbalance assessment score. The steps for obtaining the multimodal fusion emotional state points are as follows: Based on the multimodal visual behavior parameters, a visual emotion modulation vector is calculated and fused with an acoustic emotion offset vector to obtain a multimodal fused emotion state point.
2. The pet multimodal emotion recognition method based on individual perception according to claim 1, characterized in that, The steps for obtaining the personalized emotional homeostasis baseline are as follows: The acquisition is initiated when the pet is calm. The original waveform of the sound signal is recorded, the reference time information is saved, the start and end timestamps of all vocal events within a long time period are marked, the duration of each vocal event is calculated, and the interval between adjacent vocal events is counted to obtain the sound signal and vocal event records in the calm state. Based on the sound signal in the calm state and the recording of the sound event, resampling, DC removal, framing and windowing are performed. The Mel frequency cepstral coefficient values are calculated frame by frame, the pitch envelope curve is extracted frame by frame, the formant center frequency and bandwidth are extracted frame by frame, the short-time energy value is calculated frame by frame, and the frame-by-frame results are aligned with the corresponding sound event start timestamp to obtain the acoustic feature sequence and rhythm data. Based on the acoustic feature sequence and rhythm data, the mean, variance, skewness, kurtosis and quantile of Mel frequency cepstral coefficients, pitch envelope, formants and short-term energy are calculated respectively. An ordered sequence is established by combining the timestamp of the vocal event, and the periodic statistics of duration distribution and adjacent interval distribution are calculated. The acoustic feature statistics and rhythm distribution results are fused to generate a personalized emotional steady state baseline.
3. The pet multimodal emotion recognition method based on individual perception according to claim 1, characterized in that, The steps for obtaining the real-time acoustic deviation metric are as follows: Based on the personalized emotional steady-state baseline, the waveform segments of newly acquired pet vocalizations are read, and the frames are divided and windowed according to a uniform frame length and frame shift. The Mel frequency cepstral coefficients, pitch envelope, formant center frequency and bandwidth, and short-time energy are extracted frame by frame. The histogram interval count and quantile of each feature are counted in chronological order to obtain the current acoustic feature distribution. Based on the current acoustic feature distribution, align with the corresponding distribution in the personalized emotional steady-state baseline, calculate the quantile difference, skewness difference and kurtosis difference for each component, calculate the energy interval proportion difference and the frequency band interval proportion difference, and perform weighted summation according to a preset weight table to generate a real-time acoustic deviation metric.
4. The pet multimodal emotion recognition method based on individual perception according to claim 1, characterized in that, The steps for obtaining the acoustic emotion offset vector are as follows: Based on the real-time acoustic deviation metric, the components of the real-time acoustic deviation metric are transformed according to the arousal axis coefficient and the pleasure axis coefficient, the range is clipped and normalized, and the positioning in the multi-dimensional space of arousal and pleasure is completed, generating an acoustic emotion offset vector.
5. The pet multimodal emotion recognition method based on individual perception according to claim 1, characterized in that, The steps for obtaining the multimodal visual behavior parameters are as follows: Based on the acoustic emotion offset vector, visual signal segments are extracted synchronously during the start and end time of vocalization, the trajectory of key points at the tail is detected, and the number of tail wagging cycles per unit time is calculated to obtain the tail wagging frequency. At the same time, the position of key points of the auricle relative to the head is detected and the auricle offset angle is calculated to form multimodal visual behavior parameters.
6. The pet multimodal emotion recognition method based on individual perception according to claim 1, characterized in that, The steps for obtaining the emotional state sequence are as follows: Based on the personalized emotional steady-state baseline and the multimodal fused emotional state points, the multimodal fused emotional state points are collected in fixed time slices within a preset time window. The number of vocalizations in each time slice is counted to obtain the instantaneous vocalization frequency. The proportion of vocalization types in each time slice is also counted to form an emotional state sequence.
7. The pet multimodal emotion recognition method based on individual perception according to claim 1, characterized in that, The steps for obtaining the emotional homeostasis imbalance assessment score are as follows: Based on the emotional state sequence, the expected acoustic behaviors in the personalized emotional steady-state baseline are aligned by time slice, the expected vocal frequencies and the proportion of expected vocal types are read, and a quadruple is assembled for each time slice to form a time-series comparison data set. Based on the time-series comparison data set, an assessment score for emotional homeostasis imbalance is calculated.
Citation Information
Patent Citations
Method for real-time generation of empathy expression of virtual human based on multimodal emotion recognition and artificial intelligence system using the method
US20250200855A1
KR20230144482A