Deep learning-based pet dog emotion recognition method and system
Through deep learning technology, multimodal data is collected and integrated, and a multi-level emotion labeling framework and multimodal feature fusion are constructed, which solves the problems of subjectivity and environmental adaptability of traditional pet dog emotion recognition and achieves high-precision pet emotion recognition and health monitoring.
Patent Information
- Application Number
- CN202510949216.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Traditional pet dog emotion recognition technology relies on subjective judgment and lacks objective quantitative standards. It is difficult to accurately identify complex emotional states, and the recognition accuracy decreases when the environment changes.
A deep learning-based method is used to collect multimodal data, build a three-level emotion labeling framework, integrate multimodal features through a cascaded SEblock array, and inject adversarial sample data generated by a stress scenario simulator during the training phase to train the deep learning model.
It achieves high-precision recognition of pet emotions and can maintain stable recognition performance in complex and changeable practical application scenarios, improving the comprehensiveness and accuracy of emotion recognition and providing a scientific basis for pet health monitoring and behavioral intervention.
Smart Images

Figure CN120708251A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pet emotion recognition, and in particular to a pet dog emotion recognition method and system based on deep learning. Background Art
[0002] Traditional methods of identifying emotions in pet dogs rely primarily on the subjective judgment of owners and the experience of professional veterinarians, lacking objective, quantitative standards or evidence. In recent years, with the rapid development of artificial intelligence (AI) and the booming pet industry, emotion recognition in pet dogs has become an emerging research field. Some studies have begun using computer vision to analyze pets' facial expressions or using single sensors to collect physiological parameters to infer their emotional states.
[0003] However, existing pet dog emotion recognition technology has the following shortcomings: First, most methods only focus on single-modal data, ignoring the multidimensionality of pet emotional expression, making it difficult to accurately identify complex emotional states; second, existing models perform poorly when dealing with environmental changes and special scenarios (such as veterinary clinics or pet beauty salons, which may trigger pet stress reactions), making it difficult to meet the requirements of model robustness and refined personalization of emotion recognition. Summary of the Invention
[0004] The main purpose of this invention is to provide a pet dog emotion recognition method based on deep learning, aiming to solve the technical problems in the prior art.
[0005] The present invention proposes a pet dog emotion recognition method based on deep learning, comprising: Collect pet dynamic data, pet physiological data, and situational data to obtain a structured data set; Constructing a three-level emotion annotation framework, annotating the structured dataset through a cross-validation annotation mechanism to obtain an annotated dataset; Extracting multimodal features based on the labeled data set, and integrating the multimodal features through a cascaded SEblock array to obtain multimodal fusion features; Injecting adversarial sample data generated by a stress scenario simulator during the training phase, performing deep learning model training based on the adversarial sample data and the multimodal fusion features, and obtaining a pet emotion recognition model; Emotion recognition is performed on the pet to be identified using the pet emotion recognition model to obtain a pet emotion recognition result.
[0006] Preferably, the step of annotating the structured data set by a cross-validation annotation mechanism to obtain an annotated data set includes: Annotating the pet's dynamic data using an annotation tool to obtain basic emotion tags, wherein the basic emotion tags include anxiety, pleasure, and fear; Associating the scenario data with the basic emotion tag through a context association algorithm to generate a context enhancement tag; Based on the physiological parameter mapping model, calculating the correlation coefficient between the pet's physiological data and the emotional state corresponding to the basic emotion label to obtain a physiological emotion weight matrix; Cross-validating the basic emotion labels, the situational enhancement labels, and the physiological emotion weight matrix through a crowdsourcing verification platform, eliminating data with a labeling confidence lower than a threshold, and obtaining a cross-validation result; A labeled data set containing multiple levels of labels is generated according to the cross-validation results.
[0007] Preferably, the step of extracting multimodal features based on the labeled data set and integrating the multimodal features through a cascaded SEblock array to obtain multimodal fusion features includes: The facial key point detection network is used to locate the facial area based on the pet's dynamic data to obtain facial expression features; Perform audio analysis based on pet dynamic data through Mel-Cepstral time-frequency transform to obtain acoustic feature vectors; The long short-term memory network is used to model the periodic fluctuation of pet physiological data and obtain physiological state encoding; Combining the facial expression features, the acoustic feature vectors, and the physiological state codes into initial multimodal features through a feature concatenation layer; The initial multimodal features are channel-weighted optimized by cascading SEblock arrays to generate multimodal fusion features.
[0008] Preferably, the step of locating the facial region based on the pet dynamic data by using a facial key point detection network to obtain facial expression features includes: performing global channel compression on the initial multimodal features through a first SEblock module in the cascaded SEblock array to generate a first compressed feature; superimposing a variety prior weight on the first compressed feature through a second SEblock module in the cascaded SEblock array to obtain a second compressed feature; Performing a residual connection on the second compressed features through the third SEblock module in the cascaded SEblock array, fusing the initial multimodal features with the secondary weighted features to obtain a preliminary fused feature; The preliminary fusion features are dimensionally compressed through a dimensionality reduction fully connected layer to generate multimodal fusion features.
[0009] Preferably, the step of performing deep learning model training based on the adversarial sample data and the multimodal fusion features to obtain a pet emotion recognition model includes: Generate training input data based on the structured dataset using a generative adversarial network, and perform temporal modeling on the behavior sequences in the training input data using a multi-head spatiotemporal Transformer to capture the emotion evolution pattern; Adjust the contribution ratio of classification loss and regression loss through dynamic weight strategy; The gradient clipping algorithm is used to limit the parameter update amplitude of the deep learning model to prevent overfitting of adversarial samples and preserve the optimal pet emotion recognition model.
[0010] Preferably, the step of generating training input data based on the structured data set by using a generative adversarial network comprises: synthesizing the instrument noise spectrum of the structured data set through a generative adversarial network to generate enhanced acoustic data; The pet's limb motion trajectory in the pet's dynamic data under the constraint state is simulated by the posture synthesis algorithm to generate adversarial visual data; Simulate the heart rate variability and skin conductance fluctuations under fear to generate adversarial physiological data; Synchronizing the adversarial visual data, the enhanced acoustic data, and the adversarial physiological data into adversarial sample data through a spatiotemporal alignment algorithm; The adversarial sample data is mixed with the multimodal fusion features in proportion to generate training input data.
[0011] Preferably, the step of performing emotion recognition on the pet to be recognized by using the pet emotion recognition model to obtain the pet emotion recognition result includes: The multimodal data collection system is used to collect real-time data of the pet to be identified to obtain data to be processed; Preprocessing the data to be processed to obtain standardized data to be identified; Performing multimodal feature extraction on the standardized data to be identified to obtain a feature vector to be identified; Inputting the feature vector to be identified into a pet emotion recognition model to obtain an emotion prediction result; Using a time series smoothing filter, filtering out short-term abnormal fluctuations in the emotion prediction results to obtain a smoothed emotion sequence; According to the smoothed emotion sequence, an emotion state analyzer is used to generate a pet emotion recognition result including basic emotions, compound emotion components and emotion intensity.
[0012] This application also provides a pet dog emotion recognition system based on deep learning, including: The acquisition module is used to collect pet dynamic data, pet physiological data, and situational scene data to obtain a structured data set; An annotation module is used to build a three-level emotion annotation framework and annotate the structured dataset through a cross-validation annotation mechanism to obtain an annotated dataset; An extraction module, configured to extract multimodal features based on the labeled data set, and integrate the multimodal features through a cascaded SEblock array to obtain multimodal fusion features; A training module is used to inject adversarial sample data generated by a stress scenario simulator during the training phase, perform deep learning model training based on the adversarial sample data and the multimodal fusion features, and obtain a pet emotion recognition model; The recognition module is configured as a pet emotion recognition model obtained through training by the training module, and is used to perform emotion recognition on the pet to be identified and obtain a pet emotion recognition result.
[0013] The present invention also provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned pet dog emotion recognition method based on deep learning are implemented.
[0014] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned pet dog emotion recognition method based on deep learning.
[0015] The beneficial effects of the present invention are as follows: The present invention proposes a pet dog emotion recognition method based on deep learning. Through the innovative application of multimodal data fusion and deep learning algorithms, it achieves high-precision recognition of pet emotions and has significant beneficial effects. First, the present invention collects and integrates pet dynamic data, physiological data and situational scene data, constructs a comprehensive structured data set, breaks through the limitations of traditional single-modality recognition, and can capture the complex expressions of pet emotions from a multi-dimensional perspective, greatly improving the comprehensiveness and accuracy of emotion recognition, enabling the system to distinguish subtle emotional changes and transition states, and providing a more comprehensive data foundation for pet behavior interpretation; secondly, the three-level emotion labeling framework designed by the present invention establishes an objective and systematic emotion labeling system through a layered progressive mechanism of basic emotion labeling, contextual association enhancement and physiological parameter mapping, significantly improving the quality of training data, solving the problems of strong subjectivity and poor consistency of traditional labeling methods, and providing high-quality supervision signals for deep learning models; thirdly, the present invention introduces a cascaded SEblock array to achieve effective fusion of multimodal features. , highlighting key features and suppressing irrelevant information through the channel attention mechanism, greatly enhancing the feature expression ability, enabling the model to more accurately capture the correlation pattern of pet facial expressions, sound characteristics and physiological states, and improving the accuracy of emotion recognition; in addition, the present invention injects adversarial sample data generated by the stress scene simulator during the training process, constructs a recognition model with strong generalization ability, enables the system to maintain stable recognition performance in complex and changeable actual application scenarios, and effectively solves the problem of decreased recognition accuracy of traditional models in new environments and special scenarios; in the specific application field of pet emotion recognition, the artificial intelligence algorithm design of the present invention is particularly prominent. The application of multi-headed spatiotemporal Transformer enables the model to effectively capture the temporal evolution law of pet emotions, and the design of the dynamic weight loss function balances the dual goals of emotion classification and physiological parameter regression. The innovative application of the algorithm features of the present invention greatly improves the performance upper limit of the recognition model, enabling the system to accurately identify the complex emotional states of pets in real time, providing a scientific basis for pet health monitoring and behavioral intervention, while building a smoother bridge for emotional communication between owners and pets, improving the human-pet interaction experience, and not only promoting the development of pet welfare science, but also providing a technical reference for similar multimodal emotional computing fields, demonstrating the huge potential and value of artificial intelligence technology in specific application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Schematic diagram of a method flow according to an embodiment of the present invention.
[0017] Figure 2 FIG. 1 is a schematic diagram of a system structure according to an embodiment of the present invention.
[0018] Figure 3This is a schematic diagram of the internal structure of a computer device according to an embodiment of the present application.
[0019] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0020] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0021] like Figure 1 As shown, the present application provides a pet dog emotion recognition method based on deep learning, comprising: S1. Collect pet dynamic data, pet physiological data, and scenario data to obtain a structured data set; S2. Construct a three-level emotion annotation framework, annotate the structured dataset through a cross-validation annotation mechanism, and obtain an annotated dataset; S3. Extracting multimodal features based on the labeled data set, and integrating the multimodal features through a cascaded SEblock array to obtain multimodal fusion features; S4. Injecting adversarial sample data generated by the stress scenario simulator during the training phase, performing deep learning model training based on the adversarial sample data and the multimodal fusion features, and obtaining a pet emotion recognition model; S5. Perform emotion recognition on the pet to be identified using the pet emotion recognition model to obtain a pet emotion recognition result.
[0022] As described in the above steps S1-S5, the present invention collects and integrates pet dynamic data, physiological data and situational scene data, constructs a comprehensive structured data set, breaks through the limitations of traditional single-modality recognition, and can capture the complex expressions of pet emotions from a multi-dimensional perspective, greatly improving the comprehensiveness and accuracy of emotion recognition, enabling the system to distinguish subtle emotional changes and transition states, and providing a more comprehensive data basis for pet behavior interpretation; secondly, the three-level emotion labeling framework designed by the present invention establishes an objective and systematic emotion labeling system through a layered progressive mechanism of basic emotion labeling, contextual association enhancement and physiological parameter mapping, significantly improving the quality of training data, solving the problems of strong subjectivity and poor consistency of traditional labeling methods, and providing high-quality supervision signals for deep learning models; thirdly, the present invention introduces a cascaded SEblock array to realize multimodal features The effective fusion of the channel attention mechanism highlights key features and suppresses irrelevant information, greatly enhancing the feature expression ability, enabling the model to more accurately capture the correlation pattern of pet facial expressions, sound characteristics and physiological states, and improving the accuracy of emotion recognition; in addition, the present invention injects adversarial sample data generated by the stress scene simulator during the training process to construct a recognition model with strong generalization ability, so that the system can maintain stable recognition performance in complex and changeable actual application scenarios, and effectively solves the problem of decreased recognition accuracy of traditional models in new environments and special scenarios; in the specific application field of pet emotion recognition, the artificial intelligence algorithm design of the present invention is particularly prominent. The application of the multi-headed spatiotemporal Transformer enables the model to effectively capture the temporal evolution of pet emotions, and the design of the dynamic weight loss function balances the dual goals of emotion classification and physiological parameter regression. The innovative application of the algorithm features of the present invention greatly improves the performance upper limit of the recognition model, enabling the system to accurately identify the complex emotional states of pets in real time, providing a scientific basis for pet health monitoring and behavioral intervention, while building a smoother bridge for emotional communication between owners and pets, improving the human-pet interaction experience, and not only promoting the development of pet welfare science, but also providing a technical reference for similar multimodal emotional computing fields, demonstrating the huge potential and value of artificial intelligence technology in specific application scenarios.
[0023] In one embodiment, the structured data set in step S1 specifically includes: S11, pet dynamic data, the pet dynamic data including pet visual data, pet sound data and pet behavior data; S12, pet physiological data, the pet physiological data including pet electrocardiogram data, pet respiratory rate data and pet skin conductivity data; S13. Scenario data, including light intensity data, ambient noise data, and the number of interactive objects.
[0024] As described in the above steps S11-S13, the present invention first comprehensively collects pet dynamic data, pet physiological data and situational scene data, so as to construct a complete structured data set. These data together constitute the multi-dimensional feature expression of the pet's emotional state, providing a rich and comprehensive data basis for subsequent emotion recognition. Pet visual data mainly captures the pet's facial expressions, body postures and other visual information through high-resolution camera equipment. During the acquisition process, it is necessary to ensure that the image clarity and frame rate are not less than 30fps to ensure that the pet's subtle facial expression changes and rapid action reactions can be captured. The collected visual data is stored in the form of RGB image sequences or video streams. Each frame of the image contains key information such as the pet's facial area and body posture. In the data processing stage, these original image data will be processed by algorithms such as facial area detection and key point positioning to extract facial expression features. For example, when the visual data of a golden retriever is collected, the spatial coordinates of 16 key points such as its eyes, ears, and mouth can be located through the facial key point detection algorithm. These coordinate data form a description of the pet's facial expression. The pet sound data is a variety of sound signals emitted by pets recorded by a high-sensitivity microphone array, including barking, whimpering, panting and other sound performances. In the specific acquisition process of this embodiment, the sampling rate of the microphone is set to 44.1kHz to ensure that the full spectrum range of pet sounds can be covered. The collected sound data is initially in the form of a time-domain audio waveform, which needs to be converted into a frequency-domain feature representation through the Mel-Cepstral time-frequency transform algorithm. The final result is an acoustic feature matrix containing multiple frames, each of which extracts multi-dimensional MFCC feature coefficients to form an acoustic feature matrix. These feature data can effectively characterize the acoustic characteristics of pet sounds, such as pitch and timbre. Pet behavior data is a record of behavioral information such as the pet's movement trajectory and activity pattern. It is obtained through an inertial measurement unit (IMU) sensor worn on the pet or video image processing. The IMU sensor can capture changes in the pet's movement state. The obtained behavior data is stored in the form of a time series, recording the pet's movement acceleration and angular velocity changes in three-dimensional space, and finally obtaining a feature vector describing the pet's behavior pattern.
[0025] Pet ECG signal data is collected by pet ECG monitoring equipment, which can capture the detailed features of the ECG waveform. The initial form of these ECG data is a time-domain electrical signal waveform, which needs to be processed by algorithms such as R-wave detection and heart rate variability analysis to extract cardiac activity characteristics such as heart rate and heart rate variability. Changes in cardiac activity characteristics reflect the activation state of the sympathetic nervous system. Pet respiratory rate data is collected by infrared thermal imaging or respiratory belts and other equipment to record changes in the pet's respiratory rhythm. It is stored in the form of a respiratory waveform and calculated through a peak detection algorithm. Features such as respiratory frequency and depth are calculated. The respiratory frequency is calculated by the number of respiratory cycles detected in unit time, and the respiratory depth is represented by the peak-to-valley difference of the respiratory waveform. For example, when a dog is relaxed, its breathing rate may be 15-20 times / minute, and its breathing depth is stable; when it feels fear, its breathing rate may increase to 30-40 times / minute, and its breathing becomes shallow and rapid. Pet skin conductivity data is collected through electrodes attached to the pet's paw pads or other low-hair areas. Skin conductivity reflects the activity of skin sweat glands and the activation level of the sympathetic nervous system. It is an important indicator for assessing emotional arousal. The data is stored in the form of time series, recording changes in skin conductance levels and skin conductance responses.
[0026] Scenarios data is an objective record of the pet's environment, mainly including light intensity data, ambient noise data and the number of interactive objects. Light intensity data is collected by light sensors, which measure the brightness level of ambient light and record changes in ambient light intensity. In the data processing process, the present invention takes into account the diurnal variation of light intensity and calculates characteristics such as the rate of light change. For example, when light intensity suddenly increases from normal indoor levels to strong light conditions, it may cause a stress response in pets. Ambient noise data is collected by ambient sound monitoring equipment, which records the sound intensity and spectrum distribution in the environment. The present invention has a good understanding of the environmental noise data. During the data processing process, the noise intensity of different frequency bands was analyzed, and the frequency and duration of noise events were calculated. The differences in the sensitivity of pet ears to sounds of different frequencies were taken into account. The number of interactive objects refers to the number of people or other animals that interacted with the pet. These data were obtained through video image analysis and recorded the changes in the number of interactive objects around the pet. During the data processing process, the present invention identified the identity and position of different interactive objects and calculated characteristics such as the frequency and intensity of interaction. For example, when the number of people around a less social dog increases from 1 to 5, it may be observed that it changes from a relaxed state to a tense or vigilant state.
[0027] In one embodiment, the step S2 of annotating the structured dataset using a cross-validation annotation mechanism to obtain an annotated dataset includes: S21. Labeling the pet's dynamic data using a labeling tool to obtain basic emotion labels, wherein the basic emotion labels include anxiety, joy, and fear; S22, associating the scenario data with the basic emotion tag through a context association algorithm to generate a context enhancement tag; S23. Calculating the correlation coefficient between the pet's physiological data and the emotional state corresponding to the basic emotion label based on the physiological parameter mapping model to obtain a physiological emotion weight matrix; S24, cross-validating the basic emotion label, the situational enhancement label, and the physiological emotion weight matrix through a crowdsourcing verification platform, eliminating data with a labeling confidence level lower than a threshold, and obtaining a cross-validation result; S25. Generate a labeled data set containing multiple layers of labels according to the cross-validation result.
[0028] As described in the above steps S21-S25, in the present invention, the annotation tool described in step S21 refers to a software platform for annotating video and audio data, which can support manual annotators to encode and classify the facial expressions, body postures, and sound features of pets. In the specific annotation process, the annotator will observe the dynamic behavior data of the pet, including changes in facial micro-expressions, changes in the posture of the ears and tail, changes in body posture, and changes in the tone, volume, and timbre in the sound data. Based on these observations, the annotator will annotate each data segment as a basic emotion category, namely anxiety, pleasure, or fear. For example, when the annotator observes that a pet dog has its ears tilted back, its eyes wide open, the corners of its mouth tightened, and it makes a low barking sound, it will annotate the current emotion category as anxiety, pleasure, or fear. The behavior is labeled as "fear" emotion; when the pet is observed to have a relaxed body posture, a wagging tail, a relaxed facial expression, and cheerful barking, it will be labeled as "joyful" emotion; when the pet is observed to pace back and forth, lick its body, and make short, high-frequency barking sounds, it will be labeled as "anxious" emotion. The labeling tool will record the label at each time point, and the labeler will also add a confidence score to the label to indicate the degree of confidence in the label. The context association algorithm in step S22 is used to analyze the relationship between environmental factors and emotional expression. Specifically, the context scene data (including light intensity data, environmental noise data, and the number of interactive objects) is first time-aligned with the basic emotion label, and then the correlation coefficient between each context factor and the emotion label is calculated. The context association algorithm uses a conditional probability distribution model to calculate the probability of a certain emotion appearing under specific context conditions. The algorithm formula is: ; in, For the situation Pets express emotions The conditional probability of Showing Emotions for Pets When the scene appears The conditional probability of Express emotions The prior probability of Indicates the situation The prior probability of, through the above formula algorithm, can identify which environmental factors are strong triggers of specific emotions. Based on the above association analysis, the algorithm will add contextual modifiers to the original basic emotion labels to form contextual enhancement labels, such as "light-induced anxiety", "pleasure under multi-person interaction" or "fear caused by sudden noise", etc. The obtained contextual enhancement labels not only record the emotion type, but also capture the environmental inducements that lead to the emotion, thereby providing richer semantic information. The physiological parameter mapping model in step S23 is a multivariate statistical model used to quantify the correspondence between physiological signals and emotional states. Specifically, in step S23, the present invention first extracts time domain and frequency domain features of the pet's electrocardiogram signal data, respiratory rate data and skin conductivity data, and then calculates the correlation strength of these features with each emotion category through regression analysis. Through the calculation of the physiological parameter mapping model, the correlation between each physiological parameter and each emotion category is quantified as a correlation coefficient, and the obtained coefficients together constitute the physiological emotion weight matrix. The calculated correlation coefficients are organized into the physiological emotion weight matrix for subsequent emotional state evaluation and verification. The crowdsourcing verification platform in step S24 is a platform that integrates the professional knowledge of many professional animal behaviorists, veterinarians and trainers, and is used to verify the automatically generated labeling results. The cross-validation process in step S24 adopts a multi-expert voting mechanism. Each data sample is evaluated by at least 3 independent experts. The experts will evaluate the consistency of three types of labeling results: whether the basic emotion label accurately reflects the actual emotional state of the pet; whether the association between environmental factors and emotions in the situational enhancement label is reasonable; whether the correlation coefficient in the physiological emotion weight matrix is consistent with the known animal behavior theory. For each sample, the expert will give a confidence score and set a confidence threshold. Sample labels below the confidence threshold are considered unreliable and will be eliminated or Marked as needing to be re-labeled, the multi-level label structure in step S25 is a hierarchical emotion representation method, which includes three levels: the first level is the basic emotion category (anxiety, pleasure, fear); the second level is the complex emotion modified by the situation (such as "light-induced anxiety"); the third level is the detailed emotional state description with intensity quantification and physiological association. The structural design of the labeled dataset follows the JSON format. Each sample contains the following fields: sample ID, timestamp, basic emotion label, situational enhancement label, emotion intensity value (floating point number between 0-1), physiological parameter association intensity vector (the association strength of each physiological parameter with the emotion), and confidence score. The multi-level label generation process uses the decision tree algorithm to integrate the cross-validated data into a hierarchical label system. The construction of the decision tree is based on the information gain criterion, and features with strong distinguishing ability are preferentially selected as split nodes. Finally, a complete multi-level labeled dataset is formed, which includes both discrete emotion categories and continuous emotion intensity and the influence of environmental factors.
[0029] In one embodiment, the step S3 of extracting multimodal features based on the labeled data set and integrating the multimodal features through a cascaded SEblock array to obtain multimodal fusion features includes: S31, locating the facial region based on the pet dynamic data using a facial key point detection network to obtain facial expression features; S32, performing audio analysis based on the pet dynamic data by Mel-Cepstral time-frequency transform to obtain an acoustic feature vector; S33, modeling the periodic fluctuation of the pet's physiological data through a long short-term memory network to obtain a physiological state code; S34, combining the facial expression feature, the acoustic feature vector, and the physiological state code into an initial multimodal feature through a feature concatenation layer; S35. Optimize the channel weights of the initial multimodal features by cascading the SEblock array to generate multimodal fusion features.
[0030] As described in the above steps S31-S35, the present invention uses a facial key point detection network in step S31 to detect and locate the key feature points of the pet's face. The network input is an image frame sequence of the pet's facial area, and the output is the spatial coordinates of the facial key points. The network first uses a cascaded convolutional layer to extract the multi-scale features of the image, and then generates a facial heat map through a fully convolutional network. Finally, the precise position of the key points is determined by coordinate regression. In a specific embodiment, for the detection object pet of the present invention, 68 facial key points are detected, including eye contours (12 points), nose contours (9 points), mouth contours (20 points), ear contours (12 points), and facial contours (15 points). For canine pets, 56 facial key points are detected, including eye contours (12 points), nose contours (9 points), mouth contours (15 points), ear contours (10 points), and face contours (10 points). The detection accuracy of facial key points reaches the pixel level, and the average detection error is controlled within the range of 2-3 pixels. After obtaining the coordinates of facial key points, the present invention further extracts facial expression features, including geometric features and appearance features. Geometric features mainly describe the spatial relationship between key points, such as the distance between the eyes, the degree of eye opening, the degree of mouth opening, the tilt angle of the ear, etc. The appearance features are extracted through texture descriptors such as local binary patterns and directional gradient histograms to capture subtle facial texture changes. The geometric features and appearance features are spliced into a high-dimensional directional feature. The Mel-Cepstral time-frequency transform described in step S32 of the present invention is a technology specifically used for sound signal processing. It combines the advantages of Mel frequency scale and cepstrum analysis and can effectively extract the spectral characteristics of sound signals. The processing flow first pre-processes the original audio signal, including noise reduction, framing and windowing. Each frame of the pre-processed signal is sent to the short-time Fourier transform module to calculate its power spectrum. The power spectrum is then transformed by the Mel filter bank to convert the linear frequency scale into a Mel frequency scale that is more in line with human auditory perception. Next, the Mel spectrum is logarithmic and subjected to discrete cosine transform to obtain Mel frequency cepstrum coefficients (MFCCs). The specific expression is: ; in, Indicates the MFCC coefficients, Indicates the Mel filter output, is the Mel filter index, is the total number of Mel filters. In this embodiment, the first 13 MFCC coefficients are selected, and their first-order and second-order difference coefficients are calculated to form a 39-dimensional basic feature vector. For pet sound analysis, acoustic features including fundamental frequency, harmonic noise ratio, jitter rate and flicker rate are also calculated. They can reflect the pitch, clarity and stability of pet sounds. Finally, all these acoustic features are combined into a complete acoustic feature vector for subsequent emotion analysis; in step S33, the present invention uses a long short-term memory network (LSTM) to capture the long-term dependencies in the pet physiological data sequence data. For pet physiological data modeling, the input of LSTM is the time series of pet electrocardiogram signal data, respiratory rate data and skin conductivity data; the LSTM network in this embodiment adopts a bidirectional structure, taking into account past and future context information at the same time. The network includes two layers of LST M, 128 hidden units per layer, and an input sequence length of 60 seconds (physiological data collected at a sampling rate of 20 Hz); the physiological data is first normalized and then segmented into sequences of fixed length according to time windows. The LSTM network learns to capture the temporal dynamic characteristics of physiological signals, such as heart rate variability patterns, respiratory rate change trends, and skin conductivity fluctuations. The final hidden state of the network is used as a physiological state code, which is a 256-dimensional vector (the hidden state concatenation of the forward LSTM and the backward LSTM) that contains a compact representation of the pet's physiological state. In step S34, the present invention combines facial expression features, acoustic feature vectors, and physiological state codes into initial multimodal features through a feature concatenation layer. The feature concatenation layer is a simple and effective multimodal fusion method that directly concatenates feature vectors of different modalities in the feature dimension. The concatenation operation can be expressed as: ; in, Represents the initial multimodal features after splicing, with a dimension of , is the facial expression feature dimension, is the acoustic feature vector dimension, Encoding dimensions for physiological states, Indicates facial expression features; represents the acoustic eigenvector, Indicates physiological state code, Represents a vector splicing operation; in a specific line of sight, the facial expression feature dimension is 512, the acoustic feature vector dimension is 128, the physiological state encoding dimension is 256, and the initial multimodal feature dimension after splicing is 896. Since the scales of different modal features in direct splicing may be inconsistent, the contributions of different modalities should be different, and the interaction between modalities is not modeled. Therefore, before splicing, the present invention inputs each modal feature into a linear projection layer, maps it to the same feature space, and performs L2 normalization to ensure that the scales of different modal features are consistent; in step S35, the present invention optimizes the channel weights of the initial multimodal features through a cascaded SEblock array to generate multimodal fusion features. The SEblock (Squeeze-and-ExcitationBlock) is a channel attention mechanism for adaptively calibrating the importance of channel features. The cascaded SEblock array set by the present invention is to combine multiple SEblock modules in series to form a deep processing pipeline, which can gradually refine the channel weights and adapt to feature representations at different levels.
[0031] In one embodiment, the step S35 of locating the facial region based on the pet dynamic data using the facial key point detection network to obtain facial expression features includes: S351, performing global channel compression on the initial multimodal features through the first SEblock module in the cascaded SEblock array to generate a first compressed feature; S352: superimpose a variety priori weight on the first compression feature through the second SEblock module in the cascaded SEblock array to obtain a second compression feature; S353: Performing a residual connection on the second compressed features through the third SEblock module in the cascaded SEblock array, fusing the initial multimodal features with the secondary weighted features to obtain a preliminary fused feature; S354. Perform dimension compression on the preliminary fusion features through a dimensionality reduction fully connected layer to generate multimodal fusion features.
[0032] As described in steps S351-S354 above, in step S351 of the present invention, the workflow of the first SEblock module includes two stages: compression and excitation. In the compression stage, global average pooling is performed on each channel feature, compressing the spatial dimensions into a single value to obtain a channel descriptor. In the excitation stage, a two-layer fully connected network learns the nonlinear relationship between channels and generates channel weights. Finally, the channel weights are applied to the original features to obtain weighted features. The first SEblock module processes the entire initial multimodal feature without distinguishing modal sources. Its main purpose is to suppress noisy channels and enhance channels with high information content. In step S352 of the present invention, the second SEblock module introduces breed prior knowledge into its basic structure. Different pet breeds have differences in facial expressions, vocal characteristics, and physiological reactions, and these differences affect the patterns of emotional expression. For example, the facial expressions of flat-faced dog breeds (such as Pugs) are relatively difficult to identify, while the ear movements of upright-eared dog breeds (such as German Shepherds) are more expressive of emotions. To incorporate this prior knowledge, the second SEblock module introduces a breed encoding vector , and concatenate it with the channel descriptor of the first compressed feature. Then, the channel weight is calculated based on the concatenated feature. Finally, the variety-aware channel weight is applied to the first compressed feature to obtain the second compressed feature. The specific expression is: ; in, To obtain the second compression feature, is the channel dimension, is the sigmoid activation function, is the ReLU activation function, is the weight matrix of the first fully connected layer, is the weight matrix of the second fully connected layer, For the Descriptors for each channel, is the first compressed feature; the variety encoding vector It is obtained by querying a breed embedding table, which contains the codes of 50 common pet breeds. The introduced breed code vector enables the model to adjust the importance of feature channels for pets of different breeds, thereby improving the adaptability of the model. In step S353 of the present invention, the third SEblock module introduces a residual connection mechanism to solve the gradient vanishing problem in deep feature extraction while retaining the original feature information. The residual connection is implemented by adding the second compressed feature to the initial multimodal feature at the element level, and then the residual feature is further optimized through the third SEblock for channel weight optimization. The significance of step S353 is to fuse the results of multiple rounds of channel weight adjustment with the original feature, which not only retains the original information but also incorporates multi-level channel importance evaluation. In step S354 of the present invention, the preliminary fusion feature is dimensionally compressed through a dimensionality reduction full connection layer to generate a multimodal fusion feature. The preliminary fusion feature has a high dimension and there is redundant information, so dimensionality compression is required. The dimensionality reduction full connection layer is a linear transformation with a weight matrix and a bias vector, and finally a multimodal fusion feature is obtained. Specifically, the dimension of the initial fused features is 896, which is reduced to 256 dimensions after the dimensionality reduction fully connected layer. The dimensionality reduction not only reduces the computational complexity, but also fuses multimodal information through the learning weight matrix to generate a more compact and information-rich representation.
[0033] In one embodiment, the step S4 of performing deep learning model training based on the adversarial sample data and the multimodal fusion features to obtain a pet emotion recognition model includes: S41. Generate training input data based on the structured data set through a generative adversarial network; S42, performing temporal modeling on the behavior sequence in the training input data through a multi-head spatiotemporal Transformer to capture the emotion evolution pattern; S43. Adjust the contribution ratio of classification loss and regression loss through dynamic weight strategy; S44. Limit the parameter update amplitude of the deep learning model through the gradient clipping algorithm to prevent overfitting of the adversarial sample and save the optimal pet emotion recognition model.
[0034] As described in steps S41-S44 above, the GAN in step S41 of the present invention is a deep learning architecture consisting of a generator and a discriminator. The generator is responsible for generating realistic samples, while the discriminator is responsible for distinguishing between real samples and generated samples. In the pet emotion recognition task, the GAN first learns the data distribution characteristics in the structured dataset and then generates samples that are similar to the real data but contain specific variations. The obtained samples can simulate scenarios that may occur in reality but are not fully expressed in the original dataset. The multi-head spatiotemporal Transformer described in step S42 of the present invention is an extension of the standard Transformer architecture for processing sequential data with spatiotemporal relationships. In the pet emotion recognition task, the pet's emotional state is not static, but a process that changes dynamically over time. Therefore, it is necessary to perform temporal modeling on the pet behavior sequence. The workflow of the multi-head spatiotemporal Transformer includes: first, the training input data is divided into time segments of fixed length, each segment contains multimodal features, and then temporal information is added to each time step through position encoding. Then, the multi-head self-attention mechanism is used to calculate the association weights between different time steps to capture long-range temporal dependencies. Finally, the contextual information is integrated through a feedforward neural network to generate a feature representation containing temporal semantics. The multi-head mechanism of the multi-head spatiotemporal Transformer allows the model to focus on information in multiple subspaces simultaneously, thereby being able to capture different types of temporal patterns, such as short-term emotional fluctuations, medium-term emotional trends, and long-term emotional baseline changes. In the pet emotion recognition task in step S43 of the present invention, the model needs to simultaneously complete two tasks: discrete emotion classification (such as judging whether it is anxiety, pleasure, or fear) and physiological parameter regression prediction (such as predicting heart rate level, respiratory rate, etc.). Therefore, the total loss function contains two parts: classification loss and regression loss. The dynamic weight strategy is a method that dynamically adjusts the weights of these two losses during the training process. Its core idea is to pay more attention to basic classification capabilities in the early stages of training, and gradually increase the weight of regression loss as training progresses, thereby improving the model's prediction accuracy for physiological parameters. Through the dynamic adjustment strategy, the model can gradually balance the classification ability and regression ability during the training process, and ultimately achieve accurate prediction of emotional labels and physiological parameters; wherein, the loss function of the pet emotion recognition model is expressed as: ; ; ; ; in, is the total loss function, is the dynamic weight, is the discrete sentiment classification loss, is the physiological parameter regression loss, is the initial classification loss weight, is the minimum classification loss weight, is the weight attenuation coefficient, is the training round, is the sample index value, is the total number of samples, is the sentiment category index value, is the total number of emotion categories, is the smoothed label, For the samples belong to the category The predicted probability of is the physiological parameter index value, is the total number of physiological parameters, For the The weight parameters corresponding to the physiological parameters are: The model predicts The first sample The parameter value of each physiological parameter, For the The first sample The labeled values of physiological parameters are obtained. In step S44, the present invention limits the parameter update amplitude of the deep learning model through the gradient clipping algorithm to prevent overfitting of adversarial samples and save the optimal pet emotion recognition model. Gradient clipping is a regularization technology for preventing gradient explosion and model overfitting. It is particularly suitable for the training scenarios of the present invention containing adversarial samples. In the gradient clipping algorithm, when the norm of the calculated gradient vector exceeds the preset threshold, the gradient vector will be scaled down so that its norm does not exceed the threshold. Through gradient clipping, the update amplitude of the model parameters is limited to a reasonable range, avoiding excessive adjustment due to extreme features in the adversarial samples, thereby improving the generalization ability of the model and its adaptability to new data. During the training process, the model performance is regularly evaluated on the validation set, and the model parameters with the best validation performance are saved, and finally the optimal pet emotion recognition model is obtained.
[0035] In one embodiment, the step S41 of generating training input data based on the structured dataset by using a generative adversarial network includes: S411. Synthesizing the instrument noise spectrum of the structured data set using a generative adversarial network to generate enhanced acoustic data. S412, simulating the pet's limb motion trajectory in the pet's dynamic data in the constrained state by a posture synthesis algorithm to generate adversarial visual data; S413, simulating the heart rate variability and skin conductivity fluctuations under fear emotion to generate adversarial physiological data; S414, synchronizing the adversarial visual data, the enhanced acoustic data, and the adversarial physiological data into adversarial sample data through a spatiotemporal alignment algorithm; S415: Mix the adversarial sample data with the multimodal fusion features in proportion to generate training input data.
[0036] As described in steps S411-S415 above, in step S411 of the present invention, the instrument noise spectrum synthesis refers to the process of simulating the noise generated by various instruments and equipment and superimposing it on the original pet sound data according to a specific spectral distribution. Specifically, the original pet sound data is first subjected to a short-time Fourier transform to obtain a sound spectrum in the time-frequency domain; then, a generative adversarial network is used to generate spectral features that simulate instrument noise, such as the humming sound of medical equipment, interference sound from electronic equipment, or mechanical noise in the environment; finally, the generated noise spectrum is weightedly mixed with the original sound spectrum and converted back into a time domain signal through an inverse short-time Fourier transform to form enhanced acoustic data. The obtained enhanced acoustic data enables the model to effectively recognize the sound signals emitted by pets in noisy environments, thereby improving the model's emotion recognition accuracy in complex acoustic environments. In step S412 of the present invention, the present invention first extracts a sequence of skeletal key points from the pet's dynamic data, and then performs a geometric transformation on these key points through a spatial transformer network (STN) to generate a representation of the possible action responses of the pet under different constraint states. For example, for pets confined in cages or restricted by leashes, their body movements are significantly restricted. The posture synthesis algorithm simulates the possible posture changes a pet might make under these restricted conditions and generates corresponding visual data. The synthesized visual data includes subtle visual cues of the pet's emotions under restraint, such as changes in ear position, tail wagging patterns, or eye gaze behavior, thereby helping the model learn to recognize how pets express emotions under different restraints. In step S413, the present invention primarily simulates fluctuations in heart rate variability and skin conductivity under fear to generate adversarial physiological data. Heart rate variability (HRV) measures the variation in the time intervals between consecutive heartbeats, while skin conductivity reflects the activity level of the sympathetic nervous system. When in a state of fear, a pet's heart rate variability decreases (heartbeat interval variation decreases), while skin conductance increases. The process of generating counteracting physiological data includes: first, establishing a mathematical model of the physiological response to fear based on medical literature. This model describes the decreasing pattern of heart rate variability and the increasing characteristics of skin conductance under the fear state; then, introducing a random perturbation factor to simulate individual differences and environmental influences; and finally, generating time-series physiological signal data, including the changing pattern of RR periods in electrocardiogram (ECG) data and the dynamic changes of skin conductance level (SCL). The synthesized physiological data enables the model to better understand the physiological indicator characteristics under fear and improve the accuracy of fear recognition. The spatiotemporal alignment algorithm described in step S414 of the present invention is a multimodal data synchronization technology used to ensure that data from different sources and sampling rates are aligned and consistent in the time dimension.Specifically, each modal data is first timestamped, then the optimal alignment path between the different modal data is calculated using a dynamic time warping method. Finally, the temporal resolution of each modal data is adjusted through interpolation or downsampling methods to ensure complete alignment in the temporal dimension. Furthermore, the spatiotemporal alignment algorithm also considers the causal relationship between the different modal data, such as the time delay between a visually observed pet's sudden startle behavior and a sudden increase in heart rate. Through the above-mentioned precise spatiotemporal alignment, the generated adversarial sample data can accurately reflect the temporal correlation of multimodal signals during emotional changes, thereby helping the model learn cross-modal consistency features of emotional expression. In step S415, the present invention adopts a proportional mixing strategy through the mixing process, that is, according to a preset mixing ratio, the adversarial sample data and the multimodal fusion features are weighted and combined. This proportional mixing method can effectively control the impact of adversarial samples on model training. By adjusting the ratio value, a balance can be achieved between the model's adversarial robustness and basic recognition performance. When the ratio value is small, training focuses more on basic recognition ability, while when the ratio value is large, it emphasizes the model's ability to resist adversarial interference.
[0037] In one embodiment, the step S5 of performing emotion recognition on the pet to be recognized by using the pet emotion recognition model to obtain the pet emotion recognition result includes: S51, collecting real-time data of the pet to be identified through a multimodal data collection system to obtain data to be processed; S52, preprocessing the data to be processed to obtain standardized data to be identified; S53, performing multimodal feature extraction on the standardized data to be identified to obtain a feature vector to be identified; S54, inputting the feature vector to be identified into a pet emotion recognition model to obtain an emotion prediction result; S55, filtering the short-term abnormal fluctuations in the emotion prediction results through a time series smoothing filter to obtain a smoothed emotion sequence; S56. Generate a pet emotion recognition result including basic emotions, compound emotion components and emotion intensity through an emotion state analyzer according to the smoothed emotion sequence.
[0038] As described in the above steps S51-S56, in step S51 of the present invention, the multimodal data acquisition system refers to a comprehensive data acquisition platform composed of multiple sensor devices, which is used to simultaneously obtain multi-dimensional information such as vision, sound, behavior and physiology of the pet. The device collected in step S1 can be the same device. The collected multimodal data is sent to the data processing unit in real time through wireless transmission technology to form a complete data set to be processed. In step S52 of the present invention, data preprocessing is the process of converting the original sensor data into a standard format suitable for deep learning model processing, including operations such as data cleaning, format conversion, time alignment and data standardization. First, data cleaning is performed on each modality data to remove obvious noise and outliers. For visual data, a median filtering algorithm is used to remove noise in the video frame, and inter-frame interpolation is performed to supplement the frames that may be lost; for sound data, a bandpass filter (100Hz-10kHz) is used to filter out ultra-low frequency and ultra-high frequency noise that are inaudible to the human ear, and background noise is suppressed by spectral subtraction; for physiological data, wavelet transform is applied to remove baseline drift and high-frequency interference. Secondly, format conversion is performed to uniformly convert raw data of different formats into a standard data format, such as converting video data into a tensor sequence, audio data into a Mel-spectrogram, and physiological data into a normalized time series. Then, the data of different modalities are aligned in the time dimension through a timestamp alignment algorithm to ensure the consistency of each modal data at the time point. Finally, each modal data is standardized so that its distribution characteristics are suitable for the processing requirements of the deep learning model. In step S53, the present invention first performs feature extraction on the visual data, uses a pre-trained convolutional neural network (CNN) model to extract the facial expression features of the pet, and uses an optical flow algorithm to calculate the motion features between consecutive frames to capture the dynamic behavior information of the pet. Furthermore, the facial expression feature extraction process includes three steps: facial area detection, key point positioning, and expression encoding. The 2048-dimensional facial expression feature vector is extracted through the first five convolutional layers of the ResNet-50 network. Secondly, feature extraction is performed on the sound data, and the Mel-frequency cepstral coefficient (MFCC) algorithm is used to extract the acoustic features of the sound, while calculating the time domain features such as the energy and fundamental frequency of the sound. The MFCC feature extraction process first pre-emphasizes the input audio, then frames and windows it. A fast Fourier transform (FFT) is performed on each frame to obtain a spectrum. This spectrum is then converted to the Mel scale, the logarithm of the Mel spectrum is taken, and finally a discrete cosine transform (DCT) is performed to obtain MFCC coefficients. For physiological data, a wavelet transform is used to extract frequency domain features of the ECG signal. Heart rate variability metrics (such as SDNN, RMSSD, and pNN50) are calculated, and respiratory rate and skin conductance level variation patterns are extracted. The extracted modal features are combined through feature concatenation to form a multimodal feature vector.In order to reduce the feature dimension and improve the feature quality, principal component analysis (PCA) is applied to perform feature dimensionality reduction, and the principal components with an explained variance ratio of 95% are retained to obtain the final feature vector to be identified. The feature vector to be identified obtained in step S53 is input into the pet emotion recognition model obtained in step S4, and then the model outputs the emotion category probability distribution and the emotion intensity prediction value. The emotion prediction result includes the probability distribution of basic emotion categories (such as anxiety, pleasure, fear, etc.) and the corresponding emotion intensity value (range 0-1), and also includes the predicted values of physiological parameters (such as heart rate, respiratory rate, etc.). The above prediction results constitute the original emotion prediction results. In step S55, the present invention uses a time series smoothing filter to filter the short-term abnormal fluctuations in the emotion prediction result. After the time series smoothing filtering process, the original emotion prediction result is converted into a smoother and more coherent smooth emotion sequence, which effectively reduces the interference of random fluctuations in the model prediction process on the emotion recognition result. In step S56 of the present invention, the emotion state parser is a multi-level emotion parsing system for converting the smooth emotion sequence into a pet emotion recognition result with rich semantic information. First, the emotional state parser maps smoothed emotional sequences into a three-dimensional emotional space based on the Valence-Arousal-Dominance (VAD) emotion model. Valence represents the positivity of an emotion, arousal represents the level of activation, and dominance represents the sense of control. Through this mapping, the pet's emotional state is represented as a point in the three-dimensional emotional space, facilitating further analysis and quantification. Secondly, the basic emotion category is determined based on the position in the VAD space. Basic emotions include anxiety, joy, and fear, and each basic emotion corresponds to a specific region in the VAD space. At the same time, the distance between the emotion vector and each basic emotion prototype vector is calculated to determine the emotion intensity. The closer the distance, the stronger the emotion. For the recognition of compound emotions, the emotion state parser uses fuzzy logic method to calculate the membership of the emotion vector to each basic emotion. The basic emotion combination with a membership higher than the threshold constitutes a compound emotion. The membership calculation uses Gaussian membership function. Finally, the emotion state parser generates a detailed emotion recognition result report, including the dominant basic emotion category, compound emotion composition (the proportion of each basic emotion), emotion intensity value (scalar in the range of 0-1) and the temporal stability evaluation of the emotion state. The multi-level emotion analysis method can provide a richer and more detailed description of the pet's emotional state than simple emotion classification, and can provide valuable emotion reference information.
[0039] The combined steps of the preceding S5 phase constitute a complete pet emotion recognition application process. From multimodal data acquisition to the final output of emotion results, each step is implemented with clear data processing logic and algorithms. For example, when identifying the emotional state of a domestic Labrador retriever during a veterinary clinic visit, the multimodal data acquisition system first collects relevant data: a camera captures video of the dog's facial expressions, ear posture, and body movements; a microphone records its occasional low whimpers; and a physiological sensor worn around the neck detects an increase in heart rate from a normal 70 beats / minute to 120 beats / minute, respiratory rate from 18 beats / minute to 35 beats / minute, and skin conductance from a baseline value of 2.1 μS to 4.8 μS. These raw data are wirelessly transmitted to the processing unit, forming the initial dataset to be processed. During the data preprocessing phase, the video data undergoes a median filter to remove noise caused by unstable lighting. The audio data undergoes a bandpass filter to remove high-frequency medical equipment noise in the clinic environment. The physiological data undergoes a wavelet transform to remove baseline drift caused by the dog's movement. Subsequently, data from different modalities were precisely aligned using timestamps and Z-score normalized separately to ensure that the distribution characteristics of each type of data met the model input requirements. For example, after standardization, the raw heart rate data of 120 beats / minute was converted to a standardized value of 1.83 (assuming a mean heart rate of 70 and a standard deviation of 27.3). In the feature extraction stage, significant facial expression features were extracted from the video data, such as slightly squinting eyes (eye openness reduced by 20%), posterior ear pressure (ear angle changed by 32 degrees), and tightened corners of the mouth. 13-dimensional MFCC features were extracted from the sound data, showing characteristics of enhanced low-frequency energy and decreased fundamental frequency of the sound. Features such as a significant decrease in heart rate variability (SDNN value dropped from 35ms to 12ms) and an increase in the rapid response component of skin conductance were extracted from the physiological data. After feature concatenation and PCA dimensionality reduction, the multimodal features formed a 128-dimensional feature vector to be identified. These feature vectors were input into a pre-trained pet emotion recognition model. Through multi-level feature fusion and time series modeling, the model output the corresponding emotion prediction results: the probability of anxiety is 0.68, the probability of fear is 0.25, the probability of pleasure is 0.07, and the emotion intensity is 0.75 (range 0-1). Since the dogs' emotions fluctuated during the diagnosis and treatment process, the prediction results at certain moments showed short-term abnormal fluctuations, such as a sudden prediction of a happy emotion (probability 0.85) lasting for 2 seconds, which was obviously inconsistent with the overall situation.Using a time series smoothing filter (with a smoothing factor of α = 0.15), these short-term abnormal fluctuations are effectively smoothed, resulting in a more coherent emotional sequence. The corrected emotional probabilities become: Anxiety 0.72, Fear 0.21, and Pleasure 0.07, with emotional intensity fluctuating within the range of 0.7-0.8. Finally, the emotional state parser, based on the smoothed emotional sequence, maps the emotional state into VAD space, resulting in a Valence value of -0.6 (negative emotion), Arousal value of 0.75 (high activation), and Dominance value of -0.3 (medium-low sense of control). Based on these parameters, the parser determines that the dog's dominant basic emotion is Anxiety, which also contains some fear, forming a composite emotion with anxiety accounting for 75%, fear accounting for 25%, and an emotional intensity of 0.73 (medium-high intensity). The final pet emotion recognition results showed that the Labrador dog was in a moderate to high intensity anxiety-fear complex emotional state in the veterinary clinic environment, mainly manifested by multimodal characteristics such as facial tension, pressure behind the ears, low whimpering, increased heart rate and enhanced skin conductance. The emotion recognition results provide veterinarians and pet owners with objective and accurate emotional reference information, which helps to adjust diagnosis and treatment methods and improve the pet's medical experience.
[0040] like Figure 2 As shown, the present invention also provides a pet dog emotion recognition system based on deep learning, comprising: The acquisition module is used to collect pet dynamic data, pet physiological data, and situational scene data to obtain a structured data set; An annotation module is used to build a three-level emotion annotation framework and annotate the structured dataset through a cross-validation annotation mechanism to obtain an annotated dataset; An extraction module, configured to extract multimodal features based on the labeled data set, and integrate the multimodal features through a cascaded SEblock array to obtain multimodal fusion features; A training module is used to inject adversarial sample data generated by a stress scenario simulator during the training phase, perform deep learning model training based on the adversarial sample data and the multimodal fusion features, and obtain a pet emotion recognition model; The recognition module is configured as a pet emotion recognition model obtained through training by the training module, and is used to perform emotion recognition on the pet to be identified and obtain a pet emotion recognition result.
[0041] The present invention overcomes the limitations of a single modality by simultaneously collecting dynamic visual / sound data, physiological signals, and environmental scene information (such as electrocardiogram, noise intensity, etc.). Compared with traditional methods that rely solely on audio or images (which have problems with environmental interference or individual differences), the present invention optimizes channel weights through a cascaded SEblock array, effectively integrating facial expressions, acoustic features, and physiological state encoding. The triple data collaboration of the present invention also further enhances the integrity of emotional representation.
[0042] This invention employs a cross-validation mechanism combining basic emotion labels, context-enhanced labels, and a physiological weight matrix to address the subjectivity inherent in traditional manual labeling. Specifically, this invention utilizes a crowdsourcing platform to eliminate low-confidence data (e.g., the issue of insufficient labeled samples mentioned above) and, combined with a priori weighting adjustments for breeds, effectively mitigates the interference caused by differences in emotional expression among pet breeds. Compared to existing single-labeling methods (e.g., sound classification alone), this framework makes the labeling system more context-sensitive.
[0043] During the training phase, adversarial samples such as instrument noise spectra and simulated limb movement trajectories are injected, and overfitting is prevented through gradient clipping, which significantly improves the generalization ability of the model in complex environments (such as high noise and occlusion scenes). Adversarial training not only improves the verification accuracy, but also realizes the synchronous generation of multimodal adversarial data through the spatiotemporal alignment algorithm, further enhancing the model's adaptability to sudden stress scenarios.
[0044] A dynamic weighting strategy for classification and regression losses is employed, initially focusing on discrete emotion classification and gradually increasing the supervision of physiological parameter regression. Compared to fixed-weight models (such as CNNs), this adaptive mechanism improves verification accuracy while ensuring a strong correlation between physiological indicators (such as heart rate variability) and emotional state.
[0045] By modeling behavioral sequences using a multi-head spatiotemporal Transformer and combining it with a post-processing temporal smoothing filter, this approach effectively identifies continuous changes in emotional states (such as the transition from fear to anxiety). Compared to static acoustic feature analysis, this method can capture dynamic features such as tail wagging frequency and pupil constriction, significantly improving the accuracy of complex emotions (such as fear combined with aggression).
[0046] This invention can quantify the emotional state of pets, providing a quantitative basis for pet health monitoring (such as stress response warning) and behavior correction training, reducing the misdiagnosis rate of pet medical care, and at the same time, through emotional intensity analysis, providing data support for the formulation of personalized interaction plans.
[0047] It should be noted that each module and unit in the pet dog emotion recognition system based on deep learning corresponds one-to-one to the steps in the pet dog emotion recognition method based on deep learning.
[0048] like Figure 3 As shown, the present application also provides a computer device, which can be a server, and its internal structure can be as shown in FIG. Figure 3As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store all data required for the process of the pet dog emotion recognition method based on deep learning. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the pet dog emotion recognition method based on deep learning is implemented.
[0049] Those skilled in the art will understand that Figure 3 The structure shown in is merely a block diagram of a portion of the structure related to the present application solution and does not constitute a limitation on the computer device to which the present application solution is applied.
[0050] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, any one of the above-mentioned pet dog emotion recognition methods based on deep learning is implemented.
[0051] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and RAM bus dynamic RAM (RDRAM).
[0052] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.
[0053] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A pet dog emotion recognition method based on deep learning, characterized in that: include: Collect pet dynamic data, pet physiological data, and situational data to obtain a structured data set; Constructing a three-level emotion annotation framework, annotating the structured dataset through a cross-validation annotation mechanism to obtain an annotated dataset; Extracting multimodal features based on the labeled data set, and integrating the multimodal features through a cascaded SEblock array to obtain multimodal fusion features; Injecting adversarial sample data generated by a stress scenario simulator during the training phase, performing deep learning model training based on the adversarial sample data and the multimodal fusion features, and obtaining a pet emotion recognition model; Emotion recognition is performed on the pet to be identified using the pet emotion recognition model to obtain a pet emotion recognition result.
2. The pet dog emotion recognition method based on deep learning according to claim 1 is characterized in that: The step of annotating the structured data set by a cross-validation annotation mechanism to obtain an annotated data set includes: Annotating the pet's dynamic data using an annotation tool to obtain basic emotion tags, wherein the basic emotion tags include anxiety, pleasure, and fear; Associating the scenario data with the basic emotion tag through a context association algorithm to generate a context enhancement tag; Based on the physiological parameter mapping model, calculating the correlation coefficient between the pet's physiological data and the emotional state corresponding to the basic emotion label to obtain a physiological emotion weight matrix; Cross-validating the basic emotion labels, the situational enhancement labels, and the physiological emotion weight matrix through a crowdsourcing verification platform, eliminating data with a labeling confidence lower than a threshold, and obtaining a cross-validation result; A labeled data set containing multiple levels of labels is generated according to the cross-validation results.
3. The pet dog emotion recognition method based on deep learning according to claim 1 is characterized in that: The step of extracting multimodal features based on the labeled data set and integrating the multimodal features through a cascaded SEblock array to obtain multimodal fusion features includes: The facial key point detection network is used to locate the facial area based on the pet's dynamic data to obtain facial expression features; Perform audio analysis based on pet dynamic data through Mel-Cepstral time-frequency transform to obtain acoustic feature vectors; The long short-term memory network is used to model the periodic fluctuation of pet physiological data and obtain physiological state encoding; Combining the facial expression features, the acoustic feature vectors, and the physiological state codes into initial multimodal features through a feature concatenation layer; The initial multimodal features are channel-weighted optimized by cascading SEblock arrays to generate multimodal fusion features.
4. The pet dog emotion recognition method based on deep learning according to claim 1 is characterized in that: The step of locating the facial region based on the pet dynamic data by using the facial key point detection network to obtain facial expression features includes: performing global channel compression on the initial multimodal features through a first SEblock module in the cascaded SEblock array to generate a first compressed feature; superimposing a variety prior weight on the first compressed feature through a second SEblock module in the cascaded SEblock array to obtain a second compressed feature; Performing a residual connection on the second compressed features through the third SEblock module in the cascaded SEblock array, fusing the initial multimodal features with the secondary weighted features to obtain a preliminary fused feature; The preliminary fusion features are dimensionally compressed through a dimensionality reduction fully connected layer to generate multimodal fusion features.
5. The pet dog emotion recognition method based on deep learning according to claim 1 is characterized in that: The step of performing deep learning model training based on the adversarial sample data and the multimodal fusion features to obtain a pet emotion recognition model includes: Generate training input data based on the structured dataset using a generative adversarial network, and perform temporal modeling on the behavior sequences in the training input data using a multi-head spatiotemporal Transformer to capture the emotion evolution pattern; Adjust the contribution ratio of classification loss and regression loss through dynamic weight strategy; The gradient clipping algorithm is used to limit the parameter update amplitude of the deep learning model to prevent overfitting of adversarial samples and preserve the optimal pet emotion recognition model.
6. The pet dog emotion recognition method based on deep learning according to claim 1, characterized in that: The step of generating training input data based on the structured data set by using a generative adversarial network comprises: synthesizing the instrument noise spectrum of the structured data set through a generative adversarial network to generate enhanced acoustic data; The pet's limb motion trajectory in the pet's dynamic data under the constraint state is simulated by the posture synthesis algorithm to generate adversarial visual data; Simulate the heart rate variability and skin conductance fluctuations under fear to generate adversarial physiological data; Synchronizing the adversarial visual data, the enhanced acoustic data, and the adversarial physiological data into adversarial sample data through a spatiotemporal alignment algorithm; The adversarial sample data is mixed with the multimodal fusion features in proportion to generate training input data.
7. The pet dog emotion recognition method based on deep learning according to claim 1 is characterized in that: The step of performing emotion recognition on the pet to be recognized by using the pet emotion recognition model to obtain the pet emotion recognition result includes: The multimodal data collection system is used to collect real-time data of the pet to be identified to obtain data to be processed; Preprocessing the data to be processed to obtain standardized data to be identified; Performing multimodal feature extraction on the standardized data to be identified to obtain a feature vector to be identified; Inputting the feature vector to be identified into a pet emotion recognition model to obtain an emotion prediction result; Using a time series smoothing filter, filtering out short-term abnormal fluctuations in the emotion prediction results to obtain a smoothed emotion sequence; According to the smoothed emotion sequence, an emotion state analyzer is used to generate a pet emotion recognition result including basic emotions, compound emotion components and emotion intensity.
8. A pet dog emotion recognition system based on deep learning, characterized in that: include: The acquisition module is used to collect pet dynamic data, pet physiological data, and situational scene data to obtain a structured data set; An annotation module is used to build a three-level emotion annotation framework and annotate the structured dataset through a cross-validation annotation mechanism to obtain an annotated dataset; An extraction module, configured to extract multimodal features based on the labeled data set, and integrate the multimodal features through a cascaded SEblock array to obtain multimodal fusion features; A training module is used to inject adversarial sample data generated by a stress scenario simulator during the training phase, perform deep learning model training based on the adversarial sample data and the multimodal fusion features, and obtain a pet emotion recognition model; The recognition module is configured as a pet emotion recognition model obtained through training by the training module, and is used to perform emotion recognition on the pet to be identified and obtain a pet emotion recognition result.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-modal face emotion recognition method and device
CN114399818A
Method for constructing mental health care virtual scene by using GAN (Generative Adversarial Network)
CN118919025A
AI-based pet emotion recognition system
CN119049086A
Dialogue analysis-oriented graph confrontation emotion recognition method and system in Internet of Things environment
CN119474329A
Pet pacifying method
CN119646600A
Cited By
Sparse reward environment optimization learning identification method and system based on demonstration data enhancement
CN121157051A
Personality perception pet multi-mode emotion recognition method
CN121331171A
Pet neck-hanging camera Vlog generation method, device, equipment and medium
CN121334466A
A pet neck camera Vlog generation method, device, equipment and medium
CN121334466B
Pet behavior recognition and emotion detection method
CN121415134A