Animal language conversion method, emotion recognition method and model training method and device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2025-08-08
- Publication Date
- 2026-08-07
AI Technical Summary
Existing technologies cannot deeply understand the complex emotional aspects of animals, nor can they achieve deep, real-time emotional understanding and interactive communication between humans and animals. They suffer from problems such as limited emotional translation, lack of multimodal fusion analysis, insufficient temporal detection, lack of adaptive learning and optimization mechanisms, and insufficient real-time performance.
By acquiring multimodal data from animals, including sound, behavior, and physical characteristics, preprocessing and fusing the data, and utilizing deep learning and generative adversarial networks for emotion recognition and semantic mapping, animal language can be converted into human language.
It improves the accuracy of understanding and responding to animal emotions, enhances communication between humans and animals, enables deeper emotional exchange and understanding, and improves the accuracy and efficiency of cross-species communication.
Smart Images

Figure CN122535946A_ABST
Abstract
Description
Animal language conversion methods, emotion recognition methods, and model training methods and devices
[0001] This disclosure claims priority to Chinese Patent Application No. 202411793938.3, filed on December 6, 2024, entitled “Animal Language Conversion Method, Apparatus, Electronic Device and Storage Medium”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the field of artificial intelligence technology, specifically to the fields of machine learning, deep learning, and natural language processing. Background Technology
[0003] Current technologies on the market attempt to interpret the emotional world of animals through basic animal vocalization and behavior translation devices and pet emotion analysis tools utilizing artificial intelligence image recognition technology. However, these methods are often limited to a superficial interpretation of animal behavior. These technologies cannot delve into the complex emotional levels of animals, nor can they achieve deep, real-time emotional understanding and interactive communication between humans and animals. Summary of the Invention
[0004] This disclosure provides an animal language conversion method, an emotion recognition method, and a model training method and apparatus.
[0005] According to one aspect of this disclosure, an animal language conversion method is provided, comprising:
[0006] Acquire multimodal data related to animals; the multimodal data includes animal sound data, animal behavior data, and animal physical characteristics data;
[0007] The multimodal data is preprocessed to obtain fused multimodal data;
[0008] The animal's current emotion is identified based on the fused multimodal data to obtain the animal's emotion recognition result;
[0009] The emotion recognition results are semantically mapped and translated to convert animal language into human language, resulting in a language conversion result.
[0010] According to one aspect of this disclosure, a training method for an animal emotion recognition model in an animal language conversion model is provided, comprising:
[0011] Acquire multimodal data related to animals;
[0012] The multimodal data is preprocessed to obtain fused multimodal data;
[0013] The fused multimodal data is input into the animal emotion recognition model to identify the animal's current emotion and obtain the animal's emotion recognition result.
[0014] Based on the emotion recognition results, the animal emotion recognition model is trained; the emotion recognition results output by the animal emotion recognition model are used for semantic mapping and language translation to convert animal language into human language and obtain language conversion results.
[0015] According to another aspect of this disclosure, an animal emotion recognition method is provided, comprising:
[0016] Acquire multimodal data related to animals;
[0017] The multimodal data is preprocessed to obtain fused multimodal data;
[0018] The animal's current emotion is identified based on the fused multimodal data to obtain the animal's emotion recognition result.
[0019] According to another aspect of this disclosure, a method for training an animal emotion recognition model is provided, comprising:
[0020] Acquire multimodal data related to animals;
[0021] The multimodal data is preprocessed to obtain fused multimodal data;
[0022] The fused multimodal data is input into the animal emotion recognition model to identify the animal's current emotion and obtain the animal's emotion recognition result.
[0023] Based on the emotion recognition results, the animal emotion recognition model is trained.
[0024] According to another aspect of this disclosure, an animal language conversion device is provided, comprising:
[0025] The first acquisition module is used to acquire multimodal data related to animals; the multimodal data includes animal sound data, animal behavior data, and animal physical characteristics data;
[0026] The first preprocessing module is used to preprocess the multimodal data to obtain fused multimodal data;
[0027] The first recognition module is used to recognize the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result;
[0028] The conversion module is used to perform semantic mapping and language translation on the emotion recognition results to convert animal language into human language and obtain the language conversion result.
[0029] According to another aspect of this disclosure, a training apparatus for an animal emotion recognition model in an animal language conversion model is provided, comprising:
[0030] The second acquisition module is used to acquire multimodal data related to animals;
[0031] The second preprocessing module is used to preprocess the multimodal data to obtain fused multimodal data;
[0032] The second recognition module is used to input the fused multimodal data into the animal emotion recognition model to recognize the animal's current emotion and obtain the animal's emotion recognition result.
[0033] The training module is used to train the animal emotion recognition model based on the emotion recognition results; the emotion recognition results output by the animal emotion recognition model are used for semantic mapping and language translation to convert animal language into human language and obtain language conversion results.
[0034] According to another aspect of this disclosure, an animal emotion recognition device is provided, comprising:
[0035] The first acquisition module is used to acquire multimodal data related to animals;
[0036] The first preprocessing module is used to preprocess the multimodal data to obtain fused multimodal data;
[0037] The first recognition module is used to identify the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result.
[0038] According to another aspect of this disclosure, a training apparatus for an animal emotion recognition model is provided, comprising:
[0039] The second acquisition module is used to acquire multimodal data related to animals;
[0040] The second preprocessing module is used to preprocess the multimodal data to obtain fused multimodal data;
[0041] The second recognition module is used to input the fused multimodal data into the animal emotion recognition model to recognize the animal's current emotion and obtain the animal's emotion recognition result.
[0042] The training module is used to train the animal emotion recognition model based on the emotion recognition results.
[0043] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0044] At least one processor; and
[0045] The memory is communicatively connected to the at least one processor; wherein,
[0046] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.
[0047] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.
[0048] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.
[0049] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0050] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0051] Figure 1 is a flowchart illustrating an animal language conversion method according to an embodiment of the present disclosure.
[0052] Figure 2 is a schematic flowchart of the process of acquiring multimodal data according to an embodiment of the present disclosure.
[0053] Figure 3 is a schematic flowchart of preprocessing multimodal data according to an embodiment of the present disclosure.
[0054] Figure 4 is a flowchart illustrating the process of obtaining the emotion recognition result of an animal according to an embodiment of the present disclosure.
[0055] Figure 5 is a flowchart illustrating the process of obtaining multimodal feature vectors according to an embodiment of the present disclosure.
[0056] Figure 6 is a flowchart illustrating the process of obtaining animal emotion recognition results according to an embodiment of the present disclosure.
[0057] Figure 7 is a flowchart illustrating a training method for an animal emotion recognition model according to an embodiment of the present disclosure.
[0058] Figure 8 is a flowchart illustrating the process of fusing fused features at the same time step using a mid-level attention network according to an embodiment of the present disclosure.
[0059] Figure 9 is a schematic diagram of a fusion network structure for fusing multimodal features at the same time step according to an embodiment of the present disclosure.
[0060] Figure 10 is a schematic flowchart of training an animal emotion recognition model through adversarial training according to an embodiment of the present disclosure.
[0061] Figure 11 is another flowchart illustrating the training of an animal emotion recognition model through adversarial training according to an embodiment of the present disclosure.
[0062] Figure 12 is a schematic flowchart of training an animal emotion recognition model based on emotion recognition results according to an embodiment of the present disclosure.
[0063] Figure 13 is a main framework diagram of an animal emotion recognition model according to an embodiment of the present disclosure.
[0064] Figure 14 is a schematic flowchart of a pseudo-label-based animal emotion recognition model according to an embodiment of the present disclosure.
[0065] Figure 15 is a schematic flowchart of generating pseudo-tags for animal emotions according to an embodiment of the present disclosure.
[0066] Figure 16 is another schematic flowchart illustrating the generation of pseudo-labels for animal emotions according to an embodiment of the present disclosure.
[0067] Figure 17 is a flowchart illustrating an optimized animal emotion recognition model according to an embodiment of the present disclosure.
[0068] Figure 18 is a schematic diagram of a process for converting animal language into human language according to an embodiment of the present disclosure.
[0069] Figure 19 is a flowchart illustrating an animal emotion recognition method according to an embodiment of the present disclosure.
[0070] Figure 20 is a flowchart illustrating a training method for an animal emotion recognition model in an animal language conversion model according to an embodiment of the present disclosure.
[0071] Figure 21 is a schematic flowchart of updating sentiment tags according to an embodiment of the present disclosure.
[0072] Figure 22 is a schematic diagram of the structure of an animal emotion recognition device according to an embodiment of the present disclosure.
[0073] Figure 23 is a schematic diagram of the structure of a training device for an animal emotion recognition model according to an embodiment of the present disclosure.
[0074] Figure 24 is a schematic diagram of the structure of an animal language conversion device according to an embodiment of the present disclosure.
[0075] Figure 25 is a schematic diagram of the structure of a training device for an animal emotion recognition model in a language conversion model according to an embodiment of the present disclosure.
[0076] Figure 26 is a block diagram of an electronic device used to implement any of the methods of the embodiments of this disclosure. Detailed Implementation
[0077] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0078] The terms “first,” “second,” etc., used in this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0079] It should be noted that, unless it is explicitly stated that there is a sequential order of execution between different operations, or that there is a sequential order of execution between different operations in terms of technical implementation, the execution order between multiple operations may not be significant, and multiple operations may be executed simultaneously.
[0080] Currently, technologies related to human-animal communication on the market can be mainly divided into the following two types:
[0081] The first approach uses simple animal vocalization and behavior translation devices, which primarily rely on a "voiceprint database + simple algorithm" to translate animal emotions. A typical example is some commercially available pet behavior and sound recognition devices. These devices use relatively simple sensors to capture animal sounds and some typical movements, then match them to a pre-built emotion database for simple emotion mapping, such as recognizing a dog's bark as a request for food or anger.
[0082] The second approach is based on AI (Artificial Intelligence) image recognition pet emotion analysis tools, which introduce image processing and AI technology to help users understand their pets' emotional responses. For example, some companies capture animal facial expressions and movements in camera images and then combine them with existing deep learning-trained models to analyze different facial expressions and identify specific expressions such as laughter, anger, grievance, and confusion.
[0083] The two methods mentioned above have several limitations in understanding and translating animal emotions: First, emotion translation is singular, relying on pre-set voiceprint data and behavioral classifications, and cannot continuously track emotional changes, resulting in insufficient accuracy in complex scenarios. Second, the lack of multimodal fusion analysis and over-reliance on a single information source limits the comprehensiveness and accuracy of emotion translation. Third, insufficient temporal detection and the lack of ability to track the continuity of emotional states lead to sluggish perception of emotional changes. Furthermore, the lack of adaptive learning and optimization mechanisms makes it difficult to optimize and iterate when faced with unknown emotional patterns, limiting the system's flexibility and scalability. Finally, the inability to perform edge computing results in insufficient real-time performance, affecting instant interaction between humans and animals. These limitations collectively contribute to the inadequacy of existing technologies in achieving deep, real-time cross-species emotional communication.
[0084] To address at least one of the aforementioned technical problems, this disclosure provides an animal emotion recognition method. The animal emotions in this disclosure can include transient emotions expressed by the animal, such as anger or resentment, as well as the animal's needs, such as a request for food. Therefore, the animal emotions in this disclosure primarily refer to converting animal-specific expressions into expressions that humans can understand, in order to facilitate the understanding of animal behavior and needs.
[0085] Referring to Figure 1, which is a schematic diagram of the steps of the animal language conversion method in this embodiment of the present disclosure, the method can be applied to the server side and includes:
[0086] Step S101: Obtain multimodal data related to animals.
[0087] In implementation, the multimodal data includes animal sound data, animal behavior data, and animal physical characteristic data. The types of modalities can be increased or decreased according to actual needs, and this disclosure does not limit this.
[0088] Specifically, acquiring multimodal data related to animals refers to collecting information from different modalities through multiple sensors to comprehensively understand the animal's state and emotions. This includes: "Animal vocal data," which refers to various animal calls captured by audio sensors, reflecting the animal's emotions and needs; "Animal behavioral data," which is the animal's movements (such as facial expressions and body movements) and postures recorded by video cameras, helping to analyze its behavioral patterns and emotional expressions; and "Animal vital signs data," which refers to physiological indicators of the animal monitored by physiological sensors, such as heart rate and body temperature, providing physiological evidence of the animal's emotional state.
[0089] In this way, by acquiring multimodal data of animals, including vocal, behavioral, and physical data, it is possible to comprehensively capture animal emotions and behaviors, thereby providing more accurate sentiment analysis and behavioral understanding, enhancing communication and interaction between humans and animals, and improving animal welfare and human responsiveness to animal behavior.
[0090] In implementation, multimodal data can be collected by a terminal device. Data for each modality can be collected by a corresponding sensor. The same terminal device can integrate at least one sensor, but this disclosure does not limit this. After collecting the multimodal data, the terminal device can send it to a server. The collected multimodal data can be stored in the server so that the server can retrieve the multimodal data and then execute the animal language conversion method and / or animal emotion recognition method provided in this disclosure.
[0091] Step S102: Preprocess the multimodal data to obtain fused multimodal data.
[0092] Specifically, preprocessing multimodal data can include a series of processing steps on raw data collected from different sources (such as audio, video, and physiological sensors). For example, this may include noise reduction, normalization, and feature extraction to facilitate analysis and understanding. "Fused multimodal data" refers to integrating these preprocessed data into a unified dataset. This process involves time alignment (ensuring all data is synchronized in time), feature fusion (merging features from different modalities into a comprehensive feature vector), and data synchronization (handling differences in timestamps and sampling rates between different data streams). Through this preprocessing and fusion, a comprehensive data view can be provided, making subsequent sentiment recognition and language translation more accurate and efficient. Feature fusion here can include concatenating data from different modalities, ensuring that the fused multimodal data retains the independence of each modality for subsequent processing.
[0093] By preprocessing multimodal data such as animal sounds, behaviors, and physical signs, and then integrating these data into a unified dataset, the consistency and usability of the data can be improved. This leads to more accurate emotion recognition and language translation, enhances the system's comprehensive understanding of animal behavior and emotional states, and improves the efficiency and effectiveness of cross-species communication.
[0094] In practice, at least one of the aforementioned preprocessing operations can also be performed by the terminal device.
[0095] Step S103: Identify the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result.
[0096] Specifically, after obtaining the fused multimodal data, the current emotion of the animal is identified based on this data. This involves using a comprehensive dataset integrating information such as animal vocalizations, behaviors, and physical characteristics, and employing machine learning and deep learning techniques to analyze and judge the animal's emotional state. Here, "emotion recognition" refers to identifying the animal's emotional state, such as anxiety, excitement, or relaxation, by analyzing the features extracted from this multimodal data. To obtain the animal's emotion recognition results, the fused data can be input into a pre-trained emotion recognition model. This model compares the features of known emotional states to output the animal's emotion recognition result, thus providing a basis for subsequent language conversion and human-computer interaction. This animal emotion recognition model can be pre-trained separately for animal emotion recognition. The training process and the reasoning process for recognizing animal emotions are consistent; the training of the animal emotion recognition model will be explained later.
[0097] In this way, by analyzing the fused multimodal data, including animal sounds, behaviors, and physiological signs, it is possible to accurately identify the animal's current emotional state. This process not only improves the accuracy of understanding and responding to animal emotions but also enhances communication between humans and animals, enabling humans to better interpret animal needs and emotions, thereby improving animal welfare and the intimacy of the human-pet relationship.
[0098] Step S104: Semantic mapping and language translation are performed on the emotion recognition results to convert animal language into human language and obtain the language conversion result.
[0099] Specifically, "semantic mapping and language translation of emotion recognition results" refers to the process of converting animal emotional states, such as anxiety, excitement, or happiness, obtained through the analysis of multimodal data, into language expressions that humans can understand. This process involves using pre-trained language models and deep learning techniques to establish a correspondence between animal emotional features and corresponding expressions in human language—a process known as "semantic mapping." Subsequently, these mapping results are converted into specific text or speech output, achieving "language translation." Ultimately, the "language translation result" refers to the animal's emotions and intentions being transformed into human language, enabling human users to understand animal "language" more intuitively and conveniently, thereby achieving effective cross-species communication.
[0100] In this way, by semantically mapping and linguistically translating the results of emotion recognition, it is possible to transform the nonverbal communication of animals into language that humans can understand. This process not only breaks down communication barriers between humans and animals, but also greatly enhances human understanding of animal emotions and needs, making human-animal interactions more harmonious. It also provides new perspectives and tools for animal welfare and behavioral research.
[0101] This disclosure enables comprehensive capture and accurate identification of animal emotions by acquiring multimodal data. Then, through semantic mapping and language translation technologies, the animal's emotional state and intentions are converted into human-understandable language, greatly enhancing communication between humans and animals, improving the accuracy of animal emotion understanding and the real-time nature of interaction, and providing humans with a completely new way to communicate with animals. In other words, this solution can accurately identify the current emotional state of an animal and convert it into human language, thereby achieving deeper emotional exchange and understanding between animals and humans, and improving the accuracy and efficiency of cross-species communication.
[0102] In some optional embodiments, acquiring animal-related multimodal data includes: collecting sound wave information emitted by the animal to obtain animal sound data; collecting the animal's body language and movement changes to obtain animal behavior data; and collecting the animal's physical and biological indicators to obtain animal vital signs data.
[0103] Specifically, audio acquisition devices capture animal vocalizations to obtain "animal sound data." Simultaneously, visual sensors such as video cameras collect data on animal body language and movement changes, forming "animal behavior data," which reflects the animal's activity and nonverbal behavior. Furthermore, physiological sensors monitor "animal vital signs data," such as heart rate and body temperature; these physical and biological indicators provide crucial information for understanding the animal's physiological state and emotions. By integrating this multimodal data, it is possible to comprehensively capture the animal's communication methods and emotional state, laying the foundation for further data analysis and emotion recognition.
[0104] To facilitate understanding of the embodiments of this disclosure, an example is provided below, referring to Figure 2, which is a schematic flowchart of the multimodal data acquisition process in an embodiment of this disclosure. First, an audio acquisition device is used to capture the animal's vocalizations in real time. Then, the audio acquisition device sends the captured sound data to a data processing module. The data processing module can use appropriate audio filters to process noise and reduce background interference. For example, for a dog bark, its pitch, amplitude, duration, frequency, and breakpoint changes are sampled together to capture all the information of the bark. The data processing module can be located on the terminal device or on a server.
[0105] For animal behavior data collection, this can be achieved by using video equipment (such as high-definition cameras and infrared cameras) to capture animal limb movements and behavioral expressions, such as tail wagging, jumping, lying down, and other posture analysis, supplemented by specific animal physical characteristics (ears erect, pupils dilated, etc.). Then, the captured video / image data is sent to the data processing module.
[0106] For animal vital sign data collection, heart rate and body temperature can be collected through high-precision contact or non-contact body temperature detection sensors, and the collected vital sign data can be sent to the data processing module.
[0107] In this way, by collecting the sound wave information emitted by animals, body language and movement changes, as well as physical and biological indicators, relatively comprehensive animal sound data, behavioral data, and physical signs data can be obtained. The integration of this multi-dimensional information provides a rich and accurate data foundation for in-depth understanding and analysis of animal emotions and behaviors, thereby enabling humans to more accurately interpret animals' communication intentions and physiological states, strengthening non-verbal communication between humans and animals, improving the effectiveness of animal care and training, and opening up new avenues for animal health monitoring and behavioral research.
[0108] In some optional embodiments, the multimodal data is preprocessed to obtain fused multimodal data, including:
[0109] Denoising is performed on the multimodal data to clean the data, resulting in cleaned multimodal data.
[0110] The cleaned multimodal data is normalized to obtain normalized multimodal data;
[0111] The normalized multimodal data is time-series aligned and fused to obtain fused multimodal data.
[0112] Specifically, denoising multimodal data for data cleaning refers to using signal processing techniques to remove noise and interference from audio and video data to improve data quality, resulting in cleaned multimodal data. Next, "normalizing the cleaned multimodal data" means converting data from different sources and scales into a unified format or scale so that machine learning models can process it more effectively; this step yields normalized multimodal data. Finally, "temporally aligning and fusing the normalized multimodal data" involves aligning data from different modalities in time (e.g., through timestamps) and merging them into a unified dataset. This step ensures temporal consistency and integrates information from different sensors, ultimately resulting in fused multimodal data, which is a combination of data from different modalities. The fused multimodal data provides an accurate and comprehensive data foundation for subsequent emotion recognition and behavior analysis.
[0113] To facilitate understanding of the embodiments of this disclosure, an example is provided below, referring to Figure 3, which is a schematic flowchart of the preprocessing of multimodal data in an embodiment of this disclosure. First, the sound data, image data, and body temperature and heart rate data in the multimodal data are preprocessed and normalized. Data preprocessing includes noise reduction and data cleaning, which is used to remove invalid parts from the audio and visual information. For example, noise that may be generated by humans (wind, speech, etc.) is filtered out, making the audio signal clearer and easier to identify. In images, background motion and objects in video frames are cleaned to retain only key information, such as changes in animal behavior. Next, data normalization is performed. Whether it's audio signals, video, or vital sign information, it needs to be normalized, converted into a unified standard, and expressed as a feature vector that can be processed by machine learning algorithms.
[0114] Finally, the multimodal data undergoes temporal alignment and fusion. Data alignment is a prerequisite for data fusion, and it requires addressing the temporal and spatial differences between the acquisition of different modal signals. For example, a dog barking event often corresponds to a time difference with the moment of accompanying limb movement or body temperature fluctuation. This data must be temporally calibrated before input to ensure a consistent reference point for the multimodal input in the translation task.
[0115] In implementation, sensors of different modalities can acquire data based on their respective sampling rates. Each frame of data acquired by the terminal device can be timestamped by the terminal device. Timing alignment can be performed based on these timestamps. For example, one modality can be selected as the reference modality, and its timestamp can be used as the reference timestamp. Furthermore, a time difference threshold can be determined based on the differences in sampling rates between different modalities. Data whose timestamps of other modalities differ from the reference timestamp within this threshold range are considered data generated at the same time as the reference timestamp, thus aligning data from different modalities based on timestamps and achieving timing alignment.
[0116] Assuming that the sampling rate of animal sound data, i.e., audio data, is the highest, the audio timestamp can be used as the reference timestamp. In the video modality and the physical characteristic modality data, data surrounding this reference timestamp (i.e., the difference between the timestamps of other modalities and this reference timestamp is within a time difference threshold range) are searched to temporally align the video modality data and the physical characteristic modality data with the audio modality data, respectively. Of course, it is understood that if the sampling rates of the video modality and the physical characteristic modality are inconsistent, the time difference threshold used to search the video modality data and the time difference threshold used to search the physical characteristic modality data can be different. Therefore, regardless of the method used, as long as the acquisition time of data from different modalities can be aligned, this disclosure does not limit this approach.
[0117] Furthermore, since the sampling rates of data from different modalities differ, data synchronization can be used to align the sampling rates of different modalities during implementation. For example, if the sampling rates of the video modality and the vital signs modality are low, upsampling can be performed on the data of the video modality and the vital signs modality after temporal alignment to ensure that the number of frames of data from each modality is the same within the same time window. This facilitates subsequent identification of animal emotions.
[0118] In implementation, time synchronization mechanisms can be used to align data from different modalities in time. For example, NTP (Network Time Protocol) is a standard protocol used for time synchronization between devices in a network. In this embodiment, NTP is used to ensure precise synchronization of the system clocks for audio, video, and physiological signal acquisition, thereby ensuring global consistency of timestamps assigned to all data frames and avoiding data mismatch caused by clock drift during subsequent alignment.
[0119] During implementation, edge devices are supported for data preprocessing. This preprocessing typically includes: data cleaning, sampling rate conversion (i.e., the aforementioned data synchronization used for synchronizing sampling rates), frame energy normalization, etc. Edge devices can be understood as intelligent terminal devices deployed at the acquisition end (such as embedded acquisition boxes, smart collars, etc.). They are directly responsible for raw data acquisition and at least part of the preprocessing.
[0120] Data cleaning can include outlier detection and removal. The definition of outliers depends on the physical properties and sampling environment of each modality of data. For example, physiological signals exceeding physiological thresholds (e.g., canine heart rate >240 bpm), or abrupt changes, loss of data, extremely low energy, or values exceeding hardware limits in the sampled signal. In practice, outlier detection can be performed by setting statistical intervals, using the 3σ principle, or machine learning methods (e.g., Isolation Forest, LOF), and the detected outliers can be removed.
[0121] Preprocessing may also include timestamps, which are used to add a timestamp to each frame of data to facilitate subsequent timing alignment.
[0122] Preprocessing may also include fusion processing, which involves stitching together data from different modalities so that multimodal data collected at the same time can be integrated to express the different behaviors of animals at the same time from different modalities.
[0123] For the acquired multimodal data, breakpoint resumption and cache management are also supported. This can be implemented by marking all acquired data with high-precision timestamps. During breakpoint resumption, the timestamp queue of each modality's data is compared, and the data requiring retransmission is inserted into the correct timing position, thereby restoring multimodal data synchronization. An "alignment buffer" can also be introduced to temporarily cache incomplete multimodal data within a time window. Data is only output downstream after all modal data within the current time window is complete. Of course, if the data within the current time window is still incomplete by the next time window, it can still be transmitted downstream so that subsequent processing can continue based on the data within the current time window.
[0124] By performing denoising, normalization, temporal alignment, and fusion processing on multimodal data, the quality, consistency, and usability of the data can be significantly improved, making the features extracted from animal sounds, behaviors, and physical characteristics more accurate and reliable. This process not only enhances the accuracy of data analysis but also improves the model's precision in recognizing animal emotions and behaviors, providing a solid data foundation for achieving efficient and accurate cross-species communication.
[0125] In some embodiments, data from different modalities are spliced together according to a standardized format to obtain fused multimodal data, which is then transmitted.
[0126] For example, a single modality's data structure can include a timestamp, a modality identifier collected at that timestamp, and the data value for that modality. After time step alignment, data from multiple modalities corresponding to the same timestamp can be concatenated according to a standardized format. This standardized format data structure can include a timestamp and data from multiple modalities corresponding to that timestamp. For example, it could include animal sound data corresponding to an audio identifier, image data corresponding to a video identifier, and animal physical characteristic data corresponding to a physical characteristic identifier. Of course, due to different sampling rates, some modal data may be missing at certain timestamps. If data synchronization is performed during the preprocessing stage, it can ensure that multimodal data is as complete as possible at the same timestamp. If data synchronization is not performed during the preprocessing stage, it is permissible for some timestamps to have missing data from some modalities, and data synchronization can wait for subsequent high-level feature extraction based on a neural network model.
[0127] In some optional embodiments, the animal's current emotion is identified based on the fused multimodal data to obtain the animal's emotion recognition result, including:
[0128] A deep learning model is used to extract sound features, visual motion features, and analyze physical changes in the fused multimodal data to obtain a multimodal feature vector. This multimodal feature vector includes feature representations of different modalities, such as features of the audio modality, video modality, and physical change modality. This allows for accurate understanding of the features of each modality through generative adversarial networks and for uncovering feature relationships between different modalities.
[0129] Generative adversarial networks were used to perform sentiment analysis on multimodal features to obtain the emotion recognition results of animals.
[0130] Specifically, "using deep learning models to extract sound features, visual-motor features, and analyze vital sign changes from fused multimodal data" refers to using advanced deep learning techniques, such as Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs), to process and analyze multimodal datasets that integrate sound, visual, and physiological data. Since different modalities can individually express corresponding emotions, this process involves extracting sound features from audio signals, such as pitch, rhythm, and volume; extracting visual-motor features from video data, such as posture and behavioral patterns; and analyzing vital sign changes from physiological data, such as heart rate and body temperature fluctuations. The sum of these features from different modalities constitutes a multimodal feature vector. Each modality's features exist independently and can be analyzed separately. Integrating these features together provides a comprehensive reflection of the animal's physiological and behavioral state. Next, "using generative adversarial networks for sentiment analysis of multimodal features" refers to using a deep learning framework like Generative Adversarial Networks (GANs) through an adversarial training process, enabling the network to recognize and distinguish different emotional states. This process ultimately produces emotion recognition results for animals, transforming their complex emotions and behaviors into understandable emotion labels, providing a basis for further semantic mapping and language translation.
[0131] To facilitate understanding of the embodiments of this disclosure, an example is provided below, referring to Figure 4, which is a flowchart illustrating the process of obtaining animal emotion recognition results in an embodiment of this disclosure. After obtaining the fused multimodal data, fine-grained feature extraction is performed on the data of each modality using a deep learning model. Specifically, the deep learning model extracts voice features, visual action features, and analyzes physical characteristic changes to obtain a feature vector combination (i.e., a multimodal feature vector). Next, a large model based on an emotion recognition generative adversarial network (GAN) is used to analyze the voiceprint features, action changes, and physical characteristic fluctuations in the data to obtain emotion classification labels. For example, when a low-frequency barking of an animal is detected accompanied by limb tension and pupil dilation, the model identifies the animal as being in a state of high alert through comparative inference and further infers the possible underlying psychological activities (fear or confusion) through feature combination. Feature combination may include analyzing the correlation between features of different modalities and analyzing the correlation between features at different time steps in the time series.
[0132] In this way, by employing deep learning models to comprehensively analyze the fused multimodal data, it is possible to accurately extract features such as animal vocalizations, visual movements, and changes in physical characteristics, forming multimodal feature vectors. Then, generative adversarial networks are used to conduct in-depth sentiment analysis on these features, ultimately yielding animal emotion recognition results. This process not only improves the accuracy and depth of emotion recognition but also enhances the understanding of animal behavior and psychological states, providing strong technical support for achieving more effective human-animal communication.
[0133] Specifically, in some embodiments, if data synchronization and normalization are not performed during the preprocessing stage, the collected data can be timestamped, cleaned, and time-series aligned and fused during the preprocessing stage to obtain fused multimodal data that meets the aforementioned standardized format. Then, a deep learning model is used to extract sound features, visual motion features, and analyze physiological changes in the fused multimodal data to obtain multimodal feature vectors. This can be implemented as shown in Figure 5.
[0134] S501 uses a deep learning model to extract sound features, visual motion features, and analyze physical changes in the fused multimodal data, resulting in first audio features, first visual features, and first physical feature.
[0135] That is, the fused multimodal data includes audio signals from the audio modality, video data from the video modality, and physiological data from the vital signs modality. Feature extraction is performed on the audio signals to obtain the first audio features, feature extraction is performed on the video data to obtain the first visual features, and feature extraction is performed on the physiological data to obtain the first vital signs features.
[0136] The module for feature extraction of each modality may employ networks such as CNN, RNN, and MLP (Multi-Layer Perceptron), but this embodiment does not limit the specific network used.
[0137] During implementation, the aforementioned standardized structure stores the fused multimodal data. Data for each modality at each timestamp can be extracted from this fused multimodal data, and then feature extraction can be performed separately.
[0138] The aforementioned temporal alignment only aligns the macroscopic timing. However, to compensate for the sampling rate differences between different modalities, frame-level fine-tuning is needed through data synchronization to ensure that the frame rates of different modalities are the same, thus achieving the same sampling rate in the high-level feature space. For example, assuming the frame rate of the first audio feature is 16 frames / s, the frame rates of the first visual feature and the first vital sign feature both need to be aligned to 16 frames / s.
[0139] S502, synchronize the data of the first audio feature, the first visual feature and the first vital sign feature to obtain the second audio feature, the second visual feature and the second vital sign feature.
[0140] This data synchronization, or aligning the frame rates of different modalities, can be achieved by upsampling to increase the frame rate or downsampling to decrease the frame rate, thereby aligning the sampling rates of different modalities.
[0141] If normalization is not performed during the preprocessing stage, in order to align the signal ranges of different modes and prevent a single mode from dominating subsequent modeling, normalization can be performed through the following S503.
[0142] S503, normalizes the second audio feature, the second visual feature, and the second physical feature to obtain a multimodal feature vector.
[0143] The normalization process may include Z-score (Z-score Normalization, standard deviation standardization) normalization or Min-Max (Min-Max Scaling, normalization) normalization.
[0144] In practice, if the multimodal data is relatively simple, normalization can be completed during the preprocessing stage, that is, the original signal can be normalized. However, normalization during the preprocessing stage may result in the loss of some key information. Therefore, in order to retain more features, normalization can be performed in a higher-level feature space through S503 after data synchronization.
[0145] In this embodiment, independent features of each modality can be extracted separately, and then frame-level feature alignment can be achieved through data synchronization. This allows for a better description of the animal's state and reactions at the same time step, improving the accuracy of emotion recognition. Finally, normalization processing ensures that the signal ranges of different modalities are consistent, preventing a single modality from dominating emotion recognition and further improving the accuracy of animal emotion recognition.
[0146] In some embodiments, generative adversarial networks are used to perform sentiment analysis on multimodal features to obtain the emotion recognition results of animals. This can be implemented as shown in Figure 6, including:
[0147] S601 performs cross-modal feature fusion operations on multimodal feature vectors to obtain fused features.
[0148] The multimodal feature vector includes feature vectors for audio modality, video modality, and physical feature modality. This includes the normalized second audio feature, second visual feature, and second physical feature, as described earlier.
[0149] Feature fusion operations can model the intrinsic relationships between different modalities. This can be understood as each multimodal feature vector corresponding to a time step; feature fusion operations can model the features of different modalities within the multimodal feature vectors at the same time step, thus uncovering the correlation patterns between different modalities at the same time step.
[0150] In order to better understand animal emotions, in this embodiment of the disclosure, the temporal correlation of different time steps is modeled on the time series, as shown in step S602 below.
[0151] S602, perform temporal modeling on the fused features to obtain temporal features.
[0152] This time series feature can be obtained by using a neural network model that processes time series data.
[0153] S603 performs a classification operation on temporal features to obtain the animal's emotion recognition results.
[0154] In this embodiment, cross-modal feature fusion operations are used to model the relationships between multimodal data at the same time step, thereby mining high-level features. Then, temporal modeling is used to mine the temporal feature expressions of animals, generating temporal features suitable for classification. Finally, classification operations are used to recognize animal emotions. The entire process is gradual and progressive, improving the accuracy of animal recognition.
[0155] In some embodiments, the aforementioned deep learning model and generative adversarial network are used to construct an animal emotion recognition model. This animal emotion recognition model can be trained independently or together with the pre-trained speech model in S104. The specific data processing flow and training method for independent training of this emotion recognition model will be described later. Details will not be elaborated here. In this embodiment, the deep learning model and generative adversarial network are organically integrated into an animal emotion recognition model, which facilitates end-to-end training and improves the accuracy of animal emotion recognition.
[0156] In some embodiments, this disclosure can dynamically adjust the acquisition frequency of multimodal data. For example, if a preset event is detected based on emotion recognition results, the frequency of acquiring multimodal data can be increased; the preset event includes at least one of the following:
[0157] 1) The emotion recognition result indicates that the first type of target emotion has been identified, and the confidence level of the emotion recognition result is greater than the preset threshold.
[0158] The first category of target emotions can be understood as emotions associated with high-risk states, such as anxiety, tension, or the expression of hunger. During implementation, a list of first-category target emotions can be set up as needed. If the identified emotion is in this list, it indicates that the first-category target emotion has been identified.
[0159] The confidence score of the recognition will be output along with the emotion recognition result, and is used to express the credibility of the emotion recognition result.
[0160] When the first type of target emotion is identified and the confidence level is high, it is used to trigger an increase in the frequency of acquiring multimodal data, that is, to increase the sampling rate of acquiring multimodal data.
[0161] 2) The emotion recognition result indicates that the second type of target emotion was identified and was continuously detected n times, where n is a positive integer greater than 1.
[0162] The second type of target emotion can also be determined according to actual needs, and this embodiment does not limit it. For example, if the emotion used to express "aggression" is detected continuously n times or more, it can trigger an increase in the sampling rate.
[0163] 3) The emotional fluctuation value of the animal within the sliding time window where the emotion recognition result is located is higher than the preset value.
[0164] During implementation, based on the emotion recognition results and other emotion recognition results within the sliding time window, it is determined whether the animal's emotion fluctuation value is higher than the preset value.
[0165] This involves verifying the emotional fluctuations within a sliding window, which continuously monitors the animal's emotional recognition results, yielding multiple emotional recognition outcomes. The emotional fluctuation values of these multiple emotional recognition outcomes within the sliding window are then calculated.
[0166] During implementation, the statistical values of the classification probability distribution of the emotion recognition results within the sliding window can be calculated. The emotion recognition result itself is a classification probability distribution, recording the probability value of belonging to each emotion. Each time step corresponds to one emotion recognition result, and a sliding window includes multiple time steps. Therefore, the statistical values of the emotion recognition results at different time steps, such as variance, can be calculated as emotion fluctuation values.
[0167] When the emotional fluctuation value exceeds a preset value, it indicates severe emotional fluctuation in the animal. In such cases, the frequency of acquiring animal-related multimodal data can be increased. For example, the sampling rate of multimodal data can be increased to facilitate real-time tracking of the animal's emotional changes. In implementation, for instance, the original sampling rate could be increased from m1 times / s to m2 times / s, where m2 is greater than m1.
[0168] In summary, in this embodiment of the present disclosure, the frequency of acquiring multimodal data can be appropriately increased through an event-triggered mechanism, so as to facilitate real-time observation of the emotional changes of animals.
[0169] In other embodiments, after an animal's emotional state meets a preset event, the animal's emotional changes can be continuously detected. For the animal, if the preset event is not detected for a specified period of time after the detection ends, the frequency of acquiring multimodal data is reduced.
[0170] For example, from t s The system continuously detects preset events that satisfy the animal's emotional needs, thereby increasing the sampling rate. This continues until t i At time t, the animal's emotions no longer satisfy the preset event, and this continues until t. i +T s At time T, no preset event was detected, meaning the event lasted for a period of time T. s If no preset event is detected within a certain duration, the sampling rate is reduced to the default frequency.
[0171] Therefore, by reducing the frequency, power consumption can be saved, the number of times animal emotion recognition results are detected can be reduced, thereby saving computing and storage resources.
[0172] As described above, the animal language conversion model in this embodiment includes an animal emotion recognition model. Based on the same inventive concept, this embodiment also provides a separate training method for the animal emotion recognition model. It should be noted that the processing flow of the animal emotion recognition model described in this training method is also applicable to the part of the aforementioned animal language conversion method used to recognize animal emotions. The only difference is that the model parameters need to be adjusted during the training phase, and the training loss does not need to be calculated during the inference phase.
[0173] Specifically, the training method for the animal emotion recognition model is shown in Figure 7, and includes the following:
[0174] S701, acquire multimodal data related to animals.
[0175] In this embodiment, the animal in S701 can be the same animal as the animal in S101, or they can be different animals. The animal-related multimodal data during the training phase refers to the training samples used to train the animal emotion recognition model.
[0176] S702 preprocesses the multimodal data to obtain fused multimodal data.
[0177] The preprocessing operations are the same as those described above, and will not be repeated here.
[0178] Fusion multimodal data refers to data that includes different independent modalities, so that the information conveyed by a single modality can be analyzed separately, and the correlation patterns between different modalities can be discovered.
[0179] S703 inputs the fused multimodal data into the animal emotion recognition model to identify the animal's current emotion and obtain the animal's emotion recognition result.
[0180] S704, based on emotion recognition results, trains an animal emotion recognition model.
[0181] In other words, during the training phase, animal-related multimodal data is used as training samples. After preprocessing, fused multimodal data is obtained to optimize data quality and improve the training accuracy of the model. By performing inference and prediction on the fused multimodal data, the animal's emotion recognition result is obtained, which is then used to train the model and improve the recognition accuracy of the animal emotion recognition model.
[0182] In some embodiments, preprocessing of multimodal data to obtain fused multimodal data includes: summarizing multimodal data within the same time stamp range based on timestamps in the multimodal data to obtain fused multimodal data. Other preprocessing operations have been described above and will not be repeated here. It is important to emphasize that fused multimodal data is obtained by aligning data from different modalities at the same time stamp. This embodiment of the disclosure, through timestamp alignment, can macroscopically align multimodal data at the same time step, improving the recognition accuracy of the animal emotion recognition model.
[0183] In some embodiments, as described above, the animal emotion recognition model includes a deep learning model and a generative adversarial network. Based on this, the fused multimodal data is input into the animal emotion recognition model to identify the animal's current emotion, obtaining the animal's emotion recognition result. This can be implemented as follows:
[0184] A deep learning model is used to extract sound features, visual motion features, and analyze vital sign changes from the fused multimodal data to obtain multimodal feature vectors.
[0185] Generative adversarial networks are used to perform sentiment analysis on multimodal feature vectors to obtain the emotion recognition results of animals.
[0186] By employing deep learning models to comprehensively analyze the fused multimodal data, features such as animal vocalizations, visual movements, and physiological changes can be accurately extracted to form multimodal feature vectors. Then, generative adversarial networks are used to perform in-depth sentiment analysis on these features, resulting in animal emotion recognition. This process not only improves the accuracy and depth of emotion recognition but also enhances the understanding of animal behavior and psychological states, thereby improving the recognition accuracy of animal emotion recognition models.
[0187] In some embodiments, if data synchronization and normalization are not performed during the preprocessing stage, the collected data can be timestamped, cleaned, and time-series aligned and fused during the preprocessing stage to obtain fused multimodal data that meets the above-mentioned standardized format. Subsequently, a deep learning model is used to extract sound features, visual motion features, and analyze physiological changes in the fused multimodal data to obtain multimodal feature vectors. Specifically, this can be implemented as follows:
[0188] Step A1: Use a deep learning model to extract sound features, visual motion features, and analyze physical changes in the fused multimodal data to obtain the first audio feature, the first visual feature, and the first physical feature.
[0189] That is, the fused multimodal data includes audio signals from the audio modality, video data from the video modality, and physiological data from the vital signs modality. Feature extraction is performed on the audio signals to obtain the first audio features, feature extraction is performed on the video data to obtain the first visual features, and feature extraction is performed on the physiological data to obtain the first vital signs features.
[0190] In implementation, the deep learning model may include an audio feature extraction module, a video feature extraction module, and a vital sign feature extraction module. Each module may employ networks such as CNN, RNN, and MLP, and this disclosure does not limit the specific network used.
[0191] Accordingly, based on a deep learning model, sound feature extraction, visual motion feature extraction, and vital sign change analysis are performed on the fused multimodal data to obtain the first audio feature, the first visual feature, and the first vital sign feature, which can be implemented as follows:
[0192] Step A11: Based on the audio feature extraction module in the deep learning model, feature extraction is performed on the audio signal in the fused multimodal data to obtain the first audio feature;
[0193] Step A12: Based on the video feature extraction module in the deep learning model, feature extraction is performed on the video data in the fused multimodal data to obtain the first visual feature;
[0194] Step A13: Based on the vital sign feature extraction module in the deep learning model, the physiological data in the fused multimodal data is extracted to obtain the first vital sign feature.
[0195] During implementation, the training samples are also stored in a standardized structure, which stores the fused multimodal data. Features can then be extracted from each modality at each time stamp from this fused multimodal data.
[0196] The aforementioned temporal alignment only aligns the time at a macro level. However, to compensate for the differences in sampling rates between different modalities, frame-level fine-tuning is needed to make the frame rates of different modalities the same. For example, assuming the frame rate of the first audio feature is 16 frames / s, the frame rates of the first visual feature and the first vital sign feature also need to be aligned to 16 frames / s, thereby achieving alignment of the sampling rates of different modalities.
[0197] In this embodiment of the disclosure, features of different modalities are extracted through different modules to facilitate feature modeling of a single modality, extract important features of a single modality for emotion recognition, and improve the accuracy of emotion recognition in animals.
[0198] Step A2: Synchronize the data of the first audio feature, the first visual feature, and the first vital sign feature to obtain the second audio feature, the second visual feature, and the second vital sign feature.
[0199] This data synchronization, or aligning the sampling rates of different modalities, can be achieved by upsampling to increase the frame rate or downsampling to decrease the frame rate, thus aligning the sampling rates between different modalities.
[0200] In implementation, for the first visual feature of the image modality, downsampling can be performed through pooling operations, or upsampling can be performed through the upsampling module of the neural network model. If data synchronization is required for the first audio feature of the audio modality, the processing method for the visual modality can be referenced.
[0201] For the primary vital sign features of the physiological modality, the processing methods for the visual modality can be referenced. Alternatively, data synchronization can be achieved through traditional data interpolation algorithms or downsampling. For example, a time window can be constructed with a step size of 1 second, and a time window can include multiple time steps. Within each time step, the mean, variance, maximum, and minimum statistical values of the vital sign data are extracted as the primary vital sign features. Then, through data synchronization, these are expanded into a feature sequence containing multiple frames of data. Each frame of data can contain multidimensional statistics (such as mean heart rate, coefficient of variation, etc.).
[0202] In summary, the process of synchronizing the data of the first audio feature, the first visual feature, and the first vital sign feature to obtain the second audio feature, the second visual feature, and the second vital sign feature can be summarized as follows: the number of frames of the first audio feature, the first visual feature, and the first vital sign feature within the current time window are processed into the target number of frames to obtain the second audio feature, the second visual feature, and the second vital sign feature.
[0203] For example, the number of frames of the first audio feature within the current time window can be determined as the target number of frames, and the number of frames of the first visual feature and the first physical feature within the current time window can be processed as the target number of frames respectively to obtain the second visual feature and the second physical feature; wherein, the first audio feature is used as the second audio feature.
[0204] In this embodiment of the disclosure, data synchronization is used to align the frame rates or sampling rates of different modalities, so as to achieve fine-tuned frame-level alignment within the same time window, thereby improving the accuracy of animal emotion recognition.
[0205] If normalization is not performed during the preprocessing stage, in order to align the signal ranges of different modes and prevent a single mode from dominating subsequent modeling, normalization can be performed through the following step A3.
[0206] Step A3: Normalize the second audio feature, the second visual feature, and the second physical feature to obtain a multimodal feature vector.
[0207] The normalization process may include Z-score normalization or Min-Max normalization.
[0208] Of course, if the previous preprocessing has already completed the normalization, step A3 can be omitted.
[0209] In this embodiment, independent features of each modality can be extracted separately, and then frame-level feature alignment can be achieved through data synchronization. This allows for a better description of the animal's state and reactions at the same time step, improving the accuracy of emotion recognition. Finally, normalization processing ensures that the signal ranges of different modalities are consistent, preventing a single modality from dominating emotion recognition and further improving the accuracy of animal emotion recognition.
[0210] In some embodiments, each multimodal feature vector is a feature representation of multiple modalities at the same time step. In practice, features from multiple time steps within the same time window can be modeled to better classify animal emotion recognition results. Based on this, using a generative adversarial network to perform emotion analysis on multimodal features to obtain animal emotion recognition results can be implemented as follows:
[0211] Step B1: Perform cross-modal feature fusion operation on the multimodal feature vectors to obtain fused features.
[0212] The multimodal feature vector includes feature vectors for audio modality, video modality, and physical feature modality. This includes the normalized second audio feature, second visual feature, and second physical feature, as described earlier.
[0213] Feature fusion operations can model the intrinsic relationships between different modalities. This can be understood as each multimodal feature vector corresponding to a time step; feature fusion operations can model the features of different modalities within the multimodal feature vectors at the same time step, thus uncovering the correlation patterns between different modalities at the same time step.
[0214] In implementation, a mid-level attention network can be used to fuse the features at the same time step. As shown in Figure 8, this can be implemented as follows:
[0215] S801, the feature vector of the audio modality in the multimodal feature vector is used as the first query vector, and one of the feature vectors of the video modality and the feature vector of the physiological modality in the multimodal feature vector is used as the first key vector, and the other is used as the first value vector. The first query vector, the first key vector and the first value vector are processed based on the attention mechanism to obtain the first feature to be fused.
[0216] During implementation, if normalization has been performed in the preprocessing stage, the aforementioned first audio feature can be used as the feature vector of the audio modality, the first visual feature as the feature vector of the video modality, and the first vital sign feature as the feature vector of the physiological modality.
[0217] Without normalization during the preprocessing stage, the aforementioned normalized second audio features can be used as the feature vector for the audio modality, the normalized second visual features as the feature vector for the video modality, and the normalized second vital sign features as the feature vector for the physiological modality.
[0218] In practice, in order to further eliminate the differences between feature spaces of different modalities and / or the dimensional differences between features of different modalities, in this embodiment of the disclosure, the feature vectors of each modality can first be aligned with the feature dimensions through a modality transformer, and the feature representations of different modalities can be mapped to the same feature space to obtain the final feature vectors used: the feature vectors of the audio modality, the feature vectors of the video modality, and the feature vectors of the physiological modality.
[0219] For example, if normalization is not performed during the preprocessing stage, the normalized second audio features need to be transformed using an audio modality transformer to obtain the audio modality feature vector. Similarly, the normalized second visual features are transformed using a video modality transformer to obtain the video modality feature vector. Likewise, the normalized second vital sign features are transformed using a vital sign modality transformer to obtain the physiological modality feature vector.
[0220] Of course, if normalization is completed during the preprocessing stage, then the aforementioned first audio feature is transformed using an audio modality transformer to obtain the audio modality feature vector. Similarly, the first visual feature is transformed using a video modality transformer to obtain the video modality feature vector. Likewise, the first vital sign feature is transformed using a vital sign modality transformer to obtain the physiological modality feature vector.
[0221] In implementation, the audio modality transformer, video modality transformer, and vital sign modality transformer can each be implemented using at least one fully connected layer, a small CNN, or a small RNN, or a combination thereof. Their function is to project features from different modalities onto a unified dimension and feature space. Each modality transformer is required not only to align dimensions but also to ensure the comparability of feature distributions.
[0222] The feature vectors of the video modality and the feature vectors of the physiological modality can be used, one as the first key vector and the other as the first value vector; this disclosure does not limit this.
[0223] S802, the feature vector of the video modality in the multimodal feature vector is used as the second query vector, and one of the feature vectors of the audio modality and the physiological modality in the multimodal feature vector is used as the second key vector, and the other is used as the second value vector. The second query vector, the second key vector and the second value vector are processed based on the attention mechanism to obtain the second feature to be fused.
[0224] Similarly, for the feature vectors of the audio modality and the physiological modality, one can be used as the second key vector and the other as the second value vector; this disclosure does not limit this.
[0225] S803 uses the feature vector of the physiological modality in the multimodal feature vector as the third query vector, and uses one of the feature vectors of the audio modality and the video modality in the multimodal feature vector as the third key vector and the other as the third value vector. Based on the attention mechanism, the third query vector, the third key vector and the third value vector are processed to obtain the third feature to be fused.
[0226] Similarly, for the feature vectors of the audio modality and the video modality, one can be used as the third key vector and the other as the third value vector; this disclosure does not limit this.
[0227] During implementation, in steps S801-S803, each step can use a corresponding cross-attention module to construct the corresponding feature to be fused.
[0228] S804, the first feature to be fused, the second feature to be fused, and the third feature to be fused are weighted and summed to obtain the fused feature.
[0229] Therefore, feature vectors from different modalities are used as query vectors, and feature vectors from other modalities are fused together. Finally, a weighted sum is used to obtain the fused feature at the same time step. A cross-modal attention mechanism enables dynamic interaction of information between different modalities, improving the fusion effect. This fusion process mimics the extension of the Transformer self-attention mechanism in multimodal scenarios, allowing the fused feature to model the association model between different modalities at the same time step. The resulting fused feature can better represent key features beneficial for emotion classification, thereby improving the effectiveness of animal emotion classification.
[0230] In some embodiments, a fixed weight can be set to perform a weighted summation of the first feature to be fused, the second feature to be fused, and the third feature to be fused.
[0231] In other embodiments, to improve the feature representation capability of the fused features, weights can be dynamically generated to perform a weighted summation of multiple features to be fused. This can be implemented as follows: for any one of the first, second, and third features to be fused, the weight of any one feature to be fused is dynamically generated based on that feature and learnable parameters from the training phase. Here, "any one feature to be fused" refers to each of the first, second, and third features to be fused. Weights are dynamically generated for each feature to be fused for feature fusion.
[0232] For example, the input of any feature to be fused corresponds to a neural network layer with learnable parameters. After processing by the neural network layer, the weights corresponding to any feature to be fused are output. This neural network layer can be, for example, a gating network.
[0233] In one example, the fusion network that fuses multimodal features at the same time step is shown in Figure 9, including an audio modality transformer 901, a video modality transformer 902, and a feature template transformer 903. The preprocessed features obtained from the aforementioned preprocessing (such as the first audio feature, first video feature, and first physical feature, or the normalized second audio feature, normalized second visual feature, and normalized second physical feature) are input into the corresponding modality transformers to achieve dimensional alignment and feature space alignment, thereby obtaining the feature vectors of the audio modality, the video modality, and the physiological modality. Then, the features pass through three sets of attention modules. Attention module 904 uses the feature vector of the audio modality as the first query vector, and the feature vectors of the other two modalities as the first key vector and the first value vector, respectively. Attention module 905 uses the feature vector of the video modality as the second query vector, and the feature vectors of the other two modalities as the second key vector and the second value vector, respectively. Attention module 906 uses the feature vector of the feature modality as the third query vector, and the feature vectors of the other two modalities as the third key vector and the third value vector, respectively. Attention module 904 outputs the first feature to be fused, attention module 905 outputs the second feature to be fused, and attention module 906 outputs the third feature to be fused.
[0234] The processing procedure of each attention module can be represented by the following formula (1):
[0235] In formula (1), Q i K represents the input query vector (i.e., the corresponding query vector); j V represents the input key vector; j d represents the input Value vector; d represents the dimension of the Key vector, used for scaling to prevent excessive gradients; Softmax(.) is used to control the attention distribution so that key and effective features can be learned through the attention mechanism.
[0236] The process of weighted summation of the three features to be fused can be completed by the weighted fusion layer 907 in Figure 9. The processing of the weighted fusion layer can be described as generating dynamic weights based on the following formula (2) and completing the weighted summation operation based on formula (3):
[0237] In formula (2), α i W is the normalized weight for the i-th mode; i X represents the learnable parameters of the i-th modality during the training phase when performing a weighted summation of multiple features to be fused; i W represents the feature to be fused for the i-th modality or the feature vector of the i-th modality; the denominator is used to normalize the fusion weights of multiple modalities; therefore, Wj X represents the learnable parameters of the j-th modality during the training phase when performing a weighted summation of multiple features to be fused; j This represents the feature to be fused for the j-th modality or the feature vector of the j-th modality.
[0238] By determining the fusion weights of multiple features to be fused in this way, normalization can be used to ensure that the sum of the weights of all modalities is 1, thus making the feature fusion dynamically adaptive.
[0239] After dynamically determining the weights of each feature to be fused based on formula (2), they can be weighted and summed based on the following formula (3):
[0240] In formula (3), F t The fusion features represent time step t; M is the total number of modes, α i X is the normalized weight for the i-th mode; i This represents the i-th feature to be fused.
[0241] Therefore, by dynamically determining the weights of multiple features to be fused at the same time step, the contribution of important modalities to emotion can be enhanced, thereby improving the accuracy of animal emotion recognition.
[0242] In implementation, the learnable parameters of the weights used to weight and sum multiple features to be fused during the training phase are optimized as training progresses. During the inference phase, these learnable parameters can be fixed, with the corresponding fusion weights primarily determined by the feature vectors of the respective modalities or the features to be fused. Thus, during the inference phase, the weights for weighted fusion are adaptively determined for each modality.
[0243] In order to more accurately understand the emotions of animals, in this embodiment of the disclosure, the temporal correlation of different time steps is modeled on the time series, as shown in step B2 below.
[0244] Step B2: Perform temporal modeling on the fused features to obtain temporal features.
[0245] Temporal modeling of the fused features can be implemented as follows:
[0246] Step C1: Encode the multiple fused features of multiple time steps within the current time window to obtain multiple enhanced features of multiple time steps within the current time window.
[0247] Each time step corresponds to its own fusion feature. In implementation, feature transformations can be performed on multiple fusion features across multiple time steps to obtain multiple enhanced features, where the feature dimension of the enhanced features is smaller than that of the fusion features. That is, feature transformations are performed on the fusion features corresponding to multiple time steps within the same time window (i.e., the current time window) to obtain the enhanced features corresponding to each time step. The purpose of feature transformation is to compress dimensionality, improve the speed of sentiment inference, and enhance feature expressive power.
[0248] In implementation, the fused features can first be compressed and nonlinearly transformed by the encoder to achieve feature transformation and enhance the feature representation capability. The encoder processes the fused features at each time step separately to obtain the enhanced features corresponding to each time step, thereby obtaining multiple enhanced features for the same time window.
[0249] For example, the encoder may include an activation layer (ReLU), a layer normalization layer (LayerNorm), and a deactivation layer (Dropout) for feature transformation. The encoder's processing can be shown in equation (4): Z t =Dropout(LayerNorm(ReLU(W1·F t +b1))) (4)
[0250] In formula (4), W1 represents the ReLU weight matrix; b1 represents the ReLU bias; F t Z represents the fusion feature at time step t. t This represents the enhanced feature after feature transformation by the encoder at time step t.
[0251] Activation layers can enhance nonlinearity; layer normalization layers can accelerate convergence and improve model training efficiency; and deactivation layers can achieve random deactivation to prevent overfitting.
[0252] In addition to the encoder structure expressed by formula (4) above, other structures such as convolutional networks (e.g., 1D-Conv), Transformer Encoders, and residual fully connected networks can also be used to construct encoders. During implementation, the choice can be flexible based on the data type and temporal characteristics.
[0253] Step C2 involves modeling multiple enhanced features from multiple time steps onto the time series to obtain multiple intermediate features for the current time window.
[0254] In practice, the evolution of sentiment over time can be modeled using LSTM (Long Short-Term Memory) networks. For example, BiLSTM (Bidirectional Long Short-Term Memory) networks can be used to process the enhanced features at each time step to obtain the intermediate features corresponding to each time step, thereby obtaining multiple intermediate features for the current time window.
[0255] Furthermore, to further improve feature representation capabilities, in this embodiment, the enhanced features of each time step are processed by BiLSTM to obtain the initial time features corresponding to each time step output by BiLSTM. The initial time features corresponding to each time step are then processed by a self-attention mechanism module to obtain the intermediate features corresponding to each time step. In this self-attention mechanism module, Q (Query), K (Key), and V (Value) all originate from the output of BiLSTM. The temporal context state generated by BiLSTM for each time step can be expressed as shown in formula (5): h t =BiLSTM(Z) t (5)
[0256] In formula (5), Z t The enhanced feature at time step t; h t This represents the intermediate features of time step t output by the BiLSTM, that is, the context state of time step t.
[0257] Therefore, temporal modeling can model the temporal evolution of animal emotions, thereby improving the accuracy of animal emotion recognition.
[0258] Step C3 involves weighted summation of multiple intermediate features within the current time window to obtain the time series features.
[0259] In this embodiment, feature encoding enhances the expressive power of the fused features at each time step. By performing temporal modeling on multiple time steps within the same time window, the temporal features can reflect the temporal context, thereby improving the accuracy of animal emotion recognition. Furthermore, by using a weighted fusion method to fuse intermediate features from different time steps to obtain temporal features, the contribution of features from different time steps to the final temporal features can be adjusted, facilitating the extraction of feature representations from important time steps and improving the accuracy of animal emotion recognition.
[0260] In some implementations, a fixed weight can be used to sum the intermediate features of multiple time steps within the same time window to obtain the time series features.
[0261] To improve the expressive power of temporal features, this embodiment employs dynamic weighting to fuse intermediate features from different time steps. This can be implemented as follows:
[0262] Step D1: Based on the weighted generation network, process multiple intermediate features of the current time window to obtain multiple weights corresponding to multiple time steps within the current time window.
[0263] That is, the intermediate features of each time step can be input into the weight generation network to obtain the weights for each time step. In implementation, the weight generation network can be a softmax network, a gated network, etc., and this disclosure does not limit it.
[0264] Taking the softmax network as an example, the weights at generation time step t are shown in formula (6): α t =Softmax(W ɑ ·h t +b ɑ (6)
[0265] In formula (6), α t W represents the weight of time step t; ɑ b represents the weight coefficients of the softmax network; ɑ Indicates the bias of the softmax network; h t This represents an intermediate feature at time step t.
[0266] In practice, the weight of each time step can ultimately be a normalized value.
[0267] Step D2: Based on multiple weights, perform a weighted summation of multiple intermediate features at multiple time steps within the current time window to obtain the time series features.
[0268] The obtained time series features can be represented by the following formula (7): h final =∑α t ·h t (7)
[0269] In formula (7), h final This represents the temporal characteristics that ultimately model the global context within the current time window; α t h represents the weight of time step t; t This represents an intermediate feature at time step t.
[0270] In this embodiment, temporal modeling of the global context within the same time window yields temporal features that describe global characteristics. Furthermore, these temporal features emphasize the contributions of different time steps, thereby enhancing high-quality features for animal emotion recognition and suppressing secondary features. This improves the classification ability of the final temporal features and enhances the accuracy of emotion recognition.
[0271] Step B3 involves classifying the temporal features to obtain the animal's emotion recognition results.
[0272] The classification operation is performed using a classification module. The network structure of this classification module can be established according to actual needs, and this embodiment does not limit it.
[0273] In this embodiment, cross-modal feature fusion operations are used to model the relationships between multimodal data at the same time step, thereby mining high-level features. Then, temporal modeling is used to mine the temporal feature expressions of animals, generating temporal features suitable for classification. Finally, classification operations are used to recognize animal emotions. The entire process is gradual and progressive, improving the accuracy of animal recognition.
[0274] In this embodiment of the disclosure, the classification loss can be calculated using the emotion recognition results and the corresponding training labels, thereby optimizing the model parameters of the animal emotion recognition model.
[0275] In some embodiments, to improve the robustness of the animal emotion recognition model, adversarial training can be used to train the model. Adversarial training involves adding perturbations to force the animal emotion recognition model to process noisy inputs, thereby optimizing the model parameters and reducing its sensitivity to noise. This ensures that even with some dirty data, the model's recognition results will not be affected.
[0276] During implementation, as shown in Figure 10, the animal emotion recognition model is trained through adversarial training, including the following:
[0277] S1001, add perturbation to the fusion feature corresponding to the multimodal feature vector at the current time step to generate the fusion feature of the multimodal noisy feature.
[0278] During implementation, the perturbation needs to be added based on the perturbation strength. The perturbation strength can be generated randomly.
[0279] In this embodiment of the disclosure, in order to add perturbation in the direction of maximum loss and increase the training pressure of the animal emotion recognition model on boundary samples, the perturbation intensity can be dynamically generated to add perturbation to the fused features of the multimodal feature vectors. This can be implemented as follows:
[0280] S10011, obtain the first statistical value in the first classification probability distribution of the processing result of the previous time step at the current time step; the processing result includes the sentiment classification result obtained by performing sentiment classification on the fused features of the previous time step.
[0281] In other words, the fused features corresponding to the multimodal feature vectors from the previous time step are obtained, and sentiment classification is performed on them to obtain the first-class probability distribution. This sentiment classification result is a first-class probability distribution, which represents the probability value of different animal sentiment categories.
[0282] The first statistical value can be the maximum value in the first classification probability distribution, or the difference between the maximum and minimum values.
[0283] S10012, based on the negative correlation between the intensity and the first statistical value, the first disturbance intensity is generated.
[0284] The negative correlation can be linear or nonlinear, and can be determined according to actual needs. This disclosure does not limit this.
[0285] For example, in one possible implementation, the first disturbance intensity is determined as shown in the following formula (8): ε=ε0·(1-p conf ) p conf =max(Softmax(f(F) t-1 ))) (8)
[0286] In formula (8), ε represents the first disturbance intensity; ε0 represents the reference value of the disturbance intensity; F t-1 This represents the fusion feature from the previous time step, i.e., time step t-1. Softmax(f(F) t-1 )) indicates that for input F t-1 The classification process is performed to categorize animal emotions, resulting in the first category probability distribution; max(Softmax()) indicates that the maximum value output by Softmax() is taken as the confidence level p. conf .
[0287] S10013, based on the first perturbation intensity, the fusion features corresponding to the multimodal feature vectors, and the classification loss gradient of the fusion features corresponding to the multimodal feature vectors, generate the fusion features of the multimodal noisy features.
[0288] Multimodal noise-adding feature F′ t The data is input into the animal emotion recognition model. Skipping the process of feature fusion of multimodal data in the animal emotion recognition model, the fused features of the resulting multimodal noisy features are regarded as fused features and used for subsequent time series modeling and finally for animal emotion classification.
[0289] Finally, the fused feature F′ of the multimodal noisy features is obtained. t The process can be shown in formula (9):
[0290] In formula (9), F′ t F represents the fusion feature of multimodal noise features at time step t; t L represents the fused feature of the multimodal feature vector at the current time step; cls (F t ,y) represents the fusion feature F t The classification loss for sentiment classification can be cross-entropy loss; y represents the training labels. This represents the gradient of the classification loss of the fused features at time step t.
[0291] Therefore, the method of generating the first perturbation intensity in this way can reduce the perturbation of high-confidence samples, increase the training pressure of the animal emotion recognition model on boundary samples, and improve the robustness of the model and the accuracy of animal emotion recognition.
[0292] S1002, the fused features of multimodal noisy features are input into the animal emotion recognition model to identify the current emotion of the animal and obtain the emotion prediction result of the animal.
[0293] This can be understood as treating the fused features of multimodal noisy features as Ft participating in the subsequent processing of the animal emotion recognition model. That is, skipping the process of feature fusion operation of multimodal features by the fusion network in the animal emotion recognition model, and using the module after the fusion network in the animal emotion recognition model to continue processing the fused features of multimodal noisy features, thereby completing the animal emotion classification and obtaining the animal emotion prediction result.
[0294] It's important to distinguish between these different concepts. During the training phase, the emotion recognition result refers to the animal emotion identified based on the fusion of multimodal feature vectors before perturbation is added; the emotion prediction result refers to the animal emotion identified based on the fusion of multimodal noisy features after perturbation is added. The two are both different and related.
[0295] In this embodiment of the disclosure, adding perturbations to the fused features at time step t can enhance the robustness of the animal emotion recognition model through adversarial training.
[0296] In other embodiments, to avoid introducing perturbations prematurely and causing the model to become unstable during training, the animal emotion recognition model can be trained through adversarial training, as shown in Figure 11, including:
[0297] S1101, add perturbation to the feature vector of the target mode in the multimodal feature vector to obtain the perturbation feature of the target mode.
[0298] The types and number of modalities used can vary at different training stages. For example, in the initial stage, video modalities can be selected as the target modalities to add perturbations. As training progresses and the recognition accuracy and robustness of the animal emotion recognition model improve, audio modalities and / or physiological modalities can be gradually added as target modalities.
[0299] Among them, the feature vectors of video mode, audio mode, and physiological mode in the multimodal feature vector can all be used as the feature vectors of the target mode.
[0300] When there are multiple target modes, a perturbation is added to each target mode separately, and the method for adding perturbations to different target modes is consistent.
[0301] In implementation, taking a target mode as an example, the perturbation needs to be added based on the perturbation strength. The perturbation strength can be generated randomly.
[0302] In this embodiment of the disclosure, in order to add perturbations in the direction of maximum loss and improve the robustness of the augmentation model as much as possible, the perturbation intensity depends on the classification confidence of the target mode in the previous time step. Furthermore, the higher the classification confidence in the previous time window, the less emphasis is placed on the perturbation. Specifically, the adaptive determination of the perturbation intensity to generate perturbation features can be implemented as shown in Figure 11:
[0303] S11011, obtain the second statistical value in the second classification probability distribution of the target modality in the previous time step of the current time step; the second classification probability distribution is the classification probability distribution obtained by performing sentiment classification on the feature vector of the target modality.
[0304] In practice, the animal's emotions can be classified based on the feature vector of the target modality from the previous time step, yielding an emotion classification result. This result is a secondary classification probability distribution, representing the probability values of different emotion categories. The secondary statistical value can be the maximum value in this probability distribution, or the difference between the maximum and minimum values.
[0305] S11012, based on the negative correlation between the intensity and the second statistical value, generates the second perturbation intensity.
[0306] The negative correlation can be linear or nonlinear, and can be determined according to actual needs. This disclosure does not limit this.
[0307] For example, in one possible implementation, the second disturbance intensity is determined as shown in the following formula (10): ε1=ε0·(1-p conf )
[0308] In formula (8), ε1 represents the second disturbance intensity; ε0 represents the reference value of the disturbance intensity; This represents the classification feature of the feature vector of the target mode at the previous time step, i.e., time step t-1. Indicates input The classification process is performed to categorize animal emotions, resulting in a second-category probability distribution; max(Softmax()) represents taking the maximum value of the output of Softmax() as the confidence level p. conf .
[0309] The feature vector of the target modality (such as the feature vector of the audio modality) can itself serve as its classification feature. Alternatively, the feature vector of the target modality can be extracted using a classification model to obtain classification features. These classification features are used to predict the animal's emotion category based on the target modality. The classification model can be determined according to actual needs; existing classification models that can roughly classify an animal's emotion based on a single target modality are sufficient.
[0310] S11013, based on the second perturbation intensity, a perturbation is added to the feature vector of the target mode in the multimodal feature vector to obtain the perturbation feature of the target mode.
[0311] During implementation, the perturbation characteristics F″ of the target mode are finally obtained based on the second perturbation intensity. t The process can be shown in formula (11):
[0312] In formula (9), F″ f The perturbation characteristics of the target mode at time step t; The feature vector representing the target mode at the current time step; The classification loss for classifying animal emotions is represented by the classification features of the feature vector of the target modality in the previous time step, which can be the cross-entropy loss; y is the training label. This represents the gradient of the classification loss of the fused features at time step t.
[0313] By generating a reasonable second perturbation intensity and adding perturbations to the feature vectors of the corresponding target modality, perturbations can be introduced into the feature vectors of different modalities at appropriate stages, thereby improving the training efficiency and robustness of the animal emotion recognition model.
[0314] S1102, replace the feature vector of the target mode in the multimodal feature vector with the perturbation feature of the target mode to obtain the multimodal noise-added feature.
[0315] That is, the multimodal feature vector includes the feature vectors of the audio modality, the video modality, and the physiological modality. Assuming that the video modality is the target modality, the feature vectors of the audio modality and the physiological modality remain unchanged. By replacing the feature vector of the video modality with the perturbation feature of that video modality, the multimodal noise-added feature can be obtained.
[0316] Assuming the video modality and audio modality are the target modalities, the feature vector of the physiological modality remains unchanged. By replacing the feature vector of the video modality with the perturbation feature of the video modality, and replacing the feature vector of the audio modality with the perturbation feature of the audio modality, the multimodal noise-added features can be obtained.
[0317] S1103, input the multimodal noisy features into the animal emotion recognition model to identify the animal's current emotion and obtain the animal's emotion prediction result.
[0318] In other words, multimodal noisy features are input into the animal emotion recognition model to complete the entire processing flow of the animal emotion recognition model, including multimodal feature fusion operations and subsequent temporal modeling and classification processing, and finally obtain the emotion prediction results for multimodal noisy features.
[0319] Understandably, this embodiment still needs to distinguish between different concepts. During the training phase, the emotion recognition result refers to the animal emotion identified based on multimodal feature vectors before perturbation is added; the emotion prediction result refers to the animal emotion identified based on multimodal noisy features after perturbation is added. The two are both different and related.
[0320] In this embodiment of the disclosure, adding perturbations to the feature vector of the target modality can enhance the robustness of the animal emotion recognition model through adversarial training.
[0321] In implementation, perturbations can be added only to the feature vectors of the video modality in the initial stage to complete adversarial training. Once the stability of the animal emotion recognition model meets preset requirements, adversarial training of feature vectors from other modalities can be further added. Then, based on the method shown in Figure 10, perturbations can be added to the fused features of the multimodal feature vectors to continue adversarial training.
[0322] The model parameters of the animal emotion recognition model are optimized through adversarial training. Specifically, at the training loss level, the animal emotion recognition model is trained based on the emotion recognition results, as shown in Figure 12:
[0323] 1201. Based on the emotion recognition results and training labels, determine the first classification loss.
[0324] 1202. Based on the sentiment prediction results and training labels, the second classification loss is determined.
[0325] 1203. Based on the first classification loss and the second classification loss, the model parameters of the animal emotion recognition model are optimized.
[0326] In other words, the fused features before and after the perturbation are both inferred and predicted by the animal emotion recognition model to generate corresponding emotion recognition or emotion prediction results. Both results are compared with the corresponding training labels to determine the cross-entropy loss, thus obtaining the first classification loss and the second classification loss.
[0327] Two classification losses are used to optimize the model parameters of the animal emotion recognition model, enabling the model to not only handle fused features without perturbation, but also to correctly classify the emotions of perturbated fused features, thereby improving the classification accuracy and robustness of the animal emotion recognition model.
[0328] In other embodiments, consistency loss can be introduced during adversarial training to optimize model parameters. As shown in Figure 12:
[0329] 1204. The temporal features of the multimodal feature vectors obtained during the processing of multimodal feature vectors by the animal emotion recognition model are used for emotion classification to obtain the emotion recognition result.
[0330] 1205. The temporal features of the multimodal noise-added features obtained during the processing of multimodal noise-added features by the animal emotion recognition model are used for emotion classification to obtain emotion prediction results.
[0331] 1206. Based on the first consistency loss between the temporal features of multimodal feature vectors and the temporal features of multimodal noisy features, the model parameters of the animal emotion recognition model are optimized.
[0332] In this embodiment of the disclosure, the first consistency loss can be used to measure the model's ability to model features before and after the perturbation, thereby increasing the robustness of the model.
[0333] Furthermore, as shown in Figure 12, a second consistency loss can be determined between the emotion recognition result and the emotion prediction result; based on the second consistency loss, the model parameters of the animal emotion recognition model can be optimized.
[0334] In this model, both the emotion recognition result and the emotion prediction result are classification probability distributions, representing the probability of belonging to different animal emotion categories. Therefore, the parameters of the animal emotion recognition model can be optimized by ensuring the consistency of the final output.
[0335] In this embodiment of the disclosure, the second consistency loss can be used to measure the model’s ability to model features before and after the perturbation, thereby increasing the robustness of the model.
[0336] Figure 13 shows the main framework structure of the animal emotion recognition model provided in this embodiment. The multimodal data undergoes preprocessing, which includes data cleaning and macroscopic alignment of multimodal data with the same real-world data based on timestamps, ultimately resulting in standardized, fused multimodal data. This fused multimodal data includes audio signals, video data, and physiological data.
[0337] The audio signal in the fused multimodal data is processed by the audio feature extraction module 13011 in the deep learning network 1301 to extract features, resulting in the first audio feature. The video data in the fused multimodal data is processed by the video feature extraction module 13012 to extract features, resulting in the first visual feature. The physiological data in the fused multimodal data is processed by the vital sign feature extraction module 13013 to extract features, resulting in the first vital sign feature. The data synchronization module 1302 synchronizes the first audio feature, the first visual feature, and the first vital sign feature to align the frame rates of different modalities, resulting in the second audio feature, the second visual feature, and the second vital sign feature. Then, the second audio feature, the second visual feature, and the second vital sign feature are normalized to obtain the normalized second audio feature (hereinafter referred to as the third audio feature), the normalized second visual feature (hereinafter referred to as the third visual feature), and the normalized second vital sign feature (hereinafter referred to as the third vital sign feature). The third audio feature, third visual feature, and third physical feature are processed by a fusion network 1303 based on a cross-attention mechanism to obtain the fusion feature Ft at time step t. The fused multimodal data at each time step within the current time window are processed as described above to obtain the fusion features corresponding to each time step. The fusion features at each time step within the current time window are processed by an encoder 1304 to obtain the enhanced features for each time step. The enhanced features at each time step are input to a BiLSTM 1305 to obtain the initial temporal features for each time step. The initial temporal features at each time step are input to a self-attention mechanism module 1306 to obtain the intermediate features for each time step. The intermediate features at each time step are weighted and summed by a dynamic weighting layer 1307 to obtain the temporal feature h of the current time window. final Temporal characteristics h final Inputting the data into the classification module 1308 yields the animal's emotion recognition results.
[0338] During the inference phase, the processing flow of the animal emotion recognition model is the same as shown in Figure 13. During the training phase, this embodiment of the present disclosure further adds adversarial training. Specifically, by adding perturbations, the robustness and boundary handling capability of the animal emotion recognition model are improved. Perturbations can be added at each time step. Regardless of the method used to add perturbations, each time step will obtain a fused feature containing the perturbated multimodal noisy features. The fused features of the multimodal noisy features from these multiple time steps are processed by the encoder 1304, BiLSTM 1305, self-attention mechanism module 1306, and dynamic weighting layer 1307 in Figure 13 to obtain the perturbated temporal feature h. final ', containing the temporal characteristics h final After processing by classification module 1308, the sentiment prediction result is obtained. Therefore, the sentiment recognition result and sentiment prediction result within the same time window can be used to calculate the classification loss separately, and the temporal features h within the same time window... final and time series features h final 'It is necessary to calculate the consistency loss, which can be used to optimize the model parameters of the animal recognition model.'
[0339] Specifically, the above three types of losses are collectively referred to as adversarial losses, and the expression for adversarial losses is shown in formula (12): L total =L cls (F t ,y)+λ1·L cls (F′ t ,y)+λ2·L consistency L consistency =||f(F t )-f(F′ t )|| 2 (12)
[0340] In formula (12), L total Indicates resistance to loss; L cls (F t (y) represents the classification loss of the emotion recognition result before perturbation; L cls (F′ t (y) represents the classification loss of the sentiment prediction result obtained after perturbation; L consistency Represents the consistency loss (such as the first consistency loss or the second consistency loss); represents the weighting coefficient.
[0341] The entire model's parameters are optimized under the joint constraints of the three factors included in the adversarial loss during training. While the perturbation paths appear independent, they are actually highly coupled. This can be understood as the three branches (normal input, perturbation input, and consistency constraint) sharing parameters, with a unified optimization objective. This improves the model's robustness and generalization ability to anomalous samples and reduces misclassification of pseudo-labeled samples.
[0342] In some embodiments, since the number of labeled data for the sample data is limited, in order to effectively utilize the sample data, as shown in Figure 14, it can also be implemented as follows:
[0343] S1401, a pseudo-label for generating animal emotions.
[0344] In this context, pseudo-labels and expert labels are relative concepts. Expert labels refer to training labels obtained by accurately annotating sample data. In this embodiment, pseudo-labels mainly refer to the results output by the animal emotion recognition model after reasoning and analyzing the sample data; these pseudo-labels are also part of the training labels.
[0345] S1402, an animal emotion recognition model trained based on pseudo-labels.
[0346] In this embodiment of the disclosure, by automatically generating pseudo-labels, sample data can be effectively utilized to improve the generalization ability of the animal emotion recognition model.
[0347] In some embodiments, in order to efficiently utilize sample data with expert labels to generate pseudo-labels for animal emotions, as shown in Figure 15, the following implementation is possible:
[0348] S1501, under the condition that the animal emotion recognition model meets the preset conditions, input the sample data with expert labels into the animal emotion recognition model to obtain the predicted labels of the sample data output by the animal emotion recognition model.
[0349] To ensure the reasonable prediction of false labels, this embodiment requires the animal emotion recognition model to meet preset conditions, meaning the output of the animal emotion recognition model must have a certain degree of accuracy. In implementation, the preset conditions include at least one of the following:
[0350] 1) The animal emotion recognition model is trained through training rounds based on training samples with expert labels, until the target number of rounds is reached.
[0351] 2) The animal emotion recognition model meets the set conditions for entering the next stage of training based on the emotion recognition accuracy of the training samples with expert labels.
[0352] This setting could be that the accuracy of animal emotion recognition is greater than a preset accuracy threshold, or that the accuracy no longer increases significantly but fluctuates within a certain range. For example, the accuracy fluctuation range could be within the target range.
[0353] In summary, regardless of conditions 1) or 2), after the animal emotion recognition model is trained and optimized based on sample data with expert labels, the animal emotion recognition model has learned certain knowledge and can reasonably predict animal emotions. At this time, using the animal emotion recognition model to generate pseudo-labels can improve the quality of pseudo-label generation.
[0354] The predicted labels output by the animal emotion recognition model for the sample data represent the animal emotion classification distribution. In other words, the animal emotion recognition model outputs the animal emotion recognition result for the sample data. This predicted label includes the probability distribution of each animal emotion type.
[0355] S1502, the predicted labels and expert labels of the sample data are weighted and summed to obtain the pseudo labels of the sample data.
[0356] In this embodiment, a pseudo-label can be generated by weighted summation of predicted labels inferred from the animal emotion recognition model and expert labels. This allows for the reasonable reuse of sample data with expert labels based on the pseudo-labels, improving the utilization rate of sample data and the generalization ability of the animal emotion recognition model.
[0357] In one implementation, the two can be weighted and summed using fixed weights.
[0358] In other implementations, the generation of pseudo-labels through weighted calculation can also be implemented as shown in Figure 15:
[0359] S15021, Obtain the confidence level of the animal emotion recognition model for the emotion classification distribution of multiple training samples in the previous round.
[0360] Among them, the multiple training samples from the previous round can be all the samples used in the previous training process, or unlabeled samples without expert labels used in the previous training process.
[0361] S15022, determine the statistical value of the confidence level of the sentiment classification distribution of multiple training samples as a dynamic threshold.
[0362] For each training sample, the animal emotion recognition model infers and obtains a predicted label. This predicted label includes the probability of belonging to each animal emotion category. In practice, the maximum probability can be taken as the confidence level of the emotion classification distribution of that training sample.
[0363] For example, an unlabeled sample is input into an animal emotion recognition model, which outputs the emotion classification distribution of the unlabeled sample, that is, the probability distribution of belonging to each animal emotion category. Assuming that the emotion classification distribution is [0.15, 0.10, 0.70, 0.05], the maximum value of 0.7 is taken as the confidence level of the emotion classification distribution of the training sample.
[0364] Multiple training samples correspond to multiple confidence levels. From this, the statistical values of these multiple confidence levels can be determined. These statistical values can be understood as the mean, maximum value, etc., used to represent the overall confidence level of the animal emotion recognition model at its current stage.
[0365] For example, during implementation, the 80th percentile of the confidence distribution can be taken as the pseudo-label acceptance threshold for the current round, which is also known as the dynamic threshold.
[0366] When the dynamic threshold is determined using unlabeled samples from the previous round, it reflects the inference and prediction capabilities of the animal emotion recognition model for unlabeled samples, and can reflect the accuracy of the predicted labels generated by the animal emotion recognition model.
[0367] S15023, generates the fusion weight of expert labels and the fusion weight of predicted labels based on dynamic thresholds; where the closer to the dynamic threshold, the greater the fusion weight of the expert labels.
[0368] For example, the fusion weight of expert labels can be determined based on the following formula (13): α=exp(-β(p i -p thresh ) 2 (13)
[0369] In formula (13), α represents the fusion weight of the expert label; the complement of α, i.e. (1-α), is the fusion weight of the predicted label; p thresh β represents the dynamic threshold; p represents the hyperparameter that adjusts the steepness of the curve. i This represents the confidence level of the predicted label, which is the maximum probability among the predicted labels.
[0370] Based on formula (13), it can be deduced that α characterizes the level of trust in the expert label. When the confidence level of the predicted label is just close to the dynamic threshold, (p i -p thresh The value is close to 0, indicating that the contribution of expert labels is extremely large. Therefore, the predicted labels will be treated conservatively during the fusion process.
[0371] When the confidence level of the predicted label is much greater than the dynamic threshold, (p i -p threshWhen the confidence level is much greater than 0, α rapidly approaches 0, significantly increasing the fusion weight of the predicted labels. Therefore, the predicted labels are automatically trusted during the fusion process. If the confidence level of a predicted label is less than the threshold, the predicted label is unreliable and should not be used to generate pseudo-labels.
[0372] S15024: Based on the fusion weights of expert labels and predicted labels, the predicted labels and expert labels are weighted and summed to obtain pseudo-labels for the sample data.
[0373] The expression for generating pseudo-tags is shown in formula (14) below: y va =α·y expert_va +(1-α)·y pseudo_va (14)
[0374] In formula (14), α represents the fusion weight of the expert label; (1-α) is the fusion weight of the predicted label; y expert_va Indicates expert label; y pseudo_va Indicates predicted label; y va This indicates the generated pseudo-tag.
[0375] Therefore, when the confidence level of the predicted label approaches the dynamic threshold, it is in a critical range and easily affected by noise and the uncertainty of the animal emotion recognition model. At this point, increasing the fusion weight of expert labels can effectively suppress label drift and the generation of inferior pseudo-labels, avoiding interference to the animal emotion recognition model. As the performance of the animal emotion recognition model improves, the overall confidence distribution will shift to the right (i.e., increase), and the dynamic threshold will be adjusted upwards. Thus, the dynamic threshold allows the fused pseudo-labels to become increasingly of higher quality, gradually increasing trust in the pseudo-labels and accelerating the model's self-evolution loop. Finally, generating pseudo-labels through the method provided in this embodiment can suppress the weight of low-confidence predicted labels, effectively reducing the cumulative error caused by pseudo-label diffusion, and ultimately improving the accuracy of the animal emotion recognition model.
[0376] Building upon this foundation, pseudo-labels are first generated using labeled samples with expert tags to improve the utilization rate of labeled samples and the ability of the animal emotion recognition model to generate pseudo-labels. As the quality of the pseudo-labels generated by the animal emotion recognition model improves, predictive labels can be generated for unlabeled samples using the model. If the confidence level of this predicted label is higher than the dynamic threshold, it can be further used as a pseudo-label, thereby constructing (unlabeled sample, pseudo-label) sample pairs to further optimize the animal emotion recognition model. For example, in experiments, there were 1000 sample data points with expert tags; among the unlabeled samples, pseudo-labels were generated through prediction, with 60% having a confidence level ≥ 0.85. These high-confidence pseudo-labels can be used to optimize the model parameters.
[0377] In some embodiments, animals may express emotions in different modalities. Therefore, the type of animal emotion can be roughly identified from data in a single modality. In practice, to extract more pseudo-labels and high-quality sample data for learning, in this embodiment of the disclosure, generating pseudo-labels for animal emotions can also be implemented as shown in Figure 16:
[0378] S1601, determine the degree of classification divergence among multiple modalities in animal-related multimodal data.
[0379] The multimodal data includes audio modal data, video modal data, and physiological modal data. Each modality of data can theoretically be used independently to identify the animal's emotional category.
[0380] Classification divergence measures the degree of difference between animal emotion categories expressed by different modalities. A higher classification divergence indicates greater discrepancy and classification conflict between different modalities.
[0381] In one possible implementation, the classification divergence can be described based on the degree of difference between the feature vectors of different modalities, as shown in the following formula (15):
[0382] In formula (15), A avg The divergence index indicates the degree of classification divergence; a larger divergence index indicates a smaller degree of classification divergence. audio The audio vector to be compared can be represented by the feature vector of the audio mode output by the audio mode converter in Figure 13; similarly, f video The video vector to be compared can be represented by the feature vector of the video mode output by the video mode converter in Figure 13; similarly, f bio The physiological vector to be compared can be represented by the feature vector of the physiological mode output by the vital sign modality transformer in Figure 13. It can be understood that the audio vector to be compared, the audio vector to be compared, and the physiological vector to be compared need to be in the same feature space, and the implementation is not limited to the feature vector shown in Figure 13.
[0383] In other embodiments, the degree of classification divergence can also be represented based on the following method, specifically:
[0384] Step E1 involves determining multiple sentiment classification results based on animal-related multimodal data. Each modality of data corresponds to its own sentiment classification result.
[0385] The multimodal data includes audio modal data, video modal data, and physiological modal data. Each modal data point can be used as target modal data to perform the following operations:
[0386] Feature extraction is performed on the target modal data to obtain classification features;
[0387] Animal emotions are classified based on categorical features to obtain the emotion classification results. These results include the probability distribution of each animal's emotion.
[0388] Step E2: Determine the degree of classification discrepancy among multiple sentiment classification results in the multimodal data.
[0389] It can calculate the differences between sentiment classification results of each pair of modalities across multiple modalities and then take the mean. For example, this difference can be expressed as the difference, L2 distance, or cosine similarity.
[0390] S1602, when the classification divergence is less than the divergence threshold, merge the sentiment classification results of multiple modalities in the multimodal data to obtain pseudo-labels.
[0391] In other words, animal emotion recognition is performed separately for each modality of data, resulting in emotion classification results for each modality. In implementation, this animal emotion recognition model can be used to perform animal emotion recognition on the data for each modality, or other classification models can be used to perform animal emotion recognition on each modality separately.
[0392] In this embodiment of the disclosure, when the classification divergence is small, it indicates that the sentiment types represented by each modality in the multimodal data tend to be consistent. This multimodal data is a high-quality sample, and corresponding pseudo-labels can be automatically generated for model training.
[0393] In one possible implementation, the sentiment classification results of multiple modalities can be weighted and summed using fixed weights based on empirical values.
[0394] In other embodiments, to improve the quality of the pseudo-label, it can be generated based on dynamic weights, as shown in Figure 16:
[0395] S16021, Based on the classification confidence of multiple sentiment classification results of multiple modalities, generate multiple weighted values for multiple modalities.
[0396] That is, each mode corresponds to a weighted value.
[0397] For each modality, the classification confidence score of the sentiment classification result can be understood as the maximum probability value in that sentiment classification result. Essentially, this is consistent with the meaning of the confidence score of the sentiment classification distribution explained earlier.
[0398] For example, the process of generating weighted values for each modality is shown in formula (16): w m =Softmax(p conf (16)
[0399] In formula (16), pconf This represents the classification confidence score of the sentiment classification result for the corresponding modality. Softmax is used to represent the function that generates the weighted values. In addition to Softmax, gating networks can also be used, but this disclosure does not limit the implementation of such networks.
[0400] S16022 generates pseudo-labels for sample data representing multimodal data based on multiple weighted values and multiple sentiment classification results from multiple modalities.
[0401] For example, to improve the credibility of pseudo-tags, an activation layer can be used to generate pseudo-tags, as shown in the following formula (17):
[0402] y pseudo =Softmax(∑ m w m ·y m (17)
[0403] In formula (17), Softmax is the activation layer, and y m The sentiment classification result for the m-th modality; w m This represents the weighted value for the m-th modality. Activation layers can effectively filter out important content and improve the credibility of pseudo-tags.
[0404] In this embodiment of the disclosure, the classification confidence score is used to generate weights, which can better measure the recognition accuracy of each modality for animal emotion categories. The higher the accuracy, the greater its contribution to animal emotion classification, and the better the quality of the generated pseudo-labels.
[0405] In some embodiments, if the classification divergence is not less than a divergence threshold, the multimodal data is stored in an observation pool. After entering the observation pool, usable multimodal data can be selected as training samples based on a preset strategy, or corresponding training labels can be obtained through manual annotation. Thus, sample data that cannot automatically generate pseudo-labels is filtered out based on the classification divergence, making them available for subsequent processing.
[0406] In some embodiments, the quality of generated pseudo-labels needs to be controlled to ensure that high-quality pseudo-labels are used in the training of the animal emotion recognition model. Therefore, if the confidence level of a pseudo-label is greater than a dynamic threshold, the sample data corresponding to the pseudo-label is determined as the sample data for training the animal emotion recognition model. The dynamic threshold is determined based on a statistical value, which is the statistical value of the confidence level of the emotion recognition model for the emotion classification distribution of multiple training samples from the previous round. The method for determining the dynamic threshold has already been explained above and will not be repeated here.
[0407] The confidence level of a pseudo-label can be understood as the maximum probability value in the probability distribution of the generated pseudo-labels. When the pseudo-labels are generated based on a weighted sum of expert labels and predicted labels, the confidence level of the pseudo-label is the confidence level of the predicted label used.
[0408] By selecting pseudo-labels with high confidence levels to participate in the training of animal emotion recognition models, the influence of inferior labels on the models can be reduced, thereby improving the training efficiency of animal emotion recognition models.
[0409] In some embodiments, if the confidence level of a pseudo-label is no greater than a dynamic threshold, the sample data corresponding to the pseudo-label is stored in the observation pool. This allows for the screening of sample data with substandard pseudo-label quality, facilitating subsequent analysis and optimization of sample data and label quality.
[0410] In some embodiments, regardless of the method used to generate the pseudo-label, the pseudo-label and its corresponding sample data can be used to optimize the animal emotion recognition model, as shown in Figure 17:
[0411] S1701, determine at least one of the following training losses based on pseudo-labels: cross-entropy loss of pseudo-labels, loss between pseudo-labels and expert labels, and alignment loss between pseudo-label distribution and expert label distribution.
[0412] Specifically, for each pseudo-label, the cross-entropy loss is used to measure the cross-entropy loss between the emotional classification distribution output by the animal emotion recognition model after processing the sample data and the corresponding pseudo-label.
[0413] When the sample data has corresponding expert labels, the loss between the generated pseudo-labels and the expert labels can be introduced to constrain the output of the animal recognition model to be as consistent as possible with the professional labels, thereby improving the utilization rate of labeled samples.
[0414] The alignment loss between the pseudo-label distribution and the expert label distribution, as the name suggests, is used to constrain the probability distribution of multiple generated pseudo-labels and their corresponding multiple expert labels in the animal emotion category in the labeled sample set to be consistent.
[0415] S1702, determine the total loss based on at least one training loss, and optimize the animal emotion recognition model based on the total loss.
[0416] This disclosure discloses an animal emotion recognition model that generates a high-quality emotion classification distribution by introducing cross-entropy loss, thereby optimizing the model's performance on unlabeled data. The loss between pseudo-labels and expert labels allows the animal emotion recognition model to learn more fully from labeled data. Alignment loss enables the animal emotion recognition model to learn a high-quality label distribution, improving model convergence efficiency.
[0417] In some embodiments, the total loss includes the total cross-entropy loss of multiple sample data. The determination of the total cross-entropy loss can be implemented as shown in Figure 17:
[0418] S1703 generates multiple loss weights corresponding to multiple sample data; wherein, multiple sample data have multiple corresponding pseudo-labels.
[0419] The loss weights for each sample data can be dynamically generated. These loss weights are used to measure the contribution of the sample data to the total loss, thereby increasing the learning rate of high-quality sample data and reducing the learning of relatively poor-quality sample data.
[0420] The generation of loss weights can be implemented as follows: based on the difference between multiple pseudo-labels of multiple sample data and their corresponding expert labels, multiple loss weights corresponding to multiple sample data are generated; wherein, the difference and the loss weights have an inverse correlation. For example, it can be expressed as shown in the following formula (18):
[0421] In formula (18), w i P represents the loss weight for the i-th sample data; e (y i ) indicates an expert label; P p (y i ) represents a pseudo-label; ε represents a constant parameter with a value lower than the preset lower limit. Its value is extremely low, which means it does not affect the calculation of the loss weight and can also be used to avoid the denominator being 0.
[0422] The method of determining the loss weights ensures that the closer the pseudo-labels are to the expert labels, the higher the weights. This generates higher weights for high-quality pseudo-label sample data, enabling the animal emotion recognition model to learn superior feature modeling methods.
[0423] S1704, based on multiple loss weights corresponding to multiple sample data, performs a weighted summation of multiple cross-entropy losses of multiple sample data to obtain the total cross-entropy loss.
[0424] For example, the total cross-entropy loss is shown in the following formula (19):
[0425] In formula (19), w i y represents the loss weight of the i-th sample data. soft The distribution of animal emotion classification inferred by the animal emotion recognition model for the i-th sample data is a probability distribution; as the model parameters are adjusted, the prediction results for the same sample data may change. This indicates the generated pseudo-tag.
[0426] In this embodiment of the disclosure, reasonable loss weights are generated for different sample data, so that the total cross-entropy loss can better measure the shortcomings of the animal emotion recognition model, so as to adjust the model parameters in a targeted manner and improve the model training efficiency.
[0427] In some embodiments, the total loss includes the total alignment loss of multiple sample data, which measures the distribution difference between pseudo-labels and expert labels of multiple sample data; during the training phase, the total alignment loss is minimized by optimizing the model parameters of the animal emotion recognition model.
[0428] The distributional difference between pseudo-labels and expert labels can be represented by the KL (Kullback-Leibler Divergence) loss. Correspondingly, the total divergence loss is calculated as shown in formula (20):
[0429] In formula (20), L KL P represents the KL divergence loss; e (y k P represents the probability distribution of expert labels on the emotions of the k-th animal class; p (y k ) represents the distribution probability of pseudo-labels on the emotions of the k-th animal class; ε takes a very low value, that is, it does not affect the calculation of loss weights, and can also be used to avoid the denominator being 0.
[0430] In other embodiments, the distribution difference between pseudo-labels and expert labels can also be the JS (Jensen-Shannon Divergence) between them. The JS divergence for the distribution probability of each animal emotion category is shown in Equation (21):
[0431] In formula (21), p expert p represents the distribution probability of expert labels on the emotions of the k-th animal class; pseudo This represents the distribution probability of the pseudo-label on the emotion of the k-th animal class. When determining the total divergence loss, the JS divergence of each animal class is obtained based on formula (21), and then the JS divergence of each class is summed.
[0432] Furthermore, the distributional differences between pseudo-labels and expert labels can be combined with the aforementioned KL divergence loss and JS divergence loss, as shown in the following formula (22):
[0433] L total =λ1L KL (p expert ||p pseudo )+λ2L JS (p expert,p pseudo ) (twenty two)
[0434] In formula (22), the KL divergence loss and JS divergence loss for each animal emotion category can be calculated separately for pseudo-labels and expert labels, and then the two are weighted and summed. Among them, λ1 and λ2 are weights, which can be determined based on empirical values.
[0435] In some embodiments, to prevent the probability distribution of the k-th animal emotion category from having a zero value and causing KL divergence, the pseudo-labels and expert labels in formulas (21) and (22) can be updated for the probability distribution of the k-th animal emotion category as shown in formula (23):
[0436] In formula (23), K represents the total number of animal emotion categories, and δ is a very small positive number.
[0437] In this embodiment of the disclosure, for sample data with expert labels, by aligning the probability distributions of pseudo-labels and expert labels across multiple animal emotion categories, the output of the animal emotion recognition model for sample data with expert labels is constrained to be closer to the distribution of expert labels, which can prevent pseudo-label drift and improve the classification accuracy of the animal emotion recognition model.
[0438] In some embodiments, the total loss includes the total distance loss of multiple sample data. Determining the total distance loss can be implemented by: determining the distance between pseudo-labels and corresponding expert labels in multiple sample data; determining the statistical value of the distance between multiple sample data to obtain the total distance loss.
[0439] In practice, the distance between a label and its corresponding expert label can be represented by the L2 distance. Accordingly, the distance between them is shown in the following formula (24):
[0440] L va =||y expert_va -y pseudo_va || 2 (twenty four)
[0441] Among them, y expert_va Indicates expert label; y pseudo_va Indicates a pseudo tag; L va It indicates the distance between the two.
[0442] Distance loss can make the predicted pseudo-labels closer to the expert labels, thereby improving the classification accuracy of animal emotion recognition models.
[0443] In summary, the training of the animal emotion recognition model in this embodiment may include three stages, including:
[0444] In stage 1, the animal emotion recognition model is trained using sample data with expert labels. In this stage, the adversarial loss of adversarial training (as shown in formula (12)) can be used to optimize the model parameters.
[0445] Phase 2: Semi-supervised training using pseudo-labels to improve sample utilization.
[0446] This stage can employ adversarial loss, along with the aforementioned total loss for pseudo-labels, to optimize model parameters.
[0447] Phase 3 involves fine-tuning the training based on actual application scenarios to adapt to the emotion recognition of specific animals.
[0448] In some embodiments, after the animal emotion recognition model has been trained, high-quality samples can be selected based on the prediction results during the inference phase to optimize the model. For example, during the inference phase of the animal emotion recognition model, if a preset event is detected multiple times consecutively, the multimodal data corresponding to the preset event is acquired as new sample data: the preset event includes at least one of the following:
[0449] 1) The emotion recognition result indicates that the first type of target emotion has been identified, and the confidence level of the emotion recognition result is greater than the preset threshold;
[0450] 2) The emotion recognition result indicates that the second type of target emotion was identified and was continuously detected n times, where n is a positive integer greater than 1;
[0451] 3) The emotional fluctuation of the animal within the sliding time window where the emotion recognition result is located is higher than the preset value.
[0452] The various preset events have been explained above and will not be repeated here.
[0453] Understandably, by pre-setting events, it is possible to capture data samples of key animal emotions. These samples are used to optimize model parameters and improve the recognition accuracy of animal emotion recognition models.
[0454] For example, a pet begins displaying the "aggressive" emotion tag at 12:02:01 and continues for 3 times. The event recording window is from 12:02:01 to 12:02:10. The end time can be determined automatically by whether no "aggression" is detected for N consecutive seconds (e.g., 10 seconds), or by whether there are no new anomalies after a set "cooldown period" has passed since the most recent detection of "aggression".
[0455] At this point, due to the triggering of a preset event, the sampling rate is increased, and the collected multimodal data is labeled as "attack event samples." This multimodal data is used to construct training samples to enhance the learning intensity of the animal emotion recognition model for the "attack" category. The enhancement measures include: 1) prioritizing the addition of sample data corresponding to the preset event to the incremental training set to increase the probability of sampling learning; 2) assigning higher loss weights to high-risk categories during training to induce overfitting of the model to rare / high-risk labels; 3) initiating a fine-tuning process when necessary to shorten the model's response cycle to new data. During the training phase, corresponding pseudo-labels are generated. If the confidence level of the pseudo-label is greater than a dynamic threshold, it is added to the enhanced sample set for enhanced training; otherwise, it is added to the observation pool.
[0456] In some optional embodiments, the emotion recognition results can be semantically mapped and translated to convert animal language into human language, resulting in a language conversion result, including:
[0457] Emotional labels and voice features are extracted from the emotion recognition results, and the voice features are converted into standardized voice vectors.
[0458] A pre-trained language model is used to semantically map emotion tags to sound vectors in order to obtain the emotion intent;
[0459] A language generator is used to translate emotional intentions into language to generate corresponding human language, thus obtaining the language conversion result.
[0460] Specifically, a "pre-trained language model" refers to an artificial intelligence model that has been trained on a large amount of data and is capable of understanding and processing natural language. This model is used here to perform a "semantic mapping" between animal emotional labels—that is, the animal's emotional state obtained from the emotion recognition module—and sound vectors. This involves correlating the animal's vocal features with the emotional semantics in human language, thereby identifying the animal's emotional intentions. Next, the "language generator" is the process of converting the animal's nonverbal communication into language that humans can understand, based on these emotional intentions. This process involves converting the animal's emotions and intentions into specific text or speech output, i.e., the "language conversion result," enabling humans to intuitively understand the animal's "language" and achieve effective communication between humans and animals.
[0461] To facilitate understanding of the embodiments of this disclosure, an example is provided below, referring to Figure 18, which is a schematic flowchart of converting animal language into human language in an embodiment of this disclosure. After the emotion recognition module obtains the emotion recognition result, it provides the animal's emotion tag and vocal features, and then passes this information to the speech mapping module, which is responsible for converting the animal's vocal features into human-understandable speech expressions. Next, the human semantic output module receives the converted speech information and outputs it as a language conversion result, finally presenting these results to the user, realizing real-time conversion of animal language to human language and emotional communication. The entire process involves emotion recognition, feature extraction and mapping, and human language generation, aiming to promote effective communication between humans and animals.
[0462] In this way, by using a pre-trained language model to semantically map emotion tags to sound vectors, it is possible to accurately capture and understand the emotional intentions of animals. Then, a language generator is used to convert these emotional intentions into human language. This process not only achieves accurate interpretation of animal emotions but also transforms animal nonverbal communication into human-understandable language, greatly promoting communication and understanding between humans and animals, and improving the transparency and efficiency of animal emotional expression.
[0463] Based on the same technical concept, this disclosure also provides an animal emotion recognition method, as shown in Figure 19, including:
[0464] Step S1901: Obtain multimodal data related to animals.
[0465] Step S1902: Preprocess the multimodal data to obtain fused multimodal data.
[0466] Step S1903: Identify the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result.
[0467] Steps S1901-S1903 correspond to the aforementioned steps S101-S103. The implementation methods and effects of the same steps and their related content are as described above and will not be repeated here. In addition, the animal emotion recognition method shown in Figure 19 is also applicable to the embodiments of the aforementioned animal language conversion method other than those in Figure 1, and will not be repeated here either.
[0468] Based on the same technical concept, this disclosure also provides a method for training an animal emotion recognition model in an animal language conversion model. The animal language conversion model includes the aforementioned animal emotion recognition model, and a mapping network model for implementing step S104. The mapping network model includes the aforementioned pre-trained language model and / or language generator.
[0469] Figure 20 shows a flowchart illustrating the training method for the animal emotion recognition model in the animal language conversion model, including:
[0470] S2001, Acquire multimodal data related to animals;
[0471] S2002, preprocess the multimodal data to obtain fused multimodal data;
[0472] S2003, The fused multimodal data is input into the animal emotion recognition model to identify the animal's current emotion and obtain the animal's emotion recognition result;
[0473] S2004, Based on the emotion recognition results, train the animal emotion recognition model; the emotion recognition results output by the animal emotion recognition model are used for semantic mapping and language translation to convert animal language into human language and obtain language conversion results.
[0474] Steps S2001-S2004 correspond to steps S701-S704 in Figure 7. The implementation methods and effects of the same steps and their related content are as described above and will not be repeated here. In addition, the other embodiments of the training method for the animal emotion recognition model related to Figure 7 are also applicable to the method shown in Figure 20, and the corresponding content will not be repeated here either.
[0475] In addition to the content explained above, the present disclosure also provides the following alternative implementation methods. The following embodiments can also be applied to the embodiments related to Figures 1-20 above.
[0476] In some optional embodiments, it also includes:
[0477] If specific audio data is detected and there is no history of emotion matching, the specific audio data is labeled to obtain an updated emotion label;
[0478] The sample data is dynamically updated based on the updated sentiment tags, so that the model parameters can be adjusted according to the updated sample data.
[0479] Specifically, when a specific animal sound is detected that does not exist in the historical sentiment matching records (i.e., there is no corresponding previous sentiment label), a labeling process is triggered. Here, "labeling" refers to manually assigning a sentiment label to these specific sound data. This label describes the animal's emotional state when making the sound, thus "obtaining an updated sentiment label." Subsequently, this newly labeled sentiment label is incorporated into the sample database, achieving "dynamic updates to the sample data based on the updated sentiment label." This updated sample data is then used to adjust the model's parameters, i.e., "adjusting the model parameters based on the updated sample data," thereby optimizing and improving the model's ability to recognize the sentiment of newly appearing sound data. This ensures the system can adapt to new or uncommon animal sounds, improving recognition accuracy and the system's adaptive learning capabilities.
[0480] To facilitate understanding of the solutions in this disclosure, an example is given below: When the system encounters an unrecognizable vocal pattern or abnormal behavior, it prompts the user to input relevant tags or identification information. For instance, if the system detects a specific action + sound combination without a clear historical record of matching emotional expression, the user can annotate this phonetic symbol through the interface, such as "call for help" or "hunger." After the user manually annotates, the system adjusts the data parameters of the current audio and the corresponding tags for the behavior, thereby updating the weights of the existing model. After the user completes the annotation, the system updates the dataset in a timely manner, continuously optimizes and retrains the relevant generative model, so that the system can gradually improve the recognition rate of similar or new categories of emotional events.
[0481] In this way, by manually labeling specific sound data when there is a lack of historical emotion matching, and updating the emotion labels accordingly, the sample database can be continuously expanded and enriched, thereby dynamically updating the model parameters, enhancing the system's ability to recognize emotions in new sound data, improving the model's adaptability and accuracy, and ensuring that the cross-species communication system can continue to evolve and better understand and respond to the animal's communication intentions.
[0482] In some optional embodiments, the method further includes:
[0483] Collect multimodal data within a preset time window to obtain data on the animal's emotional changes;
[0484] Feature extraction is performed on the emotion change data to obtain emotion change features;
[0485] The sentiment labels are updated based on the difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window.
[0486] Specifically, a "preset time window" refers to a specific time period set for analyzing changes in animal emotions. During this period, "multimodal data" is collected, including information on the animal's vocalizations, behaviors, and physical characteristics, to gather data on its emotional state. Next, a "feature extraction" process identifies key information representing changes in the animal's emotions—the "emotional change features"—from this data. Then, the differences between the emotional features extracted in the current time window and those from the previous time window are compared. If these differences indicate a significant change in the animal's emotional state, the "emotional label" is updated to reflect the animal's latest emotional state. This process involves continuous monitoring and dynamic analysis of the animal's emotional state to ensure the real-time nature and accuracy of emotion recognition.
[0487] In this way, by collecting multimodal data from animals within a preset time window, it is possible to comprehensively capture the emotional changes of animals and extract features from this data to identify the characteristics of emotional changes. Secondly, by dynamically updating the emotional labels based on the differences between the current and previous time window's emotional features, real-time and accurate monitoring and response to the animal's emotional state can be achieved. This not only improves the accuracy of emotion recognition but also enhances the real-time nature and depth of communication between humans and animals, enabling humans to better understand and respond to the emotional needs of animals.
[0488] In some optional embodiments, multimodal data is collected within a preset time window to obtain data on changes in the animal's emotions, including:
[0489] Collect multimodal data within a preset time window;
[0490] Multimodal data is input into the emotional period recognition model, and the emotional change data of animals are obtained through the emotional period recognition model.
[0491] Specifically, a "preset time window" refers to a specific time period set for monitoring and analyzing changes in animal emotions. During this period, "multimodal data" is collected, including different types of information such as the animal's vocalizations, behaviors, and physical characteristics. This data is then fed into an "emotional period recognition model," an algorithm specifically designed to process and analyze time-series data, such as Long Short-Term Memory (LSTM) networks or Gated Recurrent Units (GRUs). By analyzing this multimodal data, the model identifies and extracts features related to the animal's emotional state, thereby "obtaining animal emotional change data." This data reflects the animal's emotional state within a continuous time window, providing a foundation for further emotion recognition and label updates. In this way, the solution enables dynamic tracking and accurate identification of animal emotional states.
[0492] By collecting multimodal data from animals within a preset time window and inputting this data into an emotion period recognition model, it is possible to accurately capture and analyze changes in animal emotions. This method can identify and understand the emotional dynamics of animals over continuous time periods, thus providing richer and more accurate data on emotional changes. This data not only enhances the depth and real-time performance of emotion recognition but also improves the quality and efficiency of communication between humans and animals, enabling humans to respond more subtly and promptly to animals' emotional needs and behavioral changes.
[0493] In some optional embodiments, the sentiment label is updated based on the difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window, including:
[0494] The Euclidean distance difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window is calculated to obtain the sentiment gap.
[0495] If the emotional gap exceeds the preset emotional gap threshold, the emotional label will be upgraded to obtain an updated emotional label.
[0496] Specifically, "emotional change features in the current time window" refers to emotion-related features extracted from multimodal data of animals within a specific time period, while "emotional change features in the previous time window" refers to the corresponding features in the previous time period. By calculating the "Euclidean distance difference" between the emotional features of these two time windows, which is a method for measuring the distance between two points in a multidimensional space, a quantitative "emotional gap" can be obtained.
[0497] The Euclidean distance difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window satisfies the formula: [D(W_i,W_{i-1})=\sqrt{\sum_{n=1}^{N}(x_{i,n}-x_{i-1,n})^2}], that is... Where D(W_i,W_{i-1}) and D(W i W i-1 ): Represents the Euclidean distance between the sentiment change features of the i-th time window and the (i-1)-th time window. This distance is used to quantify the degree of sentiment feature change within two consecutive time windows. x_{i,n} and x i,n Let x_{i-1,n} represent the value of the nth feature within the i-th time window. These features may include the pitch, rhythm, and amplitude of sound, the frequency of behavior, and changes in heart rate, etc. i-1,n This represents the value of the nth feature within the (i-1)th time window. `sum` represents the summation operation, and `sqrt` represents the square operation.
[0498] The emotional gap is calculated by measuring the Euclidean distance between the emotional change features in the current time window and those in the previous time window. This gap reflects the degree of change in the animal's emotional state within two consecutive time windows. If this gap exceeds a "preset emotional gap threshold," indicating a significant change in the animal's emotional state, an "emotional label upgrade" is triggered. Based on the emotional gap and change pattern, the emotional label is updated from one state (e.g., "alert") to another label that better reflects the current emotional state (e.g., "anxiety"), resulting in an "updated emotional label." This process makes emotion recognition more dynamic and accurate, enabling real-time responses to significant changes in animal emotions.
[0499] By calculating the Euclidean distance difference between the current and previous time window's emotional change features, the degree of change in an animal's emotional state can be quantified, resulting in an emotional gap. When the emotional gap exceeds a preset threshold, the emotional label is automatically upgraded to an updated label. This method makes emotion recognition more sensitive and accurate, enabling real-time capture and response to significant changes in animal emotions. This improves the quality and efficiency of communication between humans and animals, ensuring that humans can understand and respond to animals' emotional needs promptly and appropriately.
[0500] In some optional embodiments, the method further includes:
[0501] The sentiment weight corresponding to each time window is scored to obtain the sentiment weight score;
[0502] The emotional weight scores corresponding to similar emotional time windows within a continuous time period are accumulated. If the accumulated result is greater than the set upgrade threshold, the emotional label is updated.
[0503] Specifically, "emotional weighting scoring" refers to the process of quantitatively assessing the emotional state of an animal within each time window, where each time window contains multimodal data collected over a specific period. By analyzing this data, emotional features are extracted and assigned different weights based on the strength or salience of the features, thus calculating the emotional weighting score for each time window. Next, the weighting scores of time windows with similar emotional features across consecutive time periods are accumulated. If this accumulated score exceeds a "set upgrade threshold," a predefined value used to determine whether the emotional state is significant enough to trigger a label update, the emotional label is updated to reflect the change in the animal's emotional state. This process makes emotion recognition more dynamic and accurate, enabling real-time responses to continuous changes in animal emotions.
[0504] By assigning weights and scoring emotional states within each time window, this scheme can quantify and track changes in an animal's emotions over continuous time periods. The weighted scores of time windows with similar emotional characteristics are accumulated; when the accumulated score exceeds a preset escalation threshold, the emotional label is automatically updated to more accurately reflect the animal's current emotional state. This method improves the sensitivity and adaptability of emotion recognition, making communication between humans and animals more precise and timely, thereby enhancing the understanding and response to animals' emotional needs.
[0505] To facilitate understanding of the embodiments of this disclosure, please refer to Figure 21, which is a schematic flowchart of the emotion tag updating process in this embodiment. Animal barking audio and body temperature data are collected through sound acquisition and a body temperature sensor. This data is then sent to a data processing and fusion module for analysis to identify audio features and body temperature changes. The analysis results are used for emotion feedback; based on the emotional meaning of the audio and body temperature changes, the animal's emotion tag is dynamically adjusted, such as upgrading the "alert" state to "anxious." This process involves real-time monitoring and continuous emotion state assessment to ensure that changes in the animal's emotion are reflected in the tag update in a timely manner, thereby improving the accuracy of understanding and responding to the animal's emotional state.
[0506] By collecting data on animal emotional changes within continuous time windows and extracting emotional change features from this data, subtle changes in animal emotions can be monitored and quantified in real time. By comparing these change features with standard emotional features and updating the emotional label when the change exceeds a preset threshold, the dynamic changes in animal emotions can be captured more accurately. This enables continuous tracking and real-time response to animal emotional states, improving the sensitivity and accuracy of emotion recognition and enhancing the effectiveness of emotional communication between humans and animals.
[0507] The following describes an embodiment of the apparatus described in this application, which can be used to execute the animal emotion recognition method described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the animal emotion recognition method described above in this application.
[0508] Based on the same technical concept, this disclosure also provides an animal emotion recognition device 2200, as shown in Figure 22, including:
[0509] The first acquisition module 2201 is used to acquire multimodal data related to animals;
[0510] The first preprocessing module 2202 is used to preprocess the multimodal data to obtain fused multimodal data;
[0511] The first recognition module 2203 is used to identify the current emotion of the animal based on the fused multimodal data, so as to obtain the emotion recognition result of the animal.
[0512] In some embodiments, multimodal data includes animal sound data, animal behavior data, and animal physical characteristics data.
[0513] In some optional embodiments, the first preprocessing module 2202 preprocesses the multimodal data to obtain fused multimodal data, including:
[0514] Data cleaning is performed on the multimodal data to obtain cleaned multimodal data;
[0515] The cleaned multimodal data is normalized to obtain normalized multimodal data;
[0516] The normalized multimodal data is time-series aligned and fused to obtain fused multimodal data.
[0517] In some optional embodiments, the first recognition module 2203 identifies the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result, including:
[0518] A deep learning model is used to extract sound features, visual motion features, and analyze vital sign changes from the fused multimodal data to obtain multimodal feature vectors.
[0519] Generative adversarial networks are used to perform sentiment analysis on multimodal feature vectors to obtain the emotion recognition results of animals.
[0520] In some optional embodiments, the first identification module 2203 includes:
[0521] The first fusion unit is used to perform cross-modal feature fusion operations on multimodal feature vectors to obtain fused features;
[0522] The first temporal modeling unit is used to perform temporal modeling on the fused features to obtain temporal features;
[0523] The first classification unit is used to classify temporal features to obtain the animal's emotion recognition results.
[0524] In some optional embodiments, the first identification module 2203 includes:
[0525] The first extraction unit is used to extract sound features, visual motion features, and analyze physical changes from the fused multimodal data using a deep learning model, so as to obtain the first audio features, the first visual features, and the first physical feature.
[0526] The first synchronization unit is used to synchronize the data of the first audio feature, the first visual feature and the first vital sign feature to obtain the second audio feature, the second visual feature and the second vital sign feature.
[0527] The first normalization unit is used to normalize the second audio feature, the second visual feature, and the second physical feature to obtain a multimodal feature vector.
[0528] In some alternative embodiments, deep learning models and generative adversarial networks are used to build animal emotion recognition models.
[0529] In some optional embodiments, it also includes:
[0530] The frequency adjustment module is used to increase the frequency of acquiring multimodal data when a preset event is detected based on the emotion recognition results.
[0531] Preset events include at least one of the following:
[0532] The emotion recognition result indicates that the first type of target emotion has been identified, and the confidence level of the emotion recognition result is greater than the preset threshold.
[0533] The emotion recognition result indicates that the second type of target emotion was identified and was continuously detected n times, where n is a positive integer greater than 1;
[0534] The emotional fluctuation value of the animal within the sliding time window where the emotion recognition result is located is higher than the preset value.
[0535] In some optional embodiments, the frequency adjustment module is further configured to:
[0536] For animals, the frequency of acquiring multimodal data is reduced if the preset event is not detected for a specified duration after the event ends.
[0537] Based on the same technical concept, as shown in Figure 23, this disclosure also provides a training device 2300 for an animal emotion recognition model, comprising:
[0538] The second acquisition module 2301 is used to acquire multimodal data related to animals;
[0539] The second preprocessing module 2302 is used to preprocess the multimodal data to obtain fused multimodal data;
[0540] The second recognition module 2303 is used to input the fused multimodal data into the animal emotion recognition model to recognize the animal's current emotion and obtain the animal's emotion recognition result.
[0541] Training module 2304 is used to train an animal emotion recognition model based on emotion recognition results.
[0542] In some alternative embodiments, the animal emotion recognition model includes a deep learning model and a generative adversarial network;
[0543] The second identification module 2303 is used for:
[0544] A deep learning model is used to extract sound features, visual motion features, and analyze vital sign changes from the fused multimodal data to obtain multimodal feature vectors.
[0545] Generative adversarial networks are used to perform sentiment analysis on multimodal feature vectors to obtain the emotion recognition results of animals.
[0546] In some optional embodiments, the second identification module 2303 includes:
[0547] The second extraction unit is used to extract sound features, visual motion features, and analyze physical signs changes from the fused multimodal data based on a deep learning model, so as to obtain the first audio features, the first visual features, and the first physical sign features.
[0548] The second synchronization unit is used to synchronize the data of the first audio feature, the first visual feature and the first vital sign feature to obtain the second audio feature, the second visual feature and the second vital sign feature.
[0549] The second normalization unit is used to normalize the second audio feature, the second visual feature, and the second physical feature to obtain a multimodal feature vector.
[0550] In some optional embodiments, the second extraction unit is specifically used for:
[0551] The audio feature extraction module in the deep learning model extracts features from the audio signals in the fused multimodal data to obtain the first audio feature;
[0552] The video feature extraction module in the deep learning model extracts features from the video data in the fused multimodal data to obtain the first visual features;
[0553] The physiological data in the fused multimodal data is extracted using the feature extraction module in the deep learning model to obtain the first vital sign feature.
[0554] In some optional embodiments, the second synchronization unit is specifically used to: process the number of frames of the first audio feature, the first visual feature, and the first vital sign feature within the current time window into the target number of frames, thereby obtaining the second audio feature, the second visual feature, and the second vital sign feature.
[0555] In some optional embodiments, the second identification module 2303 includes:
[0556] The second fusion unit is used to perform cross-modal feature fusion operations on the multimodal feature vectors to obtain fused features;
[0557] The second temporal modeling unit is used to perform temporal modeling on the fused features to obtain temporal features;
[0558] The second classification unit is used to classify temporal features to obtain the animal's emotion recognition results.
[0559] In some alternative embodiments, the second fusion unit is configured to:
[0560] The feature vector of the audio modality in the multimodal feature vector is used as the first query vector. One of the feature vectors of the video modality and the feature vector of the physiological modality in the multimodal feature vector is used as the first key vector, and the other is used as the first value vector. The first query vector, the first key vector, and the first value vector are processed based on the attention mechanism to obtain the first feature to be fused.
[0561] The video modality feature vector in the multimodal feature vector is used as the second query vector, and one of the audio modality feature vector and the physiological modality feature vector in the multimodal feature vector is used as the second key vector, and the other is used as the second value vector. The second query vector, the second key vector and the second value vector are processed based on the attention mechanism to obtain the second feature to be fused.
[0562] The feature vector of the physiological modality in the multimodal feature vector is used as the third query vector. One of the feature vectors of the audio modality and the feature vector of the video modality in the multimodal feature vector is used as the third key vector, and the other is used as the third value vector. The third query vector, the third key vector and the third value vector are processed based on the attention mechanism to obtain the third feature to be fused.
[0563] The first, second, and third features to be fused are weighted and summed to obtain the fused features.
[0564] In some optional embodiments, the weight of any feature to be fused, for any of the first, second, and third features to be fused, is dynamically generated based on the feature to be fused and learnable parameters from the training phase.
[0565] In some optional embodiments, the second timing modeling unit is specifically used for:
[0566] Feature encoding is performed on multiple fused features at multiple time steps within the current time window to obtain multiple enhanced features at multiple time steps within the current time window;
[0567] By modeling multiple enhanced features from multiple time steps on a time series, multiple intermediate features of the current time window are obtained.
[0568] The time series features are obtained by weighted summation of multiple intermediate features within the current time window.
[0569] In some optional embodiments, the second timing modeling unit is specifically used for:
[0570] Feature transformation is performed on multiple fused features at multiple time steps to obtain multiple enhanced features, wherein the feature dimension of the enhanced features is smaller than that of the fused features.
[0571] In some optional embodiments, the second timing modeling unit is specifically used for:
[0572] The weighted generation network processes multiple intermediate features of the current time window to obtain multiple weights corresponding to multiple time steps within the current time window.
[0573] Based on multiple weights, the intermediate features of multiple time steps within the current time window are weighted and summed to obtain the time series features.
[0574] In some optional embodiments, an adversarial training module is also included for training the animal emotion recognition model through adversarial training.
[0575] In some alternative embodiments, the adversarial training module includes:
[0576] The first perturbation unit is used to add perturbation to the fusion feature corresponding to the multimodal feature vector at the current time step, and generate the fusion feature of multimodal noisy features;
[0577] The first prediction unit is used to input the fused features of multimodal noisy features into the animal emotion recognition model to identify the animal's current emotion and obtain the animal's emotion prediction result.
[0578] In some optional embodiments, the first perturbation unit is configured to: obtain a first statistical value in the first classification probability distribution of the processing result of the previous time step at the current time step; the processing result includes the sentiment classification result obtained by performing sentiment classification on the fused features of the previous time step;
[0579] The first perturbation intensity is generated based on the negative correlation between the intensity and the first statistical value;
[0580] Based on the first perturbation intensity, the fused features corresponding to the multimodal feature vectors, and the classification loss gradient of the fused features corresponding to the multimodal feature vectors, a fused feature of multimodal noisy features is generated.
[0581] In some alternative embodiments, the adversarial training module includes:
[0582] The second perturbation unit is used to add perturbation to the feature vector of the target mode in the multimodal feature vector to obtain the perturbation feature of the target mode;
[0583] The replacement unit is used to replace the feature vector of the target mode in the multimodal feature vector with the perturbation feature of the target mode, so as to obtain the multimodal noise-added feature;
[0584] The second prediction unit is used to input multimodal noisy features into the animal emotion recognition model to identify the animal's current emotion and obtain the animal's emotion prediction result.
[0585] The second disturbance unit, used for:
[0586] Obtain the second statistical value from the second classification probability distribution of the target modality in the previous time step of the current time step; the second classification probability distribution is the classification probability distribution obtained by performing sentiment classification on the feature vector of the target modality;
[0587] A second perturbation intensity is generated based on the negative correlation between the intensity and the second statistical value.
[0588] Based on the second perturbation intensity, a perturbation is added to the feature vector of the target mode in the multimodal feature vector to obtain the perturbation feature of the target mode.
[0589] In some optional embodiments, the training module includes:
[0590] The first loss determination unit is used to determine the first classification loss based on the emotion recognition results and training labels;
[0591] The second loss determination unit is used to determine the second classification loss based on the sentiment prediction results and training labels;
[0592] The training unit is used to optimize the model parameters of the animal emotion recognition model based on the first classification loss and the second classification loss.
[0593] In some optional embodiments, the training module is also used for:
[0594] The temporal features of the multimodal feature vectors obtained during the processing of multimodal feature vectors by the animal emotion recognition model are acquired. These temporal features are used for emotion classification to obtain the emotion recognition results.
[0595] The temporal features of the multimodal noisy features obtained during the processing of multimodal noisy features by the animal emotion recognition model are acquired. These temporal features of the multimodal noisy features are used for emotion classification to obtain emotion prediction results.
[0596] The model parameters of the animal emotion recognition model are optimized based on the first consistency loss between the temporal features of multimodal feature vectors and the temporal features of multimodal noisy features.
[0597] In some optional embodiments, the training module is also used for:
[0598] Determine the second consistency loss between emotion recognition results and emotion prediction results;
[0599] Based on the second consistency loss, the model parameters of the animal emotion recognition model are optimized.
[0600] In some optional embodiments, the second preprocessing module is used to: aggregate multimodal data within the same time stamp range based on the timestamps in the multimodal data to obtain fused multimodal data.
[0601] In some optional embodiments, it also includes:
[0602] The pseudo-tag generation module is used to generate pseudo-tags for animal emotions.
[0603] The training module is also used to train animal emotion recognition models based on pseudo-labels.
[0604] In some optional embodiments, the pseudo-tag generation module includes:
[0605] The first output unit is used to input sample data with expert labels into the animal emotion recognition model when the animal emotion recognition model meets the preset conditions, and obtain the predicted labels of the sample data output by the animal emotion recognition model.
[0606] The first generation unit is used to perform a weighted summation of the predicted labels and expert labels of the sample data to obtain the pseudo labels of the sample data.
[0607] In some optional embodiments, the first generating unit is used for:
[0608] Obtain the confidence score of the animal emotion recognition model for the emotion classification distribution of multiple training samples from the previous round;
[0609] Determine the statistical values of the confidence scores of the sentiment classification distributions of multiple training samples, and use them as dynamic thresholds;
[0610] The fusion weights of expert labels and predicted labels are generated based on dynamic thresholds; the closer the fusion weights are to the dynamic thresholds, the greater the fusion weights of the expert labels.
[0611] Based on the fusion weights of expert labels and predicted labels, a weighted sum of predicted and expert labels is performed to obtain pseudo-labels for the sample data.
[0612] In some optional embodiments, the preset conditions include at least one of the following:
[0613] The animal emotion recognition model is trained through training rounds using training samples with expert labels to reach the target number of rounds.
[0614] The animal emotion recognition model's emotion recognition accuracy based on training samples with expert labels meets the set conditions for entering the next stage of training.
[0615] In some optional embodiments, the pseudo-tag generation module includes:
[0616] The second output unit is used to determine the classification divergence between multiple modalities in animal-related multimodal data; the classification divergence is used to measure the difference between the categories of animal emotions expressed by different modalities.
[0617] The second generation unit is used to merge multiple sentiment classification results from multiple modalities in multimodal data to obtain pseudo-labels when the classification divergence degree is less than the divergence threshold.
[0618] In some optional embodiments, the second generating unit is used for:
[0619] Based on the classification confidence of multiple sentiment classification results across multiple modalities, multiple weighted values for multiple modalities are generated.
[0620] Based on multiple weighted values and multiple sentiment classification results, pseudo-labels for sample data representing multimodal data are generated.
[0621] In some optional embodiments, it also includes:
[0622] The storage module is used to store multimodal data in the observation pool when the classification divergence degree is not less than the divergence threshold.
[0623] In some optional embodiments, it also includes:
[0624] The filtering module is used to determine the sample data corresponding to the pseudo-label as the sample data for training the animal emotion recognition model when the confidence of the pseudo-label is greater than the dynamic threshold.
[0625] The dynamic threshold is determined based on statistical values, which are the confidence scores of the emotion recognition model for the emotion classification distribution of multiple training samples in the previous round.
[0626] In some alternative embodiments, the training module includes:
[0627] The loss determination unit is used to determine at least one of the following training losses based on pseudo-labels: cross-entropy loss of pseudo-labels, loss between pseudo-labels and expert labels, and alignment loss between pseudo-label distribution and expert label distribution;
[0628] An optimization unit is used to determine the total loss based on at least one training loss, so as to optimize the animal emotion recognition model based on the total loss.
[0629] In some optional embodiments, the total loss includes the total cross-entropy loss of multiple sample data, and the loss determination unit determines the total cross-entropy loss by including:
[0630] Generate multiple loss weights corresponding to multiple sample data; where each sample data has a corresponding multiple pseudo-label;
[0631] Based on multiple loss weights corresponding to multiple sample data, the weighted sum of multiple cross-entropy losses of multiple pseudo-labels is obtained to obtain the total cross-entropy loss.
[0632] In some optional embodiments, the loss determination unit is used to generate multiple loss weights corresponding to multiple sample data based on the degree of difference between multiple pseudo-labels and their corresponding expert labels;
[0633] Among them, the degree of difference and the loss weight are inversely correlated.
[0634] In some optional embodiments, the total loss includes the total distance loss of multiple sample data, and the loss determination unit determines the total distance loss, including:
[0635] Determine the distance between pseudo-labels and their corresponding expert labels in multiple sample data sets;
[0636] Determine the statistical values of the distances between multiple sample data to obtain the total distance loss.
[0637] In some optional embodiments, the total loss includes the total alignment loss of multiple sample data, which is used to measure the distribution difference between pseudo-labels and expert labels of multiple sample data; during the training phase, the total alignment loss is minimized by optimizing the model parameters of the animal emotion recognition model.
[0638] In some optional embodiments, it also includes:
[0639] The sample addition module is used in the inference phase of the animal emotion recognition model. When a preset event is detected multiple times consecutively, the multimodal data corresponding to the preset event is acquired as new sample data.
[0640] Preset events include at least one of the following:
[0641] The emotion recognition result indicates that the first type of target emotion was identified, and the confidence level of the identification is greater than the preset threshold.
[0642] The emotion recognition result indicates that the second type of target emotion was identified and was continuously detected n times, where n is a positive integer greater than 1;
[0643] The emotional fluctuations of the animals within the sliding time window where the emotion recognition results are located are higher than the preset value.
[0644] Based on the same technical concept, as shown in Figure 24, an animal language conversion device 2400 is also provided, comprising:
[0645] The first acquisition module 2401 is used to acquire multimodal data related to animals, including animal sound data, animal behavior data, and animal physical characteristics data.
[0646] The first preprocessing module 2402 is used to preprocess the multimodal data to obtain fused multimodal data;
[0647] The first recognition module 2403 is used to recognize the current emotion of the animal based on the fused multimodal data, so as to obtain the emotion recognition result of the animal;
[0648] The conversion module 2404 is used to perform semantic mapping and language translation on the emotion recognition results to convert animal language into human language and obtain language conversion results.
[0649] In some optional embodiments, the first acquisition module 2401 acquires animal-related multimodal data, including:
[0650] Collect sound wave information emitted by animals to obtain animal sound data;
[0651] Collect animal body language and movement changes to obtain animal behavior data;
[0652] Collect physical and biological indicators of animals to obtain animal vital signs data.
[0653] In some optional embodiments, the first preprocessing module 2402 preprocesses the multimodal data to obtain fused multimodal data, including:
[0654] Denoising is performed on the multimodal data to clean the data, resulting in cleaned multimodal data.
[0655] The cleaned multimodal data is normalized to obtain normalized multimodal data;
[0656] The normalized multimodal data is time-series aligned and fused to obtain fused multimodal data.
[0657] In some optional embodiments, the first recognition module 2403 identifies the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result, including:
[0658] A deep learning model is used to extract sound features, visual motion features, and analyze vital sign changes from the fused multimodal data to obtain multimodal feature vectors.
[0659] Generative adversarial networks are used to perform sentiment analysis on multimodal features to obtain the emotion recognition results for animals. Furthermore, the devices shown in Figures 22-25 may also include the following modules and / or units.
[0660] In some optional embodiments, the conversion module 2404 performs semantic mapping and language translation on the emotion recognition results to convert animal language into human language, obtaining a language conversion result, including:
[0661] Emotional labels and voice features are extracted from the emotion recognition results, and the voice features are converted into standardized voice vectors.
[0662] A pre-trained language model is used to semantically map emotion tags to sound vectors in order to obtain the emotion intent;
[0663] A language generator is used to translate emotional intentions into language to generate corresponding human language, thus obtaining the language conversion result.
[0664] In some optional embodiments, the apparatus further includes a first update module, configured to annotate the specific sound data to obtain an updated emotion tag if specific sound data is detected and no emotion matching history exists;
[0665] The sample data is dynamically updated based on the updated sentiment tags, so that the model parameters can be adjusted according to the updated sample data.
[0666] In some optional embodiments, the device further includes a second update module for collecting multimodal data within a preset time window to obtain data on the animal's emotional changes.
[0667] Feature extraction is performed on the emotion change data to obtain emotion change features;
[0668] The sentiment labels are updated based on the difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window.
[0669] In some optional embodiments, the second update module collects multimodal data within a preset time window to obtain data on the animal's emotional changes, including:
[0670] Collect multimodal data within a preset time window;
[0671] Multimodal data is input into the emotional period recognition model, and the emotional change data of animals are obtained through the emotional period recognition model.
[0672] In some optional embodiments, the second update module updates the sentiment tags based on the difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window, including:
[0673] The Euclidean distance difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window is calculated to obtain the sentiment gap.
[0674] If the emotional gap exceeds the preset emotional gap threshold, the emotional label will be upgraded to obtain an updated emotional label.
[0675] In some optional embodiments, the second update module is used to:
[0676] The sentiment weight corresponding to each time window is scored to obtain the sentiment weight score;
[0677] The emotional weight scores corresponding to similar emotional time windows within a continuous time period are accumulated. If the accumulated result is greater than the set upgrade threshold, the emotional label is updated.
[0678] Based on the same technical concept, as shown in Figure 25, a training device 2500 for an animal emotion recognition model in an animal language conversion model is also provided, comprising:
[0679] The second acquisition module 2501 is used to acquire multimodal data related to animals;
[0680] The second preprocessing module 2502 is used to preprocess the multimodal data to obtain fused multimodal data;
[0681] The second recognition module 2503 is used to input the fused multimodal data into the animal emotion recognition model to recognize the animal's current emotion and obtain the animal's emotion recognition result.
[0682] Training module 2504 is used to train the animal emotion recognition model based on the emotion recognition results; the emotion recognition results output by the animal emotion recognition model are used for semantic mapping and language translation to convert animal language into human language and obtain language conversion results.
[0683] It should be noted that the other modules and / or units in the training device 2300 of the animal emotion recognition model described above are also applicable to the training device 2500 of the animal emotion recognition model in the animal language conversion model.
[0684] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0685] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0686] Figure 26 illustrates a schematic block diagram of an example electronic device 2600 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0687] As shown in Figure 8, the electronic device 2600 includes a computing unit 2601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 2602 or a computer program loaded from a storage unit 2608 into a random access memory (RAM) 2603. The RAM 2603 may also store various programs and data required for the operation of the device 2600. The computing unit 2601, ROM 2602, and RAM 2603 are interconnected via a bus 2604. An input / output (I / O) interface 2605 is also connected to the bus 2604.
[0688] Multiple components in device 2600 are connected to I / O interface 2605, including: input unit 2606, such as keyboard, mouse, etc.; output unit 2607, such as various types of monitors, speakers, etc.; storage unit 2608, such as disk, optical disk, etc.; and communication unit 2609, such as network card, modem, wireless transceiver, etc. Communication unit 2609 allows device 2600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0689] The computing unit 2601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 2601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 2601 performs the various methods described above, such as animal language conversion methods, animal emotion recognition methods, and any animal emotion recognition model training methods. For example, in some embodiments, the animal language conversion methods, animal emotion recognition methods, and any animal emotion recognition model training methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 2608. In some embodiments, part or all of the computer program can be loaded and / or installed on device 2600 via ROM 2602 and / or communication unit 2609. When the computer program is loaded into RAM 2603 and executed by the computing unit 2601, one or more steps of the applet distribution described above can be performed. Alternatively, in other embodiments, computing unit 2601 may be configured by any other suitable means (e.g., by means of firmware) to perform an animal language conversion method, an animal emotion recognition method, or any animal emotion recognition model training method.
[0690] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0691] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0692] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0693] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0694] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0695] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0696] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0697] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for animal emotion recognition, comprising: Acquire multimodal data related to animals; The multimodal data is preprocessed to obtain fused multimodal data; The animal's current emotion is identified based on the fused multimodal data to obtain the animal's emotion recognition result.
2. The method according to claim 1, wherein, The multimodal data includes animal sound data, animal behavior data, and animal physical characteristics data.
3. The method according to claim 1, wherein, The preprocessing of the multimodal data to obtain fused multimodal data includes: The multimodal data is cleaned to obtain cleaned multimodal data; The cleaned multimodal data is normalized to obtain normalized multimodal data; The normalized multimodal data is then subjected to time-series alignment and fusion processing to obtain fused multimodal data.
4. The method according to claim 1, wherein, The step of identifying the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result includes: A deep learning model is used to extract sound features, visual motion features, and analyze vital sign changes from the fused multimodal data to obtain multimodal feature vectors. Generative adversarial networks are used to perform sentiment analysis on the multimodal feature vectors to obtain the emotion recognition results of the animal.
5. The method according to claim 4, wherein, The step of using a generative adversarial network to perform sentiment analysis on the multimodal features to obtain the animal's sentiment recognition results includes: Perform cross-modal feature fusion operation on the multimodal feature vectors to obtain fused features; Temporal modeling is performed on the fused features to obtain temporal features; The temporal features are classified to obtain the emotion recognition results of the animal.
6. The method according to claim 4, wherein, The process employs a deep learning model to extract sound features, visual motion features, and analyze vital sign changes from the fused multimodal data, resulting in a multimodal feature vector, including: A deep learning model is used to extract sound features, visual motion features, and analyze physical signs changes in the fused multimodal data to obtain first audio features, first visual features, and first physical sign features. The first audio feature, the first visual feature, and the first vital sign feature are synchronized to obtain the second audio feature, the second visual feature, and the second vital sign feature; The second audio feature, the second visual feature, and the second vital sign feature are normalized to obtain the multimodal feature vector.
7. The method according to claim 4, wherein, The deep learning model and the generative adversarial network are used to construct an animal emotion recognition model.
8. The method according to any one of claims 1-7, further comprising: When a preset event is detected based on the emotion recognition results, the frequency of acquiring the multimodal data is increased. The preset event includes at least one of the following: The emotion recognition result indicates that a first type of target emotion has been identified, and the confidence level of the emotion recognition result is greater than a preset threshold. The emotion recognition result indicates that the second type of target emotion has been identified and has been detected continuously n times, where n is a positive integer greater than 1; The emotional fluctuation value of the animal within the sliding time window where the emotion recognition result is located is higher than the preset value.
9. The method according to claim 8, further comprising: For the animal, if the preset event is not detected for a specified duration after the event ends, the frequency of acquiring the multimodal data is reduced.
10. A training method for an animal emotion recognition model, comprising: Acquire multimodal data related to animals; The multimodal data is preprocessed to obtain fused multimodal data; The fused multimodal data is input into the animal emotion recognition model to identify the animal's current emotion and obtain the animal's emotion recognition result. Based on the emotion recognition results, the animal emotion recognition model is trained.
11. The method according to claim 10, wherein, The animal emotion recognition model includes a deep learning model and a generative adversarial network; The step of inputting the fused multimodal data into the animal emotion recognition model to identify the animal's current emotion and obtain the animal's emotion recognition result includes: A deep learning model is used to extract sound features, visual motion features, and analyze vital sign changes from the fused multimodal data to obtain multimodal feature vectors. Generative adversarial networks are used to perform sentiment analysis on the multimodal feature vectors to obtain the emotion recognition results of the animals.
12. The method according to claim 11, wherein, The process employs a deep learning model to extract sound features, visual motion features, and analyze vital sign changes from the fused multimodal data, resulting in a multimodal feature vector, including: Based on the deep learning model, sound feature extraction, visual motion feature extraction, and vital sign change analysis are performed on the fused multimodal data to obtain first audio features, first visual features, and first vital sign features; The first audio feature, the first visual feature, and the first vital sign feature are synchronized to obtain the second audio feature, the second visual feature, and the second vital sign feature; The second audio feature, the second visual feature, and the second physical feature are normalized to obtain a multimodal feature vector.
13. The method according to claim 12, wherein, The process of extracting sound features, visual motion features, and analyzing vital sign changes based on the deep learning model from the fused multimodal data yields first audio features, first visual features, and first vital sign features, including: Based on the audio feature extraction module in the deep learning model, the audio signal in the fused multimodal data is extracted to obtain the first audio feature; Based on the video feature extraction module in the deep learning model, feature extraction is performed on the video data in the fused multimodal data to obtain the first visual feature; Based on the vital sign feature extraction module in the deep learning model, the physiological data in the fused multimodal data is used to extract features to obtain the first vital sign feature.
14. The method according to claim 12, wherein, The step of synchronizing the first audio feature, the first visual feature, and the first vital sign feature to obtain the second audio feature, the second visual feature, and the second vital sign feature includes: The number of frames of the first audio feature, the first visual feature, and the first vital sign feature within the current time window are processed into the target number of frames to obtain the second audio feature, the second visual feature, and the second vital sign feature.
15. The method according to claim 11, wherein, The step of using a generative adversarial network to perform sentiment analysis on the multimodal feature vectors to obtain the animal's sentiment recognition results includes: Perform cross-modal feature fusion operation on the multimodal feature vectors to obtain fused features; Temporal modeling is performed on the fused features to obtain temporal features; The temporal features are classified to obtain the emotion recognition results of the animal.
16. The method according to claim 15, wherein, The step of performing cross-modal feature fusion on the multimodal feature vectors to obtain fused features includes: The feature vector of the audio modality in the multimodal feature vector is used as the first query vector. One of the feature vectors of the video modality and the feature vector of the physiological modality in the multimodal feature vector is used as the first key vector, and the other is used as the first value vector. The first query vector, the first key vector, and the first value vector are processed based on the attention mechanism to obtain the first feature to be fused. The video modality feature vector in the multimodal feature vector is used as the second query vector, and one of the audio modality feature vector and the physiological modality feature vector in the multimodal feature vector is used as the second key vector, and the other is used as the second value vector. The second query vector, the second key vector and the second value vector are processed based on the attention mechanism to obtain the second feature to be fused. The feature vector of the physiological modality in the multimodal feature vector is used as the third query vector. One of the feature vectors of the audio modality and the feature vector of the video modality in the multimodal feature vector is used as the third key vector, and the other is used as the third value vector. The third query vector, the third key vector and the third value vector are processed based on the attention mechanism to obtain the third feature to be fused. The first feature to be fused, the second feature to be fused, and the third feature to be fused are weighted and summed to obtain the fused feature.
17. The method according to claim 16, wherein, For any one of the first feature to be fused, the second feature to be fused, and the third feature to be fused, the weight of the any one feature to be fused is dynamically generated based on the any one feature to be fused and the learnable parameters during the training phase.
18. The method according to claim 15, wherein, The step of performing temporal modeling on the fused features to obtain temporal features includes: Feature encoding is performed on multiple fused features of multiple time steps within the current time window to obtain multiple enhanced features of multiple time steps within the current time window; The multiple enhanced features of the multiple time steps are modeled on the time series to obtain multiple intermediate features of the current time window; The time series feature is obtained by weighted summation of multiple intermediate features of the current time window.
19. The method according to claim 18, wherein, The step of feature encoding multiple fused features from multiple time steps within the current time window to obtain multiple enhanced features from multiple time steps within the current time window includes: Feature transformation is performed on multiple fused features at multiple time steps to obtain multiple enhanced features, wherein the feature dimension of the enhanced features is smaller than the feature dimension of the fused features.
20. The method according to claim 18, wherein, The step of weighted summing of multiple intermediate features of the current time window to obtain the time series features includes: The weighted generation network processes multiple intermediate features of the current time window to obtain multiple weights corresponding to multiple time steps within the current time window. Based on the multiple weights, the multiple intermediate features of the multiple time steps within the current time window are weighted and summed to obtain the time series features.
21. The method of claim 11, further comprising: The animal emotion recognition model is trained through adversarial training.
22. The method according to claim 21, wherein, The process of training the animal emotion recognition model through adversarial training includes: Add perturbation to the fusion feature corresponding to the multimodal feature vector at the current time step to generate a fusion feature with added multimodal noise; The fused features of the multimodal noisy features are input into the animal emotion recognition model to identify the current emotion of the animal and obtain the emotion prediction result of the animal.
23. The method according to claim 22, wherein, The step of adding perturbations to the fused features corresponding to the multimodal feature vector at the current time step to generate fused features with added multimodal noise includes: Obtain the first statistical value from the first classification probability distribution of the processing result of the previous time step at the current time step; the processing result includes the sentiment classification result obtained by performing sentiment classification on the fused features of the previous time step; Based on the negative correlation between the intensity and the first statistical value, a first disturbance intensity is generated; Based on the first perturbation intensity, the fusion feature corresponding to the multimodal feature vector, and the classification loss gradient of the fusion feature corresponding to the multimodal feature vector, the fusion feature of the multimodal noisy feature is generated.
24. The method according to claim 21, wherein, The process of training the animal emotion recognition model through adversarial training includes: Add a perturbation to the feature vector of the target mode in the multimodal feature vector to obtain the perturbation feature of the target mode; The feature vector of the target mode in the multimodal feature vector is replaced with the perturbation feature of the target mode to obtain the multimodal noise-added feature; The multimodal noisy features are input into the animal emotion recognition model to identify the animal's current emotion and obtain the animal's emotion prediction result.
25. The method according to claim 24, wherein, The step of adding perturbation to the feature vector of the target mode in the multimodal feature vector to obtain the perturbation feature of the target mode includes: Obtain the second statistical value from the second classification probability distribution of the target modality in the previous time step of the current time step; the second classification probability distribution is the classification probability distribution obtained by performing sentiment classification on the feature vector of the target modality; A second disturbance intensity is generated based on the negative correlation between the intensity and the second statistical value; Based on the second perturbation intensity, a perturbation is added to the feature vector of the target mode in the multimodal feature vector to obtain the perturbation feature of the target mode.
26. The method according to claim 22 or 24, wherein, The step of training the animal emotion recognition model based on the emotion recognition results includes: Based on the emotion recognition results and training labels, determine the first classification loss; Based on the sentiment prediction results and the training labels, determine the second classification loss; Based on the first classification loss and the second classification loss, the model parameters of the animal emotion recognition model are optimized.
27. The method according to any one of claims 22-26, further comprising: The temporal features of the multimodal feature vectors obtained during the processing of the multimodal feature vectors by the animal emotion recognition model are acquired. The temporal features of the multimodal feature vectors are used for emotion classification to obtain the emotion recognition result. The temporal features of the multimodal noisy features obtained during the processing of the multimodal noisy features by the animal emotion recognition model are acquired. The temporal features of the multimodal noisy features are used for emotion classification to obtain the emotion prediction result. The model parameters of the animal emotion recognition model are optimized based on the first consistency loss between the temporal features of the multimodal feature vector and the temporal features of the multimodal noisy features.
28. The method according to any one of claims 22-26, further comprising: Determine the second consistency loss between the emotion recognition result and the emotion prediction result; Based on the second consistency loss, the model parameters of the animal emotion recognition model are optimized.
29. The method according to any one of claims 10-28, wherein, The preprocessing of the multimodal data to obtain fused multimodal data includes: Based on the timestamps in the multimodal data, multimodal data within the same timestamp range are aggregated to obtain the fused multimodal data.
30. The method according to any one of claims 10-29, further comprising: Generate pseudo-labels for the emotions of the animals; The animal emotion recognition model is trained based on the pseudo-labels.
31. The method according to claim 30, wherein, The pseudo-labels for generating the animal's emotions include: When the animal emotion recognition model meets the preset conditions, sample data with expert labels is input into the animal emotion recognition model to obtain the predicted labels of the sample data output by the animal emotion recognition model. The predicted labels and expert labels of the sample data are weighted and summed to obtain the pseudo labels of the sample data.
32. The method according to claim 31, wherein, The step of weighted summing of the predicted labels and expert labels of the sample data to obtain pseudo-labels for the sample data includes: Obtain the confidence level of the animal emotion recognition model for the emotion classification distribution of multiple training samples in the previous round; The statistical values of the confidence scores of the sentiment classification distributions of the multiple training samples are determined and used as dynamic thresholds; The expert label's fusion weight and the predicted label's fusion weight are generated based on the dynamic threshold; wherein, the closer to the dynamic threshold, the greater the expert label's fusion weight. Based on the fusion weights of the expert labels and the fusion weights of the predicted labels, a weighted sum is performed on the predicted labels and the expert labels to obtain the pseudo-labels of the sample data.
33. The method according to claim 31, wherein, The preset conditions include at least one of the following: The animal emotion recognition model is trained through training rounds based on training samples with expert labels to reach the target number of rounds. The animal emotion recognition model meets the set conditions for entering the next stage of training based on the emotion recognition accuracy of the training samples with expert labels.
34. The method according to claim 30, wherein, The pseudo-labels for generating the animal's emotions include: Determine the classification divergence degree among multiple modalities in the animal-related multimodal data; the classification divergence degree is used to measure the difference between the animal emotion categories expressed by different modalities. When the classification divergence degree is less than the divergence threshold, the pseudo-label is obtained by fusing multiple sentiment classification results from multiple modalities in the multimodal data.
35. The method according to claim 34, wherein, The pseudo-label is obtained by fusing multiple sentiment classification results from multiple modalities in the multimodal data, including: Based on the classification confidence of the multiple sentiment classification results of the multiple modalities, multiple weighted values of the multiple modalities are generated; Based on the multiple weighted values and the multiple sentiment classification results, pseudo-labels are generated for the sample data represented by the multimodal data.
36. The method of claim 34, further comprising: If the classification divergence degree is not less than the divergence threshold, the multimodal data is stored in the observation pool.
37. The method of claim 30, further comprising: If the confidence level of the pseudo-label is greater than the dynamic threshold, the sample data corresponding to the pseudo-label will be determined as the sample data for training the animal emotion recognition model. The dynamic threshold is determined based on statistical values, which are the confidence scores of the emotion recognition model for the emotion classification distribution of multiple training samples in the previous round.
38. The method according to any one of claims 30-37, wherein, The process of training the animal emotion recognition model based on the pseudo-labels includes: Based on the pseudo-labels, at least one of the following training losses is determined: the cross-entropy loss of the pseudo-labels, the loss between the pseudo-labels and the expert labels, and the alignment loss between the pseudo-label distribution and the expert label distribution; The total loss is determined based on the at least one training loss, and the animal emotion recognition model is optimized based on the total loss.
39. The method according to claim 38, wherein, The total loss includes the total cross-entropy loss of multiple sample data. Determining the total cross-entropy loss includes: Generate multiple loss weights corresponding to the multiple sample data; wherein, the multiple sample data have multiple corresponding pseudo-labels; Based on multiple loss weights corresponding to multiple sample data, the multiple cross-entropy losses of the multiple pseudo-labels are weighted and summed to obtain the total cross-entropy loss.
40. The method according to claim 39, wherein, The generation of multiple loss weights corresponding to the multiple sample data includes: Based on the degree of difference between the multiple pseudo-labels and their corresponding expert labels, multiple loss weights are generated for the multiple sample data. The degree of difference is inversely correlated with the loss weight.
41. The method according to claim 38, wherein, The total loss includes the total distance loss of multiple sample data. Determining the total distance loss includes: Determine the distance between the pseudo-labels and their corresponding expert labels in the plurality of sample data; The statistical values of the distances between the multiple sample data are determined to obtain the total distance loss.
42. The method according to claim 38, wherein, The total loss includes the total alignment loss of multiple sample data, which is used to measure the distribution difference between the pseudo-labels and expert labels of the multiple sample data. During the training phase, the total alignment loss is minimized by optimizing the model parameters of the animal emotion recognition model.
43. The method according to any one of claims 10-42, further comprising: In the inference phase of the animal emotion recognition model, if a preset event is detected multiple times consecutively, the multimodal data corresponding to the preset event is acquired as new sample data. The preset event includes at least one of the following: The emotion recognition result indicates that the first type of target emotion has been identified, and the confidence level of the identification is greater than a preset threshold. The emotion recognition result indicates that the second type of target emotion has been identified and has been detected continuously n times, where n is a positive integer greater than 1; The animal's emotional fluctuation within the sliding time window where the emotion recognition result is located is higher than a preset value.
44. An animal language conversion method, comprising: Acquire multimodal data related to animals, including animal sound data, animal behavior data, and animal physical characteristics data; The multimodal data is preprocessed to obtain fused multimodal data; The animal's current emotion is identified based on the fused multimodal data to obtain the animal's emotion recognition result; The emotion recognition results are semantically mapped and translated to convert animal language into human language, resulting in a language conversion result.
45. The method according to claim 44, wherein, The acquisition of animal-related multimodal data includes: Collect the sound wave information emitted by the animal to obtain the animal's sound data; Collect the animal's body language and movement changes to obtain the animal's behavioral data; The animal's physical and biological indicators were collected to obtain the animal's vital signs data.
46. The method of claim 44, wherein, The preprocessing of the multimodal data to obtain fused multimodal data includes: The multimodal data is denoised to perform data cleaning, resulting in cleaned multimodal data. The cleaned multimodal data is normalized to obtain normalized multimodal data; The normalized multimodal data is then subjected to time-series alignment and fusion processing to obtain fused multimodal data.
47. The method of claim 44, wherein, The step of identifying the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result includes: A deep learning model is used to extract sound features, visual motion features, and analyze vital sign changes from the fused multimodal data to obtain multimodal feature vectors. Generative adversarial networks are used to perform sentiment analysis on the multimodal features to obtain the emotion recognition results of the animals.
48. The method according to claim 44, wherein, The step of performing semantic mapping and language translation on the emotion recognition results to convert animal language into human language and obtain language conversion results includes: Emotional tags and voice features are extracted from the emotion recognition results, and the voice features are converted into standardized voice vectors; A pre-trained language model is used to semantically map the emotion tags to the sound vectors to obtain the emotion intent; A language generator is used to translate the emotional intent into language to generate the corresponding human language, thus obtaining the language conversion result.
49. The method according to any one of claims 44 to 48, wherein, The method further includes: If specific voice data is detected and there is no history of emotion matching, the specific voice data is labeled to obtain an updated emotion tag. The sample data is dynamically updated based on the updated sentiment tags so that the model parameters can be adjusted according to the updated sample data.
50. The method according to any one of claims 44 to 48, wherein, The method further includes: Collect multimodal data within a preset time window to obtain data on the animal's emotional changes; Feature extraction is performed on the emotional change data to obtain emotional change features; The sentiment labels are updated based on the difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window.
51. The method according to claim 50, wherein, The collection of multimodal data within a preset time window to obtain animal emotional change data includes: Collect multimodal data within a preset time window; The multimodal data is input into the emotional period recognition model, and the emotional change data of the animal is obtained through the emotional period recognition model.
52. The method according to claim 50, wherein, The step of updating sentiment tags based on the difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window includes: The Euclidean distance difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window is calculated to obtain the sentiment gap. If the emotional gap exceeds a preset emotional gap threshold, the emotional label will be upgraded to obtain an updated emotional label.
53. The method according to claim 52, wherein, The method further includes: The sentiment weight corresponding to each time window is scored to obtain the sentiment weight score; The emotional weight scores corresponding to similar emotional time windows within a continuous time period are accumulated. If the accumulated result is greater than the set upgrade threshold, the emotional label is updated.
54. A training method for an animal emotion recognition model in an animal language conversion model, comprising: Acquire multimodal data related to animals; The multimodal data is preprocessed to obtain fused multimodal data; The fused multimodal data is input into the animal emotion recognition model to identify the animal's current emotion and obtain the animal's emotion recognition result. Based on the emotion recognition results, the animal emotion recognition model is trained; the emotion recognition results output by the animal emotion recognition model are used for semantic mapping and language translation to convert animal language into human language and obtain language conversion results.
55. An animal emotion recognition device, comprising: The first acquisition module is used to acquire multimodal data related to animals; The first preprocessing module is used to preprocess the multimodal data to obtain fused multimodal data; The first recognition module is used to identify the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result.
56. A training device for an animal emotion recognition model, comprising: The second acquisition module is used to acquire multimodal data related to animals; The second preprocessing module is used to preprocess the multimodal data to obtain fused multimodal data; The second recognition module is used to input the fused multimodal data into the animal emotion recognition model to recognize the animal's current emotion and obtain the animal's emotion recognition result. The training module is used to train the animal emotion recognition model based on the emotion recognition results.
57. An animal language conversion device, comprising: The first acquisition module is used to acquire multimodal data related to animals, including animal sound data, animal behavior data, and animal physical characteristics data. The first preprocessing module is used to preprocess the multimodal data to obtain fused multimodal data; The first recognition module is used to recognize the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result; The conversion module is used to perform semantic mapping and language translation on the emotion recognition results to convert animal language into human language and obtain the language conversion result.
58. A training device for an animal emotion recognition model in an animal language conversion model, comprising: The second acquisition module is used to acquire multimodal data related to animals; The second preprocessing module is used to preprocess the multimodal data to obtain fused multimodal data; The second recognition module is used to input the fused multimodal data into the animal emotion recognition model to recognize the animal's current emotion and obtain the animal's emotion recognition result. The training module is used to train the animal emotion recognition model based on the emotion recognition results; the emotion recognition results output by the animal emotion recognition model are used for semantic mapping and language translation to convert animal language into human language and obtain language conversion results.
59. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-54.
60. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-54.
61. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-54.