Animal language conversion methods, devices, electronic equipment and storage media

By acquiring and preprocessing multimodal data from animals, and utilizing machine learning and deep learning technologies for emotion recognition and semantic mapping, this approach solves the problem of existing technologies being unable to deeply understand animal emotions, and enables efficient cross-species emotional communication and interaction.

CN119943059BActive Publication Date: 2025-12-02BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Application Number
CN202411793938.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-12-02
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

Existing technologies cannot deeply understand the complex emotional aspects of animals, nor can they achieve deep, real-time emotional understanding and interactive communication between humans and animals. They suffer from problems such as limited emotional translation, lack of multimodal fusion analysis, insufficient temporal detection, lack of adaptive learning and optimization mechanisms, and insufficient real-time performance.

Method used

By acquiring multimodal data from animals, including vocal, behavioral, and physical data, preprocessing and fusing them, and using machine learning and deep learning techniques for emotion recognition and semantic mapping, the emotional states of animals are converted into human language.

Benefits of technology

It achieves comprehensive capture and accurate identification of animal emotions, enhances the communication ability between humans and animals, improves the accuracy of emotional understanding and the real-time nature of interaction, and provides humans with a brand-new way to communicate with animals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943059B_ABST
    Figure CN119943059B_ABST
Patent Text Reader

Abstract

This disclosure provides a method, apparatus, electronic device, and storage medium for animal language conversion, relating to the field of artificial intelligence technology, specifically machine learning, deep learning, and natural language processing. The specific implementation involves: acquiring multimodal data related to the animal, including animal vocal data, animal behavioral data, and animal physical characteristics data; preprocessing the multimodal data to obtain fused multimodal data; identifying the animal's current emotion based on the fused multimodal data to obtain an emotion recognition result; and performing semantic mapping and language translation on the emotion recognition result to convert the animal language into human language, obtaining a language conversion result. This disclosure can accurately identify the animal's current emotional state and convert it into human language, thereby achieving deeper emotional communication and understanding between animals and humans, and improving the accuracy and efficiency of cross-species communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, specifically to the fields of machine learning, deep learning, and natural language processing, and particularly to methods, devices, electronic devices, and storage media for animal language conversion. Background Technology

[0002] Current technologies on the market attempt to interpret the emotional world of animals through basic animal vocalization and behavior translation devices and pet emotion analysis tools utilizing artificial intelligence image recognition technology. However, these methods are often limited to a superficial interpretation of animal behavior. These technologies cannot delve into the complex emotional levels of animals, nor can they achieve deep, real-time emotional understanding and interactive communication between humans and animals. Summary of the Invention

[0003] This disclosure provides a method, apparatus, electronic device, and storage medium for converting animal language.

[0004] According to one aspect of this disclosure, an animal language conversion method is provided, the method comprising:

[0005] Acquire multimodal data related to animals, including animal sound data, animal behavior data, and animal physical characteristics data;

[0006] The multimodal data is preprocessed to obtain fused multimodal data;

[0007] The animal's current emotion is identified based on the fused multimodal data to obtain the animal's emotion recognition result;

[0008] The emotion recognition results are semantically mapped and translated to convert animal language into human language, resulting in a language conversion result.

[0009] According to another aspect of this disclosure, an animal language conversion device is provided, comprising:

[0010] The acquisition module is used to acquire multimodal data related to animals, including animal sound data, animal behavior data, and animal physical characteristics data.

[0011] The preprocessing module is used to preprocess the multimodal data to obtain fused multimodal data;

[0012] An emotion recognition module is used to identify the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result;

[0013] The conversion module is used to perform semantic mapping and language translation on the emotion recognition results to convert animal language into human language and obtain the language conversion result.

[0014] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described in any of the above technical solutions.

[0018] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any one of the methods described above.

[0019] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in any one of the above technical solutions.

[0020] This disclosure provides a method, apparatus, device, and storage medium for animal language conversion. By acquiring multimodal data such as animal sounds, behaviors, and physical characteristics, this disclosure enables comprehensive capture and accurate identification of animal emotions. Then, through semantic mapping and language translation technologies, the animal's emotional state and intentions are converted into language that humans can understand, thereby greatly enhancing communication between humans and animals, improving the accuracy of animal emotional understanding and the real-time nature of interaction, and providing humans with a completely new way to communicate with animals. In other words, this solution can accurately identify the current emotional state of an animal and convert it into human language, thereby achieving deeper emotional communication and understanding between animals and humans, and improving the accuracy and efficiency of cross-species communication.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0022] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0023] Figure 1 This is a schematic diagram of the steps of the animal language conversion method in the embodiments of this disclosure;

[0024] Figure 2 This is a schematic diagram of the process for acquiring multimodal data in an embodiment of this disclosure;

[0025] Figure 3 This is a schematic diagram of the process for preprocessing multimodal data in an embodiment of this disclosure;

[0026] Figure 4 This is a schematic diagram of the process for obtaining the emotion recognition results of animals in an embodiment of this disclosure;

[0027] Figure 5 This is a schematic diagram of the process of converting animal language into human language in an embodiment of this disclosure;

[0028] Figure 6 This is a schematic diagram of the process for updating sentiment tags in an embodiment of this disclosure;

[0029] Figure 7 A schematic block diagram of the animal language conversion device in the embodiments of this disclosure;

[0030] Figure 8 This is a block diagram of an electronic device used to implement the animal language conversion method of the embodiments of this disclosure. Detailed Implementation

[0031] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0032] Currently, technologies related to human-animal communication on the market can be mainly divided into the following two types:

[0033] The first approach uses simple animal vocalization and behavior translation devices, which primarily rely on a "voiceprint database + simple algorithm" to translate animal emotions. A typical example is some commercially available pet behavior and sound recognition devices. These devices use relatively simple sensors to capture animal sounds and some typical movements, then match them to a pre-built emotion database for simple emotion mapping, such as recognizing a dog's bark as a request for food or anger.

[0034] The second approach is based on AI (Artificial Intelligence) image recognition pet emotion analysis tools, which introduce image processing and AI technology to help users understand their pets' emotional responses. For example, some companies capture animal facial expressions and movements in camera images and then combine them with existing deep learning-trained models to analyze different facial expressions and identify specific expressions such as laughter, anger, grievance, and confusion.

[0035] The two methods mentioned above have several limitations in understanding and translating animal emotions: First, emotion translation is singular, relying on pre-set voiceprint data and behavioral classifications, and cannot continuously track emotional changes, resulting in insufficient accuracy in complex scenarios. Second, the lack of multimodal fusion analysis and over-reliance on a single information source limits the comprehensiveness and accuracy of emotion translation. Third, insufficient temporal detection and the lack of ability to track the continuity of emotional states lead to sluggish perception of emotional changes. Furthermore, the lack of adaptive learning and optimization mechanisms makes it difficult to optimize and iterate when faced with unknown emotional patterns, limiting the system's flexibility and scalability. Finally, the inability to perform edge computing results in insufficient real-time performance, affecting instant interaction between humans and animals. These limitations collectively contribute to the inadequacy of existing technologies in achieving deep, real-time cross-species emotional communication.

[0036] To address the aforementioned technical problems, this disclosure provides an animal language conversion method, see [link to relevant documentation]. Figure 1 As shown, Figure 1 This is a schematic diagram illustrating the steps of the animal language conversion method in this embodiment of the disclosure. The method can be applied to the server side, and the method includes:

[0037] Step S101: Obtain multimodal data related to animals, including animal sound data, animal behavior data, and animal physical characteristics data.

[0038] Specifically, acquiring multimodal data related to animals refers to collecting different types of information through various sensors to comprehensively understand the animal's state and emotions. "Animal vocal data" refers to the various calls made by animals captured by audio sensors; these sounds can reflect the animal's emotions and needs. "Animal behavioral data" refers to the animal's body movements and postures recorded by video cameras, helping to analyze its behavioral patterns and emotional expressions. "Animal vital signs data" refers to the animal's physiological indicators monitored by physiological sensors, such as heart rate and body temperature; these indicators can provide physiological evidence of the animal's emotional state.

[0039] In this way, by acquiring multimodal data of animals, including vocal, behavioral, and physical data, it is possible to comprehensively capture animal emotions and behaviors, thereby providing more accurate emotion analysis and behavioral understanding, enhancing communication and interaction between humans and animals, and improving animal welfare and human responsiveness to animal behavior.

[0040] Step S102: Preprocess the multimodal data to obtain fused multimodal data.

[0041] Specifically, preprocessing multimodal data refers to a series of processing steps performed on raw data collected from different sources (such as audio, video, and physiological sensors). This may include noise reduction, normalization, and feature extraction to facilitate analysis and understanding. "Fused multimodal data" refers to integrating this preprocessed data into a unified dataset. This process involves time alignment (ensuring all data is synchronized in time), feature fusion (merging features from different modalities into a single comprehensive feature vector), and data synchronization (handling differences in timestamps and sampling rates between different data streams). Through this preprocessing and fusion, a comprehensive data view can be provided, making subsequent sentiment recognition and language translation more accurate and efficient.

[0042] By preprocessing multimodal data such as animal sounds, behaviors, and physical signs, and then integrating these data into a unified dataset, the consistency and usability of the data can be improved. This leads to more accurate emotion recognition and language translation, enhances the system's comprehensive understanding of animal behavior and emotional states, and improves the efficiency and effectiveness of cross-species communication.

[0043] Step S103: Identify the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result.

[0044] Specifically, after obtaining the fused multimodal data, the current emotion of the animal is identified based on this data. This involves using a comprehensive dataset integrating information such as the animal's vocalizations, behaviors, and physical characteristics, and employing machine learning and deep learning techniques to analyze and judge the animal's emotional state. Here, "emotion recognition" refers to identifying the animal's emotional state, such as anxiety, excitement, or relaxation, by analyzing the features extracted from this multimodal data. To obtain the animal's emotion recognition results, the fused data can be input into a trained emotion recognition model. This model compares the features of known emotional states to output the animal's emotion recognition result, thus providing a basis for subsequent language conversion and human-computer interaction.

[0045] In this way, by analyzing the fused multimodal data, including animal sounds, behaviors, and physiological signs, it is possible to accurately identify the animal's current emotional state. This process not only improves the accuracy of understanding and responding to animal emotions but also enhances communication between humans and animals, enabling humans to better interpret animal needs and emotions, thereby improving animal welfare and the intimacy of the human-pet relationship.

[0046] Step S104: Semantic mapping and language translation are performed on the emotion recognition results to convert animal language into human language and obtain the language conversion result.

[0047] Specifically, "semantic mapping and language translation of emotion recognition results" refers to the process of converting animal emotional states, such as anxiety, excitement, or happiness, obtained through the analysis of multimodal data, into language expressions that humans can understand. This process involves using pre-trained language models and deep learning techniques to establish a correspondence between animal emotional features and corresponding expressions in human language—a process known as "semantic mapping." Subsequently, these mapping results are converted into specific text or speech output, achieving "language translation." Ultimately, the "language conversion result" refers to the animal's emotions and intentions being transformed into human language, enabling human users to intuitively understand the animal's "language," thereby achieving effective cross-species communication.

[0048] In this way, by semantically mapping and linguistically translating the results of emotion recognition, it is possible to transform the nonverbal communication of animals into language that humans can understand. This process not only breaks down communication barriers between humans and animals, but also greatly enhances human understanding of animal emotions and needs, making human-animal interactions more harmonious. It also provides new perspectives and tools for animal welfare and behavioral research.

[0049] This disclosure provides a method, apparatus, device, and storage medium for animal language conversion. By acquiring multimodal data such as animal sounds, behaviors, and physical characteristics, this disclosure enables comprehensive capture and accurate identification of animal emotions. Then, through semantic mapping and language translation technologies, the animal's emotional state and intentions are converted into language that humans can understand, thereby greatly enhancing communication between humans and animals, improving the accuracy of animal emotional understanding and the real-time nature of interaction, and providing humans with a completely new way to communicate with animals. In other words, this solution can accurately identify the current emotional state of an animal and convert it into human language, thereby achieving deeper emotional communication and understanding between animals and humans, and improving the accuracy and efficiency of cross-species communication.

[0050] In some optional embodiments, acquiring animal-related multimodal data includes:

[0051] Collect sound wave information emitted by animals to obtain animal sound data;

[0052] Collect animal body language and movement changes to obtain animal behavior data;

[0053] Collect physical and biological indicators of animals to obtain animal vital signs data.

[0054] Specifically, audio acquisition devices capture animal vocalizations to obtain "animal sound data." Simultaneously, visual sensors such as video cameras collect data on animal body language and movement changes, forming "animal behavior data," which reflects the animal's activity and nonverbal behavior. Furthermore, physiological sensors monitor "animal vital signs data," such as heart rate and body temperature; these physical and biological indicators provide crucial information for understanding the animal's physiological state and emotions. By integrating this multimodal data, it is possible to comprehensively capture the animal's communication methods and emotional state, laying the foundation for further data analysis and emotion recognition.

[0055] To facilitate understanding of the solutions in the embodiments of this disclosure, examples are provided below, see [link to example]. Figure 2 As shown, Figure 2 This is a schematic diagram of the process for collecting multimodal data in an embodiment of this disclosure. First, an audio acquisition device is used to capture the sounds emitted by animals in real time. Then, the audio acquisition device sends the captured sound data to a data processing module. The data processing module can use appropriate audio filters to process noise and reduce background interference. For example, for a dog bark, its pitch, amplitude, duration, frequency, and breakpoint changes are sampled together to capture all the information of the bark.

[0056] For animal behavior data collection, video equipment (such as high-definition cameras and infrared cameras) is used to capture animal limb movements and behavioral expressions, such as tail wagging, jumping, and lying down, and these are analyzed in conjunction with specific animal physical characteristics (ears erect, pupils dilated, etc.). The captured video / image data is then sent to the data processing module.

[0057] For animal vital sign data collection, heart rate and body temperature are collected using high-precision contact or non-contact body temperature detection sensors, and the collected vital sign data is sent to the data processing module.

[0058] In this way, by collecting the sound wave information emitted by animals, body language and movement changes, as well as physical and biological indicators, relatively comprehensive animal sound data, behavioral data, and physical signs data can be obtained. The integration of this multi-dimensional information provides a rich and accurate data foundation for in-depth understanding and analysis of animal emotions and behaviors, thereby enabling humans to more accurately interpret animals' communication intentions and physiological states, strengthening non-verbal communication between humans and animals, improving the effectiveness of animal care and training, and opening up new avenues for animal health monitoring and behavioral research.

[0059] In some optional embodiments, the multimodal data is preprocessed to obtain fused multimodal data, including:

[0060] Denoising is performed on the multimodal data to clean the data, resulting in cleaned multimodal data.

[0061] The cleaned multimodal data is normalized to obtain normalized multimodal data;

[0062] The normalized multimodal data is time-series aligned and fused to obtain fused multimodal data.

[0063] Specifically, denoising multimodal data for data cleaning refers to using signal processing techniques to remove noise and interference from audio and video data to improve data quality and obtain cleaned multimodal data. Next, "normalizing the cleaned multimodal data" means converting data from different sources and scales into a unified format or scale so that machine learning models can process it more effectively. This step yields normalized multimodal data. Finally, "temporally aligning and fusing the normalized multimodal data" involves aligning data from different modalities in time and merging them into a unified dataset. This step ensures temporal consistency of the data and integrates information from different sensors, ultimately resulting in fused multimodal data, providing an accurate and comprehensive data foundation for subsequent emotion recognition and behavior analysis.

[0064] To facilitate understanding of the solutions in the embodiments of this disclosure, examples are provided below, see [link to example]. Figure 3 As shown, Figure 3This is a schematic diagram of the preprocessing flow for multimodal data in this embodiment of the disclosure. First, the audio data, image data, and body temperature and heart rate data in the multimodal data are preprocessed and normalized. This preprocessing includes noise reduction and data cleaning, which removes invalid parts from the audio and visual information. For example, noise that may be generated by humans (wind, speech, etc.) is filtered out, making the audio signal clearer and easier to identify. In images, background motion and objects in video frames are cleaned to retain only key information, such as changes in animal behavior. Next, data normalization is performed. Whether it's audio signals, video, or vital sign information, it needs to be normalized, converted into a unified standard, and expressed as a feature vector that can be processed by machine learning algorithms.

[0065] Finally, the multimodal data undergoes temporal alignment and fusion. Data alignment is a prerequisite for data fusion, and it requires addressing the temporal and spatial differences between the acquisition of different modal signals. For example, a dog barking event often corresponds to a time difference with the moment of accompanying limb movement or body temperature fluctuation. This data must be temporally calibrated before input to ensure a consistent reference point for the multimodal input in the translation task.

[0066] By performing denoising, normalization, temporal alignment, and fusion processing on multimodal data, the quality, consistency, and usability of the data can be significantly improved, making the features extracted from animal sounds, behaviors, and physical characteristics more accurate and reliable. This process not only enhances the accuracy of data analysis but also improves the model's precision in recognizing animal emotions and behaviors, providing a solid data foundation for achieving efficient and accurate cross-species communication.

[0067] In some optional embodiments, the animal's current emotion is identified based on the fused multimodal data to obtain the animal's emotion recognition result, including:

[0068] A deep learning model is used to extract sound features, visual motion features, and analyze vital sign changes from the fused multimodal data to obtain multimodal feature vectors.

[0069] Generative adversarial networks were used to perform sentiment analysis on multimodal features to obtain the emotion recognition results of animals.

[0070] Specifically, "using deep learning models to extract sound features, visual motion features, and analyze vital sign changes from fused multimodal data" refers to utilizing advanced deep learning techniques, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), to process and analyze multimodal datasets that integrate sound, visual, and physiological data. This process involves extracting sound features, such as pitch, rhythm, and volume, from audio signals; extracting visual motion features, such as posture and behavioral patterns, from video data; and analyzing vital sign changes, such as heart rate and body temperature fluctuations, from physiological data. These features collectively constitute a multimodal feature vector, which comprehensively reflects the animal's physiological and behavioral state. Next, "using generative adversarial networks for sentiment analysis of multimodal features" refers to using a deep learning framework like generative adversarial networks (GANs) to enable the network to recognize and distinguish different emotional states through an adversarial training process. This process ultimately produces animal emotion recognition results, transforming the animal's complex emotions and behaviors into understandable emotion labels, providing a basis for further semantic mapping and language translation.

[0071] To facilitate understanding of the solutions in the embodiments of this disclosure, examples are provided below, see [link to example]. Figure 4 As shown, Figure 4 This is a flowchart illustrating the process of obtaining animal emotion recognition results in this embodiment. After obtaining the fused multimodal data, fine-grained feature extraction is performed on the data of each modality using a deep learning model. Specifically, the deep learning model extracts voice features, visual motion features, and analyzes physical characteristic changes to obtain a feature vector combination. Next, a large model based on an emotion recognition generative adversarial network (GAN) is used to analyze the voiceprint features, motion changes, and physical characteristic fluctuations in the data to obtain emotion classification labels. For example, when a low-frequency barking of an animal is detected accompanied by limb tension and dilated pupils, the model identifies the animal as being in a state of high alert through comparative inference and further infers the possible underlying psychological activities (fear or confusion) through feature fusion.

[0072] In this way, by employing deep learning models to comprehensively analyze the fused multimodal data, it is possible to accurately extract features such as animal vocalizations, visual movements, and changes in physical characteristics, forming multimodal feature vectors. Then, generative adversarial networks are used to conduct in-depth sentiment analysis on these features, ultimately yielding animal emotion recognition results. This process not only improves the accuracy and depth of emotion recognition but also enhances the understanding of animal behavior and psychological states, providing strong technical support for achieving more effective human-animal communication.

[0073] In some optional embodiments, the emotion recognition results are semantically mapped and translated to convert animal language into human language, resulting in a language conversion result, including:

[0074] Emotional labels and voice features are extracted from the emotion recognition results, and the voice features are converted into standardized voice vectors.

[0075] A pre-trained language model is used to semantically map emotion tags to sound vectors in order to obtain the emotion intent;

[0076] A language generator is used to translate emotional intentions into language to generate corresponding human language, thus obtaining the language conversion result.

[0077] Specifically, a "pre-trained language model" refers to an artificial intelligence model that has been trained on a large amount of data and is capable of understanding and processing natural language. This model is used here to perform a "semantic mapping" between animal emotional labels—that is, the animal's emotional state obtained from the emotion recognition module—and sound vectors. This involves correlating the animal's vocal features with the emotional semantics in human language, thereby identifying the animal's emotional intentions. Next, the "language generator" is the process of converting the animal's nonverbal communication into language that humans can understand, based on these emotional intentions. This process involves converting the animal's emotions and intentions into specific text or speech output, i.e., the "language conversion result," enabling humans to intuitively understand the animal's "language" and achieve effective communication between humans and animals.

[0078] To facilitate understanding of the solutions in the embodiments of this disclosure, examples are provided below, see [link to example]. Figure 5 As shown, Figure 5 This is a schematic diagram illustrating the process of converting animal language into human language in an embodiment of this disclosure. After the emotion recognition module obtains the emotion recognition result, it provides the animal's emotion tag and vocal features, and then passes this information to the speech mapping module, which is responsible for converting the animal's vocal features into human-understandable speech expressions. Next, the human semantic output module receives the converted speech information and outputs it as a language conversion result, ultimately presenting these results to the user, realizing real-time conversion of animal language to human language and emotional communication. The entire process involves emotion recognition, feature extraction and mapping, and human language generation, aiming to promote effective communication between humans and animals.

[0079] In this way, by using a pre-trained language model to semantically map emotion tags to sound vectors, it is possible to accurately capture and understand the emotional intentions of animals. Then, a language generator is used to convert these emotional intentions into human language. This process not only achieves accurate interpretation of animal emotions but also transforms animal nonverbal communication into human-understandable language, greatly promoting communication and understanding between humans and animals, and improving the transparency and efficiency of animal emotional expression.

[0080] In some optional embodiments, the method further includes:

[0081] If specific audio data is detected and there is no history of emotion matching, the specific audio data is labeled to obtain an updated emotion label;

[0082] The sample data is dynamically updated based on the updated sentiment tags, so that the model parameters can be adjusted according to the updated sample data.

[0083] Specifically, when a specific animal sound is detected that does not exist in the historical sentiment matching records (i.e., there is no corresponding previous sentiment label), a labeling process is triggered. Here, "labeling" refers to manually assigning a sentiment label to these specific sound data. This label describes the animal's emotional state when making the sound, thus "obtaining an updated sentiment label." Subsequently, this newly labeled sentiment label is incorporated into the sample database, achieving "dynamic updates to the sample data based on the updated sentiment label." This updated sample data is then used to adjust the model's parameters, i.e., "adjusting the model parameters based on the updated sample data," thereby optimizing and improving the model's ability to recognize the sentiment of newly appearing sound data. This ensures the system can adapt to new or uncommon animal sounds, improving recognition accuracy and the system's adaptive learning capabilities.

[0084] To facilitate understanding of the solutions in this disclosure, an example is given below: When the system encounters an unrecognizable vocal pattern or abnormal behavior, it prompts the user to input relevant tags or identification information. For instance, if the system detects a specific action + sound combination without a clear historical record of matching emotional expression, the user can annotate this phonetic symbol through the interface, such as "call for help" or "hunger." After the user manually annotates, the system adjusts the data parameters of the current audio and the corresponding tags for the behavior, thereby updating the weights of the existing model. After the user completes the annotation, the system updates the dataset in a timely manner, continuously optimizes and retrains the relevant generative model, so that the system can gradually improve the recognition rate of similar or new categories of emotional events.

[0085] In this way, by manually labeling specific sound data when there is a lack of historical emotion matching, and updating the emotion labels accordingly, the sample database can be continuously expanded and enriched, thereby dynamically updating the model parameters, enhancing the system's ability to recognize emotions in new sound data, improving the model's adaptability and accuracy, and ensuring that the cross-species communication system can continue to evolve and better understand and respond to the animal's communication intentions.

[0086] In some optional embodiments, the method further includes:

[0087] Collect multimodal data within a preset time window to obtain data on the animal's emotional changes;

[0088] Feature extraction is performed on the emotion change data to obtain emotion change features;

[0089] The sentiment labels are updated based on the difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window.

[0090] Specifically, a "preset time window" refers to a specific time period set for analyzing changes in animal emotions. During this period, "multimodal data" is collected, including information on the animal's vocalizations, behaviors, and physical characteristics, to gather data on its emotional state. Next, a "feature extraction" process identifies key information representing changes in the animal's emotions—the "emotional change features"—from this data. Then, the differences between the emotional features extracted in the current time window and those from the previous time window are compared. If these differences indicate a significant change in the animal's emotional state, the "emotional label" is updated to reflect the animal's latest emotional state. This process involves continuous monitoring and dynamic analysis of the animal's emotional state to ensure the real-time nature and accuracy of emotion recognition.

[0091] In this way, by collecting multimodal data from animals within a preset time window, it is possible to comprehensively capture the emotional changes of animals and extract features from this data to identify the characteristics of emotional changes. Secondly, by dynamically updating the emotional labels based on the differences between the current and previous time window's emotional features, real-time and accurate monitoring and response to the animal's emotional state can be achieved. This not only improves the accuracy of emotion recognition but also enhances the real-time nature and depth of communication between humans and animals, enabling humans to better understand and respond to the emotional needs of animals.

[0092] In some optional embodiments, multimodal data is collected within a preset time window to obtain data on changes in the animal's emotions, including:

[0093] Collect multimodal data within a preset time window;

[0094] Multimodal data is input into the emotional period recognition model, and the emotional change data of animals are obtained through the emotional period recognition model.

[0095] Specifically, a "preset time window" refers to a specific time period set for monitoring and analyzing changes in animal emotions. During this period, "multimodal data" is collected, including different types of information such as the animal's vocalizations, behaviors, and physical characteristics. This data is then fed into an "emotional period recognition model," an algorithm specifically designed to process and analyze time-series data, such as Long Short-Term Memory (LSTM) networks or Gated Recurrent Units (GRUs). By analyzing this multimodal data, the model identifies and extracts features related to the animal's emotional state, thereby "obtaining animal emotional change data." This data reflects the animal's emotional state within a continuous time window, providing a foundation for further emotion recognition and label updates. In this way, the solution enables dynamic tracking and accurate identification of animal emotional states.

[0096] By collecting multimodal data from animals within a preset time window and inputting this data into an emotion period recognition model, it is possible to accurately capture and analyze changes in animal emotions. This method can identify and understand the emotional dynamics of animals over continuous time periods, thus providing richer and more accurate data on emotional changes. This data not only enhances the depth and real-time performance of emotion recognition but also improves the quality and efficiency of communication between humans and animals, enabling humans to respond more subtly and promptly to animals' emotional needs and behavioral changes.

[0097] In some optional embodiments, the sentiment label is updated based on the difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window, including:

[0098] The Euclidean distance difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window is calculated to obtain the sentiment gap.

[0099] If the emotional gap exceeds the preset emotional gap threshold, the emotional label will be upgraded to obtain an updated emotional label.

[0100] Specifically, "emotional change features in the current time window" refers to emotion-related features extracted from multimodal data of animals within a specific time period, while "emotional change features in the previous time window" refers to the corresponding features in the previous time period. By calculating the "Euclidean distance difference" between the emotional features of these two time windows, which is a method for measuring the distance between two points in a multidimensional space, a quantitative "emotional gap" can be obtained.

[0101] The Euclidean distance between the emotional change features in the current time window and the emotional change features in the previous time window satisfies the formula: [D(W_i,W_{i-1})=\sqrt{\sum_{n=1}^{N}(x_{i,n}-x_{i-1,n})^2}], where D(W_i,W_{i-1}) represents the Euclidean distance between the emotional change features in the i-th time window and the (i-1)-th time window, which quantifies the degree of change in emotional features within two consecutive time windows. x_{i,n} represents the value of the n-th feature in the i-th time window, which may include pitch, rhythm, amplitude of sound, frequency of behavior, heart rate changes, etc. x_{i-1,n} represents the value of the n-th feature in the (i-1)-th time window. sum represents the summation operation, and sqrt represents the square operation.

[0102] The emotional gap is calculated by measuring the Euclidean distance between the emotional change features in the current time window and those in the previous time window. This gap reflects the degree of change in the animal's emotional state within two consecutive time windows. If this gap exceeds a "preset emotional gap threshold," indicating a significant change in the animal's emotional state, an "emotional label upgrade" is triggered. Based on the emotional gap and change pattern, the emotional label is updated from one state (e.g., "alert") to another label that better reflects the current emotional state (e.g., "anxiety"), resulting in an "updated emotional label." This process makes emotion recognition more dynamic and accurate, enabling real-time responses to significant changes in animal emotions.

[0103] By calculating the Euclidean distance difference between the current and previous time window's emotional change features, the degree of change in an animal's emotional state can be quantified, resulting in an emotional gap. When the emotional gap exceeds a preset threshold, the emotional label is automatically upgraded to an updated label. This method makes emotion recognition more sensitive and accurate, enabling real-time capture and response to significant changes in animal emotions. This improves the quality and efficiency of communication between humans and animals, ensuring that humans can understand and respond to animals' emotional needs promptly and appropriately.

[0104] In some optional embodiments, the method further includes:

[0105] The sentiment weight corresponding to each time window is scored to obtain the sentiment weight score;

[0106] The emotional weight scores corresponding to similar emotional time windows within a continuous time period are accumulated. If the accumulated result is greater than the set upgrade threshold, the emotional label is updated.

[0107] Specifically, "emotional weighting scoring" refers to the process of quantitatively assessing the emotional state of an animal within each time window, where each time window contains multimodal data collected over a specific period. By analyzing this data, emotional features are extracted and assigned different weights based on the strength or salience of the features, thus calculating the emotional weighting score for each time window. Next, the weighting scores of time windows with similar emotional features across consecutive time periods are accumulated. If this accumulated score exceeds a "set upgrade threshold," a predefined value used to determine whether the emotional state is significant enough to trigger a label update, the emotional label is updated to reflect the change in the animal's emotional state. This process makes emotion recognition more dynamic and accurate, enabling real-time responses to continuous changes in animal emotions.

[0108] By assigning weights and scoring emotional states within each time window, this scheme can quantify and track changes in an animal's emotions over continuous time periods. The weighted scores of time windows with similar emotional characteristics are accumulated; when the accumulated score exceeds a preset escalation threshold, the emotional label is automatically updated to more accurately reflect the animal's current emotional state. This method improves the sensitivity and adaptability of emotion recognition, making communication between humans and animals more precise and timely, thereby enhancing the understanding and response to animals' emotional needs.

[0109] For ease of understanding of the embodiments of this disclosure, please refer to Figure 6 As shown, Figure 6 This is a schematic diagram of the process for updating emotion tags in an embodiment of this disclosure. Audio recordings of animal barking and body temperature data are collected via sound acquisition and a body temperature sensor. This data is then sent to a data processing and fusion module for analysis to identify audio features and changes in body temperature. The analysis results are used for emotion feedback; based on the emotional meaning of the audio and body temperature changes, the animal's emotion tag is dynamically adjusted, such as upgrading a "vigilant" state to "anxious." This process involves real-time monitoring and continuous emotion state assessment to ensure that changes in the animal's emotion are reflected in the tag update in a timely manner, thereby improving the accuracy of understanding and responding to the animal's emotional state.

[0110] By collecting data on animal emotional changes within continuous time windows and extracting emotional change features from this data, subtle changes in animal emotions can be monitored and quantified in real time. By comparing these change features with standard emotional features and updating the emotional label when the change exceeds a preset threshold, the dynamic changes in animal emotions can be captured more accurately. This enables continuous tracking and real-time response to animal emotional states, improving the sensitivity and accuracy of emotion recognition and enhancing the effectiveness of emotional communication between humans and animals.

[0111] The following describes an embodiment of the apparatus described in this application, which can be used to execute the animal language conversion method described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the animal language conversion method described above.

[0112] This disclosure also provides an animal language conversion device 700, such as Figure 7 As shown, it includes:

[0113] The acquisition module 701 is used to acquire multimodal data related to animals, including animal sound data, animal behavior data, and animal physical characteristics data.

[0114] The preprocessing module 702 is used to preprocess the multimodal data to obtain fused multimodal data;

[0115] The emotion recognition module 703 is used to identify the current emotion of an animal based on the fused multimodal data, so as to obtain the emotion recognition result of the animal;

[0116] The conversion module 704 is used to perform semantic mapping and language translation on the emotion recognition results to convert animal language into human language and obtain the language conversion result.

[0117] In some optional embodiments, the acquisition module 701 acquires animal-related multimodal data, including:

[0118] Collect sound wave information emitted by animals to obtain animal sound data;

[0119] Collect animal body language and movement changes to obtain animal behavior data;

[0120] Collect physical and biological indicators of animals to obtain animal vital signs data.

[0121] In some optional embodiments, the preprocessing module 702 preprocesses the multimodal data to obtain fused multimodal data, including:

[0122] Denoising is performed on the multimodal data to clean the data, resulting in cleaned multimodal data.

[0123] The cleaned multimodal data is normalized to obtain normalized multimodal data;

[0124] The normalized multimodal data is time-series aligned and fused to obtain fused multimodal data.

[0125] In some optional embodiments, the emotion recognition module 703 identifies the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result, including:

[0126] A deep learning model is used to extract sound features, visual motion features, and analyze vital sign changes from the fused multimodal data to obtain multimodal feature vectors.

[0127] Generative adversarial networks were used to perform sentiment analysis on multimodal features to obtain the emotion recognition results of animals.

[0128] In some optional embodiments, the conversion module 704 performs semantic mapping and language translation on the emotion recognition results to convert animal language into human language, obtaining a language conversion result, including:

[0129] Emotional labels and voice features are extracted from the emotion recognition results, and the voice features are converted into standardized voice vectors.

[0130] A pre-trained language model is used to semantically map emotion tags to sound vectors in order to obtain the emotion intent;

[0131] A language generator is used to translate emotional intentions into language to generate corresponding human language, thus obtaining the language conversion result.

[0132] In some optional embodiments, the apparatus further includes a first update module, configured to annotate the specific sound data to obtain an updated emotion tag if specific sound data is detected and no emotion matching history exists;

[0133] The sample data is dynamically updated based on the updated sentiment tags, so that the model parameters can be adjusted according to the updated sample data.

[0134] In some optional embodiments, the device further includes a second update module for collecting multimodal data within a preset time window to obtain data on the animal's emotional changes.

[0135] Feature extraction is performed on the emotion change data to obtain emotion change features;

[0136] The sentiment labels are updated based on the difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window.

[0137] In some optional embodiments, the second update module collects multimodal data within a preset time window to obtain data on the animal's emotional changes, including:

[0138] Collect multimodal data within a preset time window;

[0139] Multimodal data is input into the emotional period recognition model, and the emotional change data of animals are obtained through the emotional period recognition model.

[0140] In some optional embodiments, the second update module updates the sentiment tags based on the difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window, including:

[0141] The Euclidean distance difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window is calculated to obtain the sentiment gap.

[0142] If the emotional gap exceeds the preset emotional gap threshold, the emotional label will be upgraded to obtain an updated emotional label.

[0143] In some optional embodiments, the second update module is used to:

[0144] The sentiment weight corresponding to each time window is scored to obtain the sentiment weight score;

[0145] The emotional weight scores corresponding to similar emotional time windows within a continuous time period are accumulated. If the accumulated result is greater than the set upgrade threshold, the emotional label is updated.

[0146] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0147] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0148] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0149] like Figure 8As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0150] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0151] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the animal language conversion method. For example, in some embodiments, the animal language conversion method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the applet distribution described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the animal language conversion method by any other suitable means (e.g., by means of firmware).

[0152] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0153] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0154] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0155] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0156] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0157] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0158] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0159] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for converting animal language, the method comprising: Acquire multimodal data related to animals, including animal sound data, animal behavior data, and animal physical characteristics data; The multimodal data is preprocessed to obtain fused multimodal data; The animal's current emotion is identified based on the fused multimodal data to obtain the animal's emotion recognition result; The emotion recognition results are semantically mapped and translated to convert animal language into human language, resulting in a language conversion result. Collect multimodal data within a preset time window to obtain data on the animal's emotional changes; Feature extraction is performed on the emotional change data to obtain emotional change features; The sentiment labels are updated based on the difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window.

2. The method according to claim 1, wherein, The acquisition of animal-related multimodal data includes: Collect the sound wave information emitted by the animal to obtain the animal's sound data; Collect the animal's body language and movement changes to obtain the animal's behavioral data; The animal's physical and biological indicators were collected to obtain the animal's vital signs data.

3. The method according to claim 1, wherein, The preprocessing of the multimodal data to obtain fused multimodal data includes: The multimodal data is denoised to perform data cleaning, resulting in cleaned multimodal data. The cleaned multimodal data is normalized to obtain normalized multimodal data. The normalized multimodal data is then subjected to time-series alignment and fusion processing to obtain fused multimodal data.

4. The method according to claim 1, wherein, The step of identifying the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result includes: A deep learning model is used to extract sound features, visual motion features, and analyze vital sign changes from the fused multimodal data to obtain multimodal feature vectors. Generative adversarial networks are used to perform sentiment analysis on the multimodal features to obtain the emotion recognition results of the animals.

5. The method according to claim 1, wherein, The step of performing semantic mapping and language translation on the emotion recognition results to convert animal language into human language and obtain the language conversion result includes: Emotional tags and voice features are extracted from the emotion recognition results, and the voice features are converted into standardized voice vectors; A pre-trained language model is used to semantically map the emotion tags to the sound vectors to obtain the emotion intent; A language generator is used to translate the emotional intent into language to generate the corresponding human language, thus obtaining the language conversion result.

6. The method according to any one of claims 1 to 5, wherein, The method further includes: If specific voice data is detected and there is no history of emotion matching, the specific voice data is labeled to obtain an updated emotion tag. The sample data is dynamically updated based on the updated sentiment tags so that the model parameters can be adjusted according to the updated sample data.

7. The method according to claim 1, wherein, The collection of multimodal data within a preset time window to obtain animal emotional change data includes: Collect multimodal data within a preset time window; The multimodal data is input into the emotional period recognition model, and the emotional change data of the animal is obtained through the emotional period recognition model.

8. The method according to claim 1, wherein, The step of updating sentiment tags based on the difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window includes: The Euclidean distance difference between the sentiment change characteristics of the current time window and the sentiment change characteristics of the previous time window is calculated to obtain the sentiment gap. If the emotional gap exceeds a preset emotional gap threshold, the emotional label will be upgraded to obtain an updated emotional label.

9. The method according to claim 8, wherein, The method further includes: The sentiment weight corresponding to each time window is scored to obtain the sentiment weight score; The emotional weight scores corresponding to similar emotional time windows within a continuous time period are accumulated. If the accumulated result is greater than the set upgrade threshold, the emotional label is updated.

10. An animal language conversion device, comprising: The acquisition module is used to acquire multimodal data related to animals, including animal sound data, animal behavior data, and animal physical characteristics data. The preprocessing module is used to preprocess the multimodal data to obtain fused multimodal data; An emotion recognition module is used to identify the animal's current emotion based on the fused multimodal data to obtain the animal's emotion recognition result; The conversion module is used to perform semantic mapping and language translation on the emotion recognition results to convert animal language into human language and obtain language conversion results; the conversion module is also used to collect multimodal data within a preset time window to obtain animal emotion change data; to extract features from the emotion change data to obtain emotion change features; and to update the emotion labels based on the difference between the emotion change features of the current time window and the emotion change features of the previous time window.

11. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-9.

13. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • AI-based pet emotion recognition system

    CN119049086A

  • KR20190125668A

Cited By

  • A pet bidirectional translation method and system based on end-cloud cooperation and archive enhancement

    CN122655793A

  • Animal language conversion method and apparatus, animal emotion recognition method and apparatus, and model training method and apparatus

    EP4783162A1

  • Animal language conversion method and apparatus, animal emotion recognition method and apparatus, and model training method and apparatus

    WO2026118524A1