Animal language conversion method and device, electronic equipment and storage medium
By acquiring and analyzing multimodal data of animals and using machine learning and deep learning technologies to identify and translate the emotional state of animals, the problem that the existing technology cannot deeply understand animal emotions is solved, and deep-level and real-time emotional communication between humans and animals is achieved.
Patent Information
- Application Number
- CN202411793938.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-06
AI Technical Summary
The existing technology is difficult to deeply understand the complex emotional levels of animals, and it is impossible to achieve deep and real-time emotional understanding and interactive communication between humans and animals.
By obtaining multimodal data such as animal sounds, behaviors and signs, preprocessing and fusion, using machine learning and deep learning technology to identify and analyze the animal's emotional state, and then semantic mapping and linguistic translation, converting the animal's emotional state into a language that can be understood by humans.
It has achieved comprehensive capture and accurate identification of animal emotions, enhanced communication skills between humans and animals, improved the accuracy of animal emotional understanding and real-time interaction, and provided humans with a new way to communicate with animals.
Smart Images

Figure CN119943059A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as machine learning, deep learning, and natural language processing, and especially to animal language conversion methods, devices, electronic devices, and storage media. Background Art
[0002] The current technologies on the market attempt to interpret the emotional world of animals through some basic animal sound and behavior translation devices and pet emotion analysis tools using artificial intelligence image recognition technology. However, these methods are often limited to the surface interpretation of animal behavior. These technical means cannot penetrate into the complex emotional level of animals and cannot achieve deep, real-time emotional understanding and interactive communication between humans and animals. Summary of the invention
[0003] The present invention provides an animal language conversion method, device, electronic device and storage medium.
[0004] According to one aspect of the present disclosure, there is provided an animal language conversion method, the method comprising:
[0005] Acquiring multimodal data related to an animal, wherein the multimodal data includes animal sound data, animal behavior data, and animal vital sign data;
[0006] Preprocessing the multimodal data to obtain fused multimodal data;
[0007] Identifying the current emotion of the animal according to the fused multimodal data to obtain an emotion recognition result of the animal;
[0008] The emotion recognition result is semantically mapped and language translated to convert the animal language into human language to obtain a language conversion result.
[0009] According to another aspect of the present disclosure, there is provided an animal language conversion device, comprising:
[0010] An acquisition module, used to acquire multimodal data related to an animal, wherein the multimodal data includes animal sound data, animal behavior data, and animal vital sign data;
[0011] A preprocessing module, used for preprocessing the multimodal data to obtain fused multimodal data;
[0012] An emotion recognition module, used to recognize the current emotion of the animal according to the fused multimodal data to obtain an emotion recognition result of the animal;
[0013] The conversion module is used to perform semantic mapping and language translation on the emotion recognition result to convert the animal language into human language to obtain a language conversion result.
[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any method in any of the above technical solutions.
[0018] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any one of the methods described in the above technical solutions.
[0019] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program implements any one of the methods described in the above technical solutions when executed by a processor.
[0020] The present disclosure provides an animal language conversion method, device, equipment and storage medium. The present disclosure can achieve comprehensive capture and accurate identification of animal emotions by acquiring multimodal data such as animal sounds, behaviors and physical signs. Then, through semantic mapping and language translation technology, the emotional state and intention of the animal are converted into a language that humans can understand, thereby greatly enhancing the communication ability between humans and animals, improving the accuracy of animal emotional understanding and the real-time nature of interaction, and providing humans with a new way to communicate with animals. That is, this solution can accurately identify the current emotional state of the animal and convert it into human language, thereby achieving a deeper level of emotional communication and understanding between animals and humans, and improving the accuracy and efficiency of cross-species communication.
[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0023] Figure 1 is a schematic diagram of the steps of the animal language conversion method in the embodiment of the present disclosure;
[0024] Figure 2 is a schematic diagram of a process for collecting multimodal data in an embodiment of the present disclosure;
[0025] Figure 3 is a schematic diagram of a process for preprocessing multimodal data in an embodiment of the present disclosure;
[0026] Figure 4 is a schematic diagram of a process for obtaining an animal's emotion recognition result in an embodiment of the present disclosure;
[0027] Figure 5 is a schematic diagram of a process of converting animal language into human language in an embodiment of the present disclosure;
[0028] Figure 6 is a schematic diagram of a process of updating emotion tags in an embodiment of the present disclosure;
[0029] Figure 7 Principle block diagram of the animal language conversion device in the embodiment of the present disclosure;
[0030] Figure 8 It is a block diagram of an electronic device used to implement the animal language conversion method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0031] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0032] Currently, the technologies for human-animal communication on the market are mainly divided into the following two types:
[0033] The first method is to use a simple translation device for animal sounds and behaviors, which mainly realizes the animal emotion translation function based on "voiceprint database + simple algorithm". Typical representatives are some pet behavior and sound recognition devices on the market, which can use relatively simple sensors to capture the sounds and some typical actions of animals, and then match them to a pre-built emotion library for simple mapping of emotions, such as recognizing that the barking of a dog represents a request for food, anger, etc.
[0034] The second method is to use AI (Artificial Intelligence) image recognition-based pet emotion analysis tools, which introduce image processing and AI technology to help users understand the emotional reactions of pets. For example, some companies capture animal faces and movements in camera image data, and then use existing deep learning training models to analyze different facial expressions of animals and identify specific expressions such as smiles, anger, grievances, and confusion.
[0035] There are multiple limitations in the above two methods in terms of animal emotion understanding and translation: first, emotion translation is single and relies on preset voiceprint data and behavior classification, which cannot continuously track emotion changes, resulting in insufficient accuracy in complex scenarios; second, lack of multimodal fusion analysis and over-reliance on a single information source limit the comprehensiveness and accuracy of emotion translation; third, insufficient timing detection and lack of ability to track the consistency of emotional states lead to slow perception of emotion changes; in addition, lack of adaptive learning and tuning mechanism makes it difficult to optimize and iterate when facing unknown emotion patterns, limiting the flexible scalability of the system; finally, the inability to perform edge computing leads to insufficient real-time performance, affecting the instant interaction between humans and animals. These limitations together lead to the inadequacy of existing technologies in achieving deep, real-time cross-species emotional communication.
[0036] In order to solve the above technical problems, the present disclosure provides an animal language conversion method, see Figure 1 As shown, Figure 1 : is a schematic diagram of the steps of the animal language conversion method in the embodiment of the present disclosure, which can be applied to the server side, and the method includes:
[0037] Step S101, acquiring multimodal data related to an animal, wherein the multimodal data includes animal sound data, animal behavior data, and animal vital sign data.
[0038] Specifically, obtaining multimodal data related to animals means collecting different types of information through multiple sensors to fully understand the state and emotions of animals. Among them, "animal sound data" refers to the various sounds made by animals captured by audio sensors, which can reflect the emotions and needs of animals; "animal behavior data" refers to the body movements and postures of animals recorded by video cameras to help analyze their behavior patterns and emotional expressions; and "animal vital signs data" refers to the physiological indicators of animals monitored by physiological sensors, such as heart rate, body temperature, etc. These indicators can provide physiological basis for the emotional state of animals.
[0039] In this way, by obtaining multimodal data of animals, including sound, behavior and physical sign data, it is possible to fully capture the emotions and behaviors of animals, thereby providing more accurate emotional analysis and behavioral understanding, enhancing communication and interaction between humans and animals, and improving animal welfare and humans' ability to respond to animal behavior.
[0040] Step S102, preprocessing the multimodal data to obtain fused multimodal data.
[0041] Specifically, preprocessing multimodal data refers to a series of processing steps on the raw data collected from different sources (such as audio, video, and physiological sensors), which may include noise reduction, normalization, feature extraction, etc., for easy analysis and understanding. "Fused multimodal data" refers to the integration of these preprocessed data into a unified data set. This process involves time alignment (ensuring that all data is synchronized in time), feature fusion (combining features of different modalities into a comprehensive feature vector) and data synchronization (handling timestamps and sampling rate differences of different data streams). Through this preprocessing and fusion, a comprehensive data view can be provided, making subsequent emotion recognition and language translation more accurate and efficient.
[0042] In this way, by preprocessing multimodal data such as animal sounds, behaviors, and physical signs, and then fusing these data into a unified data set, the consistency and availability of the data can be improved, which in turn can make emotion recognition and language translation more accurate, enhance the system's comprehensive understanding of animal behavior and emotional state, and improve the efficiency and effectiveness of cross-species communication.
[0043] Step S103, identifying the current emotion of the animal according to the fused multimodal data to obtain an emotion recognition result of the animal.
[0044] Specifically, after obtaining the fused multimodal data, the current emotion of the animal is identified based on the fused multimodal data, which specifically uses a comprehensive data set that integrates information such as animal sounds, behaviors, and physical signs, and uses machine learning and deep learning techniques to analyze and judge the emotional state of the animal. The "emotion recognition" here refers to identifying the emotional state of the animal, such as anxiety, excitement, or relaxation, by analyzing the features extracted from these multimodal data. To obtain the emotion recognition result of the animal, for example, the fused data can be input into a trained emotion recognition model, which outputs the emotion recognition result of the animal by comparing the features of the known emotional state, so as to provide a basis for subsequent language conversion and human-computer interaction.
[0045] In this way, by analyzing the fused multimodal data, including the animal's voice, behavior, and physiological signs, it is helpful to accurately identify the animal's current emotional state. This process not only improves the accuracy of understanding and responding to animal emotions, but also enhances communication between humans and animals, allowing humans to better interpret the needs and emotions of animals, thereby improving animal welfare and the intimacy of the human-pet relationship.
[0046] Step S104, semantic mapping and language translation are performed on the emotion recognition result to convert the animal language into human language to obtain a language conversion result.
[0047] Specifically, "semantic mapping and language translation of emotion recognition results" refers to the process of converting the emotional states of animals, such as anxiety, excitement or happiness, obtained by analyzing multimodal data, into language expressions that can be understood by humans. Specifically, this process involves using pre-trained language models and deep learning technology to establish a correspondence between the emotional characteristics of animals and the corresponding expressions in human language, namely "semantic mapping". Subsequently, these mapping results are converted into specific text or voice output to achieve "language translation". Ultimately, the "language conversion result" means that the emotions and intentions of animals are converted into the form of human language, so that human users can intuitively understand the "language" of animals, thereby achieving effective communication across species.
[0048] In this way, by semantically mapping and linguistically translating the results of emotion recognition, it is possible to convert the non-verbal communication of animals into a language that humans can understand. This process not only breaks the communication barriers between humans and animals, but also greatly enhances humans' understanding of animal emotions and needs, making the interaction between humans and animals more harmonious. It also provides new perspectives and tools for animal welfare and behavior research.
[0049] The present disclosure provides an animal language conversion method, device, equipment and storage medium. The present disclosure can achieve comprehensive capture and accurate identification of animal emotions by acquiring multimodal data such as animal sounds, behaviors and physical signs. Then, through semantic mapping and language translation technology, the emotional state and intention of the animal are converted into a language that humans can understand, thereby greatly enhancing the communication ability between humans and animals, improving the accuracy of animal emotional understanding and the real-time nature of interaction, and providing humans with a new way to communicate with animals. That is, this solution can accurately identify the current emotional state of the animal and convert it into human language, thereby achieving a deeper level of emotional communication and understanding between animals and humans, and improving the accuracy and efficiency of cross-species communication.
[0050] In some optional embodiments, obtaining multimodal data related to an animal includes:
[0051] Collect the sound wave information emitted by animals to obtain animal sound data;
[0052] Collect animal body language and movement changes to obtain animal behavior data;
[0053] Collect the physical biological indicators of animals and obtain animal vital signs data.
[0054] Specifically, the sounds of animals are captured by audio collectors to collect the sound wave information emitted by animals to obtain "animal sound data"; at the same time, visual sensors such as video cameras are used to collect the body language and movement changes of animals to form "animal behavior data", which reflects the activities and non-verbal behaviors of animals; in addition, physiological sensors are used to monitor the heartbeat, body temperature and other "animal vital signs data" of animals. These physical biological indicators provide important information for understanding the physiological state and emotions of animals. By combining these multimodal data, it is possible to fully capture the communication methods and emotional states of animals, laying the foundation for further data analysis and emotion recognition.
[0055] To facilitate understanding of the solutions of the embodiments of the present disclosure, the following examples are given. Figure 2 As shown, Figure 2 : is a flow chart of collecting multimodal data in the embodiment of the present disclosure. First, an audio collector is used to capture the sounds made by animals in real time, and then the audio collector sends the captured sound data to a data processing module, through which a corresponding audio filter can be used to process noise and reduce background interference. For example, for a dog barking, its pitch, amplitude, duration, frequency, breakpoint change and other dimensional information are sampled together to capture all the information of the barking.
[0056] For animal behavior data collection, the animal's body movements and behavior performance are obtained through video equipment (such as high-definition cameras and infrared cameras), such as tail shaking, jumping movements, lying posture analysis, and supplemented by specific animal body characteristics (ears erect, pupil dilation, etc.). Then, the captured video / image data is sent to the data processing module.
[0057] For animal vital sign data collection, the heart rate and body temperature are collected through high-precision contact or non-contact body temperature detection sensors, and the collected vital sign data are sent to the data processing module.
[0058] In this way, by collecting the sound wave information, body language and movement changes, and physical biological indicators emitted by animals, we can collect more comprehensive animal sound data, behavioral data, and vital sign data. The integration of these multi-dimensional information provides a rich and accurate data basis for in-depth understanding and analysis of animal emotions and behaviors, enabling humans to more accurately interpret animals' communication intentions and physiological states, strengthening non-verbal communication between humans and animals, and improving the effectiveness of animal care and training. It also opens up new avenues for animal health monitoring and behavioral research.
[0059] In some optional embodiments, preprocessing the multimodal data to obtain fused multimodal data includes:
[0060] De-noising the multimodal data to clean the data, and obtain cleaned multimodal data;
[0061] Normalizing the cleaned multimodal data to obtain normalized multimodal data;
[0062] The normalized multimodal data are time-series aligned and fused to obtain fused multimodal data.
[0063] Specifically, denoising multimodal data for data cleaning refers to using signal processing technology to remove noise and interference in audio and video data to improve the quality of the data and obtain cleaned multimodal data. Next, "normalizing the cleaned multimodal data" means converting data from different sources and scales into a unified format or scale so that the machine learning model can process it more efficiently. This step results in normalized multimodal data. Finally, "temporal alignment and fusion of normalized multimodal data" involves aligning data of different modalities in time and merging them into a unified data set. This step ensures the consistency of the data in time and integrates information from different sensors. The final result is fused multimodal data, which provides an accurate and comprehensive data foundation for subsequent emotion recognition and behavior analysis.
[0064] To facilitate understanding of the solutions of the embodiments of the present disclosure, the following examples are given. Figure 3 As shown, Figure 3It is a flowchart of preprocessing multimodal data in an embodiment of the present disclosure. First, the sound data, image data, body temperature, and heartbeat data in the multimodal data are preprocessed and normalized, and the processing includes denoising and cleaning data, which is used to process the invalid parts in the audio and visual information. For example, the noise that may be generated by humans (wind, speech, etc.) is filtered out to make the audio signal clearer and easier to identify. In terms of images, the background motion and objects in the video frame are cleaned up to retain only key information, such as behavioral changes of animals. Next, data normalization processing is performed. Whether it is audio signals, videos, or vital signs information, they all need to be normalized, converted into a unified standard, and expressed as a feature vector that can be processed by the machine learning algorithm.
[0065] Finally, the multimodal data is time-aligned and fused. Data alignment is the premise of data fusion, which needs to solve the time and space differences between the acquisition of different modal signals. For example, a dog barking event often has a corresponding time difference with the moment of accompanying limb movement or body temperature fluctuation. These data must be time-calibrated before input so that the multimodal input in the translation task has a consistent reference point.
[0066] In this way, by denoising, normalizing, aligning and fusing multimodal data, the quality, consistency and availability of the data can be significantly improved, making the features extracted from animal sounds, behaviors and signs more accurate and reliable. This process not only enhances the accuracy of data analysis, but also improves the accuracy of the model in identifying animal emotions and behaviors, providing a solid data foundation for efficient and accurate cross-species communication.
[0067] In some optional embodiments, the current emotion of the animal is identified according to the fused multimodal data to obtain the emotion recognition result of the animal, including:
[0068] A deep learning model is used to extract sound features, visual motion features, and physical sign changes from the fused multimodal data to obtain a multimodal feature vector.
[0069] Generative adversarial networks are used to perform sentiment analysis on multimodal features to obtain animal emotion recognition results.
[0070] Specifically, "using deep learning models to extract sound features, visual motion features, and physical sign changes from the fused multimodal data" refers to the use of advanced deep learning techniques, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), to process and analyze multimodal datasets that integrate sound, visual, and physiological data. This process involves extracting sound features such as pitch, rhythm, and volume from audio signals; extracting visual motion features such as posture and behavior patterns from video data; and analyzing physical sign changes such as heart rate and body temperature fluctuations from physiological data. Together, these features form a multimodal feature vector that comprehensively reflects the physiological and behavioral states of the animal. Next, "using generative adversarial networks to perform sentiment analysis on multimodal features" refers to the use of a deep learning framework, generative adversarial networks (GANs), to enable the network to recognize and distinguish different emotional states through an adversarial training process. This process ultimately produces animal emotion recognition results, converting the animal's complex emotions and behaviors into understandable emotional labels, providing a basis for further semantic mapping and language translation.
[0071] To facilitate understanding of the solutions of the embodiments of the present disclosure, the following examples are given. Figure 4 As shown, Figure 4 It is a flowchart of obtaining the emotion recognition results of animals in the embodiment of the present disclosure. After obtaining the fused multimodal data, fine-grained feature extraction is performed on the data of each modality through a deep learning model. Specifically, the sound feature extraction, visual motion feature and physical sign change analysis are extracted through the deep learning model to obtain a feature vector combination. Next, a large model based on the emotion recognition generative adversarial network (GAN) is used to analyze the voiceprint features, motion changes and physical sign fluctuations in the data to obtain an emotion classification label. For example, when an animal's low-frequency barking accompanied by tense limbs and dilated pupils is detected, the model recognizes that the animal is in a state of high alert through comparative reasoning, and further features are combined to infer the psychological activities (fear or confusion) that may be hidden behind it.
[0072] In this way, by using a deep learning model to conduct a comprehensive analysis of the fused multimodal data, it is possible to accurately extract features such as the animal's voice, visual movements, and physical changes, forming a multimodal feature vector, and then use the generative adversarial network to conduct in-depth sentiment analysis on these features to obtain the animal's emotion recognition results. This process not only improves the accuracy and depth of emotion recognition, but also enhances the understanding of animal behavior and psychological state, providing strong technical support for more effective communication between humans and animals.
[0073] In some optional embodiments, the emotion recognition results are semantically mapped and language translated to convert the animal language into human language, and the language conversion results are obtained, including:
[0074] Extract emotion labels and sound features from emotion recognition results, and convert the sound features into standardized sound vectors;
[0075] A pre-trained language model is used to semantically map the emotion label with the sound vector to obtain the emotion intention;
[0076] A language generator is used to translate the emotional intention into language to generate the corresponding human language and obtain the language conversion result.
[0077] Specifically, a "pre-trained language model" refers to an artificial intelligence model that has been trained with a large amount of data and is able to understand and process natural language. The model is used here to "semantically map" the animal's emotional label, that is, the animal's emotional state obtained from the emotion recognition module, with the sound vector, that is, to correspond the animal's sound characteristics to the emotional semantics in human language, thereby identifying the animal's emotional intentions. Next, the "language generator" is the process of converting the animal's non-verbal communication into a language that humans can understand based on these emotional intentions. This process involves converting the animal's emotions and intentions into specific text or voice outputs, namely the "language conversion results", so that humans can intuitively understand the animal's "language" and achieve effective communication between humans and animals.
[0078] To facilitate understanding of the solutions of the embodiments of the present disclosure, the following examples are given. Figure 5 As shown, Figure 5 It is a flowchart of converting animal language into human language in the embodiment of the present disclosure. After the emotion recognition module obtains the emotion recognition result, it provides the emotion label and sound characteristics of the animal, and then passes this information to the voice mapping module, which is responsible for converting the sound characteristics of the animal into human-understandable voice expressions. Next, the human semantic output module receives the converted voice information and outputs it as a language conversion result, and finally presents these results to the user, realizing real-time conversion of animal language to human language and emotional communication. The whole process involves the recognition of emotions, the extraction and mapping of features, and the generation of human language, aiming to promote effective communication between humans and animals.
[0079] In this way, by using a pre-trained language model to perform semantic mapping between emotion tags and sound vectors, the emotional intentions of animals can be accurately captured and understood, and then these emotional intentions can be converted into human language using a language generator. This process not only achieves accurate interpretation of animal emotions, but also enables the non-verbal communication of animals to be converted into language understandable to humans, greatly promoting communication and understanding between humans and animals, and improving the transparency of animal emotional expression and the efficiency of communication.
[0080] In some optional embodiments, the method further comprises:
[0081] If specific sound data is detected and there is no emotion matching history record, the specific sound data is labeled to obtain an updated emotion label;
[0082] The sample data is dynamically updated according to the updated sentiment labels, so that the model parameters are adjusted according to the updated sample data.
[0083] Specifically, when specific animal sound data is detected, and this sound data does not exist in the historical emotion matching record, that is, there is no previous emotion label corresponding to it, a labeling process will be triggered. The "labeling" here refers to artificially assigning an emotion label to these specific sound data. This label describes the emotional state of the animal when making the sound, thereby "obtaining an updated emotion label." Subsequently, this newly labeled emotion label is included in the sample database to achieve "dynamic update of sample data according to the updated emotion label." The updated sample data is used to adjust the parameters of the model, that is, "adjusting the model parameters according to the updated sample data", so as to optimize and enhance the model's emotion recognition ability for newly emerging sound data, ensure that the system can adapt to new or uncommon animal sounds, and improve the recognition accuracy and the system's adaptive learning ability.
[0084] To facilitate understanding of the solution of the disclosed embodiment, an example is given below: when the system faces a call pattern or abnormal performance that cannot be identified, it will remind the user to enter relevant labels or identification information. For example, the system detects a specific action + sound combination, and there is no obvious emotional expression matching historical record. At this time, the user can annotate this phonetic symbol through the interface, such as: "calling for help" or "hungry" state. After the user manually labels, the system will adjust the data parameters of the current audio and the corresponding label of the behavior, thereby updating the weight of the existing model. After the user completes the labeling, the system will update the data set in a timely manner, continuously optimize and retrain the relevant generation model, so that the system can gradually improve the recognition rate of similar or new categories of emotional events.
[0085] In this way, by manually annotating when specific sound data is detected and there is a lack of emotion matching history, and updating the emotion labels accordingly, the sample database can be continuously expanded and enriched, and the model parameters can be dynamically updated to enhance the system's emotion recognition ability for new sound data, improve the model's adaptability and accuracy, and ensure that the cross-species communication system can continue to evolve and better understand and respond to animals' communication intentions.
[0086] In some optional embodiments, the method further comprises:
[0087] Collect multimodal data within a preset time window to obtain the animal's emotional change data;
[0088] Extract features from the emotion change data to obtain emotion change features;
[0089] The emotion label is updated according to the gap between the emotion change characteristics of the current time window and the emotion change characteristics of the previous time window.
[0090] Specifically, the "preset time window" refers to a specific time period set for analyzing the emotional changes of animals. During this time period, "multimodal data" is collected, including information such as the animal's voice, behavior, and physical signs, to collect data on the animal's emotional state. Then, through the "feature extraction" process, key information that can represent the animal's emotional changes, namely "emotional change features," is identified from these data. Then, the difference between the emotional features extracted in the current time window and the features of the previous time window is compared. If these differences indicate that the animal's emotional state has changed significantly, the "emotional label" will be "updated" based on this change to reflect the animal's latest emotional state. This process involves continuous monitoring and dynamic analysis of the animal's emotional state to ensure the real-time and accuracy of emotion recognition.
[0091] In this way, by collecting multimodal data of animals within a preset time window, the emotional changes of animals can be fully captured, and feature extraction of these data can be performed to identify the characteristics of emotional changes. Secondly, by dynamically updating the emotional label according to the difference between the emotional features of the current and previous time windows, real-time and accurate monitoring and response to the emotional state of animals can be achieved. This not only improves the accuracy of emotion recognition, but also enhances the real-time and depth of communication between humans and animals, allowing humans to better understand and respond to the emotional needs of animals.
[0092] In some optional embodiments, collecting multimodal data within a preset time window to obtain the emotional change data of the animal includes:
[0093] Collect multimodal data within a preset time window;
[0094] The multimodal data is input into the emotion period recognition model, and the emotion change data of the animal is obtained through the emotion period recognition model.
[0095] Specifically, the "preset time window" refers to a specific time period set for monitoring and analyzing the emotional changes of animals. During this period, "multimodal data" is collected, which includes different types of information such as the animal's voice, behavioral movements, and physical signs. These data are then input into the "emotional period recognition model", which is an algorithm specially designed to process and analyze time series data, such as long short-term memory networks (LSTM) or gated recurrent units (GRU). The model analyzes these multimodal data, identifies and extracts features related to the animal's emotional state, and thus "obtains the animal's emotional change data." These data reflect the animal's emotional state within a continuous time window, providing a basis for further emotion recognition and label updates. In this way, the solution can achieve dynamic tracking and accurate identification of the animal's emotional state.
[0096] In this way, by collecting multimodal data of animals within a preset time window and inputting these data into the emotion period recognition model, it is possible to accurately capture and analyze the emotional changes of animals. This method can identify and understand the emotional dynamics of animals in continuous time periods, thereby providing richer and more accurate emotional change data. These data not only enhance the depth and real-time nature of emotion recognition, but also improve the quality and efficiency of communication between humans and animals, allowing humans to respond to the emotional needs and behavioral changes of animals in a more detailed and timely manner.
[0097] In some optional embodiments, the emotion label is updated according to the difference between the emotion change feature of the current time window and the emotion change feature of the previous time window, including:
[0098] Calculate the Euclidean distance difference between the emotion change characteristics of the current time window and the emotion change characteristics of the previous time window to obtain the emotion gap;
[0099] If the emotion gap exceeds a preset emotion gap threshold, the emotion label is upgraded to obtain an updated emotion label.
[0100] Specifically, the "emotional change features of the current time window" refer to the emotion-related features extracted from the animal's multimodal data in a specific time period, while the "emotional change features of the previous time window" refer to the corresponding features in the previous time period. By calculating the "Euclidean distance difference" between the emotion features of these two time windows, a method of measuring the distance between two points in multidimensional space, a quantified "emotional gap" can be obtained.
[0101] The Euclidean distance difference between the emotion change feature of the current time window and the emotion change feature of the previous time window satisfies the formula: [D(W_i,W_{i-1})=\sqrt{\sum_{n=1}^{N}(x_{i,n}-x_{i-1,n})^2}], where D(W_i,W_{i-1}): represents the Euclidean distance between the emotion change features of the i-th time window and the i-1-th time window. This distance is used to quantify the degree of change of emotion features in two consecutive time windows. x_{i,n} represents the value of the n-th feature in the i-th time window. These features may include the pitch, rhythm, amplitude of the sound, the frequency of the behavior, the heart rate change of physical signs, etc. x_{i-1,n} represents the value of the n-th feature in the i-1-th time window. sum represents the summation operation, and sqrt represents the square value operation.
[0102] By calculating the Euclidean distance difference between the emotion change characteristics of the current time window and the emotion change characteristics of the previous time window, the emotion gap is obtained, which reflects the degree of change of the animal's emotional state in two consecutive time windows. If this gap exceeds the "preset emotion gap threshold", that is, a pre-set limit, indicating that the animal's emotional state has changed significantly, it will trigger the "emotion label upgrade". According to the emotion gap and change pattern, the emotion label is updated from one state (such as "alert") to another label that is more in line with the current emotional state (such as "anxiety"), thereby obtaining the "updated emotion label". This process makes emotion recognition more dynamic and accurate, and can respond to significant changes in animal emotions in real time.
[0103] In this way, by calculating the Euclidean distance difference between the emotion change characteristics of the current and previous time windows, the degree of change in the animal's emotional state can be quantified to obtain the emotion gap. When the emotion gap exceeds the preset threshold, the emotion label will be automatically upgraded to obtain an updated emotion label. This method makes emotion recognition more sensitive and accurate, and can capture and respond to significant changes in animal emotions in real time, thereby improving the quality and efficiency of communication between humans and animals, and ensuring that humans can understand and respond to the emotional needs of animals in a timely and appropriate manner.
[0104] In some optional embodiments, the method further comprises:
[0105] Score the sentiment weight corresponding to each time window to obtain the sentiment weight score;
[0106] The sentiment weight scores corresponding to similar sentiment time windows in continuous time periods are accumulated. If the accumulated result is greater than the set upgrade threshold, the sentiment label is updated.
[0107] Specifically, "emotional weight scoring" refers to the process of quantitatively evaluating the emotional state of an animal within each time window, where each time window contains multimodal data collected within a specific time period. By analyzing this data, emotional features are extracted, and different weights are assigned according to the strength or significance of the features, and then the emotional weight score for each time window is calculated. Next, the weight scores of time windows with similar emotional features in consecutive time periods are accumulated. If this accumulated score exceeds the "set upgrade threshold", a predefined value used to determine whether the emotional state is significant enough to trigger a label update, the emotional label is updated to reflect the change in the animal's emotional state. This process makes emotion recognition more dynamic and accurate, and can respond to continuous changes in animal emotions in real time.
[0108] In this way, by assigning weights and scoring the emotional state in each time window, this solution can quantify and track the emotional changes of animals in continuous time periods. The weighted scores of time windows with similar emotional characteristics are accumulated. When the accumulated result exceeds the preset upgrade threshold, the emotional label will be automatically updated to more accurately reflect the current emotional state of the animal. This method improves the sensitivity and adaptability of emotion recognition, making communication between humans and animals more accurate and timely, thereby enhancing the ability to understand and respond to the emotional needs of animals.
[0109] To facilitate understanding of the solutions of the embodiments of the present disclosure, see Figure 6 As shown, Figure 6 : is a flowchart of updating emotion labels in the embodiment of the present disclosure. The barking audio and body temperature data of the animal are collected through sound collection and body temperature sensors, and then these data are sent to the data processing and fusion module for analysis to identify audio features and body temperature changes. The analysis results are used for emotional feedback, and the animal's emotion label is dynamically adjusted according to the emotional meaning of the audio and body temperature changes, such as upgrading the "alert" state to "anxiety". This process involves real-time monitoring and continuous emotional state assessment to ensure that the animal's emotional changes can be reflected in the label update in a timely manner, thereby improving the accuracy of understanding and responding to the animal's emotional state.
[0110] In this way, by collecting the emotional change data of animals in a continuous time window and extracting the emotional change features in these data, it is possible to monitor and quantify the subtle changes in animal emotions in real time. By comparing the differences between these change features and standard emotional features and updating the emotional label when the change exceeds the preset threshold, it is possible to more accurately capture the dynamic changes in animal emotions, thereby achieving continuous tracking and real-time response to the emotional state of animals, improving the sensitivity and accuracy of emotion recognition, and enhancing the effectiveness of emotional communication between humans and animals.
[0111] The following describes an embodiment of the device of the present application, which can be used to execute the animal language conversion method in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the animal language conversion method in the above embodiment of the present application.
[0112] The present disclosure also provides an animal language conversion device 700, such as Figure 7 As shown, including:
[0113] An acquisition module 701 is used to acquire multimodal data related to animals, where the multimodal data includes animal sound data, animal behavior data, and animal vital sign data;
[0114] A preprocessing module 702 is used to preprocess the multimodal data to obtain fused multimodal data;
[0115] The emotion recognition module 703 is used to recognize the current emotion of the animal according to the fused multimodal data to obtain the emotion recognition result of the animal;
[0116] The conversion module 704 is used to perform semantic mapping and language translation on the emotion recognition result to convert the animal language into human language to obtain a language conversion result.
[0117] In some optional embodiments, the acquisition module 701 acquires multimodal data related to the animal, including:
[0118] Collect the sound wave information emitted by animals to obtain animal sound data;
[0119] Collect animal body language and movement changes to obtain animal behavior data;
[0120] Collect the physical biological indicators of animals and obtain animal vital signs data.
[0121] In some optional embodiments, the preprocessing module 702 preprocesses the multimodal data to obtain fused multimodal data, including:
[0122] De-noising the multimodal data to clean the data, and obtain cleaned multimodal data;
[0123] Normalizing the cleaned multimodal data to obtain normalized multimodal data;
[0124] The normalized multimodal data are time-series aligned and fused to obtain fused multimodal data.
[0125] In some optional embodiments, the emotion recognition module 703 recognizes the current emotion of the animal according to the fused multimodal data to obtain the emotion recognition result of the animal, including:
[0126] A deep learning model is used to extract sound features, visual motion features, and physical sign changes from the fused multimodal data to obtain a multimodal feature vector.
[0127] Generative adversarial networks are used to perform sentiment analysis on multimodal features to obtain animal emotion recognition results.
[0128] In some optional embodiments, the conversion module 704 performs semantic mapping and language translation on the emotion recognition result to convert the animal language into human language, and obtains a language conversion result, including:
[0129] Extract emotion labels and sound features from emotion recognition results, and convert the sound features into standardized sound vectors;
[0130] A pre-trained language model is used to semantically map the emotion label with the sound vector to obtain the emotion intention;
[0131] A language generator is used to translate the emotional intention into language to generate the corresponding human language and obtain the language conversion result.
[0132] In some optional embodiments, the apparatus further comprises a first updating module, which is used to label the specific sound data to obtain an updated emotion label if the specific sound data is detected and there is no emotion matching history record;
[0133] The sample data is dynamically updated according to the updated sentiment labels, so that the model parameters are adjusted according to the updated sample data.
[0134] In some optional embodiments, the device further comprises a second updating module for collecting multimodal data within a preset time window to obtain emotional change data of the animal;
[0135] Extract features from the emotion change data to obtain emotion change features;
[0136] The emotion label is updated according to the gap between the emotion change characteristics of the current time window and the emotion change characteristics of the previous time window.
[0137] In some optional embodiments, the second updating module collects multimodal data within a preset time window to obtain the animal's emotion change data, including:
[0138] Collect multimodal data within a preset time window;
[0139] The multimodal data is input into the emotion period recognition model, and the emotion change data of the animal is obtained through the emotion period recognition model.
[0140] In some optional embodiments, the second updating module updates the emotion label according to the difference between the emotion change feature of the current time window and the emotion change feature of the previous time window, including:
[0141] Calculate the Euclidean distance difference between the emotion change characteristics of the current time window and the emotion change characteristics of the previous time window to obtain the emotion gap;
[0142] If the emotion gap exceeds a preset emotion gap threshold, the emotion label is upgraded to obtain an updated emotion label.
[0143] In some optional embodiments, the second update module is used to:
[0144] Score the sentiment weight corresponding to each time window to obtain the sentiment weight score;
[0145] The sentiment weight scores corresponding to similar sentiment time windows in continuous time periods are accumulated. If the accumulated result is greater than the set upgrade threshold, the sentiment label is updated.
[0146] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0147] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0148] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0149] like Figure 8As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0150] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0151] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the animal language conversion method. For example, in some embodiments, the animal language conversion method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the applet distribution described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the animal language conversion method in any other appropriate manner (e.g., by means of firmware).
[0152] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0153] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0154] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0155] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0156] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0157] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0158] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0159] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for converting animal language, the method comprising: Acquiring multimodal data related to an animal, wherein the multimodal data includes animal sound data, animal behavior data, and animal vital sign data; Preprocessing the multimodal data to obtain fused multimodal data; Identifying the current emotion of the animal according to the fused multimodal data to obtain an emotion recognition result of the animal; The emotion recognition result is semantically mapped and language translated to convert the animal language into human language to obtain a language conversion result.
2. The method according to claim 1, wherein: The obtaining of multimodal data related to the animal comprises: Collecting sound wave information emitted by the animal to obtain the animal sound data; Collecting the body language and movement changes of the animal to obtain the animal behavior data; The physical biological indicators of the animal are collected to obtain the animal's vital sign data.
3. The method according to claim 1, wherein: The preprocessing of the multimodal data to obtain fused multimodal data includes: Performing denoising processing on the multimodal data to perform data cleaning to obtain cleaned multimodal data; Normalizing the cleaned multimodal data to obtain normalized multimodal data; The normalized multimodal data is subjected to time series alignment and fusion processing to obtain fused multimodal data.
4. The method according to claim 1, wherein: The identifying the current emotion of the animal according to the fused multimodal data to obtain the emotion recognition result of the animal includes: Using a deep learning model to extract sound features, visual motion features, and physical sign changes from the fused multimodal data to obtain a multimodal feature vector; A generative adversarial network is used to perform sentiment analysis on the multimodal features to obtain the animal's emotion recognition results.
5. The method according to claim 1, wherein: The emotion recognition result is semantically mapped and language translated to convert the animal language into human language to obtain a language conversion result, including: Extracting emotion labels and sound features from the emotion recognition results, and converting the sound features into standardized sound vectors; Using a pre-trained language model to semantically map the emotion label with the sound vector to obtain the emotion intention; The emotional intention is translated into language using a language generator to generate a corresponding human language and obtain a language conversion result.
6. The method according to any one of claims 1 to 5, wherein: The method further comprises: If specific sound data is detected and there is no emotion matching history record, the specific sound data is labeled to obtain an updated emotion label; The sample data is dynamically updated according to the updated emotion label, so that the model parameters are adjusted according to the updated sample data.
7. The method according to any one of claims 1 to 5, wherein: The method further comprises: Collect multimodal data within a preset time window to obtain the animal's emotional change data; Extracting features from the emotion change data to obtain emotion change features; The emotion label is updated according to the gap between the emotion change characteristics of the current time window and the emotion change characteristics of the previous time window.
8. The method according to claim 7, wherein: The collecting of multimodal data within a preset time window to obtain the emotional change data of the animal includes: Collect multimodal data within a preset time window; The multimodal data is input into an emotion period recognition model, and the emotion change data of the animal is obtained through the emotion period recognition model.
9. The method according to claim 7, wherein: The updating of the emotion label according to the gap between the emotion change feature of the current time window and the emotion change feature of the previous time window includes: Calculate the Euclidean distance difference between the emotion change feature of the current time window and the emotion change feature of the previous time window to obtain the emotion gap; If the emotion gap exceeds a preset emotion gap threshold, the emotion label is upgraded to obtain an updated emotion label.
10. The method according to claim 9, wherein: The method further comprises: Score the sentiment weight corresponding to each time window to obtain the sentiment weight score; The sentiment weight scores corresponding to similar sentiment time windows in continuous time periods are accumulated. If the accumulated result is greater than the set upgrade threshold, the sentiment label is updated.
11. An animal language conversion device, comprising: An acquisition module, used to acquire multimodal data related to an animal, wherein the multimodal data includes animal sound data, animal behavior data, and animal vital sign data; A preprocessing module, used for preprocessing the multimodal data to obtain fused multimodal data; An emotion recognition module, used to recognize the current emotion of the animal according to the fused multimodal data to obtain an emotion recognition result of the animal; The conversion module is used to perform semantic mapping and language translation on the emotion recognition result to convert the animal language into human language to obtain a language conversion result.
12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.
14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
System and method for recognizing voice emotion of animal
CN104700829A
Multi-modal emotion recognition method and device, electronic equipment and storage medium
CN113111855A
User emotion analysis model training method and device, electronic equipment and storage medium
CN114925159A
Multifunctional pet monitoring necklace and interactive management system thereof
CN115136903A
Intelligent sensing necklace for pets and control method of intelligent sensing necklace
CN116711654A
Cited By
Cross-species animal behavior identification method and system based on motion sensor
CN120217166A
Animal language conversion method and apparatus, animal emotion recognition method and apparatus, and model training method and apparatus
WO2026118524A1