A method and system for human-pet intercom based on multimodal interaction

By employing a multimodal interactive human-pet communication method and integrating real-time pet response data using a multimodal translation model, the problem of insufficient translation accuracy in existing technologies is solved, achieving efficient and accurate human-pet interaction and enhancing the adaptability and commercial value of the device.

CN120636418BActive Publication Date: 2026-01-06SHENZHEN ZHONGHE DAHUI ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510786448.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2026-01-06
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

Existing pet interaction devices have fragmented functions, requiring the use of multiple devices in combination. They lack adaptive learning mechanisms, leading to operational redundancy and resource waste. Furthermore, their translation accuracy and adaptability are insufficient, making it difficult to achieve natural communication. Existing technologies that cannot achieve remote two-way voice interaction suffer from low confidence translation error rates.

Method used

A multimodal interactive human-pet intercom method is adopted, which receives signals from the user device through the pet device to obtain the pet's real-time response data. The multimodal translation model is used to fuse and translate sound, motion images and environmental data, including the modular design of the input layer, encoder and decoder, to achieve cross-modal semantic fusion and dynamic translation.

Benefits of technology

It significantly improves the reliability and accuracy of translation results, enhances the adaptability and computational efficiency of human-pet interaction, reduces information loss, and has commercial potential.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636418B_ABST
    Figure CN120636418B_ABST
Patent Text Reader

Abstract

This application discloses a method and system for human-pet dialogue based on multimodal interaction, relating to the field of artificial intelligence. The method includes: receiving dialogue signals sent by a user device to a target pet; determining the animal's speech information corresponding to the dialogue signals based on a local language library, which includes the correspondence between each dialogue signal and animal speech information; acquiring real-time response data of the target pet to the dialogue signals, including the target pet's sound data, motion image data, and environmental data; inputting the real-time response data into a multimodal translation model to obtain a translation result of the target pet's response to the dialogue signals, and sending the translation result to the user device. In this embodiment, the translation result integrates sound data, motion data, and environmental data, achieving efficient data fusion, improving translation reliability, and solving the problem of high misjudgment rate with low confidence, thereby improving the accuracy and adaptability of human-pet interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a method and system for human-pet intercom based on multimodal interaction. Background Technology

[0002] In modern society, pets have become important emotional companions, and owners yearn to overcome the language barrier between species and more accurately understand their pets' emotions and behavioral intentions. However, research has revealed significant limitations in current pet interaction devices on the market. First, helmet-style translators are bulky and uncomfortable to wear, easily causing pets to resist them. Second, while mobile translation applications are portable, they require close-range operation, making it difficult to achieve remote two-way voice interaction. Third, location-based devices only integrate basic location tracking functions and lack real-time communication capabilities. Fourth, traditional communication devices are bulky, complex to deploy, require manual training to trigger preset commands, and have a narrow language database, making it difficult to meet personalized needs.

[0003] It is evident that existing pet interaction devices generally exhibit fragmented functionality, requiring users to use multiple independent devices to cover different scenarios, leading to operational redundancy and resource waste. Furthermore, existing solutions lack adaptive learning mechanisms, failing to dynamically optimize interaction logic through AI technology, severely limiting the efficiency of natural communication between humans and pets and the depth of emotional connection. Summary of the Invention

[0004] The purpose of this application is to address at least one of the aforementioned technical deficiencies.

[0005] On one hand, embodiments of this application provide a human-pet intercom method based on multimodal interaction, which is executed by a pet device and includes:

[0006] It receives dialogue signals sent by the user device to the target pet and determines the animal voice information corresponding to the dialogue signal based on the local language library, which includes the correspondence between each dialogue signal and the animal voice information.

[0007] Acquire real-time response data of the target pet to dialogue signals. The real-time response data includes the target pet's sound data, motion image data, and environmental data.

[0008] The real-time data is input into the multimodal translation model to obtain the translation results of the target pet's response to the dialogue signals, and the translation results are sent to the user's device.

[0009] Optionally, the multimodal translation model includes an input layer, an encoder, and a decoder. Real-time data is input into the multimodal translation model to obtain the translated results of the target pet's response to dialogue signals, including:

[0010] Based on the input layer, feature extraction processing is performed on the sound data, motion image data, and environmental data in the real-time data to obtain the sound features corresponding to the sound data, the image features corresponding to the motion image data, and the environmental features corresponding to the environmental data.

[0011] Based on the encoder, cross-modal semantic fusion processing is performed on the sound features and the image features to obtain the fused features;

[0012] The fused features and the environmental features are input into the decoder, so that the decoder performs dynamic translation generation processing based on the fused features and the environmental features to obtain the translation result of the target pet's response to the dialogue signal.

[0013] Optionally, the encoder includes a dynamic cross-modal attention layer, a sparse local attention layer, and adaptive noise suppression. The cross-modal semantic fusion processing of the sound features and image features based on the encoder to obtain the fused features includes:

[0014] Based on the sparse local attention layer, each local window corresponding to the sound feature is determined, and the features of each local window are integrated through skip connections to obtain the optimized sound feature.

[0015] Based on the adaptive noise suppression layer, the optimized sound features and the image features are denoised to obtain denoised sound features and denoised image features.

[0016] The dynamic cross-modal attention layer performs self-attention fusion calculation on the denoised sound features and the denoised image features to obtain the fused features.

[0017] Optionally, the step of performing noise reduction processing on the optimized sound features and image features based on the adaptive noise suppression layer to obtain denoised sound features and denoised image features includes:

[0018] Based on the adaptive noise suppression layer, the background noise regions included in the optimized sound features and the image features are determined respectively;

[0019] The features corresponding to the background noise region are masked in the optimized sound features and image features to obtain the denoised sound features and denoised image features.

[0020] Optionally, the step of performing self-attention fusion calculation on the denoised sound features and the denoised image features through the dynamic cross-modal attention layer to obtain the fused features includes:

[0021] Based on the sound features and the image features, the attention weights of the sound data and the attention weights of the motion image data are determined respectively through the dynamic cross-modal attention layer;

[0022] Based on the attention weights of the sound data and the motion image data, self-attention fusion calculation is performed on the denoised sound features and the denoised image features to obtain the fused features.

[0023] Optionally, the decoder includes a dynamic weight allocation layer and a decoding layer, and the decoder obtains the translation result of the target pet's response to the dialogue signal in the following manner:

[0024] The task label is obtained, and the dynamic weight allocation layer adjusts the attention weight based on the task label to obtain the adjusted attention weight. The adjusted attention weight includes the attention weight of the fused features and the attention weight corresponding to the environmental data.

[0025] Based on the adjusted attention weights, task optimization processing is performed on the features corresponding to the environmental data and the fused features to obtain a task-optimized feature vector.

[0026] The decoding layer maps the task-optimized feature vectors to the vocabulary space to obtain the translation result.

[0027] Optionally, the method further includes:

[0028] Receive the location viewing signal sent by the user equipment terminal for the target pet;

[0029] The system acquires the real-time accurate location information and historical movement trajectory of the target pet, and sends the real-time accurate location information and the historical movement trajectory to the user device.

[0030] Optionally, the method further includes:

[0031] Receive the environmental viewing signal sent by the user equipment terminal for the target pet;

[0032] The included cameras are automatically turned on, and image data of the environmental scene where the target pet is located is acquired through the cameras;

[0033] The image data is sent to the user equipment terminal, and the image data includes at least one of video data and image data.

[0034] Optionally, after sending the image data to the user equipment, the process further includes...

[0035] Receive the alarm signal sent by the user equipment terminal for the target pet, and initiate the alarm processing corresponding to the alarm signal;

[0036] The alarm processing includes at least one of playing a preset human distress voice, turning on the alarm indicator light, and sending a distress signal.

[0037] On the other hand, embodiments of this application provide a human-pet intercom method based on multimodal interaction, which is executed by a user device and includes:

[0038] Receive a dialogue signal trigger command for the target pet, wherein the dialogue signal trigger command includes a specific command identifier;

[0039] According to the dialogue signal trigger instruction, a dialogue signal for the target pet is sent to the pet device, and the translation result of the dialogue signal returned by the pet device is received.

[0040] The human speech data corresponding to the translation result is determined and played back.

[0041] In another aspect, embodiments of this application provide a human-pet intercom device based on multimodal interaction, wherein the pet device terminal includes the device, which includes:

[0042] The signal receiving module is used to receive dialogue signals sent by the user device to the target pet, and determine the animal voice information corresponding to the dialogue signals according to the local language library, wherein the local language library includes the correspondence between each dialogue signal and the animal voice information;

[0043] The data acquisition module is used to acquire real-time response data of the target pet to the dialogue signal, wherein the real-time response data includes the target pet's sound data, motion image data, and environmental data;

[0044] The translation result determination module is used to input the real-time response data into the multimodal translation model to obtain the translation result of the target pet's response to the dialogue signal, and send the translation result to the user device.

[0045] Furthermore, embodiments of this application provide a human-pet intercom device based on multimodal interaction, the user equipment terminal including the device, the device comprising:

[0046] The instruction receiving module is used to receive a dialogue signal trigger instruction for the target pet, wherein the dialogue signal trigger instruction includes a specific instruction identifier;

[0047] The signal sending module is used to send a dialogue signal for the target pet to the pet device terminal according to the dialogue signal triggering instruction, and to receive the translation result of the dialogue signal returned by the pet device terminal;

[0048] The voice playback module is used to determine the voice data corresponding to the translation result and play it using human language.

[0049] On the other hand, embodiments of this application provide a human-pet intercom system based on multimodal interaction. The human-pet intercom system includes a main control system module, a positioning module, an information transmission module, a memory module, a power module, a camera, a microphone, and a speaker output hole. The human-pet intercom system is used to execute any one of the methods in a human-pet intercom method based on multimodal interaction.

[0050] In another aspect, embodiments of this application provide an electronic device, including a processor and a memory:

[0051] The memory is configured to store machine-readable instructions that, when executed by the processor, cause the processor to perform any one of a multimodal human-pet intercom method.

[0052] The beneficial effects of the technical solutions provided in this application include at least the following:

[0053] In this embodiment, after a dialogue with the pet, the translation results of the pet's response data integrate sound data, motion data, and environmental data, achieving efficient fusion of sound, image, and environmental information. Compared with the translation results obtained from only a single sound data in the prior art, the translation results obtained at this time significantly improve the reliability of the translation, solve the problem of high misjudgment rate in low-confidence translation in the prior art, and significantly improve the accuracy and adaptability of human-pet interaction, possessing significant technical advantages and commercial potential.

[0054] Furthermore, the multimodal translation model in this application embodiment, through modular design, clearly delineates the steps of multimodal data processing, feature fusion, noise suppression, and task-driven generation, which enhances the correlation between modalities, reduces information loss, and improves computational efficiency and robustness, thereby achieving flexible task adaptation. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0056] Figure 1A flowchart illustrating a human-pet intercom method based on multimodal interaction provided in this application embodiment;

[0057] Figure 2 A flowchart illustrating another human-pet intercom method based on multimodal interaction provided in this application embodiment;

[0058] Figure 3 A schematic diagram of the structure of a human-pet intercom device based on multimodal interaction provided in this application embodiment;

[0059] Figure 4 A schematic diagram of another human-pet intercom device based on multimodal interaction provided in this application embodiment;

[0060] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0061] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting the invention.

[0062] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0063] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0064] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0065] Specifically, such as Figure 1 As shown, this method is executed by the pet device and may include:

[0066] Step S101: Receive dialogue signals sent by the user device to the target pet, and determine the animal voice information corresponding to the dialogue signals according to the local language library and play the animal voice information. The local language library includes the correspondence between each dialogue signal and the animal voice information.

[0067] Optionally, the pet device refers to a terminal device placed on the pet. Internally, this pet device includes a main control system module, a positioning module, an information transmission module, a memory module, and a power module. Externally, the pet device features a camera, microphone, speaker, power button (power on / off and reset), and an SOS indicator light. The main control system module is responsible for the overall operation of the device; the positioning module provides real-time accurate location tracking for the pet; the camera views the pet's environment and takes photos and records videos; the information transmission module sends and receives information; the memory module stores information, including caching and permanently storing audio and image files; the power module provides power for the device; the microphone picks up sounds from the pet; and the speaker plays audio. Furthermore, the pet device's design incorporates a loop for threading a pet collar through and securing it to the pet's neck or body, allowing the device to be worn on the pet. The user device refers to the terminal device with an app (application software) that matches the pet device installed. This terminal device includes, but is not limited to, mobile terminals and PC (Personal Computer) terminals. The target pet refers to the pet that the user wants to have a conversation with, including but not limited to animals such as dogs, cats, and birds.

[0068] Optionally, when a user wants to communicate with a target pet, they can send a dialogue signal from the user's device to the pet's device. The pet's device can receive the dialogue signal sent by the user's device and then retrieve its local language library. This local language library stores the animal's voice information corresponding to each dialogue signal. The user can then compare the local language library with the received dialogue signal to determine the corresponding animal's voice information. For example, if the target pet is a puppy, and the dialogue signal is a call for the puppy to come home, the pet's device, upon receiving the dialogue signal, can determine the dog language voice data representing the call to the puppy and play it in dog language.

[0069] Optionally, the pet device can receive encrypted dialogue signals sent by the user device via Bluetooth Low Energy (BLE) or 4G / Wi-Fi modules. These dialogue signals contain command types (such as "call" or "command") and voice content (encoded as a binary stream). The data can then be decrypted using the TLS / HTTPS protocol to extract the voice content, and then the voice can be converted into text commands using a lightweight ASR (Automatic Speech Recognition) model.

[0070] Step S102: Obtain real-time response data of the target pet to the dialogue signal. The real-time response data includes the target pet's sound data, motion image data, and environmental data.

[0071] Optionally, after the animal's voice data is played, the target pet will react in real time upon hearing the voice. The terminal device can then acquire the target pet's real-time reaction data, which includes the pet's sound data, motion image data, and environmental data. The sound data refers to the sound emitted by the target pet upon hearing the voice information, which can be received through the speaker on the pet device. The motion image data refers to the data of the target pet's actions upon hearing the voice information (such as tail wagging), which can be obtained through video recording or photographing by the camera on the pet device. The environmental data refers to the data of the target pet's surrounding environment, which can also be obtained through video recording or photographing by the camera on the pet device.

[0072] Step S103: Input the real-time reflected data into the multimodal translation model to obtain the translation result of the target pet's response to the dialogue signal, and send the translation result to the user device.

[0073] Optionally, the pet terminal device can be configured with a pre-trained multimodal translation model. In this case, the acquired real-time response data can be input into the multimodal translation model, which can translate the real-time response data to obtain a translation result. This translation result represents what the target pet wants to express after hearing the voice information.

[0074] Furthermore, after obtaining the translation result, it can be sent to the user device. Upon receiving the translation result, the user device determines the corresponding human speech data and plays it back. Specifically, the translation result can be transmitted by first translating it into text data and then sending the text data to the user device. Upon receiving the text data, the user device determines the corresponding speech data and plays it back in human language. Alternatively, the translation result can be sent to the user device as speech data, and upon receiving the speech data, the user device determines the corresponding speech data and plays it back in human language. This human speech data refers to a pre-set language, such as Chinese, English, or Spanish.

[0075] In optional embodiments of this application, the multimodal translation model includes an input layer, an encoder, and a decoder. Real-time reflected data is input into the multimodal translation model to obtain a translation result of the target pet's response to dialogue signals, including:

[0076] Based on the input layer, feature extraction processing is performed on the sound data, motion image data, and environmental data in the real-time data to obtain the sound features corresponding to the sound data, the image features corresponding to the motion image data, and the environmental features corresponding to the environmental data.

[0077] Based on the encoder, cross-modal semantic fusion processing is performed on the sound features and image features to obtain the fused features;

[0078] The fused features and environmental features are input into the decoder, which then performs dynamic translation generation based on the fused features and environmental features to obtain the translation result of the target pet's response to the dialogue signal.

[0079] Optionally, real-time data can be input into a multimodal translation model. The input layer of the multimodal translation model can perform feature extraction processing on the sound data, motion image data, and environmental data in the real-time data to obtain the sound features corresponding to the sound data, the image features corresponding to the motion image data, and the environmental features corresponding to the environmental data. For example, the sound data can be converted into a mono audio signal, and then a Mel spectrogram can be generated through short-time Fourier transform (STFT). The Mel spectrogram is then subjected to standard framing processing, and the resulting spectral tensor is used as the sound feature corresponding to the sound data. The motion image data can be cropped in real time to represent the main pet region and then normalized. The resulting image tensor is used as the image feature corresponding to the motion image data, thereby achieving data augmentation (such as random flipping, brightness adjustment, etc.). The environmental data can be converted into a 32-dimensional One-Hot vector as the environmental feature corresponding to the environmental data.

[0080] Furthermore, the extracted sound features, image features, and environmental features can be input into cross-modal semantics for fusion processing to obtain fused features. The fused features and environmental features are then input into the decoder, which performs dynamic translation generation processing based on the fused features and environmental features to obtain the translation result of the target pet's response to the dialogue signal.

[0081] In optional embodiments of this application, the encoder includes a dynamic cross-modal attention layer, a sparse local attention layer, and an adaptive noise suppression layer. Based on the encoder, cross-modal semantic fusion processing is performed on the sound features and image features to obtain the fused features, including:

[0082] The sparse local attention layer determines each local window corresponding to the sound features, and the features of each local window are integrated through skip connections to obtain the optimized sound features.

[0083] Based on the adaptive noise suppression layer, the optimized sound features and image features are denoised to obtain the denoised sound features and denoised image features.

[0084] The fused features are obtained by performing self-attention fusion calculation on the denoised sound features and the denoised image features through a dynamic cross-modal attention layer.

[0085] Optionally, the encoder in this embodiment may include a dynamic cross-modal attention layer, a sparse local attention layer, and adaptive noise suppression. The sparse local attention layer can divide the input sound features into blocks, that is, divide them into local windows according to the time series (e.g., 8 frames per window), and then calculate the local attention weights in each window to reduce the global computation. Finally, the features of each window are integrated through a skip connection to obtain the optimized sound features.

[0086] Furthermore, the adaptive noise suppression layer performs noise reduction processing on the optimized sound features and image features, and the resulting noise-reduced sound features and image features are then input into the dynamic cross-modal attention layer. The dynamic cross-modal attention layer then performs self-attention fusion calculation on the input noise-reduced sound features and image features to obtain the fused features.

[0087] In optional embodiments of this application, based on an adaptive noise suppression layer, the optimized sound features and image features are denoised to obtain denoised sound features and denoised image features, including:

[0088] The background noise regions included in the optimized sound features and image features are determined based on the adaptive noise suppression layer;

[0089] By masking the features corresponding to the background noise region in the optimized sound features and image features, the denoised sound features and denoised image features are obtained.

[0090] Optionally, when denoising sound features and image features, a lightweight CNN can be used to first detect background noise regions (such as thunderstorm sounds) in the sound features and image features and generate binary masks. Then, the features corresponding to the background noise regions are masked in the attention calculation of the optimized sound features and image features to obtain the denoised sound features and denoised image features.

[0091] In an optional embodiment of this application, a self-attention fusion calculation is performed on the denoised audio features and the denoised image features through a dynamic cross-modal attention layer to obtain the fused features, including:

[0092] Based on sound features and image features, attention weights for sound data and attention weights for motion image data are determined respectively through a dynamic cross-modal attention layer.

[0093] Based on the attention weights of the audio data and the motion image data, self-attention fusion calculation is performed on the denoised audio features and the denoised image features to obtain the fused features.

[0094] Optionally, a sound-image correlation score (0~1) can be calculated based on sound features and image features. Then, based on the obtained correlation score, self-attention calculation is performed on the sound features and image features respectively to obtain the attention weights of the sound data and the action image data. Finally, based on the attention weights of the sound data and the action image data, self-attention fusion calculation is performed on the denoised sound features and the denoised image features to achieve cross-modal attention fusion.

[0095] In an optional embodiment of this application, the decoder includes a dynamic weight allocation layer and a decoding layer. The decoder obtains the translation result of the target pet's response to the dialogue signal in the following manner:

[0096] The task labels are obtained, and the dynamic weight allocation layer adjusts the attention weights based on the task labels to obtain the adjusted attention weights. The adjusted attention weights include the attention weights of the fused features and the attention weights corresponding to the environmental data.

[0097] Based on the adjusted attention weights, task optimization processing is performed on the features corresponding to the environmental data and the fused features to obtain the task-optimized feature vector.

[0098] The decoding layer maps the task-optimized feature vectors to the vocabulary space to obtain the translation results.

[0099] Optionally, the decoder includes a dynamic weight allocation layer and a decoding layer. When the fused features are input to the decoder, the dynamic weight allocation layer can obtain the task label (such as the "go home" label), then map the task label to a query vector, and adjust the attention weights of the multimodal features guided by the task query to obtain a task-optimized feature vector. The adjustment of the attention weights of the multimodal features is determined by the following formula:

[0100]

[0101] in, For task query vectors, Modal value vector, Let be the dimension of the key vector. For normalization function, This is the modal bond vector.

[0102] Furthermore, the decoding layer captures the long-distance dependencies within the feature vectors optimized for the task based on a self-attention mechanism, and then maps the features to the vocabulary space based on a fully connected layer to obtain the final translation result.

[0103] In this application, the pre-defined multimodal translation model, through modular design, clearly delineates the steps of multimodal data processing, feature fusion, noise suppression, and task-driven generation, which enhances the correlation between modalities, reduces information loss, and improves computational efficiency and robustness, thereby achieving flexible task adaptation.

[0104] In optional embodiments of this application, the method further includes:

[0105] Receive location tracking signals sent by the user's device for the target pet;

[0106] It acquires the target pet's real-time accurate location information and historical movement trajectory, and sends the real-time accurate location information and historical movement trajectory to the user's device.

[0107] Optionally, when a user wants to know the specific location information of their pet, they can send a location viewing signal for the target pet to the pet device through the user's device. When the pet device receives the location viewing signal, it can obtain the real-time accurate location information and the pre-stored historical movement trajectory of the target pet, and then send the real-time accurate location information and historical movement trajectory to the user device.

[0108] In optional embodiments of this application, the method further includes:

[0109] Receive environmental viewing signals sent by the user's device for the target pet;

[0110] Automatically turn on the included cameras and acquire image data of the environment in which the target pet is located;

[0111] Image data is sent to the user's device. The image data includes at least one of video data and image data.

[0112] Optionally, when a user wants to know the environment around their pet, they can send an environment viewing signal for the target pet to the pet device via the user's device. When the pet device receives the environment viewing signal, it can automatically turn on its camera and then acquire image data of the environment scene where the target pet is located. This image data can be video data or image data. Finally, the image data is sent to the user device, which displays it after receiving it, so that the user can know the environment around their pet in real time.

[0113] In an optional embodiment of this application, after sending the image data to the user equipment, it further includes...

[0114] Receive alarm signals sent by the user equipment for the target pet and initiate alarm processing corresponding to the alarm signal;

[0115] The alarm processing includes at least one of the following: playing a preset human distress signal, turning on the alarm indicator light, and sending a distress signal.

[0116] Optionally, after the user device displays image data, if the user discovers that the pet is being attacked, they can send an alarm signal to the pet device via the user device. Upon receiving the alarm signal, the pet device can play a preset human distress signal, activate the alarm indicator light, and send a distress message. For example, after receiving the alarm signal, the pet device can flash the SOS light, automatically dial 110 (the police emergency number) (i.e., send a distress signal), and play a human voice saying, "I am a puppy (or kitten), I am in trouble at XX place, please come and save me!"

[0117] In this embodiment, after a dialogue with the pet, the translation results of the pet's response data integrate sound data, motion data, and environmental data, achieving efficient fusion of sound, image, and environmental information. Compared with the translation results obtained from only a single sound data in the prior art, the translation results obtained at this time significantly improve the reliability of the translation, solve the problem of high misjudgment rate in low-confidence translation in the prior art, and significantly improve the accuracy and adaptability of human-pet interaction, possessing significant technical advantages and commercial potential.

[0118] like Figure 2As shown in the embodiment of this application, a human-pet intercom method based on multimodal interaction is also provided. This method is executed by the user device and includes:

[0119] Step 301: Receive a dialogue signal trigger command for the target pet. The dialogue signal trigger command includes a specific command identifier.

[0120] Step 302: Send a dialogue signal for the target pet to the pet device according to the dialogue signal trigger instruction, and receive the translation result of the dialogue signal returned by the pet device.

[0121] Step 303: Determine the speech data corresponding to the translation result and play it using human language.

[0122] The specific implementation methods of steps 301-303 have been described in detail above, and the specific content can be found in the previous text. They will not be repeated here.

[0123] Optionally, in practical applications, users can also use a custom language library on their devices. In this case, users can use this module to collect their pet's sounds (based on recordings from the corresponding app), and then translate them into human language after observation and understanding, storing them in the custom pet language library. Alternatively, they can upload them to a cloud server to share with other users.

[0124] This application embodiment also provides a human-pet intercom system based on multimodal interaction. The human-pet intercom system includes a main control system module, a positioning module, an information transmission module, a memory module, a power module, a camera, a microphone, and a speaker output hole. The human-pet intercom system is used to execute any one of the methods in a human-pet intercom method based on multimodal interaction.

[0125] The functions of each module in a multimodal interactive human-pet intercom system have been described in detail above, and will not be repeated here.

[0126] This application provides a human-pet intercom device based on multimodal interaction, wherein the pet device terminal includes the device, and the device is as follows: Figure 3 As shown, the device may include: a signal receiving module 401, a data acquisition module 402, and a translation result determination module 403, wherein,

[0127] The signal receiving module is used to receive dialogue signals sent by the user device to the target pet, and determine the animal voice information corresponding to the dialogue signal according to the local language library. The local language library includes the correspondence between each dialogue signal and the animal voice information.

[0128] The data acquisition module is used to acquire real-time response data of the target pet to dialogue signals. The real-time response data includes the target pet's sound data, motion image data, and environmental data.

[0129] The translation result determination module is used to input real-time reflected data into the multimodal translation model to obtain the translation result of the target pet's response to the dialogue signal, and then send the translation result to the user device.

[0130] Optionally, the multimodal translation model includes an input layer, an encoder, and a decoder. The translation result determination module, when inputting real-time feedback data into the multimodal translation model to obtain the translation result of the target pet's response to the dialogue signal, is specifically used for:

[0131] Based on the input layer, feature extraction processing is performed on the sound data, motion image data, and environmental data in the real-time data to obtain the sound features corresponding to the sound data, the image features corresponding to the motion image data, and the environmental features corresponding to the environmental data.

[0132] Based on the encoder, cross-modal semantic fusion processing is performed on the sound features and image features to obtain the fused features;

[0133] The fused features and environmental features are input into the decoder, which then performs dynamic translation generation based on the fused features and environmental features to obtain the translation result of the target pet's response to the dialogue signal.

[0134] Optionally, the encoder includes a dynamic cross-modal attention layer, a sparse local attention layer, and adaptive noise suppression. The translation result determination module, when performing cross-modal semantic fusion processing on the sound and image features based on the encoder to obtain the fused features, is specifically used for:

[0135] The sparse local attention layer determines each local window corresponding to the sound features, and the features of each local window are integrated through skip connections to obtain the optimized sound features.

[0136] Based on the adaptive noise suppression layer, the optimized sound features and image features are denoised to obtain the denoised sound features and denoised image features.

[0137] The fused features are obtained by performing self-attention fusion calculation on the denoised sound features and the denoised image features through a dynamic cross-modal attention layer.

[0138] Optionally, when the translation result determination module performs noise reduction processing on the optimized sound features and image features based on the adaptive noise suppression layer to obtain the denoised sound features and denoised image features, it is specifically used for:

[0139] The background noise regions included in the optimized sound features and image features are determined based on the adaptive noise suppression layer;

[0140] By masking the features corresponding to the background noise region in the optimized sound features and image features, the denoised sound features and denoised image features are obtained.

[0141] Optionally, the translation result determination module performs self-attention fusion calculation on the denoised audio features and denoised image features through a dynamic cross-modal attention layer. When obtaining the fused features, it is specifically used for:

[0142] Based on sound features and image features, attention weights for sound data and attention weights for motion image data are determined respectively through a dynamic cross-modal attention layer.

[0143] Based on the attention weights of the audio data and the motion image data, self-attention fusion calculation is performed on the denoised audio features and the denoised image features to obtain the fused features.

[0144] Optionally, the decoder includes a dynamic weight allocation layer and a decoding layer. The decoder obtains the translation results of the target pet's response to the dialogue signals in the following way:

[0145] The task labels are obtained, and the dynamic weight allocation layer adjusts the attention weights based on the task labels to obtain the adjusted attention weights. The adjusted attention weights include the attention weights of the fused features and the attention weights corresponding to the environmental data.

[0146] Based on the adjusted attention weights, task optimization processing is performed on the environmental features and the fused features to obtain the task-optimized feature vector.

[0147] The decoding layer maps the task-optimized feature vectors to the vocabulary space to obtain the translation results.

[0148] Optionally, the device also includes an information viewing module, specifically used for:

[0149] Receive location tracking signals sent by the user's device for the target pet;

[0150] It acquires the target pet's real-time accurate location information and historical movement trajectory, and sends the real-time accurate location information and historical movement trajectory to the user's device.

[0151] Optionally, the information viewing module is also used for:

[0152] Receive environmental viewing signals sent by the user's device for the target pet;

[0153] Automatically turn on the included cameras and acquire image data of the environment in which the target pet is located;

[0154] Image data is sent to the user's device. The image data includes at least one of video data and image data.

[0155] Optionally, the device also includes an information viewing module, specifically used for:

[0156] After sending the image data to the user device, the alarm signal sent by the user device for the target pet is received, and the alarm processing corresponding to the alarm signal is initiated.

[0157] The alarm processing includes at least one of the following: playing a preset human distress signal, turning on the alarm indicator light, and sending a distress signal.

[0158] This application provides a human-pet intercom device based on multimodal interaction, wherein the user equipment terminal includes the device, and the device is as follows: Figure 4 As shown, the device may include: an instruction receiving module 501, a signal transmitting module 502, and a voice playback module 503, wherein,

[0159] The instruction receiving module is used to receive dialogue signal trigger instructions for the target pet. The dialogue signal trigger instructions include a specific instruction identifier.

[0160] The signal sending module is used to send dialogue signals for the target pet to the pet device according to the dialogue signal trigger command, and to receive the translation results of the dialogue signals returned by the pet device.

[0161] The voice playback module is used to determine the voice data corresponding to the translation result and play it back in human language.

[0162] The human-pet intercom device based on multimodal interaction in this embodiment can execute the human-pet intercom method based on multimodal interaction shown in the embodiment of this application. The implementation principle is similar and will not be described again here.

[0163] This application provides an electronic device, which includes a processor and a memory configured to store machine-readable instructions that, when executed by the processor, cause the processor to perform a human-pet intercom method based on multimodal interaction.

[0164] This application provides an electronic device, such as... Figure 5 As shown, Figure 5The illustrated electronic device includes a processor 2001 and a memory 2003. The processor 2001 and the memory 2003 are connected, for example, via a bus 2002. Optionally, the electronic device 2000 may further include a transceiver 2004. It should be noted that in practical applications, the transceiver 2004 is not limited to one type, and the structure of this electronic device 2000 does not constitute a limitation on the embodiments of this application.

[0165] Processor 2001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 2001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0166] Bus 2002 may include a pathway for transmitting information between the aforementioned components. Bus 2002 may be a PCI bus or an EISA bus, etc. Bus 2002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0167] The memory 2003 may be ROM or other type of static storage device capable of storing static information and instructions, RAM or other type of dynamic storage device capable of storing information and instructions, or EEPROM, CD-ROM or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0168] The memory 2003 stores the application code that executes the scheme of this application, and its execution is controlled by the processor 2001. The processor 2001 executes the application code stored in the memory 2003 to implement... Figure 3 and Figure 4 The illustrated embodiment provides the action of a human-pet intercom device based on multimodal interaction.

[0169] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0170] The above description is only a partial embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for pet-to-human intercom based on multi-modal interaction, characterized in that, The method is executed by a pet device end, and the method comprises: receiving a dialogue signal sent by a user device end for a target pet, and determining animal voice information corresponding to the dialogue signal according to a local language library, wherein the local language library comprises a corresponding relationship between each dialogue signal and animal voice information; obtaining real-time reflection data of the target pet for the dialogue signal, wherein the real-time reflection data comprises voice data, action image data and environment data of the target pet; inputting the real-time reflection data into a multi-modal translation model to obtain a translation result of the target pet responding to the dialogue signal, and sending the translation result to the user device; the multi-modal translation model comprises an input layer, an encoder and a decoder, and the inputting of the real-time reflection data into the multi-modal translation model to obtain the translation result of the target pet responding to the dialogue signal comprises: based on the input layer, the voice data, the action image data and the environment data in the real-time reflection data are respectively subjected to feature extraction processing to obtain voice features corresponding to the voice data, image features corresponding to the action image data and environment features corresponding to the environment data, based on the encoder, the voice features and the image features are subjected to cross-modal semantic fusion processing to obtain fused features; the fused features and the environment features are input into the decoder to enable the decoder to generate a dynamic translation according to the fused features and the environment features to obtain the translation result of the target pet responding to the dialogue signal; the encoder comprises a dynamic cross-modal attention layer, a sparse local attention layer and an adaptive noise suppression layer, and the cross-modal semantic fusion processing of the voice features and the image features by the encoder to obtain the fused features comprises: based on the sparse local attention layer, each local window corresponding to the voice features is determined, and the features of each local window are integrated through a skip connection to obtain optimized voice features; based on the adaptive noise suppression layer, the optimized voice features and the image features are subjected to noise reduction processing to obtain noise-reduced voice features and noise-reduced image features; the noise-reduced voice features and the noise-reduced image features are subjected to self-attention fusion calculation by the dynamic cross-modal attention layer to obtain the fused features.

2. The method of claim 1, wherein, based on the adaptive noise suppression layer, the optimized voice features and the image features are subjected to noise reduction processing to obtain noise-reduced voice features and noise-reduced image features, comprising: based on the adaptive noise suppression layer, the optimized voice features and the image features each comprise a background noise region; the features corresponding to the background noise region are shielded in the optimized voice features and the image features to obtain noise-reduced voice features and noise-reduced image features.

3. The method of claim 1, wherein, the noise-reduced voice features and the noise-reduced image features are subjected to self-attention fusion calculation by the dynamic cross-modal attention layer to obtain the fused features, comprising: According to the sound feature and the image feature, the dynamic cross-modal attention layer is used to determine an attention weight of the sound data and an attention weight of the motion image data respectively; According to the attention weight of the sound data and the attention weight of the motion image data, a self-attention fusion calculation is performed on the noise-reduced sound feature and the noise-reduced image feature to obtain a fused feature.

4. The method of claim 1, wherein, The decoder includes a dynamic weight distribution layer and a decoding layer, and the decoder obtains a translation result of the target pet responding to the dialogue signal in the following manner: The dynamic weight distribution layer adjusts the attention weight based on the task label to obtain an adjusted attention weight, and the adjusted attention weight includes an attention weight of the fused feature and an attention weight corresponding to the environment data; According to the adjusted attention weight, the environment feature and the fused feature are processed for task optimization to obtain a task-optimized feature vector; The decoding layer maps the task-optimized feature vector to a vocabulary space to obtain the translation result.

5. The method of claim 1, wherein, The method further includes: receiving a location viewing signal sent by the user equipment end for the target pet; obtaining real-time accurate location information and historical motion trajectory of the target pet, and sending the real-time accurate location information and the historical motion trajectory to the user equipment.

6. The method of claim 1, wherein, The method further includes: receiving an environment viewing signal sent by the user equipment end for the target pet; automatically starting the included camera, and obtaining image data of an environment scene where the target pet is located through the camera; sending the image data to the user equipment end, wherein the image data includes at least one of video data and picture data.

7. The method of claim 6, wherein, After the image data is sent to the user equipment end, the method further includes receiving an alarm signal sent by the user equipment end for the target pet, and starting an alarm processing corresponding to the alarm signal; The alarm processing includes at least one of playing a preset human help-seeking voice, turning on an alarm indicator light, and sending a help-seeking signal.

8. A pet-to-pet intercom system based on multi-modal interaction, characterized in that, The human-pet intercom system includes a master control system module, a positioning module, an information transmission module, a memory module, a power module, a camera, a microphone, and a speaker sound hole, and is used to execute the method in any one of claims 1-7. The human-pet intercom system includes a master control system module, a positioning module, an information transmission module, a memory module, a power module, a camera, a microphone, and a speaker sound hole, and is used to execute the method in any one of claims 1-7.

Citation Information

Patent Citations

  • Pet identity recognition method and device and storage medium

    CN118942124A

  • Pet voice translation method and system, electronic equipment and storage medium

    CN119626263A