Man-pet talkback method and system based on multi-modal interaction
By using a multimodal translation model in pet interaction devices to integrate various pet data and generate highly reliable translation results, the problems of functional fragmentation and low interaction efficiency in existing technologies are solved, and more efficient natural communication between humans and pets is achieved.
Patent Information
- Application Number
- CN202510786448.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Existing pet interaction devices have fragmented functions, and users need to use multiple independent devices in combination to cover different scenarios. This results in serious operational redundancy and resource waste. In addition, there is a lack of adaptive learning mechanisms, making it impossible to dynamically optimize the interaction logic. This limits the efficiency of natural communication between humans and pets and the depth of emotional connection.
This paper presents a multimodal interaction-based method for human-pet intercom communication. The pet device receives conversation signals from the user's device and uses a multimodal translation model to integrate the pet's sound data, motion image data, and environmental data to generate a translation result, which is then sent back to the user's device. This method includes an input layer, an encoder, and a decoder, and utilizes cross-modal semantic fusion and self-attention fusion to improve the reliability of the translation results.
It significantly improves the reliability of translation results, solves the problem of high misjudgment rate of low-confidence translation in existing technologies, improves the accuracy and adaptability of human-pet interaction, and has significant technical advantages and commercial potential.
Smart Images

Figure CN120636418A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence, and in particular to a method and system for human-pet intercom based on multimodal interaction. Background Art
[0002] In modern society, pets have become important emotional companions. Owners are eager to overcome species-specific language barriers and more accurately understand their pets' emotions and behavioral intentions. However, research has revealed significant limitations in pet interaction devices currently available on the market. First, helmet-mounted translators are bulky and uncomfortable to wear, which can easily lead to pet resistance. Second, mobile translation apps, while portable, rely on close-range operation, making it difficult to achieve long-distance two-way voice interaction. Third, positioning devices only integrate basic location tracking functions and lack real-time communication capabilities. Fourth, traditional communication devices are bulky, complex to deploy, require manual training to trigger preset commands, and have a narrow language library, making them difficult to meet personalized needs.
[0003] As can be seen, existing pet interaction devices generally exhibit fragmented functionality, requiring users to stack multiple independent devices to cover different scenarios, resulting in redundant operations and wasted resources. Furthermore, existing solutions lack adaptive learning mechanisms, making it impossible to dynamically optimize interaction logic through AI technology, severely limiting the efficiency of natural communication and the depth of emotional connection between humans and pets. Summary of the Invention
[0004] The purpose of this application is to solve at least one of the above technical deficiencies.
[0005] On the one hand, an embodiment of the present application provides a method for human-pet intercom based on multimodal interaction, which is executed by a pet device and includes: Receive a conversation signal sent by a user device to a target pet, and determine the animal voice information corresponding to the conversation signal based on a local language library, wherein the local language library includes a correspondence between each conversation signal and the animal voice information; Acquire real-time response data of the target pet to the conversation signal, the real-time response data including sound data, motion image data and environmental data of the target pet; The real-time reflection data is input into the multimodal translation model to obtain the translation result of the target pet's response to the dialogue signal, and the translation result is sent to the user device.
[0006] Optionally, the multimodal translation model includes an input layer, an encoder, and a decoder, which inputs real-time reflection data into the multimodal translation model to obtain a translation result of the target pet responding to the conversation signal, including: Based on the input layer, feature extraction processing is performed on the sound data, action image data and environmental data in the real-time reflection data respectively to obtain the sound features corresponding to the sound data, the image features corresponding to the action image data and the environmental features corresponding to the environmental data. Performing cross-modal semantic fusion processing on the sound features and the image features based on the encoder to obtain fused features; The fused features and the environmental features are input into the decoder, so that the decoder performs dynamic translation generation processing according to the fused features and the environmental features to obtain a translation result of the target pet's response to the dialogue signal.
[0007] Optionally, the encoder includes a dynamic cross-modal attention layer, a sparse local attention layer, and adaptive noise suppression, and the cross-modal semantic fusion processing of the sound features and the image features based on the encoder to obtain the fused features includes: Determining each local window corresponding to the sound feature based on the sparse local attention layer, and integrating the features of each local window through a skip connection to obtain an optimized sound feature; Based on the adaptive noise suppression layer, performing noise reduction processing on the optimized sound features and the image features to obtain noise-reduced sound features and noise-reduced image features; The dynamic cross-modal attention layer performs self-attention fusion calculation on the denoised sound features and the denoised image features to obtain fused features.
[0008] Optionally, performing noise reduction processing on the optimized sound features and the image features based on the adaptive noise suppression layer to obtain noise-reduced sound features and noise-reduced image features includes: determining, based on the adaptive noise suppression layer, background noise regions respectively included in the optimized sound feature and the image feature; The features corresponding to the background noise area are shielded from the optimized sound features and the image features to obtain the noise-reduced sound features and the noise-reduced image features.
[0009] Optionally, performing self-attention fusion calculation on the denoised sound features and the denoised image features through the dynamic cross-modal attention layer to obtain fused features includes: Determining, by the dynamic cross-modal attention layer, an attention weight of the sound data and an attention weight of the action image data according to the sound features and the image features; According to the attention weight of the sound data and the attention weight of the action image data, a self-attention fusion calculation is performed on the denoised sound features and the denoised image features to obtain fused features.
[0010] Optionally, the decoder includes a dynamic weight allocation layer and a decoding layer, and the decoder obtains a translation result of the target pet's response to the dialogue signal in the following manner: Obtaining a task label, and the dynamic weight allocation layer adjusting the attention weight based on the task label to obtain an adjusted attention weight, wherein the adjusted attention weight includes the attention weight of the fused feature and the attention weight corresponding to the environmental data; performing task optimization processing on the features corresponding to the environmental data and the fused features according to the adjusted attention weights to obtain a task-optimized feature vector; The decoding layer maps the task-optimized feature vector to a vocabulary space to obtain the translation result.
[0011] Optionally, the method further includes: Receiving a location check signal sent by the user device to the target pet; The real-time accurate location information and the historical movement trajectory of the target pet are obtained, and the real-time accurate location information and the historical movement trajectory are sent to the user device.
[0012] Optionally, the method further includes: receiving an environment checking signal sent by the user device to the target pet; Automatically turning on the included camera and acquiring image data of the environment scene where the target pet is located through the camera; The image data is sent to the user equipment end, where the image data includes at least one of video data and picture data.
[0013] Optionally, after sending the image data to the user device, the method further includes: receiving an alarm signal sent by the user device to the target pet, and initiating an alarm process corresponding to the alarm signal; The alarm processing includes at least one of playing a preset human distress voice, turning on an alarm indicator light, and sending a distress signal.
[0014] On the other hand, an embodiment of the present application provides a method for human-pet intercom based on multimodal interaction, which is executed by a user device and includes: receiving a conversation signal trigger instruction for a target pet, wherein the conversation signal trigger instruction includes a specific instruction identifier; sending a conversation signal for the target pet to a pet device end according to the conversation signal trigger instruction, and receiving a translation result for the conversation signal returned by the pet device end; Determine the human voice data corresponding to the translation result and play the voice data.
[0015] On the other hand, an embodiment of the present application provides a human-pet intercom device based on multimodal interaction, which is included in a pet device end, and the device includes: a signal receiving module, configured to receive a conversation signal sent by a user device to a target pet, and determine the animal voice information corresponding to the conversation signal based on a local language library, wherein the local language library includes a correspondence between each conversation signal and the animal voice information; a data acquisition module for acquiring real-time reflection data of the target pet in response to the dialogue signal, wherein the real-time reflection data includes voice data, motion image data, and environmental data of the target pet; The translation result determination module is used to input the real-time reflection data into the multimodal translation model, obtain the translation result of the target pet's response to the dialogue signal, and send the translation result to the user device.
[0016] On the other hand, an embodiment of the present application provides a human-pet intercom device based on multimodal interaction, which is included in a user device, and includes: An instruction receiving module is used to receive a conversation signal trigger instruction for a target pet, wherein the conversation signal trigger instruction includes a specific instruction identifier; a signal sending module, configured to send a conversation signal for the target pet to a pet device end according to the conversation signal trigger instruction, and receive a translation result for the conversation signal returned by the pet device end; The voice playing module is used to determine the voice data corresponding to the translation result and play the voice data in human language.
[0017] On the other hand, an embodiment of the present application provides a human-pet intercom system based on multimodal interaction, which includes a main control system module, a positioning module, an information transmission module, a memory module, a power module, a camera, a microphone and a speaker outlet. The human-pet intercom system is used to execute any one of the human-pet intercom methods based on multimodal interaction.
[0018] In another aspect, an embodiment of the present application provides an electronic device, including a processor and a memory: The memory is configured to store machine-readable instructions, which, when executed by the processor, cause the processor to perform any one of the methods for human-pet intercom based on multimodal interaction.
[0019] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least: In an embodiment of the present application, after a conversation with a pet, the translation result of the pet's feedback data integrates sound data, action data and environmental data, realizing efficient fusion of sound, image and environmental information. The translation result obtained at this time is compared with the translation result obtained based only on a single sound data in the prior art, which significantly improves the translation reliability, solves the problem of high misjudgment rate of low-confidence translation in the prior art, significantly improves the accuracy and adaptability of human-pet interaction, and has significant technical advantages and commercial potential.
[0020] In addition, the multimodal translation model in the embodiment of the present application adopts a modular design, which clearly divides the steps of multimodal data processing, feature fusion, noise suppression and task-driven generation, enhances the correlation between modalities while reducing information loss, and improves computational efficiency and robustness, thereby achieving flexible task adaptation. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] Figure 1 A flowchart of a method for human-pet intercom based on multimodal interaction provided in an embodiment of the present application; Figure 2 A flowchart of another method for human-pet intercom based on multimodal interaction provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of a human-pet intercom device based on multimodal interaction provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of another human-pet intercom device based on multimodal interaction provided in an embodiment of the present application; Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present invention.
[0024] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.
[0025] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0026] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0027] Specifically, such as Figure 1 As shown, the method is executed by the pet device side, and the method may include: Step S101: receiving a conversation signal sent by a user device to a target pet, determining the animal voice information corresponding to the conversation signal according to a local language library and playing the animal voice information. The local language library includes a correspondence between each conversation signal and the animal voice information.
[0028] Optionally, the "pet device" refers to a terminal device located on the pet's side. This device internally includes a main control system module, a positioning module, an information transmission module, a memory module, and a power module. The pet device's exterior features a camera, microphone, speaker, power buttons (on / off and reset), and an SOS signal light. The main control system module is responsible for operating the entire device system; the positioning module is used to accurately locate the pet in real time; the camera is used to view the pet's environment and take photos and record videos; the information transmission module is used to send and receive information; the memory module is used for information storage, including caching and permanent storage of audio and image files; the power module is used to provide energy for device operation; the microphone is used to receive sounds emitted by the pet; and the speaker is used to play back the sounds. Furthermore, the pet device's exterior structure is designed with a loop through which a pet collar is passed and secured to the pet's neck or body, thereby enabling the pet device to be worn by the pet. The user device refers to a terminal device installed with an APP (Application Software) that matches the pet device, which includes but is not limited to mobile terminals, PC (Personal Computer) terminals, etc.; the target pet refers to the pet that the user wants to communicate with, which includes but is not limited to dogs, cats, birds and other animals.
[0029] Optionally, when a user wishes to engage in a conversation with a target pet, they can send a conversation signal from their user device to the pet device. The pet device can then receive the conversation signal sent by the user device for the target pet and retrieve a local language library containing the animal voice information corresponding to each conversation signal. The library can then be compared with the received conversation signal to determine the animal voice information corresponding to the conversation signal. For example, if the target pet is a puppy and the conversation signal is a call for the puppy to come home, the pet device can determine, based on the local language library, the dog language voice data representing the call for the puppy to come home and play the voice message in the dog language.
[0030] Optionally, the pet device can receive encrypted conversation signals sent by the user device through a low-energy Bluetooth (BLE) or 4G / Wi-Fi module. The conversation signal contains the command type (such as "call", "command") and voice content (encoded as a binary stream). The data can then be decrypted using the TLS / HTTPS protocol, the voice content can be extracted, and then the voice can be converted into text commands through a lightweight ASR (automatic speech recognition) model.
[0031] Step S102: obtaining the target pet's real-time response data to the dialogue signal, where the real-time response data includes the target pet's sound data, motion image data, and environment data.
[0032] Optionally, after the animal voice data is played, the target pet will react in real time after hearing the voice. At this time, the terminal device can obtain the target pet's real-time reaction data, which includes the target pet's sound data, motion image data, and environmental data. Among them, the sound data refers to the sound data emitted by the target pet after hearing the voice information, which can be collected through the speaker outlet arranged on the pet device end; the motion image data refers to the data of the action generated by the target pet after hearing the voice information (such as the action data of wagging the tail), which can be obtained by recording or taking photos through the camera arranged on the pet device end; the environmental data refers to the data of the target pet's surrounding environment, which can also be obtained by recording or taking photos through the camera arranged on the pet device end.
[0033] In step S103, the real-time reflection data is input into the multimodal translation model to obtain a translation result of the target pet's response to the dialogue signal, and the translation result is sent to the user device.
[0034] Optionally, the pet terminal device can be configured with a pre-trained multimodal translation model. In this case, the acquired real-time reflection data can be input into the multimodal translation model. The multimodal translation model can translate the real-time reflection data to obtain a translation result, which represents what the target pet wants to express after hearing the voice information.
[0035] Furthermore, after obtaining the translation result, the translation result can be sent to the user device. After receiving the translation result, the user device determines the human voice data corresponding to the translation result and plays the voice. When transmitting the translation result, the translation result can be first translated into text data, which is then sent to the user device. After receiving the text data, the user device determines the corresponding voice data and plays it in the human language. Alternatively, the translation result can be sent to the user device in the form of voice data. After receiving the voice data, the user device determines the corresponding voice data and plays it in the human language. The human voice data refers to a pre-set language, such as Chinese, English, or Spanish.
[0036] In an optional embodiment of the present application, the multimodal translation model includes an input layer, an encoder, and a decoder. Real-time reflection data is input into the multimodal translation model to obtain a translation result of the target pet's response to the dialogue signal, including: Based on the input layer, the sound data, action image data and environmental data in the real-time reflection data are respectively subjected to feature extraction processing to obtain the sound features corresponding to the sound data, the image features corresponding to the action image data and the environmental features corresponding to the environmental data. Based on the encoder, cross-modal semantic fusion processing is performed on the sound features and image features to obtain the fused features; The fused features and environmental features are input into the decoder, so that the decoder performs dynamic translation generation processing based on the fused features and environmental features to obtain a translation result in which the target pet responds to the dialogue signal.
[0037] Optionally, the real-time response data can be input into a multimodal translation model. The input layer of the multimodal translation model can perform feature extraction on the sound data, motion image data, and environmental data in the real-time response data, respectively, to obtain sound features corresponding to the sound data, image features corresponding to the motion image data, and environmental features corresponding to the environmental data. For example, the sound data can be converted into a monophonic audio signal, then a short-time Fourier transform (STFT) is used to generate a mel-spectrogram. The mel-spectrogram is then subjected to standard framing, and the resulting spectrum tensor is used as the sound features corresponding to the sound data. The pet's main body area is cropped from the motion image data in real time and normalized. The resulting image tensor is used as the image features corresponding to the motion image data, thereby achieving data enhancement (such as random flipping and brightness adjustment). The environmental data can be converted into a 32-dimensional one-hot vector as the environmental features corresponding to the environmental data.
[0038] Furthermore, the extracted sound features, image features, and environmental features can be input into cross-modal semantics for fusion processing to obtain fused features, and the fused features and environmental features can be input into the decoder. The decoder then performs dynamic translation generation processing based on the fused features and environmental features to obtain the translation result of the target pet's response to the dialogue signal.
[0039] In an optional embodiment of the present application, the encoder includes a dynamic cross-modal attention layer, a sparse local attention layer, and an adaptive noise suppression layer. The encoder performs cross-modal semantic fusion processing on the sound features and image features to obtain fused features, including: Based on the sparse local attention layer, each local window corresponding to the sound feature is determined, and the features of each local window are integrated through jump connections to obtain the optimized sound features; Based on the adaptive noise suppression layer, the optimized sound features and image features are subjected to noise reduction processing to obtain noise-reduced sound features and noise-reduced image features; The dynamic cross-modal attention layer is used to perform self-attention fusion calculation on the denoised sound features and the denoised image features to obtain the fused features.
[0040] Optionally, the encoder in the embodiment of the present application may include a dynamic cross-modal attention layer, a sparse local attention layer and adaptive noise suppression. The sparse local attention layer can perform block processing on the input sound features, that is, divide them into local windows according to the time series (such as 8 frames per window), and then calculate the local attention weight in each window, thereby reducing the global calculation amount, and finally integrate the features of each window through skip connection to obtain optimized sound features.
[0041] Furthermore, the adaptive noise suppression layer performs denoising on the optimized sound features and image features, obtains the denoised sound features and denoised image features, and then inputs them into the dynamic cross-modal attention layer. The dynamic cross-modal attention layer then performs self-attention fusion calculation on the input denoised sound features and denoised image features to obtain the fused features.
[0042] In an optional embodiment of the present application, based on the adaptive noise suppression layer, noise reduction processing is performed on the optimized sound features and image features to obtain noise-reduced sound features and noise-reduced image features, including: determining, based on the adaptive noise suppression layer, background noise regions included in the optimized sound features and image features; The features corresponding to the background noise area are shielded in the optimized sound features and image features to obtain the noise-reduced sound features and noise-reduced image features.
[0043] Optionally, when denoising sound features and image features, a lightweight CNN can be used to first detect background noise areas (such as thunderstorm sounds) in the sound features and image features and generate binary masks. Then, in the attention calculation of the optimized sound features and image features, the features corresponding to the background noise areas are masked to obtain the denoised sound features and denoised image features.
[0044] In an optional embodiment of the present application, a dynamic cross-modal attention layer is used to perform self-attention fusion calculation on the denoised sound features and the denoised image features to obtain fused features, including: Based on the sound features and image features, the attention weights of the sound data and the attention weights of the action image data are determined respectively through the dynamic cross-modal attention layer; According to the attention weight of the sound data and the attention weight of the action image data, self-attention fusion calculation is performed on the denoised sound features and the denoised image features to obtain the fused features.
[0045] Optionally, a sound-image correlation score (0-1) can be calculated based on the sound features and image features, and then self-attention calculation can be performed on the sound features and image features respectively based on the obtained correlation score to obtain the attention weight of the sound data and the attention weight of the action image data. Finally, based on the attention weight of the sound data and the attention weight of the action image data, self-attention fusion calculation can be performed on the denoised sound features and the denoised image features to achieve cross-modal attention fusion.
[0046] In an optional embodiment of the present application, the decoder includes a dynamic weight allocation layer and a decoding layer. The decoder obtains the translation result of the target pet's response to the dialogue signal in the following manner: Obtain the task label. The dynamic weight allocation layer adjusts the attention weight based on the task label to obtain the adjusted attention weight. The adjusted attention weight includes the attention weight of the fused feature and the attention weight corresponding to the environmental data. According to the adjusted attention weights, the features corresponding to the environmental data and the fused features are processed for task optimization to obtain the feature vectors for task optimization. The decoding layer maps the task-optimized feature vector to the vocabulary space to obtain the translation result.
[0047] Optionally, the decoder includes a dynamic weight allocation layer and a decoding layer. When the fused features are input to the decoder, the dynamic weight allocation layer can obtain the task label (such as the "go home" label), then map the task label to a query vector, and use the task query as a guide to adjust the attention weight of the multimodal features to obtain a task-optimized feature vector. The adjustment of the attention weight of the multimodal features is determined by the following formula: in, is the task query vector, modal value vector, is the dimension of the key vector, is the normalization function, is the modal key vector.
[0048] Furthermore, the decoding layer captures the long-distance dependencies within the task-optimized feature vector based on the self-attention mechanism, and then maps the features to the vocabulary space based on the fully connected layer to obtain the final translation result.
[0049] In this application, the preset multimodal translation model adopts a modular design, which clearly divides the steps of multimodal data processing, feature fusion, noise suppression and task-driven generation, enhances the correlation between modalities while reducing information loss, and improves computational efficiency and robustness, thereby achieving flexible task adaptation.
[0050] In an optional embodiment of the present application, the method further includes: Receive a location check signal sent by a user device for a target pet; Obtain the real-time and accurate location information and historical movement trajectory of the target pet, and send the real-time and accurate location information and historical movement trajectory to the user's device.
[0051] Optionally, when the user wants to know the specific location information of the pet, a location check signal for the target pet can be sent to the pet-end device through the user-end device. When the pet-end device receives the location check signal, it can obtain the real-time and accurate location information of the target pet and the pre-stored historical movement trajectory, and then send the real-time and accurate location information and historical movement trajectory to the user device.
[0052] In an optional embodiment of the present application, the method further includes: Receiving an environment viewing signal sent by a user device to a target pet; Automatically turning on the included camera and acquiring image data of the environment scene where the target pet is located through the camera; The image data is sent to the user equipment end, where the image data includes at least one of video data and picture data.
[0053] Optionally, when the user wants to know the environment around the pet, the user can send an environment viewing signal for the target pet to the pet-side device through the user-side device. When the pet-side device receives the environment viewing signal, it can automatically turn on the included camera, and then obtain image data of the environment scene where the target pet is located through the camera. The image data can be video data or picture data. Finally, the image data is sent to the user device side, and the user device side displays the image data after receiving it, so that the user can know the environment around the pet in real time.
[0054] In an optional embodiment of the present application, after sending the image data to the user device, the following steps are also included: Receive an alarm signal sent by a user device for a target pet, and initiate an alarm process corresponding to the alarm signal; The alarm processing includes at least one of playing a preset human distress voice, turning on an alarm indicator light, and sending a distress signal.
[0055] Optionally, after the user device displays the image data, if the user discovers that their pet is being attacked, they can send an alarm signal from the user device to the pet device. Upon receiving the alarm signal, the pet device can play a preset human distress call, turn on an alarm indicator, and send a distress message. For example, upon receiving the alarm signal, the pet device can flash an SOS light, automatically dial 110 (i.e., send a distress signal), and play a human voice message saying, "I am a puppy (or kitten). I am in trouble in XX place. Please come and save me!"
[0056] In an embodiment of the present application, after a conversation with a pet, the translation result of the pet's feedback data integrates sound data, action data and environmental data, realizing efficient fusion of sound, image and environmental information. The translation result obtained at this time is compared with the translation result obtained based only on a single sound data in the prior art, which significantly improves the translation reliability, solves the problem of high misjudgment rate of low-confidence translation in the prior art, significantly improves the accuracy and adaptability of human-pet interaction, and has significant technical advantages and commercial potential.
[0057] like Figure 2 As shown, the embodiment of the present application also provides a method for human-pet intercom based on multimodal interaction, which is executed by a user device and includes: Step 301: Receive a conversation signal trigger instruction for a target pet, where the conversation signal trigger instruction includes a specific instruction identifier.
[0058] Step 302: sending a conversation signal for the target pet to the pet device according to the conversation signal trigger instruction, and receiving a translation result for the conversation signal returned by the pet device.
[0059] Step 303: determine the voice data corresponding to the translation result and play the voice data in human language.
[0060] The specific implementation of steps 301 to 303 has been described in detail in the previous text. The specific content can be found in the previous text and will not be repeated here.
[0061] Optionally, in actual applications, the user device also has a custom language library. In this case, the user can use this module to collect the pet's voice (based on the corresponding APP recording), and then translate it into human language after observation and understanding and store it in the custom pet language library. It can also be uploaded to the cloud server for sharing with other users.
[0062] An embodiment of the present application also provides a human-pet intercom system based on multimodal interaction, which includes a main control system module, a positioning module, an information transmission module, a memory module, a power module, a camera, a microphone, and a speaker outlet. The human-pet intercom system is used to execute any one of the human-pet intercom methods based on multimodal interaction.
[0063] Among them, the functions of each module in a human-pet intercom system based on multimodal interaction have been described in detail in the previous article. The specific content can be found in the previous article and will not be repeated here.
[0064] The embodiment of the present application provides a human-pet intercom device based on multimodal interaction, which is included in the pet device end. Figure 3 As shown, the device may include: a signal receiving module 401, a data acquisition module 402 and a translation result determination module 403, wherein: A signal receiving module is configured to receive a conversation signal sent by a user device to a target pet and determine the animal voice information corresponding to the conversation signal based on a local language library, wherein the local language library includes a correspondence between each conversation signal and the animal voice information; A data acquisition module is used to acquire real-time response data of the target pet to the dialogue signal, wherein the real-time response data includes sound data, motion image data and environmental data of the target pet; The translation result determination module is used to input the real-time reflection data into the multimodal translation model, obtain the translation result of the target pet's response to the dialogue signal, and send the translation result to the user device.
[0065] Optionally, the multimodal translation model includes an input layer, an encoder, and a decoder. When the translation result determination module inputs the real-time reflection data into the multimodal translation model and obtains the translation result of the target pet's response to the dialogue signal, it is specifically used to: Based on the input layer, feature extraction processing is performed on the sound data, action image data and environmental data in the real-time reflection data to obtain the sound features corresponding to the sound data, the image features corresponding to the action image data and the environmental features corresponding to the environmental data. Based on the encoder, cross-modal semantic fusion processing is performed on the sound features and image features to obtain the fused features; The fused features and environmental features are input into the decoder, so that the decoder performs dynamic translation generation processing based on the fused features and environmental features to obtain a translation result in which the target pet responds to the dialogue signal.
[0066] Optionally, the encoder includes a dynamic cross-modal attention layer, a sparse local attention layer, and adaptive noise suppression. When the translation result determination module performs cross-modal semantic fusion processing on the sound features and image features obtained by the encoder to obtain the fused features, it is specifically used to: Based on the sparse local attention layer, each local window corresponding to the sound feature is determined, and the features of each local window are integrated through jump connections to obtain the optimized sound features; Based on the adaptive noise suppression layer, the optimized sound features and image features are subjected to noise reduction processing to obtain noise-reduced sound features and noise-reduced image features; The dynamic cross-modal attention layer is used to perform self-attention fusion calculation on the denoised sound features and the denoised image features to obtain the fused features.
[0067] Optionally, when the translation result determination module performs noise reduction processing on the optimized sound features and image features based on the adaptive noise suppression layer to obtain the noise-reduced sound features and the noise-reduced image features, it is specifically configured to: determining, based on the adaptive noise suppression layer, background noise regions included in the optimized sound features and image features; The features corresponding to the background noise area are shielded in the optimized sound features and image features to obtain the noise-reduced sound features and noise-reduced image features.
[0068] Optionally, the translation result determination module performs self-attention fusion calculation on the denoised sound features and the denoised image features through a dynamic cross-modal attention layer. When the fused features are obtained, they are specifically used to: Based on the sound features and image features, the attention weights of the sound data and the attention weights of the action image data are determined respectively through the dynamic cross-modal attention layer; According to the attention weight of the sound data and the attention weight of the action image data, self-attention fusion calculation is performed on the denoised sound features and the denoised image features to obtain the fused features.
[0069] Optionally, the decoder includes a dynamic weight allocation layer and a decoding layer. The decoder obtains the translation result of the target pet's response to the conversation signal in the following manner: Obtain the task label. The dynamic weight allocation layer adjusts the attention weight based on the task label to obtain the adjusted attention weight. The adjusted attention weight includes the attention weight of the fused feature and the attention weight corresponding to the environmental data. According to the adjusted attention weights, the environmental features and the fused features are processed for task optimization to obtain the task-optimized feature vector; The decoding layer maps the task-optimized feature vector to the vocabulary space to obtain the translation result.
[0070] Optionally, the device further includes an information viewing module, specifically configured to: Receive a location check signal sent by a user device for a target pet; Obtain the real-time and accurate location information and historical movement trajectory of the target pet, and send the real-time and accurate location information and historical movement trajectory to the user's device.
[0071] Optionally, the information viewing module is also used to: Receiving an environment viewing signal sent by a user device to a target pet; Automatically turning on the included camera and acquiring image data of the environment scene where the target pet is located through the camera; The image data is sent to the user equipment end, where the image data includes at least one of video data and picture data.
[0072] Optionally, the device further includes an information viewing module, specifically configured to: After sending the image data to the user device, receiving an alarm signal sent by the user device for the target pet, and starting an alarm process corresponding to the alarm signal; The alarm processing includes at least one of playing a preset human distress voice, turning on an alarm indicator light, and sending a distress signal.
[0073] The embodiment of the present application provides a human-pet intercom device based on multimodal interaction, which is included in the user device. Figure 4 As shown, the device may include: an instruction receiving module 501, a signal sending module 502 and a voice playing module 503, wherein: An instruction receiving module is used to receive a conversation signal trigger instruction for a target pet, wherein the conversation signal trigger instruction includes a specific instruction identifier; A signal sending module is used to send a conversation signal for a target pet to the pet device end according to a conversation signal trigger instruction, and receive a translation result for the conversation signal returned by the pet device end; The voice playback module is used to determine the voice data corresponding to the translation result and play the voice in human language.
[0074] A human-pet intercom device based on multimodal interaction in this embodiment can execute a human-pet intercom method based on multimodal interaction shown in the embodiment of this application. Its implementation principle is similar and will not be repeated here.
[0075] An embodiment of the present application provides an electronic device, and the electronic device in the embodiment of the present application includes: a processor; and a memory, the memory being configured to store machine-readable instructions, which, when executed by the processor, causes the processor to execute a human-pet intercom method based on multimodal interaction.
[0076] The present application embodiment provides an electronic device, such as Figure 5 As shown, Figure 5 The electronic device shown includes a processor 2001 and a memory 2003. The processor 2001 and the memory 2003 are connected, for example, via a bus 2002. Optionally, the electronic device 2000 may further include a transceiver 2004. It should be noted that in practical applications, the number of transceivers 2004 is not limited to one, and the structure of the electronic device 2000 does not constitute a limitation on the embodiments of the present application.
[0077] Processor 2001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 2001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0078] The bus 2002 may include a path for transmitting information between the above components. The bus 2002 may be a PCI bus or an EISA bus, etc. The bus 2002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0079] The memory 2003 may be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, a CD-ROM or other optical disk storage, an optical disc storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0080] The memory 2003 is used to store the application code for executing the solution of the present application, and the execution is controlled by the processor 2001. The processor 2001 is used to execute the application code stored in the memory 2003 to implement Figure 3 and Figure 4 The illustrated embodiment provides an action of a human-pet intercom device based on multimodal interaction.
[0081] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0082] The above descriptions are only partial embodiments of the present invention. It should be pointed out that ordinary technicians in this technical field can make several improvements and modifications without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for human-pet intercom based on multimodal interaction, characterized in that: The method is executed by a pet device, and includes: receiving a conversation signal sent by a user device to a target pet, and determining animal voice information corresponding to the conversation signal based on a local language library, wherein the local language library includes a correspondence between each conversation signal and animal voice information; Acquiring real-time reflection data of the target pet in response to the dialogue signal, wherein the real-time reflection data includes voice data, motion image data, and environmental data of the target pet; The real-time reflection data is input into a multimodal translation model to obtain a translation result of the target pet's response to the dialogue signal, and the translation result is sent to the user device.
2. The method according to claim 1, characterized in that The multimodal translation model includes an input layer, an encoder, and a decoder. Inputting the real-time reflection data into the multimodal translation model to obtain a translation result of the target pet's response to the dialogue signal includes: Based on the input layer, feature extraction processing is performed on the sound data, action image data and environment data in the real-time reflection data to obtain sound features corresponding to the sound data, image features corresponding to the action image data and environment features corresponding to the environment data. Performing cross-modal semantic fusion processing on the sound features and the image features based on the encoder to obtain fused features; The fused features and the environmental features are input into the decoder, so that the decoder performs dynamic translation generation processing according to the fused features and the environmental features to obtain a translation result of the target pet's response to the dialogue signal.
3. The method according to claim 2, characterized in that The encoder includes a dynamic cross-modal attention layer, a sparse local attention layer, and adaptive noise suppression. The encoder performs cross-modal semantic fusion processing on the sound features and the image features to obtain fused features, including: Determining each local window corresponding to the sound feature based on the sparse local attention layer, and integrating the features of each local window through a skip connection to obtain an optimized sound feature; Based on the adaptive noise suppression layer, performing noise reduction processing on the optimized sound features and the image features to obtain noise-reduced sound features and noise-reduced image features; The dynamic cross-modal attention layer performs self-attention fusion calculation on the denoised sound features and the denoised image features to obtain fused features.
4. The method according to claim 3, characterized in that The step of performing noise reduction processing on the optimized sound features and the image features based on the adaptive noise suppression layer to obtain noise-reduced sound features and noise-reduced image features includes: determining, based on the adaptive noise suppression layer, background noise regions respectively included in the optimized sound feature and the image feature; The features corresponding to the background noise area are shielded from the optimized sound features and the image features to obtain the noise-reduced sound features and the noise-reduced image features.
5. The method according to claim 3, characterized in that The self-attention fusion calculation is performed on the denoised sound features and the denoised image features by the dynamic cross-modal attention layer to obtain fused features, including: Determining, by the dynamic cross-modal attention layer, an attention weight of the sound data and an attention weight of the action image data according to the sound features and the image features; According to the attention weight of the sound data and the attention weight of the action image data, a self-attention fusion calculation is performed on the denoised sound features and the denoised image features to obtain fused features.
6. The method according to claim 2, characterized in that The decoder includes a dynamic weight allocation layer and a decoding layer. The decoder obtains the translation result of the target pet's response to the dialogue signal in the following manner: Obtaining a task label, and the dynamic weight allocation layer adjusting the attention weight based on the task label to obtain an adjusted attention weight, wherein the adjusted attention weight includes the attention weight of the fused feature and the attention weight corresponding to the environmental data; performing task optimization processing on the environmental features and the fused features according to the adjusted attention weights to obtain a task-optimized feature vector; The decoding layer maps the task-optimized feature vector to a vocabulary space to obtain the translation result.
7. The method according to claim 1, characterized in that The method further comprises: Receiving a location check signal sent by the user device to the target pet; The real-time accurate location information and the historical movement trajectory of the target pet are obtained, and the real-time accurate location information and the historical movement trajectory are sent to the user device.
8. The method according to claim 1, characterized in that The method further comprises: receiving an environment checking signal sent by the user device to the target pet; Automatically turning on the included camera and acquiring image data of the environment scene where the target pet is located through the camera; The image data is sent to the user equipment end, where the image data includes at least one of video data and picture data.
9. The method according to claim 8, characterized in that After sending the image data to the user equipment, the method further includes: receiving an alarm signal sent by the user device to the target pet, and initiating an alarm process corresponding to the alarm signal; The alarm processing includes at least one of playing a preset human distress voice, turning on an alarm indicator light, and sending a distress signal.
10. A human-pet intercom system based on multimodal interaction, characterized in that: The human-pet intercom system includes a main control system module, a positioning module, an information transmission module, a memory module, a power module, a camera, a microphone and a speaker outlet. The human-pet intercom system is used to execute any one of the methods in claims 1-9.
Citation Information
Patent Citations
Image information input method, electronic equipment and computer readable storage medium
CN111881315A
Pet identification method and device, equipment and storage medium
CN113673487A
Human and animal situation language intelligent communication system
CN117711369A
Pet identity recognition method and device and storage medium
CN118942124A
Pet voice translation method and system, electronic equipment and storage medium
CN119626263A
Cited By
A pet bidirectional translation method and system based on end-cloud cooperation and archive enhancement
CN122655793A