Voice dialogue method and device based on animal behaviors and wearable equipment
By collecting and analyzing animal movement data and environmental sounds in real time, personalized dialogue information is generated, which solves the problem of lack of emotional communication in existing animal interaction products, provides a more realistic and intelligent interactive experience, and reduces device power consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING XINGZHE WUJIANG INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, animal-user interaction products lack emotional communication and cannot recognize animal behavior and environmental sounds, resulting in insufficient accuracy and realism in the interaction.
By collecting animal movement data and environmental sounds in real time, analyzing behavioral states and parsing speech information, dialogue information based on speech information and behavioral states is generated, and dialogue triggering and termination mechanisms are introduced, combined with personalized models and active and passive interaction methods.
It enables more realistic and intelligent animal-user interaction, improves the accuracy and emotional connection of the interaction, reduces device power consumption, and extends battery life.
Smart Images

Figure CN121963731A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of smart wearable device technology, and in particular relates to a voice dialogue method, device and wearable device based on animal behavior. Background Technology
[0002] In recent years, with the rapid development of artificial intelligence, voice recognition, cloud computing, and the Internet of Things, smart products for animals have gradually increased. For example, various smart animal feeders, remote monitoring cameras, animal voice toys, and smart voice-interactive robots have appeared on the market. These products have improved the way users communicate with their animals to some extent.
[0003] However, these existing technologies primarily focus on "functional interactions" (such as feeding, location tracking, camera recording, and voice playback), failing to truly achieve "emotional communication" between animals and users. Most products lack motion sensing or environmental detection capabilities, making it impossible to recognize the animal's surroundings and behavioral state, thus affecting the accuracy and authenticity of the interaction. For example, even when an animal is resting or eating, random voice prompts or accidental interactions may still be played, disrupting the genuine companionship atmosphere. Summary of the Invention
[0004] This application provides a voice dialogue method, device, and wearable device based on animal behavior, aiming to solve the problems of accuracy and authenticity of interaction caused by the inability of existing animal behavior-based voice dialogue methods to combine with real-world situations.
[0005] To address the aforementioned technical problems, in a first aspect, this application provides a voice dialogue method based on animal behavior, the method comprising: Real-time collection of animal movement data and ambient sounds around the animals; The motion data is analyzed to obtain the animal's behavioral state; The ambient sound is analyzed to obtain the speech information contained in the ambient sound; When a dialogue trigger command is detected, a dialogue program is executed; the dialogue trigger command is used to indicate that the acquired environmental sound, including dialogue trigger voice information and / or the animal's behavioral state, is the target behavioral state; the dialogue program is used to generate first dialogue information based on the acquired voice information and the animal's behavioral state.
[0006] Secondly, this application provides a voice dialogue device based on animal behavior, the device comprising: The data acquisition module is configured to collect animal movement data and ambient sounds around the animal in real time. The first processing module is configured to analyze the motion data to obtain the animal's behavioral state; The second processing module is configured to parse the ambient sound to obtain the speech information contained in the ambient sound; The third processing module is configured to execute a dialogue program when a dialogue trigger instruction is detected; the dialogue trigger instruction is used to indicate that the acquired ambient sound includes dialogue trigger voice information and / or the animal's behavioral state as the target behavioral state; the dialogue program is used to generate first dialogue information based on the acquired voice information and the animal's behavioral state.
[0007] Thirdly, this application provides a wearable device worn on an animal, the wearable device comprising: at least one motion sensor, at least one audio sensor, at least one processor, and at least one memory, wherein... The at least one motion sensor is configured to collect animal motion data in real time; The at least one audio collector is configured to collect environmental sounds around the animal in real time and encode the acquired environmental sounds to obtain corresponding audio data. The at least one memory stores computer-readable instructions, which are executed by the at least one processor to enable the wearable device to perform the following: The motion data is analyzed to obtain the animal's behavioral state; The ambient sound is analyzed to obtain the speech information contained in the ambient sound; When a dialogue trigger command is detected, a dialogue program is executed; the dialogue trigger command is used to indicate that the acquired environmental sound, including dialogue trigger voice information and / or the animal's behavioral state, is the target behavioral state; the dialogue program is used to generate first dialogue information based on the acquired voice information and the animal's behavioral state.
[0008] Fourthly, this application provides a storage medium storing computer-readable instructions that are executed by one or more processors to implement the animal behavior-based voice dialogue method described above.
[0009] This application provides a voice dialogue method, device, wearable device, and storage medium based on animal behavior. By analyzing the animal's behavioral state and audio information in the environment in real time, it generates dialogue information based on voice information and behavioral state when a dialogue trigger command is detected. This solves the problems of stiff interaction and lack of context awareness in the prior art. It has the ability to generate dialogue information based on voice information and behavioral state when a dialogue trigger command is detected by analyzing the animal's behavioral state and voice information in the environment in real time, thus providing a more realistic and intelligent interactive experience. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 The embodiments provided in this application illustrate the main flow diagram of a voice dialogue method based on animal behavior; Figure 2 for Figure 1 The diagram shows a sub-process of the embodiment shown. Figure 3 for Figure 1 The diagram shows a sub-process of the embodiment shown. Figure 4 for Figure 1 The diagram shows a sub-process of the embodiment shown. Figure 5 A schematic block diagram of a voice dialogue device based on animal behavior provided in this application; Figure 6 A schematic block diagram of a wearable device provided in this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0015] This application provides a voice dialogue method based on animal behavior. This method is applicable to wearable devices. In the following method embodiments, for ease of description, the execution subject of each step of the method is a wearable device as an example, but this does not constitute a specific limitation.
[0016] like Figure 1As shown in this exemplary embodiment, the voice dialogue method based on animal behavior may include the following steps S110-S140: S110: Real-time acquisition of animal movement data and ambient sounds around the animals.
[0017] First, real-time data collection of animal movement and surrounding environmental sounds is performed. Movement data can be collected by attaching small sensors to the animal. It's important to understand that movement data refers to real-time information about the animal's physical activity acquired through sensor devices, such as acceleration, angular velocity, and changes in position. This data reflects the animal's movement patterns and intensity. Ambient sounds refer to various sound wave signals present in the space surrounding the animal, including but not limited to human speech, other animal calls, natural environmental sounds, and mechanical noises, which are collected using devices such as microphones.
[0018] For example, inertial measurement units such as accelerometers and gyroscopes in motion sensors can be used to acquire the animal's acceleration and angular velocity information. Ambient sound can be captured by placing microphone arrays within the animal's activity area or on a wearable device to record the acoustic environment surrounding the animal. Real-time acquisition of these data streams provides the foundation for subsequent analysis and dialogue generation.
[0019] S120. Analyze the motion data to obtain the animal's behavioral status.
[0020] Motion data is analyzed to obtain the animal's behavioral state. Various algorithms can be used for motion data analysis. For example, thresholds can be set to distinguish between stationary and moving states, or more complex behavioral patterns, such as running, jumping, and scratching, can be identified by analyzing the frequency and amplitude characteristics of the motion data. These analytical results are mapped to predefined behavioral state labels, such as sleeping, walking, eating, drinking, running, and playing.
[0021] S130. Analyze the ambient sound to obtain the speech information contained in the ambient sound.
[0022] Environmental sound analysis involves extracting the speech information contained within the sound. This analysis can include speech recognition technology, which converts human speech into text; it can also include voiceprint recognition technology, which identifies the source of a specific sound; or it can determine the type of sound by analyzing its spectral characteristics, such as animal calls, music, or environmental noise. This extracted audio information provides a basis for understanding the environment in which animals live.
[0023] S140. When a dialogue trigger command is detected, the dialogue program is executed.
[0024] The dialogue trigger instruction is used to indicate that the acquired environmental sound, including dialogue trigger voice information and / or the animal's behavioral state, is the target behavioral state; the dialogue program is used to generate first dialogue information based on the voice information in the audio information and the animal's behavioral state.
[0025] Furthermore, when a dialogue trigger command is detected, the dialogue program is executed. The detection of dialogue trigger commands can be based on various methods. A dialogue trigger command refers to a specific signal or condition that initiates the voice dialogue program, such as detecting specific voice keywords, gestures, or specific behavioral patterns of an animal. For example, it can be set to trigger when the volume of ambient sound exceeds a specific threshold, or when the animal's movement data exhibits a specific pattern within a specific time period. Alternatively, it can be triggered by specific signals sent from external devices. Once a command that meets the conditions is detected, the preset dialogue program is initiated.
[0026] Finally, the dialogue program refers to a pre-defined algorithm and logical flow used to receive input information, process it, and generate corresponding dialogue output. This module selects or generates appropriate dialogue content based on the currently parsed voice information (e.g., recognizing the user's "hello") and the animal's behavioral state (e.g., the animal is in a "playing" state). For example, if the animal is in a "resting" state and soft music is detected in the ambient sound, the dialogue program might generate a soothing voice message. If the animal is in an "excited" state and the user is calling its name, the dialogue program might generate an encouraging interactive voice message.
[0027] Imagine a home environment where a user equips their dog with a wearable device. This device continuously collects the dog's movement data and ambient sounds in real time. For example, when the dog roams freely in the living room, the device's motion sensors record its running, jumping, and tail wagging movements, while the microphone array captures sounds from conversations between the user and family members, television playback, and occasional barks from the dog.
[0028] The processor in the device continuously analyzes the collected data. First, motion data is analyzed to determine the dog's behavioral state. For example, if the motion data shows high-intensity displacement and acceleration changes in a short period, the system may identify its behavior as "active" or "playful." If the motion data shows prolonged periods of low-intensity or no displacement, it may be identified as "resting" or "sleeping." Simultaneously, ambient sound is analyzed to determine the audio information it contains. For example, using voice recognition technology, the system may recognize a user's spoken message such as, "Puppy, come play!"
[0029] When the system detects a dialogue trigger command, such as recognizing a user's voice command, "Puppy, come play!", the dialogue program is executed. At this point, the dialogue program considers both the currently parsed voice information (the content of the voice command "Puppy, come play!") and the animal's behavioral state (e.g., the animal is currently in an "active" state). Based on this information, the dialogue program generates initial dialogue information. For example, it can generate a voice simulating an excited bark from an animal, accompanied by a synthesized voice message, "The user is calling me to play, I'm so happy!", which is then played through the device's built-in speaker.
[0030] For example, when human voice information is detected in the ambient sound and specific data is present in the motion data, such as acceleration exceeding a certain value or vertical displacement exceeding a certain distance, and the analysis indicates that the animal's behavioral state is that of being picked up, then the dialogue program is executed. Furthermore, the dialogue program can interact based on subsequent voice information.
[0031] In this way, the method can analyze and obtain the animal's behavioral state and the interaction information of the interacting object based on the collected motion and sound data. Based on the animal's actual behavioral state and the voice information in the environment, it generates dialogue content with context awareness and emotional relevance, thereby achieving a more natural and realistic interaction.
[0032] In this embodiment, a dialogue program is proposed to generate dialogue information. However, in this process, the dialogue program may not end automatically, resulting in wasted resources and unnatural interaction. For example, it may continue to run even after the dialogue has ended or when the ambient sound is below the threshold for a long time, causing ineffective consumption of computing resources and a decline in user experience.
[0033] In this regard, such as Figure 3 As shown, in some embodiments of this example, step S150 is also included: S150. When a dialogue end command is detected or no ambient sound exceeding the set volume threshold is detected within a set time, the dialogue program ends.
[0034] The dialogue end instruction is a signal or command used to explicitly indicate the termination of a dialogue process. This instruction can be sent by the user through an external device (e.g., a smartphone application, a remote control) or automatically generated internally by the system based on specific logic (e.g., reaching a preset limit on the number of dialogue rounds, recognizing a specific ending keyword). The absence of ambient sound exceeding a set volume threshold within a set time serves as an automatic mechanism for determining the end of the dialogue.
[0035] Specifically, the system continuously monitors the volume of ambient sounds. If the average or peak volume of the ambient sounds does not reach a preset volume threshold within a preset time window, the conversation is considered to have ended. Alternatively, it can incorporate speech recognition technology; if no valid human speech or animal sounds are recognized within a set time, the conversation is considered to have ended even in the presence of low-volume background noise.
[0036] Terminating a dialog program means ending its operation, thereby freeing up system resources. This can include stopping all child processes of the dialog program, such as speech recognition, behavior analysis, and dialogue information generation, or putting the dialog program into a dormant state, waiting for a new trigger command.
[0037] For example, suppose you are having a conversation with an animal. When the user sends a "end conversation" command via a smartphone app connected to the device, or directly says "end, goodbye" in voice, the device receives and recognizes this command as an end-of-conversation instruction, and then terminates the current conversation. Alternatively, in another scenario, after the animal and user have had a conversation, the user leaves, and the animal becomes quiet. In this case, the device continuously monitors ambient sound. If the average volume of ambient sound remains below 35 decibels (a set volume threshold) for 10 consecutive seconds (a set time), the device determines that the conversation has ended naturally and automatically stops the conversation.
[0038] The system collects animal movement data and ambient sounds in real time, analyzing this data to determine the animal's behavioral state and audio information. When a dialogue trigger command is detected, the dialogue program starts and generates initial dialogue information based on the voice information in the audio and the animal's behavioral state. To ensure intelligent management and resource optimization of the dialogue program, this solution introduces a dialogue program termination mechanism. The system continuously monitors for explicit dialogue end commands, allowing users to actively intervene and terminate the dialogue, thus ensuring the controllability of the interaction process. Simultaneously, the system intelligently monitors ambient sounds. If the ambient sound volume remains below a preset threshold for a set period of time, it indicates that there may be no effective dialogue interaction in the current environment, and the system will automatically determine that the dialogue has ended naturally. Once either of the above conditions is met—either a dialogue end command is detected or no ambient sound exceeding the set volume threshold is detected within the set time—the system immediately terminates the dialogue program. This mechanism ensures that the dialogue program runs when necessary and terminates promptly when there is no need for interaction, effectively avoiding unnecessary resource consumption and reducing device power consumption.
[0039] like Figure 2As shown, in some embodiments of this example, the dialogue program is used to generate first dialogue information based on the voice information in the audio information and the animal's behavioral state, specifically including steps S141 to S142: S141. Generate an initial dialogue scheme based on the acquired voice information and the animal's behavioral state; S142. Input the initial dialogue scheme into the pre-trained animal personalization model to generate the first dialogue information.
[0040] Among them, the personalized animal model is trained based on historical movement data and animal calls from audio information.
[0041] A "pre-trained personalized animal model" refers to an intelligent model trained on a large amount of data, capable of understanding and simulating specific animal behaviors and vocalization patterns. This model can be a deep learning model, such as a recurrent neural network (RNN), long short-term memory network (LSTM), or a Transformer model, or it can be a model built based on statistical learning or expert systems. It constructs an animal personality profile (such as activity level, attachment level, appetite index, etc.) based on historical behavioral data, and continuously updates the animal's personalized characteristic parameters by combining historical behavioral and vocal data. This utilizes the animal's long-term accumulated movement data (e.g., activity level, sleep patterns, play habits, etc.) and its unique vocalization data (e.g., the pitch, frequency, and duration of barks, meows, and purrs in different contexts). Through this historical data, the model can learn the animal's unique behavioral patterns, emotional expressions, and responses to specific stimuli, thus making its generated dialogue more consistent with the animal's personality.
[0042] This application's solution effectively enhances the personalization of dialogue by combining real-time contextual awareness with individual historical experience. The final dialogue information not only accurately responds to the animal's current immediate state but also incorporates its unique "personality," making the interaction between humans and animals more realistic and emotionally profound.
[0043] like Figure 4 As shown in an exemplary embodiment, the animal behavior-based voice dialogue method of this application further includes step S160: S160. When the target external device is detected to be within the set range, a second dialogue message is actively generated.
[0044] "Detecting a target external device within the set range" means that the system can identify and confirm that a preset external device has entered the preset geographical or communication distance range.
[0045] Specifically, the detection of target external devices within a set range can be achieved through various methods. For example, short-range wireless communication technologies such as Bluetooth, Wi-Fi Direct, and NFC can be used to determine whether the target external device has entered the set range by monitoring its signal strength or connection status; when the signal strength reaches a preset threshold or a connection is successfully established, the device is considered to have entered the set range. Furthermore, positioning technologies such as GPS and UWB can be used to obtain the distance between the target external device and the device in real time; when the calculated distance is less than or equal to a preset distance, the device is also considered to have entered the set range.
[0046] "Proactive generation of secondary dialogue information" refers to the system autonomously creating and outputting dialogue content without receiving external instructions when specific conditions are met (such as detecting a target external device within a set range). This feature aims to enhance the initiative and timeliness of interaction, making communication between animals and users more natural. In specific implementation, multiple dialogue templates or phrases can be preset, and secondary dialogue information can be generated by selecting and combining them from a preset library based on contextual information such as the detected external device type, time, and environment. For example, in a specific application scenario, when a user returns home from outdoors, the wearable device detects that the user's mobile phone's Bluetooth is automatically connected. When the user is detected to be home, the collar automatically triggers proactive voice output, such as "The user is back! I'm a little tired from running!" This feature, as a condition for triggering proactive dialogue, aims to ensure the timeliness and relevance of the interaction.
[0047] By introducing a technical solution that proactively generates second dialogue information when a target external device is detected within a set range, this application effectively solves the problem of lack of initiative and timeliness in the interaction of existing technologies, thereby providing a more intelligent and humanized companion service. Furthermore, by combining proactive and passive dialogue, the overall interactivity is improved.
[0048] like Figure 5 As shown, this application also proposes an embodiment of a voice dialogue device 100 based on animal behavior. The device includes a data acquisition module 101, a first processing module 102, a second processing module 103, and a third processing module 104. The data acquisition module 101 is configured to acquire animal movement data and ambient sounds around the animal in real time. The first processing module 102 is configured to analyze the movement data to obtain the animal's behavioral state. The second processing module 103 is configured to parse the ambient sounds to obtain the speech information contained within them. The third processing module 104 is configured to execute a dialogue program when a dialogue trigger command is detected. The dialogue trigger command indicates that the acquired ambient sounds including dialogue trigger speech information and / or the animal's behavioral state are the target behavioral state. The dialogue program generates first dialogue information based on the acquired speech information and the animal's behavioral state.
[0049] In an optional embodiment, the device further includes a fourth processing module configured to actively generate second dialogue information when a target device is detected to be within a set range.
[0050] In an optional embodiment, the device further includes a fifth processing module configured to terminate the dialogue program when a dialogue end command is detected or when no ambient sound exceeding a set volume threshold is detected within a set time.
[0051] In an optional embodiment, the device further includes a sixth processing module configured to generate an initial dialogue scheme based on the acquired voice information and the animal's behavioral state; input the initial dialogue scheme into a pre-trained animal personalization model to generate first dialogue information; wherein the animal personalization model is trained based on historical movement data and animal calls in audio information.
[0052] like Figure 6 As shown, this application also proposes a wearable device 20, which includes at least one motion sensor, at least one audio acquisition unit 209, at least one processor, and at least one memory.
[0053] In this device, at least one motion sensor is configured to collect animal motion data in real time; at least one audio collector 209 is configured to collect ambient sounds around the animal in real time and encode the acquired ambient sounds to obtain corresponding audio data; at least one memory stores computer-readable instructions, which are executed by at least one processor to enable the wearable device 20 to perform the following method: Analyze motion data to obtain the animal's behavioral status; Analyze ambient sounds to obtain the audio information contained within them; When a dialogue trigger command is detected, the dialogue program is executed; the dialogue program is used to generate the first dialogue information based on the voice information in the audio information and the animal's behavioral state.
[0054] Specifically, the wearable device 20 is an electronic collar worn around the neck of an animal. The motion sensor is a six-axis sensor 205, and the audio acquisition unit 209 includes a microphone array 208, a speaker 207, and an audio converter 204 (ADC converter, DAC converter). The six-axis sensor 205 is an inertial measurement unit that integrates a three-axis accelerometer and a three-axis gyroscope to measure motion parameters in six degrees of freedom. In some other embodiments, the wearable device can also be an electronic device worn on other parts of the animal's body, such as its legs, head, or torso. Other motion sensors, such as three-axis or nine-axis sensors, can also be used. By analyzing and processing the real-time motion data and ambient sound collected by the six-axis sensor 205 and the microphone array 208 in a multimodal data fusion manner, the animal's behavioral state and environmental audio information can be accurately identified. This achieves the effect of generating context-appropriate dialogue content based on behavioral state and voice information only when a dialogue trigger command is detected. Specifically, the wearable device 20 is worn directly on the animal to ensure accurate collection of motion data; the six-axis sensor 205 monitors the animal's acceleration, angular velocity and other motion parameters in real time to analyze behavioral states such as resting, playing or eating; the microphone array 208 captures ambient sound and extracts voice information from the audio information through speech recognition technology; when a dialogue trigger command is detected, the dialogue program integrates the behavioral state and voice information to generate the first dialogue information, avoiding triggering interaction in inappropriate situations, and the generated dialogue information can be played through the speaker 207.
[0055] In one embodiment, the wearable device 20 includes a SOC chip 200, which integrates a main MCU 201 and a low-power MCU 202. The main MCU 201 is connected to a communication unit, which includes a WIFI module 2031 and a Bluetooth module 2032. A microphone 208 and a speaker 207 are connected to an audio converter 204 via I2S and then to the MCU. The microphone 208 is also connected to a comparator amplifier 206 and, together with a six-axis sensor 205, to the low-power MCU 202.
[0056] The low-power MCU 202 includes at least one processor, which is used to send a wake-up command to the main MCU 201 to switch the main MCU 201 from sleep state to running state when the main MCU 201 is in sleep state and the motion data exceeds a set value or the ambient sound exceeds a set volume value.
[0057] The main MCU 201 includes at least one processor for analyzing the motion data to obtain the animal's behavioral state; parsing the environmental sound to obtain the audio information contained in the environmental sound; and executing a dialogue program when a dialogue trigger command is detected; the dialogue trigger command is used to indicate that the acquired environmental sound includes dialogue trigger voice information and / or the animal's behavioral state as the target behavioral state; the dialogue program is used to generate first dialogue information based on the acquired voice information and the animal's behavioral state.
[0058] The wearable device 20 can also actively generate second dialogue information when it detects that the communication unit is automatically connected to the target terminal and the target terminal is no more than a set distance from the wearable device 20.
[0059] This wearable device can use STMicroelectronics' STM32L0 series microcontroller as the low-power MCU and the STM32F4 series microcontroller as the main MCU. When the device is worn normally on the animal, the main MCU enters a deep sleep mode to save power. During this time, the STM32L0 low-power MCU continues to operate, continuously reading accelerometer data at a low sampling rate and calculating its root mean square (RMS) value. When this value continuously exceeds a preset "activity threshold" (e.g., an RMS value greater than 0.1g for 5 consecutive seconds), or when the ambient sound energy detected by the low-power microphone continuously exceeds a preset "sound threshold" (e.g., a sound pressure level greater than 35dB for 3 consecutive seconds), the low-power MCU sends a wake-up command to the main MCU via a GPIO interrupt signal. In this embodiment, the acceleration threshold is 0.1g, and the ambient sound level is set to 35dB. Understandably, when the acceleration is less than 0.1g or the ambient noise is less than 35dB, the main MCU will be controlled to re-enter sleep mode. In sleep mode, the average power consumption of the entire device is less than 1mW.
[0060] This application effectively solves the problem of excessive power consumption caused by continuous operation of the main processor in wearable devices. By introducing a dual architecture of a main MCU and a low-power MCU, the low-power MCU wakes up the main MCU by monitoring data thresholds, achieving low-power operation. The low-power MCU is responsible for preliminary motion and sound monitoring during the main MCU's sleep period, waking up the main MCU only when a potential interaction request is detected, thus avoiding unnecessary long-term operation of the main MCU. This significantly reduces the overall power consumption of the device and extends battery life.
[0061] In one embodiment, the wearable device 20 is connected to an external server via a communication unit. The wearable device 20 can control the communication unit to establish a connection with the server when a dialogue trigger command is detected. The server is equipped with a pre-trained animal personalization model. This model generates an initial dialogue scheme based on acquired voice information and the animal's behavioral state. The initial dialogue scheme is then input into the pre-trained animal personalization model to generate the first dialogue information. The animal personalization model is trained based on historical movement data and animal calls from historical environmental sounds.
[0062] The communication unit is a hardware module that enables data exchange between the wearable device and external devices. Its function is to achieve wireless data transmission between the wearable device and the target server, which is crucial for remote processing and receiving of dialogue information. This communication unit can be a short-range wireless communication module, such as a Bluetooth or Wi-Fi module, used to connect to nearby gateway devices or smartphones, and then communicate with the target server via the internet. Alternatively, the communication unit can be a cellular mobile communication module, such as a 4G or 5G module, allowing the wearable device to communicate directly with the target server via a mobile network without intermediate devices. The external device, including the target server, is a remote computer system with powerful computing and storage capabilities, used to receive data sent by the wearable device and perform complex computational tasks. Its role is to host and run complex personalized animal models, process audio information and behavioral state data from the wearable device, and generate initial dialogue information, thereby compensating for the insufficient computing resources of the local device. The target server can be a server cluster based on a cloud computing architecture.
[0063] This application also proposes a storage medium, in an exemplary embodiment, storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set, when executed by a processor of a computer device, implements the animal behavior-based voice dialogue method of any of the above embodiments.
[0064] It should be understood that, in the embodiments of this application, the processor may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0065] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the animal behavior-based voice dialogue method described above.
[0066] In summary, the animal behavior-based voice dialogue method, apparatus, and wearable device provided in this application analyze the animal's behavioral state and audio information in the environment in real time. Upon detecting a dialogue trigger command, it generates dialogue information based on voice information and behavioral state, providing a more realistic and intelligent interactive experience. A personalized model trained on the target animal's historical movement data and vocalizations is introduced, making the generated dialogue information more closely match the individual animal characteristics. Furthermore, both active and passive dialogue triggering mechanisms are employed, further improving the flexibility of the interaction. The wearable device adopts a dual architecture of a main MCU and a low-power MCU. The low-power MCU monitors data thresholds to wake up the main MCU, achieving low-power operation. The wearable device interacts with a cloud server via a dual-mode WIFI / Bluetooth communication unit, utilizing the server's computing power to run the personalized model and generate dialogue information, compensating for local computing limitations and possessing a greater ability for continuous learning.
[0067] It is understood that this application has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of this application. Furthermore, based on the teachings of this application, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of this application. Therefore, this application is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of this application.
Claims
1. A voice dialogue method based on animal behavior, characterized in that, The method includes: Real-time collection of animal movement data and ambient sounds around the animals; The motion data is analyzed to obtain the animal's behavioral state; The ambient sound is analyzed to obtain the speech information contained in the ambient sound; When a dialogue trigger command is detected, a dialogue program is executed; the dialogue trigger command is used to indicate that the acquired environmental sound, including dialogue trigger voice information and / or the animal's behavioral state, is the target behavioral state; the dialogue program is used to generate first dialogue information based on the acquired voice information and the animal's behavioral state.
2. The method according to claim 1, characterized in that, The method further includes: when the target device is detected to be within a set range, actively generating second dialogue information.
3. The method according to claim 1, characterized in that, The method further includes: terminating the dialogue program when a dialogue end command is detected or when no ambient sound exceeding a set volume threshold is detected within a set time.
4. The method according to claim 1, characterized in that, The dialogue program is used to generate first dialogue information based on the acquired voice information and the animal's behavioral state, including: An initial dialogue plan is generated based on the acquired voice information and the animal's behavioral state; The initial dialogue scheme is input into a pre-trained animal personalization model to generate the first dialogue information; wherein the animal personalization model is trained based on historical movement data and animal calls in historical environmental sounds.
5. A voice dialogue device based on animal behavior, characterized in that, The device includes: The data acquisition module is configured to collect animal movement data and ambient sounds around the animal in real time. The first processing module is configured to analyze the motion data to obtain the animal's behavioral state; The second processing module is configured to parse the ambient sound to obtain the speech information contained in the ambient sound; The third processing module is configured to execute a dialogue program when a dialogue trigger instruction is detected; the dialogue trigger instruction is used to indicate that the acquired ambient sound includes dialogue trigger voice information and / or the animal's behavioral state as the target behavioral state; the dialogue program is used to generate first dialogue information based on the acquired voice information and the animal's behavioral state.
6. A wearable device, said wearable device being worn on an animal, characterized in that, include: The system comprises at least one motion sensor, at least one audio acquisition unit, at least one memory, a main MCU, and a low-power MCU, wherein both the main MCU and the low-power MCU include at least one processor. The at least one motion sensor is configured to collect animal motion data in real time; The at least one audio collector is configured to collect environmental sounds around the animal in real time and encode the acquired environmental sounds to obtain corresponding audio data. The at least one memory stores computer-readable instructions, which are executed by at least one processor of the low-power MCU to enable the wearable device to: When the main MCU is in sleep mode, if the motion data exceeds a set value or the ambient sound exceeds a set volume value, a wake-up command is sent to the main MCU to switch the main MCU from sleep mode to running mode. The computer-readable instructions are executed by at least one processor of the main power MCU, enabling the wearable device to: The motion data is analyzed to obtain the animal's behavioral state; The audio data is parsed to obtain the speech information contained in the ambient sound. When a dialogue trigger command is detected, a dialogue program is executed; the dialogue trigger command is used to indicate that the acquired environmental sound, including dialogue trigger voice information and / or the animal's behavioral state, is the target behavioral state; the dialogue program is used to generate first dialogue information based on the acquired voice information and the animal's behavioral state.
7. The wearable device according to claim 6, characterized in that, The wearable device further includes a communication unit for communicating with an external device, the external device including a target terminal. The computer-readable instructions are executed by the at least one processor, enabling the wearable device to also achieve: When it is detected that the communication unit automatically connects to the target terminal and the target terminal is no more than a set distance from the wearable device, a second dialogue message is actively generated.
8. The wearable device according to claim 6, characterized in that, The computer-readable instructions are executed by the at least one processor, enabling the wearable device to also perform: The dialogue program ends when a dialogue end command is detected or when no ambient sound exceeding a set volume threshold is detected within a set time.
9. The wearable device according to claim 6, characterized in that, The wearable device further includes a communication unit for communicating with an external device, including a target server. The computer-readable instructions are executed by the at least one processor, enabling the wearable device to also perform the following: When a dialogue trigger command is detected, the control communication unit establishes a connection with the target server to send voice information to the target server; In response to sending voice information to the target server, the system receives first dialogue information returned by the target server; wherein the target server is configured to: An initial dialogue scheme is generated based on the acquired voice information and the animal's behavioral state; The initial dialogue scheme is input into a pre-trained animal personalization model to generate the first dialogue information; wherein the animal personalization model is trained based on historical movement data and animal calls in historical environmental sounds.
10. The wearable device according to claim 6, characterized in that, The wearable device includes an electronic collar worn around the neck of an animal.