An electronic pet behavior triggering method based on sound perception

By constructing sound feature vectors using a microphone array for multi-category analysis and self-noise control, the problem of insufficient comprehensive analysis of multi-dimensional acoustic features and self-noise interference in electronic pet devices is solved, resulting in more natural and stable behavioral responses.

CN122455009APending Publication Date: 2026-07-24BEIJING AH LATIN INTERNATIONAL CULTURAL DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING AH LATIN INTERNATIONAL CULTURAL DEVELOPMENT CO LTD
Filing Date
2026-05-18
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing electronic pet devices lack comprehensive analysis of multi-dimensional acoustic characteristics, making it difficult to distinguish different sound categories and reflect their differentiated impact on internal states. Furthermore, self-noise interference affects the accuracy and stability of behavior triggering.

Method used

Sound signals are collected through a microphone array, sound feature vectors are constructed, multi-category analysis is performed, and dynamic updates are combined with electronic pet status parameters to generate priority behavior commands, and self-noise control is performed during execution.

Benefits of technology

It achieves accurate recognition of user emotions, command intentions, and sound events, improving the naturalness, intelligence, and robustness of electronic pet behavior responses, reducing self-noise interference, and enhancing the accuracy of behavior triggering and overall system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122455009A_ABST
    Figure CN122455009A_ABST
Patent Text Reader

Abstract

The application discloses a sound perception-based electronic pet behavior triggering method, and relates to the technical field of intelligent interaction.The method comprises the following steps: collecting environmental sound signals through a microphone array and constructing a sound feature vector, determining sound intensity levels, sound source directions, and user voice classes, environmental event sound classes, and third-party biological sound classes; analyzing different sound classes to obtain user instruction intentions, user emotion labels, and sound event types; dynamically updating in combination with current emotion parameters and energy parameters of the electronic pet; generating priority behavior instructions based on the above information and sending the priority behavior instructions to an action execution module; and monitoring the short-term energy of the electronic pet in real time during the execution process and performing self-noise control.The technical problems of single sound processing, rigid state updating, serious self-noise interference, and lack of continuity of behavior response in the life state of the existing electronic pet are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent interaction technology, and in particular to a method for triggering electronic pet behavior based on sound perception. Background Technology

[0002] With the rapid development of smart terminals and artificial intelligence technologies, electronic pets are gradually becoming important interactive companion devices. Existing electronic pets typically collect environmental sound signals through microphones and trigger preset behavioral actions based on speech recognition or simple sound classification results, thereby responding to user commands or changes in the environment.

[0003] However, in practical applications, the existing technology still has the following shortcomings: On the one hand, existing solutions focus on keyword recognition of voice content, usually only matching and triggering fixed behaviors based on specific commands, lacking comprehensive analysis and utilization of multi-dimensional acoustic features such as emotional information, intensity changes, and speech rate characteristics in sound signals, resulting in electronic pets' behavioral responses lacking emotional continuity and natural consistency. On the other hand, existing technologies often employ a relatively fragmented or singular processing method for sound signals from different sources (such as user voice, environmental sounds, and third-party biological sounds), making it difficult to effectively distinguish sound categories and reflect their differentiated impact on the internal state of electronic pets, thus resulting in insufficient rationality and targeting of behavioral decisions. Furthermore, during the execution of behavioral commands, the electronic pet's own self-noise, such as the sound of the motor running, structural friction, and the output sound of the sound-generating device, is easily received by the microphone acquisition device. Existing technologies generally lack effective detection and control mechanisms for this type of self-noise, which in turn interferes with the continuous sound analysis results, affecting the accuracy of behavior triggering and the overall stability of interaction. Therefore, there is an urgent need for a new method to trigger electronic pet behavior, which can perform unified and precise feature extraction and category analysis on multiple types of sound input, combine dynamic updates of the electronic pet's internal state parameters, and achieve effective self-noise control during behavior execution, so as to significantly improve the intelligence, stability, and user companionship experience of electronic pet behavior response. Summary of the Invention

[0004] To address the aforementioned issues, this invention proposes a sound-based electronic pet behavior triggering method. By constructing sound feature vectors through a microphone array, it achieves accurate classification and analysis of user voice, environmental event sounds, and third-party biological sounds. Furthermore, it dynamically integrates and updates sound intensity levels, user command intent, user emotion tags, and sound event types with the electronic pet's emotion parameters and energy parameters to generate priority behavior commands and implement self-noise adaptive control during execution.

[0005] To achieve the above objectives, this invention provides a method for triggering electronic pet behavior based on sound perception, comprising the following steps: By using a microphone array deployed on an electronic pet device to collect ambient sound signals in real time, and performing analog-to-digital conversion, preprocessing, and acoustic feature extraction on the ambient sound signals, a sound feature vector is constructed. Primary sound information is determined based on sound feature vectors, including sound intensity level, sound source location information, and sound category; the sound category includes user voice, environmental event sound, and third-party biological sound. The environmental sound signal is analyzed and processed based on the sound category to obtain the sound category analysis result, including the user's command intent, the user's emotion label, and the sound event type. Specifically, when the sound category is the user's voice category, the user's command intent and the user's emotion label are determined. When the sound category is the environmental event sound category or the third-party biological sound category, the sound event type is determined. Obtain the current status parameters of the electronic pet, including emotion parameters and energy parameters; update the emotion parameters based on the user's emotion tags, and update the energy parameters based on the sound intensity level and sound event type to obtain the updated status parameters; Based on the initial sound information, sound category parsing results, and updated status parameters, behavioral commands are generated and sent to the electronic pet's action execution module in priority order. The action execution module executes the behavior command and collects the electronic pet's own sound signal during the execution process, and calculates the short-term energy value within the current time window. When the short-term energy value is greater than the preset self-noise judgment energy threshold, self-noise control processing is performed. When the short-term energy value is not greater than the self-noise judgment energy threshold, the execution of the behavior command is maintained or resumed.

[0006] The one or more technical solutions provided in this invention have at least the following technical effects or advantages: By extracting features and performing multi-category analysis on sound signals, accurate identification of user emotions, user command intentions, and sound event types can be achieved; on this basis, through the dynamic correlation and update of sound information and the internal state parameters of the electronic pet, the emotion parameters and energy parameters can change in real time with the sound input, thereby driving the electronic pet to exhibit reasonable behavior in different situations that conforms to its "life state," which is different from traditional voice-controlled toys that only execute command mappings; at the same time, the introduction of real-time detection and control of self-noise during the execution of behavioral commands effectively suppresses the noise generated by the electronic pet's own actions, reduces interference with sound analysis, avoids false triggering or execution chaos, and improves the accuracy of behavior triggering and the overall system stability. Compared with existing technologies, this invention significantly improves the naturalness, intelligence and robustness of electronic pet behavior response through multi-category sound perception, dynamic linkage of the pet's internal state, and self-noise control during action execution.

[0007] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a flowchart illustrating a sound-based electronic pet behavior triggering method provided in an embodiment of this application. Detailed Implementation

[0010] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0011] Example 1, as Figure 1 As shown, this application provides a method for triggering electronic pet behavior based on sound perception, wherein the method includes: S1: Real-time acquisition of ambient sound signals through a microphone array deployed on the electronic pet device, and analog-to-digital conversion, preprocessing and acoustic feature extraction of the ambient sound signals to construct a sound feature vector; Furthermore, the construction of the sound feature vector includes: The ambient sound signals are collected in real time by a microphone array deployed on the electronic pet device, and the ambient sound signals are processed by analog-to-digital conversion to obtain digital sound signals; Preprocessing of digital audio signals includes pre-emphasis, framing, windowing, and noise reduction. Acoustic features are extracted from each preprocessed digital audio signal frame, including Mel frequency cepstral coefficients, short-time energy, short-time zero-crossing rate, fundamental frequency, and sound source horizontal angle. The acoustic features corresponding to each digital audio signal frame are combined in chronological order to construct an audio feature vector.

[0012] Specifically, the system first collects ambient sound signals in real time using a microphone array deployed on the electronic pet device. The microphone array consists of multiple microphones arranged in a preset geometric structure, such as a circular array or a linear array. The collected ambient sound signals are continuous analog signals, which undergo analog-to-digital conversion: the analog signal is sampled at a preset sampling frequency, and the amplitude of each sampling point is converted into a digital value according to a preset quantization precision, thus obtaining a digital sound signal. The preferred sampling frequency is 16kHz, and the preferred quantization precision is 16-bit or 24-bit. After acquiring the digital audio signal, it undergoes preprocessing, including pre-emphasis, framing, windowing, and noise reduction. Pre-emphasis enhances the high-frequency components of the signal to improve spectral distribution; framing divides the continuous digital audio signal into multiple short frames, each 20–30 milliseconds long, making the signal within each frame approximately stable; windowing weights each frame to reduce spectral leakage, using a Hamming window; noise reduction employs spectral subtraction to subtract the estimated environmental noise power spectrum (obtained by detecting periods without speech activity) from the power spectrum of the noisy signal, yielding the denoised signal frame. After noise reduction, acoustic features are extracted from each preprocessed digital audio signal frame, including: Mel-frequency cepstral coefficients (MFCCs), obtained by converting the signal to the frequency domain using a fast Fourier transform, followed by Mel filter banks, logarithmic operations, and discrete cosine transforms, typically with 12 to 13 dimensions; short-time energy, obtained by calculating the sum of squares of the signal amplitude in each frame; short-time zero-crossing rate, obtained by counting the number of times the signal waveform in each frame crosses zero; fundamental frequency, extracted using autocorrelation or cepstral methods; and the horizontal angle of the sound source, calculated using the time difference between the received signals from each channel of the microphone array, combined with the geometric layout of the microphone array, through a geometric positioning algorithm to determine the incident direction of the sound source, thus obtaining the horizontal angle relative to the front of the electronic pet. The acoustic features corresponding to each digital audio signal frame are combined in chronological order to construct a sound feature vector, which is used to characterize the comprehensive characteristics of environmental sound in the time and frequency dimensions, and serves as the input for subsequent sound analysis processing.

[0013] S2: Determine primary sound information based on sound feature vectors, including sound intensity level, sound source location information, and sound category; the sound category includes user voice, environmental event sound, and third-party biological sound. Furthermore, the determination of primary sound information based on sound feature vectors includes: Short-time energy is extracted based on sound feature vectors, and the short-time energy of consecutive frames is statistically processed to calculate the average energy value within a preset time window. The average energy value is then compared with a preset energy threshold range to determine the sound intensity level. Obtain the horizontal angle of the sound source from the sound feature vector and use it as the sound source location information; Determining sound categories based on sound feature vectors includes: A dual threshold detection method based on short-time energy and short-time zero-crossing rate is used to classify sound signals into speech signals or non-speech signals. For voice signals, they are classified as user voice; when multiple consecutive frames are identified as voice frames, the signal segment consisting of the consecutive voice frames is identified as a valid voice segment. For non-speech signals, their acoustic features are matched with a preset environmental event feature template library and a biological sound feature template library, respectively. Based on the matching results, the sound category to which they belong is determined to be either environmental event sound or third-party biological sound.

[0014] Specifically, primary sound information, including sound intensity level, sound source location information, and sound category, is determined based on sound feature vectors. First, short-time energy is extracted from the sound feature vectors. The short-time energy of consecutive frames is statistically processed to calculate the average energy value within a preset time window. This average energy value is then compared with a preset energy threshold range to determine the sound intensity level, which is divided into five levels. Next, the horizontal angle of the sound source is read from the sound feature vector and output as the sound source location information; Then, the sound category is determined based on the sound feature vector. This determination of the sound category is divided into two levels: The first level distinguishes between speech signals and non-speech signals. Specifically, it extracts the short-time energy E and short-time zero-crossing rate Z of each signal frame from the sound feature vector. A dual-threshold speech activity detection method based on short-time energy and short-time zero-crossing rate is used to determine whether it is a speech signal or a non-speech signal. When E is greater than a preset energy threshold and Z is within a preset zero-crossing rate range, it is determined to be a speech signal; otherwise, it is determined to be a non-speech signal. When multiple consecutive frames are determined to be speech frames, they are identified as valid speech segments. The second level is used to further subdivide sound categories. For speech signals, they are uniformly classified as user speech; for non-speech signals, they are further distinguished as environmental event sounds or third-party biological sounds. Specifically, for non-speech signals, their acoustic features are extracted and matched with templates in a pre-built environmental event feature template library and a biological sound feature template library. The sound category is determined based on the matching results. The environmental event feature template library includes: doorbell sounds, whose energy envelope shows a rapid rise followed by exponential decay, and whose high-frequency components in the Mel frequency cepstral coefficients are greater than their low-frequency components; footsteps, whose short-time energy sequence shows periodic spikes, and whose corresponding fundamental frequency is in the low-frequency range; and collision sounds, whose short-time energy shows instantaneous spikes, a short-time zero-crossing rate higher than 100%, and a duration less than 50 milliseconds. The biological sound feature template library includes: dog barks, with a fundamental frequency in the low-to-mid frequency range (e.g., 100Hz-300Hz) and a regular decrease in adjacent high-frequency coefficients of the Mel-frequency cepstral coefficients; cat meows, with a fundamental frequency in the mid-frequency range (e.g., 200Hz-500Hz) and a fundamental frequency that fluctuates periodically over time; and bird calls, with a fundamental frequency in the high-frequency range (e.g., 1000Hz-4000Hz) and a peak density in the short-time energy sequence that is higher than that of speech signals. By comparing the acoustic features of the current non-speech signal with the feature template library, it is determined whether the sound category corresponding to the current non-speech signal is environmental event sound or third-party biological sound.

[0015] S3: Based on the sound category, the environmental sound signal is parsed and processed to obtain the sound category parsing result, including the user's command intent, the user's emotion label, and the sound event type; wherein, when the sound category is the user's voice category, the user's command intent and the user's emotion label are determined, and when the sound category is the environmental event sound category or the third-party biological sound category, the sound event type is determined. Furthermore, obtaining the sound category analysis result includes: When the sound category is user voice, the acoustic features of the current voice signal are matched with each keyword audio template in the preset command intent template library. The command intent corresponding to the keyword audio template with a matching score higher than the preset similarity threshold and the maximum value is determined as the user command intent. The fundamental frequency, short-time energy, and short-time zero-crossing rate are extracted from the sound feature vector, and the speech rate is determined by combining the effective speech segment duration. Based on the preset emotion judgment rules, the user's emotion label is determined. When the sound category is environmental event sound or third-party biological sound, the acoustic features of the current sound signal are matched with the feature template library of the corresponding category, and the specific sound event type is determined based on the matching result.

[0016] Specifically, based on the sound category determined in step S2, the environmental sound signal is analyzed to obtain the sound category analysis result; When the voice category is user voice, the user's command intent and user emotion tag are obtained respectively: To obtain user command intent, firstly, the acoustic features corresponding to the valid speech segment are matched with keyword audio templates in a pre-built command intent template library. This matching process uses dynamic time warping to obtain the matching score for each template by calculating the distance or similarity of the matching path. Different user command intents in the command intent template library correspond to several keyword audio templates. For example, a calling intent corresponds to keyword audio templates such as "come here" or "come quickly," an interaction intent corresponds to keyword audio templates such as "shake hands" or "raise hand," and a feeding intent corresponds to keyword audio templates such as "eat" or "feed." Then, the template with the highest matching score exceeding a threshold is selected from the matching scores, and the command intent corresponding to this template is output as the user command intent. If the matching scores of all templates are below the threshold, it is determined that the valid speech segment does not contain a valid command. To obtain user emotion labels, firstly, the fundamental frequency, short-time energy, and short-time zero-crossing rate are extracted from the sound feature vector. The duration of the speech is then calculated based on the number of frames and frame length of the effective speech segment to estimate the speech rate. Next, according to preset emotion determination rules, the fundamental frequency, short-time energy sequence, short-time zero-crossing rate sequence, and speech rate are mapped to corresponding emotion labels. Specifically: when the fundamental frequency is more than 20% higher than the user's neutral fundamental frequency and the short-time energy fluctuates drastically (i.e., the variance of the short-time energy is greater than a preset threshold), it is determined to be a happy emotion; when the fundamental frequency is more than 15% lower than the neutral fundamental frequency and the short-time energy is stable (i.e., the variance is less than a preset threshold) and the speech rate is slow, it is determined to be a sad emotion; when the short-time energy experiences a sudden surge (i.e., the rate of change of energy between adjacent frames exceeds a threshold), and the fundamental frequency increases while the short-time zero-crossing rate increases, it is determined to be an angry emotion; when the fundamental frequency is within ±10% of the neutral fundamental frequency, and the short-time energy and zero-crossing rate are stable while the speech rate is within a normal range, it is determined to be a calm emotion. Among them, the user's neutral fundamental frequency is obtained by collecting stable speech during the user registration phase and then statistically analyzing it. When the sound category is either environmental event sound or third-party biological sound, the corresponding sound event type is obtained. Specifically, the acoustic features of the current non-speech signal are matched with the corresponding environmental event feature template library or biological sound feature template library, and the specific sound event type is output based on the matching result, including one or more of the following: doorbell sound, footsteps, collision sound, dog bark, cat meow, and bird chirping. When the sound category is user voice, the sound event type is empty.

[0017] S4: Obtain the current status parameters of the electronic pet, including emotion parameters and energy parameters; update the emotion parameters based on the user's emotion tags, and update the energy parameters based on the sound intensity level and sound event type to obtain the updated status parameters; Furthermore, the obtained updated state parameters include: Obtain the current status parameters of the electronic pet, including emotion parameters and energy parameters; Based on the sound feature vector, the feature derivation quantities are calculated, including fundamental frequency offset, fundamental frequency fluctuation amplitude, short-time energy fluctuation, short-time zero-crossing rate change, and fundamental frequency normalization value. When the input is a valid speech segment, the emotion change is calculated based on the user's emotion label and feature derivation, and the emotion parameters are updated based on the emotion change; when the input is a non-speech signal, the emotion parameters remain unchanged; when there is no speech input, the emotion parameters are updated to the baseline emotion parameter value. Intensity coefficients are constructed based on sound intensity levels, and event coefficients are constructed based on sound event types and characteristic derivations. When the input is a valid speech segment, the energy parameter is updated based on the intensity coefficient; when the input is a non-speech signal, the energy parameter is updated based on the intensity coefficient and the event coefficient; when there is no sound input, the energy parameter is updated in a linear decay manner over time. The updated emotion and energy parameters are stored for use in the generation of subsequent behavioral instructions.

[0018] Specifically, the current state parameters of the electronic pet are obtained, including emotion parameters and energy parameters. The emotion parameters are used to characterize the current level of pleasure or depression of the electronic pet, and the energy parameters are used to characterize the current level of energy of the electronic pet.

[0019] First, extraction is performed for different input types: when the input is a valid speech segment, the fundamental frequency sequence corresponding to that valid speech segment is extracted. Short-time energy sequence Short-time zero-crossing rate sequence and speaking speed When the input is a non-speech signal, the corresponding fundamental frequency sequence is extracted. Short-time energy sequence Short-time zero-crossing rate sequence And the Mel frequency cepstral coefficient sequence; Based on the above characteristics, feature derivation quantities are calculated, including: the fundamental frequency offset for the effective speech segment. ,in The mean of the fundamental frequency. The user-neutral baseband; and the baseband normalization value for non-speech signals. ,in The reference fundamental frequency for system calibration; also includes the fundamental frequency fluctuation amplitude. Short-term energy fluctuation ,in Short-time energy mean; short-time zero-crossing rate change. ,in The mean of the zero-crossing rates; I. Update of Emotional Parameters: When the input is a valid speech segment, for different user emotion tags, the emotion change is calculated based on the feature derivation quantity according to the feature combination method consistent with the emotion determination in step S3, and the emotion change quantity is calculated based on the feature derivation quantity. ,include: When the user's emotion label is happy ; When a user's emotion tag is "sad" ; When a user's emotion label is anger ; When the user's emotion label is calm ,in, The baseline emotion parameter values ​​set for the electronic pet during initialization. The current sentiment parameter; The emotion parameters are updated based on the magnitude of emotion changes. ; When the input is a non-speech signal, the emotion parameters remain unchanged; When there is no voice input, the emotion parameters are updated to the baseline emotion parameter values.

[0020] 2. Energy parameter update: Update according to the sound intensity level determined in step S2 and the sound event type determined in step S3.

[0021] First, calculate the corresponding average energy value for the current valid speech segment or non-speech signal. And based on the energy threshold range corresponding to the sound intensity level determined in step S2, construct the intensity coefficient: ,in These are the lower and upper bounds of the energy threshold range corresponding to the sound intensity level, respectively. Based on the sound event type determined in step S3, construct event coefficients by combining feature derivation quantities. : For doorbell sounds (energy surge followed by attenuation + MFCC high-frequency dominance). ;in, The total dimension of MFCC, This is the starting index for the high-frequency components. Here, k is the cutoff index for low-frequency components, and k is the summation index. The first of the Mel frequency cepstral coefficient sequence First-order cepstral coefficients; For footsteps (periodic energy spikes + low frequencies). , where P is the main peak value of the autocorrelation function of the short-time energy sequence (characterizing the periodic intensity). For collision sounds (instantaneous spike + high zero-crossing rate + short duration). ; For dog barking (mid-low frequency + MFCC harmonic reduction). ,in The first of the Mel frequency cepstral coefficient sequence cepstral coefficients, The first of the Mel frequency cepstral coefficient sequence First-order cepstral coefficients; For cat meows (mid-frequency + fundamental frequency periodic fluctuations), the event coefficient ; For birdsong (high frequency + dense energy peaks), the event coefficient ;in The number of short-time energy peaks per unit time (peak density). The energy parameters are updated based on the intensity coefficient and event coefficient when the input is a valid speech segment. When it is a non-speech signal When there is no sound input, the energy parameter is updated according to a linear decay over time, calculated using the following formula: ;in, For the updated energy parameters, The energy parameters before the update. For time intervals, The time coefficient is determined based on the natural energy decay rate during system operation. Finally, the updated energy and emotion parameters are stored for use in generating subsequent behavioral instructions.

[0022] S5: Based on the primary sound information, sound category parsing results, and updated status parameters, generate behavior commands and send the behavior commands to the electronic pet's action execution module in priority order; Furthermore, the generation behavior instructions include: Based on the initial sound information, sound category parsing results, and updated status parameters, trigger conditions are determined sequentially according to a preset priority order, and action instructions are generated, including: When the energy parameter is lower than the preset low battery threshold, a low battery sleep behavior instruction is generated, including stopping movement and entering a sleep posture; When the user's intent is recognized, a corresponding intent-based action instruction is generated; When the user's emotion label is identified as "sad" or "angry", a soothing behavior instruction is generated; When the type of sound event is identified, a corresponding event response behavior instruction is generated; Based on the updated emotional and energy parameters, the action parameters for intentional behavioral instructions, soothing behavioral instructions, and event response behavioral instructions are determined, including turning action parameters, movement action parameters, vocalization action parameters, and limb action parameters. Behavior commands are sent to the electronic pet's action execution module in priority order to trigger the electronic pet to perform the corresponding behavior. When there are multiple behavior commands, the higher priority behavior command interrupts the lower priority behavior command, and behavior commands with the same priority are executed in the order they are generated.

[0023] Specifically, based on the primary sound information determined in step S2, the analysis results of the sound category in step S3, and the status parameters of the electronic pet updated in step S4, the behavior commands are judged and generated in the following priority order: the low battery response has the highest priority, and once triggered, all subsequent conditions are ignored; otherwise, the corresponding behavior commands are generated by matching the user's command intent, user emotion label, and sound event type in that order.

[0024] (a) Low battery response (highest priority) When the energy parameter is lower than the preset low battery threshold, a low battery sleep behavior command is generated: stop moving and enter sleep mode, maintaining only the sound detection function. In this state, the device will only wake up and regenerate the behavior command according to the priority rules when a sound intensity level greater than or equal to 4 is detected or the user's calling intention is recognized. After execution, the device will enter sleep mode again.

[0025] (ii) Behavioral instructions based on user intent (high priority) When a user's intent is identified, a corresponding intent-based action instruction is generated based on that intent. Calling Intent: Turn towards the sound source and move toward the sound source, while triggering an interactive vocal action; Interactive intent: To perform the action of raising hands and wagging tail; Feeding intention: Turn towards the user and perform a feeding simulation.

[0026] (iii) Behavioral instructions based on user emotion tags (medium priority) When the user's emotion label is "sad" or "angry", a soothing behavior instruction is generated: turn towards the user and move towards the user, trigger a soothing vocal action, and then execute the head-rubbing action and the snuggling action in sequence.

[0027] (iv) Behavioral instructions based on sound event type (low priority) When the sound category is environmental event sound or third-party biological sound, corresponding event response behavior instructions are generated based on the specific event type, including: Doorbell sound: Turn towards the door and trigger an alarm sound action; Footsteps: Turn towards the direction of the sound source and look around; Collision sound: Turn towards the sound source and perform avoidance actions, including pulling back the body and blinking; Dog barking: Turn towards the source of the sound and adopt an alert posture, including pricking up ears and tensing the body; Cat meow: Turn towards the sound source and perform a probe action, while simultaneously triggering a probing sound action; Birdsong: Turn towards the sound source and perform head movements, including raising and turning the head.

[0028] The behavioral commands include parameters for turning, moving, vocalizing, and limb movements. These are combined with the updated emotion parameter M' and energy parameter E', and compared to the factory-set baseline emotion parameter value Mref and baseline energy parameter value Eref. Based on the comparison results, the corresponding level is selected from the preset action level set. (1) Determination of steering action parameters The turning angle is taken as the target direction angle, which is the horizontal angle of the sound source determined in step S2.

[0029] (2) Determination of movement parameters The movement distance is the spatial distance between the current position and the target position; The movement speed is determined based on the combination of M' and the baseline emotion parameter value Mref, and E' and the baseline energy parameter value Eref: If M' > Mref and E' > Eref, use the fast file. If M' > Mref and E' ≤ Eref, use the medium speed gear; If M' ≤ Mref and E' > Eref, use the slow speed mode; If M' ≤ Mref and E' ≤ Eref, only the turning action is performed, and no movement action is performed; the movement speed is 0.

[0030] (3) Determination of vocalization parameters The tone of voice is determined by the type of voice corresponding to the behavior command. The soothing voice uses a preset low tone, the interactive and probing voices use a preset mid tone, and the alert voice uses a preset high tone. Each preset tone is a fixed voice parameter set at the factory when the electronic pet leaves the factory. The duration of the sound output is the actual duration of the triggering sound segment. The volume of the sound is determined by the combination of M' and Mref, and E' and Eref: If M' > Mref and E' > Eref, use the high volume setting; If M' > Mref and E' ≤ Eref, use the medium volume setting; If M' ≤ Mref, use the low volume setting.

[0031] (4) Determination of limb movement parameters The types of actions are defined by behavioral instructions, including erecting ears, wagging tail, peeking out, retracting, blinking, raising head, and turning head. The amplitude of the movement is determined based on the comparison between M' and Mref: if M' > Mref, a large amplitude is used; if M' ≤ Mref, a small amplitude is used. The movement frequency applies only to tail wagging; for other limb movements, no movement frequency parameter is set, and they are executed according to a single action command. The movement frequency is determined based on the comparison between E' and Eref: if E' > Eref, a high frequency setting is used; if E' ≤ Eref, a low frequency setting is used. The duration of the action is the duration of the triggering sound segment.

[0032] The generated action instructions are sent to the action execution module for execution in order of priority. When there are multiple action instructions, they are executed in order of priority. Higher priority action instructions can interrupt lower priority action instructions. Action instructions with the same priority are executed in the order they were generated. After execution, the system returns to standby state.

[0033] S6: The action execution module executes the behavior command and collects the electronic pet's own sound signal during the execution process, and calculates the short-term energy value within the current time window; when the short-term energy value is greater than the preset self-noise judgment energy threshold, self-noise control processing is performed; when the short-term energy value is not greater than the self-noise judgment energy threshold, the execution of the behavior command is maintained or resumed.

[0034] Furthermore, the action execution module executes behavioral commands and collects the electronic pet's own sound signals during execution, calculating the short-term energy value within the current time window; when the short-term energy value is greater than a preset self-noise judgment energy threshold, self-noise control processing is performed; when the short-term energy value is not greater than the self-noise judgment energy threshold, the execution of the behavioral commands is maintained or resumed, including: During the execution of behavioral commands by the action execution module, the electronic pet's own sound signals are collected, and the short-term energy value within the current time window is calculated; When the short-term energy value exceeds the preset self-noise detection energy threshold, self-noise control processing is performed, including: When the action instruction includes a vocalization, pause the vocalization. When the action command includes a movement action, the movement speed is adjusted to the next lower speed level corresponding to the current speed level, and the movement action is stopped when the current speed level is the lowest. When a behavioral instruction includes a limb movement with a movement frequency parameter, the movement frequency is adjusted to the next lower frequency level corresponding to the current frequency level, and the limb movement is stopped when the current frequency level is the lowest. When multiple action instructions are executed simultaneously, only the action instruction with the highest priority is retained, and the execution of the other action instructions is suspended. After self-noise control processing, when the short-term energy value meets the condition of not exceeding the self-noise judgment energy threshold, the suspended or reduced behavior action is restored, and the original behavior instruction is continued to be executed.

[0035] Specifically, during the execution of behavioral instructions by the action execution module, the sound signals generated by the electronic pet itself are collected in real time. These sound signals are a mixture of sound signals generated by the motor operation, structural movement and sound generation device, and the short-time energy value Es within the current time window is calculated. The short-term energy value Es is compared with a preset self-noise reference value En, and the execution of the behavior command is controlled according to the comparison result: when Es≤En, the normal execution of the current behavior command is maintained; when Es>En, self-noise control processing is performed. This self-noise control processing makes corresponding adjustments for different action types contained in the behavior command, including: If the current action instruction includes a vocalization action, then pause the execution of the vocalization action; If the current action command includes a movement action, then adjust the movement speed to the next lower speed level corresponding to the current speed level; if the current speed level is already the lowest, then stop the movement action. If the current action command includes a limb movement, for limb movements with a frequency parameter (such as tail wagging), adjust the frequency of the limb movement to the next lower frequency level of the current frequency level; if the current frequency level is already the lowest, stop the limb movement; for limb movements without a frequency parameter, continue to execute according to the single action command. When there are multiple action instructions that are executed simultaneously, only the action instruction with the highest priority is retained, and the execution of the other action instructions is suspended. After performing self-noise control processing, the short-time energy value Es is recalculated. When Es≤En is satisfied, the paused or reduced behavior is resumed, and the original behavior instructions are executed again.

[0036] Through the above self-noise control processing, the electronic pet can dynamically adjust its behavior when generating significant self-noise, reduce interference with sound acquisition, and automatically resume normal behavior after the noise decreases, ensuring the continuity and stability of the interaction.

[0037] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0038] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for triggering electronic pet behavior based on sound perception, characterized in that, The method includes the following steps: S1: Real-time acquisition of ambient sound signals through a microphone array deployed on the electronic pet device, and analog-to-digital conversion, preprocessing and acoustic feature extraction of the ambient sound signals to construct a sound feature vector; S2: Determine primary sound information based on sound feature vectors, including sound intensity level, sound source location information, and sound category; the sound category includes user voice, environmental event sound, and third-party biological sound. S3: Based on the sound category, the environmental sound signal is parsed and processed to obtain the sound category parsing result, including the user's command intent, the user's emotion label, and the sound event type; wherein, when the sound category is the user's voice category, the user's command intent and the user's emotion label are determined, and when the sound category is the environmental event sound category or the third-party biological sound category, the sound event type is determined. S4: Obtain the current status parameters of the electronic pet, including emotion parameters and energy parameters; update the emotion parameters based on the user's emotion tags, and update the energy parameters based on the sound intensity level and sound event type to obtain the updated status parameters; S5: Based on the primary sound information, sound category parsing results, and updated status parameters, generate behavior commands and send the behavior commands to the electronic pet's action execution module in priority order; S6: The action execution module executes the behavior command and collects the electronic pet's own sound signal during the execution process, and calculates the short-term energy value within the current time window; when the short-term energy value is greater than the preset self-noise judgment energy threshold, self-noise control processing is performed; when the short-term energy value is not greater than the self-noise judgment energy threshold, the execution of the behavior command is maintained or resumed.

2. The electronic pet behavior triggering method as described in claim 1, characterized in that, The construction of the sound feature vector includes: The ambient sound signals are collected in real time by a microphone array deployed on the electronic pet device, and the ambient sound signals are processed by analog-to-digital conversion to obtain digital sound signals; Preprocessing of digital audio signals includes pre-emphasis, framing, windowing, and noise reduction. Acoustic features are extracted from each preprocessed digital audio signal frame, including Mel frequency cepstral coefficients, short-time energy, short-time zero-crossing rate, fundamental frequency, and sound source horizontal angle. The acoustic features corresponding to each digital audio signal frame are combined in chronological order to construct an audio feature vector.

3. The electronic pet behavior triggering method as described in claim 2, characterized in that, The determination of primary sound information based on sound feature vectors includes: Short-time energy is extracted based on sound feature vectors, and the short-time energy of consecutive frames is statistically processed to calculate the average energy value within a preset time window. The average energy value is then compared with a preset energy threshold range to determine the sound intensity level. Obtain the horizontal angle of the sound source from the sound feature vector and use it as the sound source location information; Determining sound categories based on sound feature vectors includes: A dual threshold detection method based on short-time energy and short-time zero-crossing rate is used to classify sound signals into speech signals or non-speech signals. For voice signals, they are classified as user voice; when multiple consecutive frames are identified as voice frames, the signal segment consisting of the consecutive voice frames is identified as a valid voice segment. For non-speech signals, their acoustic features are matched with a preset environmental event feature template library and a biological sound feature template library, respectively. Based on the matching results, the sound category to which they belong is determined to be either environmental event sound or third-party biological sound.

4. The electronic pet behavior triggering method as described in claim 3, characterized in that, The obtained sound category analysis results include: When the sound category is user voice, the acoustic features of the current voice signal are matched with each keyword audio template in the preset command intent template library. The command intent corresponding to the keyword audio template with a matching score higher than the preset similarity threshold and the maximum value is determined as the user command intent. The fundamental frequency, short-time energy, and short-time zero-crossing rate are extracted from the sound feature vector, and the speech rate is determined by combining the effective speech segment duration. Based on the preset emotion judgment rules, the user's emotion label is determined. When the sound category is environmental event sound or third-party biological sound, the acoustic features of the current sound signal are matched with the feature template library of the corresponding category, and the specific sound event type is determined based on the matching result.

5. The electronic pet behavior triggering method as described in claim 4, characterized in that, The updated state parameters obtained include: Obtain the current status parameters of the electronic pet, including emotion parameters and energy parameters; Based on the sound feature vector, the feature derivation quantities are calculated, including fundamental frequency offset, fundamental frequency fluctuation amplitude, short-time energy fluctuation, short-time zero-crossing rate change, and fundamental frequency normalization value. When the input is a valid speech segment, the emotion change is calculated based on the user's emotion label and feature derivation, and the emotion parameters are updated based on the emotion change; when the input is a non-speech signal, the emotion parameters remain unchanged; when there is no speech input, the emotion parameters are updated to the baseline emotion parameter value. Intensity coefficients are constructed based on sound intensity levels, and event coefficients are constructed based on sound event types and characteristic derivations. When the input is a valid speech segment, the energy parameter is updated based on the intensity coefficient; when the input is a non-speech signal, the energy parameter is updated based on the intensity coefficient and the event coefficient; when there is no sound input, the energy parameter is updated in a linear decay manner over time. The updated emotion and energy parameters are stored for use in the generation of subsequent behavioral instructions.

6. The electronic pet behavior triggering method as described in claim 5, characterized in that, The generation behavior instructions include: Based on the initial sound information, sound category parsing results, and updated status parameters, trigger conditions are determined sequentially according to a preset priority order, and action instructions are generated, including: When the energy parameter is lower than the preset low battery threshold, a low battery sleep behavior instruction is generated, including stopping movement and entering a sleep posture; When the user's intent is recognized, a corresponding intent-based action instruction is generated; When the user's emotion label is identified as "sad" or "angry", a soothing behavior instruction is generated; When the type of sound event is identified, a corresponding event response behavior instruction is generated; Based on the updated emotional and energy parameters, the action parameters for intentional behavioral instructions, soothing behavioral instructions, and event response behavioral instructions are determined, including turning action parameters, movement action parameters, vocalization action parameters, and limb action parameters. Behavior commands are sent to the electronic pet's action execution module in order of priority to trigger the electronic pet to perform the corresponding behavior. When there are multiple behavior commands, the higher priority behavior command interrupts the lower priority behavior command, and behavior commands with the same priority are executed in the order they are generated.

7. The electronic pet behavior triggering method as described in claim 6, characterized in that, The action execution module executes behavioral commands and collects the electronic pet's own sound signals during execution, calculating the short-term energy value within the current time window. When the short-term energy value is greater than a preset self-noise judgment energy threshold, self-noise control processing is performed; when the short-term energy value is not greater than the self-noise judgment energy threshold, the execution of the behavioral commands is maintained or resumed, including: During the execution of behavioral commands by the action execution module, the electronic pet's own sound signals are collected, and the short-term energy value within the current time window is calculated; When the short-term energy value exceeds the preset self-noise detection energy threshold, self-noise control processing is performed, including: When the action instruction includes a vocalization, pause the vocalization. When the action command includes a movement action, the movement speed is adjusted to the next lower speed level corresponding to the current speed level, and the movement action is stopped when the current speed level is the lowest. When a behavioral instruction includes a limb movement with a movement frequency parameter, the movement frequency is adjusted to the next lower frequency level corresponding to the current frequency level, and the limb movement is stopped when the current frequency level is the lowest. When multiple action instructions are executed simultaneously, only the action instruction with the highest priority is retained, and the execution of the other action instructions is suspended. After self-noise control processing, when the short-term energy value meets the condition of not exceeding the self-noise judgment energy threshold, the suspended or reduced behavior action is restored, and the original behavior instruction is continued to be executed.