Pet voice interaction device and pet voice interaction method
By designing a pet voice interactive device, using deep learning to recognize effective pet calls and realize voice recognition and response, the existing pet training equipment is solved, and real-time effective interaction with pets and cost reduction is achieved.
Patent Information
- Application Number
- CN202411945188.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-06
AI Technical Summary
Existing pet training equipment prevents pet barking inhumane ways, and the beeping training method is not effective. Smart pet interactive devices on the market have problems of high costs and relying on cellular networks.
Design a pet voice interactive device, including a sound detection and acquisition module, a sound deep learning and processing module and an intelligent voice assistant module, to recognize effective pet calls through deep learning and realize voice recognition and response.
Real-time and effective interaction with pets, such as remote pet fun, training pets, finding pets and pet companionship, avoiding inhumane training methods and reducing equipment costs.
Smart Images

Figure CN119943060A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent pet education, and in particular to a pet voice interaction device and a pet voice interaction method. Background Art
[0002] With the development of pet feeding levels, "rich" pet raising has become a mainstream trend. Pets have gradually become an important role of emotional companionship in the family, and pet life needs have extended to emotional attention.
[0003] However, pets are not humans and cannot communicate with language. They can only convey their emotions through unique ways, such as the "woof" of dogs or the "meow" of cats. These pet calls convey emotions in the following ways: sometimes they need to be comforted, such as when the pet is alone at home; sometimes they need to be trained, such as the barking of dogs that disturbs people's quiet life at night; and sometimes they need to be guided, such as when the pet is accidentally lost.
[0004] Based on these needs, a wide variety of smart devices have appeared on the market. One of them is the pet trainer commonly seen in the market, which is used to stop the annoying barking of pets, mainly through training methods such as electric shock, vibration or buzzer sound. Among these methods, the device that uses electric shock stimulation to stop animals from barking generates high-frequency electrostatic stimulation of about 15,000 volts and a current of about 10um on the neck of the pet when a high-voltage pulse is output. However, since the barking of pets is made through the vocal cords, this method and its device do not solve the problem symptomatically, but instead stop barking in an inhumane way. The buzzer sound method, although more humane, does not have a good training effect.
[0005] In addition, there are products such as voice calling modules for pet trackers on the market, most of which need to interact with pets through cellular networks and remote real-time intercom. However, this method not only requires a hardware module with cellular voice communication, but also requires regular subscription to high-cost voice call packages. Summary of the invention
[0006] Based on the shortcomings of existing products, the present invention proposes a pet voice interaction device, comprising: a sound detection and collection module, a sound deep learning and processing module and an intelligent voice assistant module; the sound detection and collection module is used to detect and collect sounds in the environment, so as to collect effective calls of pets in the environment, and transmit the effective calls of pets to the sound deep learning and processing module; the sound deep learning and processing module is used to analyze and deeply learn the collected effective calls of pets, and extract the call audio information of the effective calls of pets; the intelligent voice assistant module is used to perform voice recognition based on the call audio information, and match voice wake-up terms and play voice response terms.
[0007] Furthermore, the sound detection and collection module includes a sound detection module and a sound collection module. The sound detection module is used to monitor the sound in the environment in real time. When the amplitude of the sound in the environment exceeds a preset threshold, the sound detection module is used to wake up the sound collection module to collect sound from the surrounding environment.
[0008] Furthermore, the sound collection module samples the sounds in the environment at a predetermined sampling frequency within a set time interval, and records the intensity or amplitude of the sound waves of each sample to form an original audio file.
[0009] Furthermore, the original audio file is compressed and saved as a correspondence between audio amplitude and digital time series, and is sent to the buffer of the sound detection and acquisition module for decoding and conversion into an array with a corresponding relationship. The audio is then converted from the time domain to the frequency domain through a short-time Fourier transform to obtain a spectrogram. The spectrogram is input into a sound source classifier, a CNN convolutional neural network model architecture is used to extract feature maps, and a linear classifier model is used to predict whether the collected sound belongs to the pet sound category based on the frequency, rhythm and amplitude characteristics of the corresponding pet sound category, so as to determine whether it is a valid call of the pet.
[0010] Furthermore, the sound deep learning and processing module includes an audio analysis module, a deep learning algorithm unit, a voiceprint storage unit and an audio encoding unit; the audio analysis module is responsible for pre-processing the effective calls of the pet wearing this device, and then sending it to the deep learning algorithm unit; the feature map output by the convolutional network in the deep learning algorithm unit is cut into independent frames and input into the recurrent network; each of the frames corresponds to a single step of the original audio spectrum, and the number of each frame and the duration of each frame are selected when designing the model as hyperparameters; for each frame, the recurrent network is connected to a linear classifier to predict the probability of each character in the vocabulary, derive the correct character sequence, and convert it into effective voice information transmitted by the pet; the audio encoding unit transcribes the converted effective voice information to form a wake-up word information text, and transmits it to the intelligent voice assistant module.
[0011] Furthermore, the intelligent voice assistant module includes an intelligent voice recognition module, an intelligent voice response module and a voice broadcast module; the intelligent voice recognition module is used to decode the wake-up word information text, convert it into the control command and send it to the intelligent voice response module; the intelligent voice response module is used to manage and deploy the audio database files, and locate the corresponding audio file according to the received control command and send it to the audio broadcast module; the audio broadcast module is responsible for broadcasting the audio file to achieve real-time and effective interaction with the effective calls of the pet.
[0012] Furthermore, the pet voice interaction device also includes at least one of a power module, a wireless radio frequency module, a pet posture sensing module, a control module and an indication module; the power module is used to power the device; the wireless radio frequency module unit is used to connect the pet voice interaction device to the cloud and / or the terminal user's APP; the pet posture sensing unit includes a motion sensing sensor, which is used to detect the body vibration generated when the pet makes the effective call of the pet, as an auxiliary judgment for accurately identifying the effective call of the pet; the control module is used to perform corresponding configuration operations on the pet voice interaction device; the indication module is used to provide effective feedback on different operations or event processing results.
[0013] To solve the above problems, the present invention also provides a pet voice interaction method, including: detecting and collecting sounds in the environment to collect effective calls of pets in the environment; analyzing and deep learning the collected effective calls of pets to extract audio information of the effective calls of pets; performing voice recognition based on the call audio information, and matching voice wake-up terms and playing voice response terms.
[0014] Furthermore, the method includes at least one of the following steps: real-time monitoring of the sounds in the environment, and when the amplitude of the sounds in the environment exceeds a preset threshold, collecting the sounds of the surrounding environment; sampling the sounds in the environment at a predetermined sampling frequency within a set time interval, and recording the intensity or amplitude of the sound waves of each sample to form an original audio file; compressing the original audio file, saving it with the corresponding relationship between the audio amplitude and the digital time series, decoding it, converting it into an array with a corresponding relationship, and then converting the audio from the time domain to the frequency domain through a short-time Fourier transform to obtain a spectrogram; extracting a feature spectrum based on the spectrogram, and predicting whether the collected sound is a pet sound category based on the frequency, rhythm and amplitude characteristics of the corresponding pet sound category, and determining whether it is whether it is a valid call of the pet; pre-processing the valid call of the pet; cutting the feature map into independent frames and inputting them into the recurrent network; each of the frames corresponds to a single step of the original audio spectrum, and the number of each of the frames and the duration of each of the frames are selected when designing the model as hyperparameters; for each of the frames, the recurrent network connects the linear classifier to predict the probability of each character in the vocabulary, derives the correct character sequence, and thus converts it into the valid voice information transmitted by the pet; transcribes the converted valid voice information to form a wake-up word information text; decodes the wake-up word information text and converts it into the control command; according to the control command, locates the corresponding audio file; broadcasts the audio file to achieve real-time and effective interaction with the valid call of the pet.
[0015] Furthermore, the method also includes at least one of the following steps: connecting to the cloud and / or the terminal user's APP; detecting the body vibration generated when the pet makes the effective call of the pet as an auxiliary judgment for accurately identifying the effective call of the pet; and providing effective feedback on different operations or event processing results.
[0016] The pet voice interaction device provided by the present invention has a corresponding structure and connection relationship, and can start sound collection under preset conditions, and extract effective pet sounds based on the sounds collected in the environment, and then perform deep learning and conversion processing on the original pet sound source, and convert it into an effective voice recognition entry to drive the intelligent voice recognition module to play matching audio response information, which can realize real-time and effective interaction with pets, such as remote pet teasing, pet training, (remote) pet search and pet companionship.
[0017] The pet voice interaction method provided by the present invention includes: detecting and collecting sounds in the environment to collect effective calls of pets in the environment; analyzing and deeply learning the collected effective calls of pets to extract audio information of the effective calls of pets; performing voice recognition based on the call audio information, and matching voice wake-up terms and playing voice response terms. Based on the sounds collected in the environment, effective pet sounds can be extracted, and the original pet sound sources can be deeply learned and processed to convert them into effective voice recognition terms to drive the intelligent voice recognition module to play matching audio response information. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic diagram (block diagram) of the hardware principle of the device according to the embodiment of the present invention;
[0019] Figure 2 is a schematic diagram of signal processing logic in an embodiment of the present invention;
[0020] Figure 3 2 is a schematic diagram of the principle of using short-time Fourier transform to achieve audio conversion from time domain to frequency domain in an embodiment of the present invention;
[0021] Figure 4 is a schematic diagram of the principle of using a convolutional neural network to determine whether the original audio is a valid pet sound in an embodiment of the present invention;
[0022] Figure 5 It is a schematic diagram of the principle of obtaining the corresponding wake-up word information text according to the audio in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0024] The embodiment of the present invention provides a pet voice interaction device, including: Figure 1 The sound detection and acquisition module (unlabeled), the sound deep learning and processing module (unlabeled) and the intelligent voice assistant module (unlabeled, dotted box on the right) are shown.
[0025] The sound detection and collection module is used to detect and collect sounds in the environment to collect effective calls of pets in the environment and transmit the effective calls of pets (as raw information) to the sound deep learning and processing module. The sound detection and collection module may also have a corresponding storage module (not shown) and an artificial intelligence module (not shown).
[0026] The sound deep learning and processing module is used to analyze and deeply learn the collected effective pet calls, and extract the audio information of the effective pet calls.
[0027] The intelligent voice assistant module is used to perform voice recognition based on the audio information of the call, and to match the voice wake-up terms and play the voice response terms.
[0028] like Figure 1 The sound detection and collection module includes a sound detection module (not labeled) and a sound collection module (not labeled). The sound detection module is used to monitor the sound in the environment (surrounding sound wave signals) in real time. When the amplitude of the sound (sound wave signal) in the environment exceeds the preset threshold, the sound detection module is used to wake up the sound collection module to collect the sound of the surrounding environment. The sound collection module mainly distinguishes the sound of the pet from a complex environment and extracts it.
[0029] Specifically, Figure 2 In one embodiment, the system (e.g., the software of the hardware system) is first "initialized", and then "environmental sound detection" is performed. When the amplitude of the sound (sound wave signal) in the environment exceeds a preset threshold. The specific threshold can be: the volume is set to 40 decibels, and the frequency is set to 100 Hz. At this time, if Figure 2 If the detection result of "exceeding the sound information threshold" is "Yes", the corresponding sound source is judged to be an "effective sound source". At this time, the sound collection module can be awakened to collect sounds from the surrounding environment. On the contrary, when the detection result of "exceeding the sound information threshold" is "No (N0)", that is, the corresponding sound source is judged to be "noise", the pet voice interaction device can not perform subsequent operation steps.
[0030] like Figure 2 In the embodiment of the present invention, "preprocessing" can be performed first (first identifying pet sounds, such as dog barking, from environmental sounds) and then "pet sound recognition" (semantic recognition) can be performed.
[0031] like Figure 2 , when performing "pet sound recognition", existing research on pet sounds can be used. The corresponding pets can be common cats and dogs. The calls of pets, such as dog calls and cat calls, have specific sound frequencies, amplitudes and rhythms, which can be identified by algorithms. For example, the base frequency of a dog's call is 400-500Hz, and the sub-band is about 900Hz. The frequency range of the sound emitted by a cat is from 400-1200Hz, and the specific frequency, rhythm and amplitude are different. By setting the algorithm, it can be identified according to the corresponding preset information, so as to determine whether the corresponding sound is emitted by a pet, and whether the pet is a dog or a cat (to distinguish it from the noise, human voice and other sounds in the environment). More importantly, since the sound of each pet has its own characteristics (different timbre, mainly different specific spectrum of the pet's call), the corresponding pet voice interaction device can have a built-in intelligent pet voice recognition and learning function module, and can also be designed to perform training and learning at the same time, so that after a period of training and the model is formed, the corresponding pet voice interaction device can more accurately identify the correct effective sound of a certain pet.
[0032] like Figure 2 The next step of "Pet Sound Recognition" is to judge the result, that is, "Is it a pet sound?" If the result is "Yes", it will enter the "Output Pet Sound Audio File" step. If the result is "No", it will return to the above step, that is, "Environmental Sound Detection".
[0033] Regarding the step of "outputting pet sound audio files", in an embodiment of the present invention, the sound collection module samples the sounds (sound wave signals) in the environment at a predetermined sampling frequency within a set time interval, and records the intensity or amplitude of the sound waves of each sample to form an original audio file.
[0034] Specifically, in one embodiment, the time interval that can be set can be, for example, 3S, that is, 3S is selected as the sound source. In addition, a predetermined sampling frequency is set, and the sampling frequency can usually be set to between 200Hz and 4KHz at the beginning (the sampling rate can be set by the corresponding software according to the corresponding algorithm matching). As above, after setting the time interval and predetermining the sampling frequency range and using it, it can be optimized after a period of time (optimizing the time interval and the sampling frequency range), so as to form a range that is more suitable for each pet, and form a more accurate and targeted recognition effect.
[0035] Please refer to Figure 2 , in "Output pet sound audio file", enter the "Audio preprocessing" step. Among them, outputting pet sound audio files can be performed in the sound detection and acquisition module, and audio preprocessing can be performed in the sound deep learning and processing module.
[0036] Please refer to Figure 3 , the original audio file is compressed and saved with the corresponding relationship between audio amplitude and digital time series (specifically, it can be saved in a file), and sent to the buffer of the sound detection and acquisition module for decoding and converted into an array with a corresponding relationship (specifically, it can be sent to the buffer for decoding and converted into a Numpy array with a corresponding relationship, or other arrays with a corresponding relationship). Then, the audio is converted from the time domain to the frequency domain through the short-time Fourier transform (the purpose at this time is to shorten the duration of the sound signal to a small time period, and then apply the Fourier transform to each small time period to determine the frequency contained in the segment), and obtain the spectrogram (specifically, the Fourier transform results of all time periods can be merged, and the data can be passed to the draw method for drawing, and presented as a waveform graph. Finally, the librosa function can be used and SpecAugment can be used for enhancement. The spectrogram formed at this time is a Mel scale spectrogram). The principle of the short-time Fourier transform that converts the original inter-sound signal into frequency information can be referred to. Figure 3 .
[0037] Please refer to Figure 4 , Figure 4 First, the above process of converting "raw audio data" into "spectrogram" is shown, namely Figure 3 The corresponding process then inputs the spectrogram into the sound source classifier, uses the CNN convolutional neural network model architecture to extract the feature spectrum (feature spectrum), and uses the linear classifier model to predict whether the collected sound belongs to the pet sound category based on the frequency, rhythm and amplitude characteristics of the corresponding pet sound category, and determine whether it is a valid pet call. Figure 4 The displayed judgment results may be car horns, dog barking (effective pet sounds), drilling holes (environmental noise), etc. This process distinguishes the sounds of pets from complex environments, including recognition from the frequency, rhythm and amplitude of the sounds.
[0038] In an embodiment of the present invention, the sound deep learning and processing module can be a module that transmits the original information to the voice deep learning and processing module, and performs deep learning processing on the pet's voice signal through the voice processing algorithm module of the module, so as to analyze the effective information conveyed by the pet's call, extract and transform the effective information, form the pet voice, and send it to the intelligent voice assistant module synchronously. That is, the sound deep learning and processing module includes the voice processing algorithm module. The voice processing algorithm module is used to use a deep learning algorithm for effective pet calls (dog barks), analyze the information expressed in the sound, and convert it into language.
[0039] In this embodiment, non-pet barking sounds (such as non-dog barking sounds) are filtered out in the sound detection and collection module, that is, corresponding sound filtering operations can be performed to filter out non-pet valid barking sounds (such as noise) and the like.
[0040] Specifically, the sound deep learning and processing module includes an audio analysis module ( Figure 1 Not shown), deep learning algorithm unit ( Figure 1 Not shown), voiceprint storage unit ( Figure 1 ) and the audio encoding unit ( Figure 1 not shown).
[0041] The audio analysis module is responsible for pre-processing the effective calls of the pet wearing this pet voice interaction device (the pet's audio information), and then sending it to the deep learning algorithm unit (the audio analysis module mainly analyzes the sound of the pet wearing this product).
[0042] Please continue to refer to Figure 2 , after the “audio preprocessing” step, the embodiment of the present invention enters the “deep learning algorithm” step.
[0043] The hardware of this embodiment may include a deep learning algorithm unit (not shown). Figure 2 The "deep learning algorithm" step shown can be performed in the deep learning algorithm unit. The specific deep learning algorithm unit may include a deep learning algorithm CTC (connectionist temporal classification) model unit (or an audio deep learning module, which is mainly used to interpret the extracted pet audio files, interpret what useful information is conveyed by the pet's calls, and then convert it into information text). The feature map output by the convolutional network in the deep learning algorithm CTC (model) unit is cut into independent frames, refer to Figure 5 The corresponding principle is Figure 5 The "slicing audio into frame sequences" step in , and input into the recurrent network (Recurrent Neural Network, RNN), that is Figure 5In the "frame sequence input to RNN step". Each frame corresponds to a single step of the original audio spectrum (the step can be determined according to requirements). The number of frames and the duration of each frame are selected as hyperparameters when designing the model. For each frame, the recurrent network connects the linear classifier to predict the probability of each character in the vocabulary, derive the correct character sequence, and convert it into effective voice information transmitted by the pet.
[0044] The audio encoding unit transcribes the converted valid voice information to form the wake-up word information text. Please refer to Figure 5 The steps of "RNN outputs character probability", "Filters the maximum character probability", "Merges duplicate characters", and "Removes blanks" in the above figure.
[0045] Please refer to Figure 2 , the above process of deriving the correct character sequence and forming the wake-up word information text, namely Figure 2 The steps of "Pet language encoding" and "Wake-up command word conversion" are as follows.
[0046] After the wake-up word information text is formed, it can be transmitted to the intelligent voice assistant module (specifically, it is transmitted to the intelligent voice recognition module of the intelligent voice assistant module described below, please refer to the subsequent content).
[0047] It should be noted that for pet dogs, which have the highest rate of being kept at home, they usually express their emotions to the outside world through barking in daily life. Generally speaking, the low sound emitted by dogs represents warning, while the high-pitched sound represents friendliness. In addition, they will also adjust the tone, duration, frequency, etc. of their barking to convey different information. Studies have found that there are about 170 kinds of dog barking, and it can be confirmed that dogs all over the world are using a unified "dog language". Ten common meanings of dog barking: long humming means wanting to play or want to eat; short whining means anger and disappointment; crisp barking means there is a need; barking with teeth showing means to attack you; wolf howling means the dog is lonely; howling means being frightened; rhythmic barking means the dog is in a happy mood; loud barking means it wants to drive away strangers; low whining means it is uncomfortable; barking with the head turned away means feeling uneasy; direct barking means feeling angry. In addition, you can also refer to the relevant content of the patent application with publication number CN1478269A (device and method for determining the emotions of a dog based on the characteristics of barking), which discloses determining various emotions of a dog based on the characteristics of barking. The embodiment of the present invention provides a pet voice interaction device for obtaining the wake-up word information text corresponding to the effective sound of the pet through corresponding settings.
[0048] Please refer to Figure 2 After the "wake-up command conversion", the corresponding "wake-up command word decoding" is performed. Specifically, after receiving the wake-up word information text of the analyzed pet voice, the intelligent voice assistant module decodes it into a "control command" (please refer to Figure 2 ) and perform audio database search and positioning, which is "matching audio files" (please refer to Figure 2 ), select the corresponding response audio to interact with your pet in real time and effectively.
[0049] Specifically, the intelligent voice assistant module includes an intelligent voice recognition module ( Figure 1 Intelligent response unit (not marked), intelligent voice response module ( Figure 1 The intelligent answering unit is shown in the figure, not marked) and the voice broadcast module ( Figure 1 The unit shown in the figure is a voice announcement unit, not marked).
[0050] The intelligent voice recognition module is used to decode the wake-up word information text, convert it into a control command and send it to the intelligent voice response module (such as Figure 5 The information text obtained by the process analysis is displayed as "lonely". This term needs to be decoded into corresponding computer instruction information that can control the voice response module, and then the response module can call the corresponding audio file to play it).
[0051] Please refer to Figure 2 In "Extract Audio File", the intelligent voice response module is used to manage and allocate audio database files, and locate the corresponding audio file and send it to the audio broadcast module according to the received control command.
[0052] Please refer to Figure 2 In the "Broadcast Audio File", the audio broadcast module is responsible for broadcasting the corresponding audio file, which can realize real-time and effective interaction with the pet's effective calls (pet interactive calls).
[0053] The audio database in the embodiment of the present invention may include audio files, which may mainly include recording files of the owner's own voice and other music files for entertaining pets, etc. Among them, the recording files of the owner's own voice are obtained in the following ways: the owner can record the device and store it locally in the pet voice interaction device; the owner can use the mobile phone to record and store it in the cloud, and the pet voice interaction device can directly call the cloud audio, and this audio can be updated in real time; the owner can also respond in real time through the cloud, and the pet voice interaction device end calls the audio data (recording file) of the real-time response in the cloud.
[0054] The pet voice interaction device provided in the embodiment of the present invention may also include a power module (not shown), a wireless radio frequency module (see Figure 1 Wireless module, not marked), pet posture sensing module (not shown), control module ( Figure 1 The control unit is shown in the figure, not marked) and the indicator module ( Figure 1 At least one of the units shown as indicating units, not labeled).
[0055] Among them, the power module is used to power the pet voice interaction device.
[0056] The wireless radio frequency module unit is used to connect the pet voice interactive device to the cloud and / or the end user's APP. Through the settings of the cloud and / or the end user's APP, the corresponding pet voice interactive device can rely on the cloud big database information to achieve real-time update and optimization of the voiceprint database and audio database, thereby enhancing the deep learning algorithm unit of the pet voice interactive device based on the advantages of online big data analysis and algorithms.
[0057] The pet posture sensing unit includes a motion sensing sensor, which is used to detect the body vibration generated when the pet makes an effective pet call, as an auxiliary judgment for accurately identifying the effective pet call.
[0058] The control module is used to perform corresponding configuration operations on the pet voice interaction device.
[0059] The indication module is used to provide effective feedback on different operations or event processing results.
[0060] The pet voice interaction device provided in the embodiment of the present invention can be specifically designed as a structure for a pet to wear, such as a pet collar or other pet wearable product.
[0061] The pet voice interaction device provided by the embodiment of the present invention has the above-mentioned structure and its connection relationship. It can start sound collection under preset conditions, and extract effective pet sounds based on the sounds collected in the environment, and then perform deep learning and conversion processing on the original pet sound source, and convert it into effective voice recognition entries to drive the intelligent voice recognition module to play matching audio response information, which can realize real-time and effective interaction with pets, such as remote pet teasing, pet training, (remote) pet search and pet companionship.
[0062] An embodiment of the present invention also provides a pet voice interaction method, including: detecting and collecting sounds in an environment to collect effective calls of pets in the environment; analyzing and deep learning the collected effective calls of pets to extract audio information of the effective calls of pets; performing voice recognition based on the call audio information, and matching voice wake-up terms and playing voice response terms.
[0063] The pet voice interaction method provided by the embodiment of the present invention also includes at least one of the following steps: real-time monitoring of the sounds in the environment, and when the amplitude of the sounds in the environment exceeds a preset threshold, collecting the sounds of the surrounding environment; within a set time interval, using a predetermined sampling frequency, sampling the sounds in the environment, and recording the intensity or amplitude of the sound waves of each sample to form an original audio file; compressing the original audio file, saving it with the corresponding relationship between the audio amplitude and the digital time series, decoding it, converting it into an array with a corresponding relationship, and then converting the audio from the time domain to the frequency domain through a short-time Fourier transform to obtain a spectrogram; extracting a feature spectrum based on the spectrogram, and predicting the collected sound based on the frequency, rhythm and amplitude characteristics of the corresponding pet sound category. Whether the sound belongs to the pet sound category, determine whether it is an effective pet call; pre-process the effective pet calls; cut the feature map into independent frames and input them into the recurrent network; each frame corresponds to a single step of the original audio spectrum, and the number of frames and the duration of each frame are selected when designing the model as hyperparameters; for each frame, the recurrent network connects the linear classifier to predict the probability of each character in the vocabulary, derive the correct character sequence, and thus convert it into effective voice information transmitted by the pet; transcribe the converted effective voice information to form a wake-up word information text; decode the wake-up word information text and convert it into a control command; according to the control command, locate the corresponding audio file; broadcast the audio file to achieve real-time and effective interaction with the effective calls of the pet.
[0064] The pet voice interaction method of the embodiment of the present invention also includes at least one of the following steps: powering; connecting to the cloud and / or the terminal user's APP; detecting the body vibrations generated when the pet makes an effective pet call as an auxiliary judgment for accurately identifying the effective pet call; and providing effective feedback on different operations or event processing results.
[0065] For more information about the pet voice interaction method according to an embodiment of the present invention, please refer to the corresponding content of the aforementioned pet voice interaction device embodiment.
[0066] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A pet voice interaction device, characterized in that: include: Sound detection and acquisition module, sound deep learning and processing module and intelligent voice assistant module; The sound detection and collection module is used to detect and collect sounds in the environment, so as to collect effective calls of pets in the environment, and transmit the effective calls of pets to the sound deep learning and processing module; The sound deep learning and processing module is used to analyze and deeply learn the collected effective calls of the pet, and extract the audio information of the effective calls of the pet; The intelligent voice assistant module is used to perform voice recognition based on the calling audio information, and to match voice wake-up terms and play voice response terms.
2. The pet voice interaction device according to claim 1, characterized in that: The sound detection and collection module includes a sound detection module and a sound collection module. The sound detection module is used to monitor the sound in the environment in real time. When the amplitude of the sound in the environment exceeds a preset threshold, the sound detection module is used to wake up the sound collection module to collect the sound of the surrounding environment.
3. The pet voice interactive device as claimed in claim 2, characterized in that: The sound collection module samples the sounds in the environment at a predetermined sampling frequency within a set time interval, and records the intensity or amplitude of the sound wave of each sample to form an original audio file.
4. The pet voice interaction device as claimed in claim 3, characterized in that: The original audio file is compressed and saved as a corresponding relationship between audio amplitude and digital time series, and is sent to the buffer of the sound detection and acquisition module for decoding and conversion into an array with a corresponding relationship. The audio is then converted from the time domain to the frequency domain through a short-time Fourier transform to obtain a spectrogram. The spectrogram is input into a sound source classifier, a CNN convolutional neural network model architecture is used to extract feature spectra, and a linear classifier model is used to predict whether the collected sound belongs to the pet sound category based on the frequency, rhythm and amplitude characteristics of the corresponding pet sound category, so as to determine whether it is a valid call of the pet.
5. The pet voice interaction device as claimed in claim 4, characterized in that: The sound deep learning and processing module includes an audio analysis module, a deep learning algorithm unit, a voiceprint storage unit and an audio encoding unit; The audio analysis module is responsible for pre-processing the effective calls of the pet wearing the device and then sending them to the deep learning algorithm unit; The feature map output by the convolutional network in the deep learning algorithm unit is cut into independent frames and input into the recurrent network; each frame corresponds to a single step of the original audio spectrum, and the number of each frame and the duration of each frame are selected as hyperparameters when designing the model; for each frame, the recurrent network is connected to a linear classifier to predict the probability of each character in the vocabulary, derive the correct character sequence, and convert it into effective voice information transmitted by the pet; The audio encoding unit transcribes the converted valid voice information to form a wake-up word information text, and transmits it to the intelligent voice assistant module.
6. The pet voice interaction device as claimed in claim 5, characterized in that: The intelligent voice assistant module includes an intelligent voice recognition module, an intelligent voice response module and a voice broadcast module; The intelligent voice recognition module is used to decode the wake-up word information text, convert it into the control command and send it to the intelligent voice response module; The intelligent voice response module is used to manage and deploy audio database files, and locate the corresponding audio file according to the received control command and send it to the audio broadcast module; The audio broadcast module is responsible for broadcasting the audio file to achieve real-time and effective interaction with the effective calls of the pet.
7. The pet voice interaction device according to any one of claims 1 to 6, characterized in that: The pet voice interaction device further includes at least one of a power module, a wireless radio frequency module, a pet posture sensing module, a control module and an indication module; The power module is used to supply power to the device; The wireless radio frequency module unit is used to connect the pet voice interaction device to the cloud and / or the terminal user's APP; The pet posture sensing unit includes a motion sensing sensor for detecting body vibrations generated when the pet emits the pet's effective call, as an auxiliary judgment for accurately identifying the pet's effective call; The control module is used to perform corresponding configuration operations on the pet voice interaction device; The indication module is used to provide effective feedback for different operations or event processing results.
8. A pet voice interaction method, characterized in that: include: Detect and collect sounds in the environment to collect effective calls of pets in the environment; Analyze and deeply learn the collected effective calls of the pet, and extract the audio information of the effective calls of the pet; Voice recognition is performed based on the calling audio information, and voice wake-up term matching and voice response term playback are performed.
9. The pet voice interaction method according to claim 8, characterized in that: Also includes at least one of the following steps: Monitor the sounds in the environment in real time. When the amplitude of the sounds in the environment exceeds the preset threshold, collect the sounds of the surrounding environment. Within a set time interval, a predetermined sampling frequency is used to sample the sounds in the environment, and the intensity or amplitude of the sound waves of each sample is recorded to form an original audio file; The original audio file is compressed, saved as a corresponding relationship between audio amplitude and digital time series, decoded, and converted into an array with a corresponding relationship, and then converted from the time domain to the frequency domain through short-time Fourier transform to obtain a spectrogram; a feature spectrum is extracted according to the spectrogram, and based on the frequency, rhythm and amplitude characteristics of the corresponding pet sound category, whether the collected sound belongs to the pet sound category is predicted, and whether it is a valid call of the pet is determined; Pre-processing the effective calls of the pet; Cut the feature map into independent frames and input them into the recurrent network; Each of the frames corresponds to a single step of the original audio spectrum, and the number of frames and the duration of each frame are selected as hyperparameters when designing the model; for each frame, the recurrent network is connected to the linear classifier to predict the probability of each character in the vocabulary, derive the correct character sequence, and thus convert it into effective voice information transmitted by the pet; Transcribing the converted valid voice information to form a wake-up word information text; Decoding the wake-up word information text and converting it into the control command; According to the control command, locate the corresponding audio file; The audio file is broadcasted to achieve real-time and effective interaction with the effective calls of the pet.
10. The pet voice interaction method according to claim 8, characterized in that: Also includes at least one of the following steps: APP that connects to the cloud and / or end users; Detecting the body vibration generated when the pet makes the effective call of the pet as an auxiliary judgment for accurately identifying the effective call of the pet; Provide effective feedback on different operations or event processing results.
Citation Information
Patent Citations
Device and method for judging dog's feeling from cry vocal character analysis
CN1478269A
Cited By
Pet language translation method and system based on audio learning
CN120780863A
Audio-based pet emotion recognition method and system
CN121122332A