Pet language translation method and system based on audio learning
By constructing a standardized pet voice library and combining it with audio learning methods, the problem of inaccurate matching in pet translators has been solved, achieving accurate translation and training of pet language.
Patent Information
- Application Number
- CN202511196288.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-08-26
AI Technical Summary
Existing pet translation devices cannot accurately match the vocalizations of pets of different types and ages, leading to mistranslations or incorrect translations, and are unable to verify the accuracy of their translations.
A standardized pet sound library is constructed. Through audio learning methods, the target sound signal is played and a feedback operation is triggered. Pet sounds are collected in real time and matched with the standardized sound library to output semantic translation results of pet needs.
It achieves accurate and reasonable pet language translation, reduces the probability of mistranslation or incorrect translation, and provides training logic that can be stably recognized and replicated.
Smart Images

Figure CN120780863B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of animal training, and in particular to a pet language translation method and system based on audio learning. BACKGROUND
[0002] At present, pets have become indispensable members in many families, and the pet-keeping population continues to expand worldwide. With the upgrading of pet-keeping needs, people no longer satisfy with simple feeding, walking and other basic care, but yearn for deeper interaction with pets. People hope to understand the expression of pet needs and expect pets to respond accurately to their instructions. This demand for two-way communication between pets and owners has driven the continuous development of pet interaction and training related technologies.
[0003] Although some pet translation machine products have appeared on the current market, trying to realize the "translation" of pet language through sound collection and intelligent analysis, there are still many limitations in actual application. However, most of these products rely on simple recognition of pet natural sound, but it is difficult to establish a standardized sound corresponding logic. It is found in research that the core problem of existing pet translation machines is "inability to prove the correctness of translation", for example, different types and ages of pets, even if expressing the same needs (such as hunger), their sound tone and frequency may have significant differences, which leads to the fact that existing translation machines cannot accurately match and are prone to misinterpretation or mistranslation. SUMMARY
[0004] To solve the above technical problems, the present application provides a pet language translation method and system based on audio learning.
[0005] In a first aspect, the present application provides a pet language translation method based on audio learning, comprising: constructing a pet standardized sound library; the pet standardized sound library includes a plurality of sound signals, each sound signal being associated with a pet demand; when a preset scene is triggered, playing a target sound signal and triggering a feedback operation corresponding to the target sound signal; the target sound signal is a sound signal in the pet standardized sound library and related to the preset scene; in the learning stage, when a first sound signal emitted by a pet is detected to match a sound signal in the pet standardized sound library, triggering a feedback operation corresponding to the matching sound signal; real-time collection of pet sound in the current environment, in response to the pet sound matching a sound signal in the pet standardized sound library, outputting a pet demand corresponding to the matching sound signal; the pet demand corresponds to the semantic translation result of the pet sound.
[0006] Optionally, each sound signal in the pet standardized sound library has different acoustic characteristic parameters.
[0007] Optionally, the acoustic feature parameter comprises at least one of a fundamental frequency, a frequency bandwidth, a formant, a time domain feature, an amplitude, a spectral centroid, a mel-frequency cepstral coefficient.
[0008] Optionally, the process of matching the pet sound with the sound signal in the pet standardized sound library comprises: extracting an acoustic feature parameter of the pet sound; performing a cosine similarity calculation on the acoustic feature parameter of the pet sound and acoustic feature parameters of a plurality of sound signals in the pet standardized sound library; and if there is a sound signal with a cosine similarity calculation result greater than a set threshold, the matching is successful.
[0009] Optionally, the set threshold adopts a dynamic updating mode, and the dynamic updating mode comprises: inputting the pet sound into a pre-trained classification matching model to output a target adjustment parameter; wherein the classification matching model is used to identify background interference in the pet sound; and based on the target adjustment parameter, a basic threshold is optimized to obtain the set threshold.
[0010] Optionally, the playing of the target sound signal when the preset scene is triggered comprises: obtaining a work-rest mapping table; wherein the work-rest mapping table comprises a mapping relationship between a plurality of time nodes and pet work-rest rules; in response to a current time node matching a time node in the work-rest mapping table, a corresponding target sound signal is extracted from the pet standardized sound library for playing; wherein the current time node matching the time node in the work-rest mapping table represents triggering the preset scene.
[0011] Optionally, the playing of the target sound signal when the preset scene is triggered comprises: obtaining a behavior feature of the pet; based on the behavior feature of the pet, predicting a demand tendency of the pet; wherein the predicted demand tendency of the pet represents triggering the preset scene; and based on the predicted demand tendency of the pet, a corresponding target sound signal is extracted from the pet standardized sound library for playing.
[0012] Optionally, the playing of the target sound signal when the preset scene is triggered comprises: obtaining position information of the pet and position information of M marker points; M is a positive integer greater than 2; based on the position information of the pet and the position information of the M marker points, determining a relative position relationship between the pet and the M marker points; based on the relative position relationship between the pet and the M marker points, a corresponding target sound signal is extracted from the pet standardized sound library for playing; wherein the relative position relationship between the pet and the M marker points is used to trigger the preset scene.
[0013] Optionally, the feedback operation comprises at least one of controlling the feeder to feed, controlling the waterer to dispense water, controlling the toy device to activate the toy, and triggering the user's communication device to prompt the user about the pet demand.
[0014] In a second aspect, the present application provides a pet language translation system based on audio learning, comprising: a construction module configured to construct a pet standardized sound library; the pet standardized sound library comprises a plurality of sound signals, each sound signal being associated with a pet demand; an audio learning module configured to play a target sound signal and trigger a feedback operation corresponding to the target sound signal when a preset scene is triggered; the target sound signal is a sound signal in the pet standardized sound library and related to the preset scene; a feedback reinforcement module configured to trigger a feedback operation corresponding to a matched sound signal in the pet standardized sound library when a first sound signal emitted by a pet is detected to match the sound signal in the pet standardized sound library during a learning stage; and a translation module configured to collect pet sounds in a current environment in real time, output a pet demand corresponding to a matched sound signal in the pet standardized sound library in response to the pet sound matching the sound signal, and the pet demand corresponding to a semantic translation result of the pet sound.
[0015] The present application provides a pet language translation method based on audio learning, which core is to construct a pet standardized sound library, which is equivalent to defining a language system for pets, and then in the environment where the pet is located, a sound signal in the pet standardized sound library is extracted and played in response to a preset scene trigger, and a control response is performed based on the pet demand corresponding to the sound signal, so that the pet can learn and know what physical feedback it can get by emitting sound in the current language environment, and then when a first sound signal emitted by the pet is detected to match a sound signal in the pet standardized sound library during a learning stage, a feedback operation corresponding to the matched sound signal is triggered to reinforce the feedback. Finally, in the application process, pet sounds in the current environment are collected in real time, and a pet demand corresponding to a matched sound signal in the pet standardized sound library is output in response to the pet sound matching the sound signal, and the pet demand corresponds to a semantic translation result of the pet sound. As can be seen, the above method can provide a stable and reproducible training logic, and can realize accurate and reasonable pet language translation and reduce the probability of misinterpretation or misinterpretation. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 A step flowchart of a pet language translation method based on audio learning provided by the embodiments of the present application;
[0017] Figure 2A flow chart of steps of another pet language translation method based on audio learning provided by an embodiment of the present application;
[0018] Figure 3 A module block diagram of a pet language translation system based on audio learning provided by an embodiment of the present application;
[0019] Figure 4 A module block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0020] In the following description, for purposes of explanation and not limitation, specific details are set forth, such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known systems, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary detail.
[0021] It is found in research that the core problem existing in the existing pet translation machine is "inability to prove the correctness of translation", for example, different types and ages of pets, even if expressing the same demand (such as hunger), the tone and frequency of their sound may differ significantly, which leads to the fact that the existing translation machine cannot accurately match and is prone to mis-translation or wrong translation.
[0022] In view of the above problems, the present application proposes the following embodiments to solve the above technical problems.
[0023] Please refer to Figure 1 The present application provides a pet language translation method based on audio learning, which specifically comprises steps 101-104.
[0024] Step 101: Construct a pet standardized sound library.
[0025] The pet standardized sound library includes a plurality of sound signals, and each sound signal is associated with a pet demand.
[0026] The pet standardized sound library described above can be constructed according to the collected existing pet sound data. According to the determined pet sound data, the pet demand associated with each sound signal is determined. At the same time, the sound signal is stored.
[0027] For example, in the pet standardized sound library, the following corresponding relationship is set:
[0028] The sound signal A corresponds to the demand of the pet: hunger;
[0029] The sound signal B corresponds to the demand of the pet: thirst;
[0030] Sound signal C, corresponding to the pet's demand: playing toys;
[0031] Sound signal D, corresponding to the pet's demand: going out to play.
[0032] The above is only an example of part of the pet standardized sound library, and the specific sound signals and corresponding relationships contained in the specific pet standardized sound library are not limited.
[0033] Step 102: When the preset scene is triggered, the target sound signal is played, and the feedback operation corresponding to the target sound signal is triggered.
[0034] The target sound signal is a sound signal in the pet standardized sound library and related to the preset scene.
[0035] The preset scene can be indicative of the pet eating, the pet playing, the pet drinking, etc. When these scenes are triggered, the sound signal related to the current scene (i.e. the target sound signal) is extracted from the pet standardized sound library for playing. Then, the feedback operation corresponding to the target sound signal is triggered, which corresponds to the preset scene.
[0036] For example, when the preset scene is the pet eating, the sound signal A related to the current scene of the pet eating (such as the sound signal indicative of hunger) is extracted for playing, and then the feeder can be controlled to feed. Further, the pet can understand in this environment that when the sound signal A appears, food can be obtained.
[0037] The above process is a continuous learning and perception process of the pet based on audio, and is also a learning stage of the pet.
[0038] Step 103: In the learning stage, when the first sound signal emitted by the pet is detected to match the sound signal in the pet standardized sound library, the feedback operation corresponding to the matched sound signal is triggered.
[0039] That is, in the continuous learning stage, if the pet itself triggers a sound signal matching the sound signal in the pet standardized sound library, the feedback operation corresponding to the matched sound signal is triggered to realize feedback reinforcement in the learning stage of the pet.
[0040] Step 104: Real-time collection of pet sound in the current environment, in response to the pet sound matching the sound signal in the pet standardized sound library, outputting the pet demand corresponding to the matched sound signal.
[0041] The pet demand corresponds to the semantic translation result of the pet sound.
[0042] That is, after the playing learning of the sound signals in the pet standardized sound library is completed through the foregoing steps, in the actual translation stage, the pet sound in the current environment is collected in real time, and when it is determined that the pet sound matches the sound signal in the pet standardized sound library, the pet demand corresponding to the matched sound signal is output.
[0043] For example, a text prompt can be sent to the terminal device of the pet owner, which represents the translation result of the current pet sound. Of course, the pet demand can also be directly played through the microphone, and the present application is not limited in this regard.
[0044] Considering that the language system is a carrier tool for communication, which maps the content that needs to be expressed (real existence), but the language itself is defined, and there is no such language system definition between the growing environment of pets and different types, which further leads to the fact that the existing translation machine cannot accurately match and is prone to misinterpretation or misinterpretation. Based on this, the pet language translation method based on audio learning provided in the embodiments of the present application is to build a set of pet standardized sound library, which is equivalent to defining a set of language system for pets. Then, in the environment where the pet is located, the sound signal in the pet standardized sound library is extracted and played in response to a preset scene trigger, and a control response is performed based on the pet demand corresponding to the sound signal, so that the pet can learn and know in the current language environment what kind of physical feedback can be obtained by emitting a sound. Then, in the learning stage, when the first sound signal emitted by the pet is detected to match the sound signal in the pet standardized sound library, the feedback operation corresponding to the matched sound signal is triggered, so as to strengthen the feedback. Finally, in the application process, the pet sound in the current environment is collected in real time, and when the pet sound matches the sound signal in the pet standardized sound library, the pet demand corresponding to the matched sound signal is output; the pet demand corresponds to the semantic translation result of the pet sound. As can be seen, the above method can provide a stable recognition and reproducible training logic, which can realize accurate and reasonable pet language translation and reduce the probability of misinterpretation or misinterpretation.
[0045] Optionally, when the preset scene trigger is triggered, the target sound signal is played, and after the preset time of playing the target sound signal, the feedback operation corresponding to the target sound signal is triggered.
[0046] The above-mentioned preset time can be set to 1-3 seconds. That is, the reaction time of the pet to the sound is provided. In other words, the preset time can give the pet itself a learning feedback time, for example, when the pet hears the target sound signal, it will react for a while, and then the corresponding physical feedback is provided so that the pet can effectively familiarize the association between the sound and the feedback response.
[0047] Optionally, each sound signal in the pet standardized sound library has different acoustic characteristic parameters.
[0048] That is, the process of matching the voice signal of the pet with the voice signal in the pet standardized voice library can be based on the matching of the acoustic characteristic parameters.
[0049] Optionally, the acoustic characteristic parameters include at least one of a fundamental frequency, a frequency bandwidth, a formant, a time domain feature, an amplitude, a spectral centroid, and a mel-frequency cepstral coefficient.
[0050] It can be understood that the acoustic characteristic parameters possessed by each voice signal can be one of the above examples, or a combination of at least two. For example, the acoustic characteristic parameters of the voice signal can include the fundamental frequency and the amplitude.
[0051] The above-listed acoustic characteristic parameters are explained as follows. The fundamental frequency is the most basic frequency in the voice signal, which determines the pitch of the voice. For example, the "whimpering sound" emitted by the pet usually has a low fundamental frequency, while the "excited barking sound" has a high fundamental frequency. The pet can quickly perceive the basic tonality of the voice through the difference in the fundamental frequency.
[0052] The frequency bandwidth refers to the frequency range contained in the voice signal, that is, the difference between the highest frequency and the lowest frequency. For example, the "barking sound" of the pet has a wide frequency bandwidth, while the "slight whispering sound" has a narrow frequency bandwidth.
[0053] The formant is the frequency area where the energy is concentrated due to resonance during the propagation of the voice, which determines the key features of the voice timbre.
[0054] The time domain feature mainly describes the change characteristics of the voice in the time dimension, covering the duration, rhythm, waveform change, etc. of the voice. For example, the "short warning sound" of the pet has a short duration and a rapid rhythm in the time domain, while the "coy humming sound" has a longer duration and a more gentle waveform. The pet can judge the emotion conveyed by the voice through these time dimension features.
[0055] The amplitude reflects the energy size of the voice signal, which directly corresponds to the loudness of the voice. The larger the amplitude, the louder the voice. For example, the "excited barking" of the pet has a large amplitude, while the "dejected low humming" has a small amplitude. The pet can perceive the strength of the voice through the amplitude, and further judge the urgency of the voice, etc.
[0056] The spectral centroid is the center of gravity position of the energy distribution of the voice spectrum, which can reflect the "brightness" of the voice. The higher the spectral centroid, the brighter and sharper the voice sounds; otherwise, it is dull. For example, the "sharp squeaking sound" of the pet has a high spectral centroid, while the "low growling sound" has a low spectral centroid, which will affect the pet's perception tendency of the voice.
[0057] Mel-frequency cepstral coefficients (MFCCs) are feature parameters extracted based on the perceptual characteristics of sound, which can effectively capture the spectral envelope information of sound and have good noise resistance. In pet sound applications, it can better fit the auditory perceptual characteristics of pets and improve the accuracy of sound recognition and association.
[0058] Optionally, the feedback operation includes at least one of controlling the feeder to feed, controlling the water dispenser to dispense water, controlling the toy device to start the toy, and triggering the communication device of the user to prompt the user about the pet demand.
[0059] Example one, when the pet demand represented by the target sound signal is "hunger", the feedback operation is to control the feeder to feed. Specifically, the feedback operation control instruction is sent to the feeder to open the discharge port and release a set amount of pet food (the set amount can be set according to the weight of the pet), so that the pet gradually associates the sound signal with "eating".
[0060] Example two, when the pet demand represented by the target sound signal is "thirst", the feedback operation is to control the water dispenser to dispense water. Then, the feedback operation control instruction is sent to the water dispenser to open the water outlet and release a set amount of water, so that the pet gradually associates the sound signal with "drinking water".
[0061] Example three, when the pet demand represented by the target sound signal is "toy", the feedback operation is to control the toy device to start the toy. Then, the feedback operation control instruction is sent to the toy device to start the toy. Specifically, the toy device can be a cat toy, which can swing after being started. The toy can also be a ball launching device, which can launch a ball unless the ball is launched. The above feedback control makes the pet gradually associate the sound signal with "playing with toys".
[0062] Example four, when the pet demand represented by the target sound signal is "going out to play", the feedback operation is to trigger the communication device of the user to prompt the user about the pet demand. For example, the user's communication device is reminded by a short message or an application to prompt the pet owner to take the pet out to play. The above feedback control makes the pet gradually associate the sound signal with "going out to play".
[0063] Please refer to Figure 2 Optionally, the process of matching the pet sound with the sound signal in the pet standardized sound library can specifically include steps 201-203.
[0064] Step 201: Extracting acoustic feature parameters of the pet sound.
[0065] The acoustic feature parameters can refer to the description in the foregoing embodiments, which will not be repeated here.
[0066] Step 202: Perform cosine similarity calculation on the acoustic feature parameters of the pet sound and the acoustic feature parameters of the plurality of sound signals in the pet standardized sound library.
[0067] Step 203: If there is a sound signal with a cosine similarity calculation result greater than the set threshold, the matching is successful.
[0068] The set threshold can be 0.85, 0.8, 0.7, etc., and the specific value is not limited.
[0069] Specifically, in the learning phase, when the pet dog emits a first sound signal, the sound is collected by the microphone, and the acoustic feature parameter A is extracted. Then, the acoustic feature parameter A is calculated with the acoustic feature parameters of the plurality of sound signals in the standardized sound library, and the cosine similarity between the acoustic feature parameter A and the acoustic feature parameter B is calculated to be 0.88. Since 0.88>0.85 (set threshold), the comparison is successful, that is, the pet dog's current sound expression is identified as hunger demand, and the feedback behavior of the feeder is triggered to feed, thereby strengthening the learning feedback.
[0070] In actual application, the pet sound is collected by the microphone, and the acoustic feature parameter A is extracted. Then, the acoustic feature parameter A is calculated with the acoustic feature parameters of the plurality of sound signals in the standardized sound library, and the cosine similarity between the acoustic feature parameter A and the acoustic feature parameter B is calculated to be 0.88. Since 0.88>0.85 (set threshold), the comparison is successful, that is, the pet dog's current sound expression is identified as hunger demand, and the pet demand is output to represent the translation result of the pet sound.
[0071] It should be noted that the cosine similarity calculation can effectively measure the similarity of two acoustic feature parameters, improve the recognition accuracy of the pet's real demand, and complete the comparison of the pet sound and the standardized sound library in a short time, quickly determine the pet demand and trigger the corresponding feedback.
[0072] In addition, if the pet sound is not matched successfully, the unknown type of pet sound can be collected and uploaded to the server for subsequent in-depth analysis to identify new features.
[0073] Optionally, the set threshold adopts a dynamic updating mode, and the dynamic updating mode includes: inputting the pet sound into a pre-trained classification matching model to output a target adjustment parameter; wherein the classification matching model is used to identify the background interference in the pet sound; based on the target adjustment parameter, the basic threshold is optimized to obtain the set threshold.
[0074] In an embodiment, the classification matching model is used to identify the background interference level in the pet sound, such as a quiet environment, the background interference level is 0, a lower noise environment, the background interference level is 1, a medium noise environment, the background interference level is 2, and a higher noise environment, the background interference level is 3. Different background interference levels correspond to different target adjustment parameters.
[0075] For example, the basic threshold can be set to 0.7. The target adjustment parameter corresponding to the background interference level 0 can be 0, and the corresponding set threshold is 0.7+0=0.7. The target adjustment parameter corresponding to the background interference level 1 can be 0.05, and the corresponding set threshold is 0.7+0.05=0.75. The target adjustment parameter corresponding to the background interference level 2 can be 0.08, and the corresponding set threshold is 0.7+0.08=0.78. The target adjustment parameter corresponding to the background interference level 3 can be 0.1, and the corresponding set threshold is 0.7+0.1=0.8. The target adjustment parameter corresponding to the background interference level 3 can be 0.12, and the corresponding set threshold is 0.7+0.12=0.82.
[0076] It should be noted that strong interference can distort the acoustic characteristics of pet sounds, which may overlap with parameters such as frequency and amplitude of pet calls, resulting in deviation of the extracted pet sound feature vector. Therefore, by increasing the threshold, those sounds that can still maintain high similarity even under strong interference (i.e. pet calls whose core features are not severely damaged) can be screened out, reducing such misjudgments.
[0077] Of course, in another embodiment, the classification matching model can be used to identify the background interference type in the pet sound, such as identifying television sound interference, traffic interference, thunderstorm interference, etc., and different target adjustment parameters are set for different background interference types for adjustment.
[0078] In summary, the embodiments of the present application can improve the recognition anti-interference ability in complex environments by using a dynamic updating method for the set threshold, enhance the adaptability of the model to pet sound interference, and reduce the probability of false negatives and false positives.
[0079] The preset scene triggering strategy provided by the embodiments of the present application is described below.
[0080] In an embodiment, when the preset scene is triggered, the target sound signal is played, which can specifically include: obtaining a rest mapping table; wherein the rest mapping table includes a mapping relationship between a plurality of time nodes and pet rest habits; in response to the current time node matching the time node in the rest mapping table, the corresponding target sound signal is extracted from the pet standardized sound library for playing; wherein the current time node matches the time node in the rest mapping table, indicating that the preset scene is triggered.
[0081] For example, the above schedule mapping table can be constructed according to historical data. For example, according to the schedule of a pet dog, such as eating from 7:00 to 9:00 in the morning, playing with toys from 2:00 to 3:00 in the afternoon, and going out to play from 6:00 to 8:00 in the evening. The schedule mapping table can be constructed according to the schedule.
[0082] Specifically, when it is determined that the current time node is 7:30 in the morning, according to the schedule mapping table, the target sound signal representing the pet's demand as "hunger" can be extracted from the pet standardization sound library for playing.
[0083] When it is determined that the current time node is 2:00 in the afternoon, according to the schedule mapping table, the target sound signal representing the pet's demand as "toy" can be extracted from the pet standardization sound library for playing.
[0084] Since pets have a strong time perception instinct, such as expecting to eat at a fixed time, in the embodiments of the present application, the sound signal is bound to the schedule, so that the pet is in a "time-sound-demand" cycle, the memory of the sound signal at each time node is strengthened, the association between the sound and the demand is more profound, the training effect is more excellent, and the pet can make more focused learning and training on the sound in a specific time period. can avoid the cognitive confusion of pets caused by random or irregular sounds, so that the pet can make more focused learning and training on the sound in a specific time period.
[0085] In another embodiment, playing a target sound signal can specifically include: obtaining a behavior feature of the pet; predicting a demand tendency of the pet based on the behavior feature of the pet; wherein the predicted demand tendency of the pet represents triggering a preset scene; and extracting a corresponding target sound signal from a pet standardization sound library based on the predicted demand tendency of the pet for playing.
[0086] Generally, the behavior feature of the pet can show the current demand tendency of the pet. For example, the pet scratches the edge of the food bowl with its claws, which can be determined as the pet having a "eating" demand. For example, the pet scratches the door with its claws, which can be determined as the pet having a "going out to play" demand tendency.
[0087] Further, when it is identified that the pet scratches the edge of the food bowl with its claws, it can be determined that the pet has a "eating" tendency, and a target sound signal representing the pet's demand as "hunger" can be extracted from the pet standardization sound library for playing.
[0088] It should be noted that the above method can quickly respond when the pet actively shows the demand behavior, avoid missing the best training opportunity, make the training more suitable for the real-time state of the pet, and strengthen the immediate linkage memory of "behavior-sound-feedback". Moreover, the method can adapt to individual behavior differences of the pet, improve the training flexibility, and strengthen the pet's active participation, and improve the interactivity of the training.
[0089] In another embodiment, the playing of the target sound signal triggered by the preset scene can specifically include: obtaining the position information of the pet and the position information of the M markers; M is a positive integer greater than 2; determining the relative position relationship between the pet and the M markers based on the position information of the pet and the position information of the M markers; extracting the corresponding target sound signal from the pet standardized sound library based on the relative position relationship between the pet and the M markers to play; wherein the relative position relationship between the pet and the M markers is used to trigger the preset scene.
[0090] The above-mentioned markers can specifically refer to the point of the "feeder", the point of the "water feeder", the point of the "toy device", and the point of the "front door of the home". The above-mentioned relative position relationship can refer to the relative distance (such as the straight-line distance) between the pet and the M markers.
[0091] When it is detected that the straight-line distance between the pet position and the point of the "feeder" is less than 0.3 meters (i.e., the relative position relationship between the pet and the point of the "feeder"), the target sound signal associated with "hunger" is extracted from the pet standardized sound library to play. The above-mentioned distance value can be set according to the actual environment, such as 0.5 meters, 0.1 meters, etc.
[0092] Of course, in this embodiment, other playing trigger conditions can also be set, such as when it is detected that the straight-line distance between the pet position and the point of the "feeder" is less than 0.3 meters for 5 times, the target sound signal associated with "hunger" is extracted from the pet standardized sound library to play. The purpose of setting this playing trigger condition is to ensure that the pet currently indeed has the need to eat. If it is detected that the straight-line distance between the pet position and the point of the "feeder" is less than 0.3 meters for 5 times, it can represent that the pet repeatedly wanders around the "feeder".
[0093] As can be seen, the above-mentioned manner can be adapted to the scene, strengthen the pet's cognition and memory of the space, and strengthen the instant linkage memory of "space-sound-feedback".
[0094] Through the above manner, the training of the pet language can be realized. In addition, it needs to be noted that the training time can be set as a fixed period, such as the pet can be continuously trained for one month or three months as a fixed period, and the sound signals in the pet standardized sound library are played at intervals within the fixed period, so that each sound signal reaches the set playing times. For example, the set playing times can be 50 times or 100 times to ensure that the pet can effectively learn and train.
[0095] In an embodiment, considering that the sound of different types of pets is greatly different, different pet standardized sound libraries can be constructed for different types of pets.
[0096] For example, different pet standardized sound libraries are constructed for pet cats and pet dogs. The pet standardized sound library for pet cats can be constructed according to the known sound data of the pet cats collected. The pet standardized sound library for pet dogs can be constructed according to the known sound data of the pet dogs collected.
[0097] In actual application, different pet standardized sound libraries can be obtained according to different types of pets selected.
[0098] It should be noted that the exclusive standardized sound library constructed for different types of pets can fully adapt to the sound characteristics of various pets, significantly improve the accuracy and acceptance of the sound signal perceived by the pets, and construct a sound environment suitable for different pets, so as to realize efficient and accurate training of pets of different types, and further improve the universality and adaptability of the training method.
[0099] Of course, in another embodiment, different pet standardized sound libraries can also be constructed for pets of different breeds of the same type.
[0100] For example, the breeds of pet dogs include huskies and pugs. The pet standardized sound library for pet dogs of the husky breed can be constructed according to the known sound data of the pet dogs of the husky breed collected. The pet standardized sound library for pet dogs of the pug breed can be constructed according to the known sound data of the pet dogs of the pug breed collected.
[0101] In actual application, different pet standardized sound libraries can be obtained according to different breeds of pets selected.
[0102] It should be noted that the exclusive standardized sound library constructed for different breeds of pets of the same type can accurately adapt to the sound differences between the breeds, and can further enhance the accuracy of training and the stability of training results.
[0103] Please refer to Figure 3Based on the same inventive concept, the application provides a pet language translation system 300 based on audio learning, comprising: a construction module 301, configured to construct a pet standardized sound library; the pet standardized sound library comprises a plurality of sound signals, each sound signal being associated with a pet demand; an audio learning module 302, configured to play a target sound signal and trigger a feedback operation corresponding to the target sound signal when a preset scene is triggered; the target sound signal is a sound signal in the pet standardized sound library and related to the preset scene; a feedback reinforcement module 303, configured to, in a learning stage, trigger a feedback operation corresponding to a matched sound signal when detecting that a first sound signal emitted by a pet matches a sound signal in the pet standardized sound library; a translation module 304, configured to collect pet sounds in a current environment in real time, and output a pet demand corresponding to the matched sound signal in response to the pet sounds matching a sound signal in the pet standardized sound library; the pet demand corresponds to a semantic translation result of the pet sounds.
[0104] Please refer to Figure 4 Based on the same inventive concept, the application provides a module frame of an electronic device 400 applying the above method. The electronic device 400 comprises at least one processor 401 (only one is shown in the figure), a memory 402, a computer program 403 stored in the memory 402 and executable on the at least one processor 401, and the processor 401 implements the steps of the method in any of the preceding embodiments when executing the computer program 403. Figure 4 The electronic device 400 can be a device associated with a pet, such as a pet smart collar, a pet foot ring, a pet monitoring device, and a smart pet door. It can also be a household device, such as a household security monitoring device, and the application does not limit it.
[0105] Those skilled in the art can understand that
[0106] The electronic device 400 is only an example and does not limit the electronic device 400, which can include more or fewer components than shown, or combine certain components, or different components. For example, positioning sensors, cameras, speakers, microphones, and the like. Figure 4
[0107] The processor 401 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0108] The memory 402 can be an internal storage unit of the electronic device 400, such as a hard disk or a memory of the electronic device 400 in some embodiments. The memory 402 can also be an external storage device of the electronic device 400, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 400 in other embodiments. Further, the memory 402 can include both the internal storage unit and the external storage device of the electronic device 400.
[0109] It should be noted that the above system, device, etc. are based on the same concept as the method embodiments of the present application, and the modules designed by the system and the steps performed by the device and the resulting technical effects can be referred to the method embodiments part, which will not be repeated here.
[0110] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the above described functions. Each functional unit or module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of each functional unit or module are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the unit or module in the system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0111] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to realize the steps in the above-mentioned various method embodiments.
[0112] The embodiment of the present application provides a computer program product. When the computer program product is run on a mobile terminal, the mobile terminal is caused to execute the steps in the above-mentioned various method embodiments.
[0113] The integrated unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the present application realizes all or part of the processes in the above-mentioned embodiment methods, which can be completed by instructing related hardware through a computer program. The computer program can be stored in a computer readable storage medium. The computer program is executed by a processor to realize the steps in the above-mentioned various method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium at least includes any entity or system capable of carrying the computer program code to the photographing system / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal and a software distribution medium. For example, a U disk, a mobile hard disk, a magnetic disk or an optical disk.
[0114] In the above-mentioned embodiments, the description of each embodiment has its own focus. The parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0115] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in the present application can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized by hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0116] In addition, in the description of the present application and the appended claims, the terms "first", "second", "third" and the like are only used for differentiation and cannot be understood as indicating or implying relative importance.
[0117] In the embodiments provided by the present application, it should be understood that the disclosed system / device and method can be implemented in other manners. For example, the embodiments of the system / device described above are merely illustrative, and the division of the modules or units can be different from the above. For example, one or more units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be implemented by using some interfaces, and the indirect couplings or communication connections can be implemented in electronic, mechanical or other forms.
[0118] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.
[0119] The above-described embodiments are merely used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent ones; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A pet language translation method based on audio learning, characterized in that, include: Construct a standardized pet sound library, which includes a variety of sound signals, each of which is associated with a specific pet need; When a preset scenario is triggered, a target sound signal is played, and a feedback operation corresponding to the target sound signal is triggered after a preset time of playing the target sound signal; the target sound signal is a sound signal in the pet standardized sound library that is related to the preset scenario; During the learning phase, when the first sound signal emitted by the pet is detected to match the sound signal in the pet standardized sound library, a feedback operation corresponding to the matched sound signal is triggered. The system collects pet sounds in the current environment in real time, and responds by matching the pet sounds with sound signals in the standardized pet sound library, outputting the pet's needs based on the matched sound signals. The semantic translation result of the pet's voice corresponds to the pet's needs; The step of playing a target sound signal when a preset scenario is triggered includes: obtaining a schedule mapping table; wherein the schedule mapping table includes multiple time points and mapping relationships between pet schedules; in response to a current time point matching a time point in the schedule mapping table, extracting the corresponding target sound signal from the pet's standardized sound library and playing it; wherein a match between the current time point and a time point in the schedule mapping table indicates that a preset scenario has been triggered, or... This includes: acquiring the location information of a pet and the location information of M marker points; M is a positive integer greater than 2; determining the relative positional relationship between the pet and the M marker points based on the location information of the pet and the location information of the M marker points; extracting the corresponding target sound signal from the pet's standardized sound library and playing it based on the relative positional relationship between the pet and the M marker points; wherein, the relative positional relationship between the pet and the M marker points is used to trigger a preset scene.
2. The pet language translation method based on audio learning according to claim 1, characterized in that, In the standardized pet sound library, each sound signal has different acoustic characteristic parameters.
3. The pet language translation method based on audio learning according to claim 2, characterized in that, The acoustic characteristic parameters include at least one of the following: fundamental frequency, frequency bandwidth, formant, time-domain characteristics, amplitude, spectral centroid, and Mel-frequency cepstral coefficients.
4. The pet language translation method based on audio learning according to claim 2, characterized in that, The process of matching the pet's sound with the sound signals in the standardized pet sound library includes: Extract the acoustic feature parameters of the pet's sounds; The acoustic feature parameters of the pet's sound are compared with the acoustic feature parameters of various sound signals in the standardized pet sound library using cosine similarity calculation. If there is a sound signal whose cosine similarity calculation result is greater than the set threshold, then the match is successful.
5. The pet language translation method based on audio learning according to claim 4, characterized in that, The set threshold is dynamically updated, and the dynamic update method includes: The pet sound is input into a pre-trained classification and matching model, which outputs target adjustment parameters; wherein, the classification and matching model is used to identify background interference in the pet sound; Based on the target adjustment parameters, the basic threshold is optimized to obtain the set threshold.
6. The pet language translation method based on audio learning according to claim 1, characterized in that, The feedback operations include at least one of the following: controlling the feeder to feed, controlling the waterer to dispense water, controlling the toy device to start the toy, and triggering the user's communication device to prompt the user about the pet's needs.
7. A pet language translation system based on audio learning, characterized in that, include: Build modules are used to create a standardized pet sound library; The standardized pet sound library includes a variety of sound signals, and each sound signal is associated with a pet's needs. The audio learning module is used to play a target sound signal when a preset scenario is triggered, and to trigger a feedback operation corresponding to the target sound signal after a preset time of playing the target sound signal; the target sound signal is a sound signal in the pet standardized sound library that is related to the preset scenario; The feedback reinforcement module is used to trigger a feedback operation corresponding to the matched sound signal when the first sound signal emitted by the pet is detected to match the sound signal in the pet standardized sound library during the learning phase. The translation module is used to collect pet sounds in the current environment in real time, and in response to the matching of the pet sounds with the sound signals in the standardized pet sound library, outputs the pet's needs with the matched sound signals; the pet's needs correspond to the semantic translation results of the pet sounds; The audio learning module is specifically used to acquire a schedule mapping table; wherein the schedule mapping table includes multiple time points and mapping relationships with pets' schedule patterns; in response to a current time point matching a time point in the schedule mapping table, the corresponding target sound signal is extracted from the pet's standardized sound library and played; wherein a match between the current time point and a time point in the schedule mapping table indicates the triggering of a preset scenario, or... Specifically, it is used to obtain the location information of the pet and the location information of M marker points; M is a positive integer greater than 2; based on the location information of the pet and the location information of the M marker points, the relative positional relationship between the pet and the M marker points is determined; based on the relative positional relationship between the pet and the M marker points, the corresponding target sound signal is extracted from the pet standardized sound library and played; wherein, the relative positional relationship between the pet and the M marker points is used to trigger a preset scene.
Citation Information
Patent Citations
Infant or pet nursing method and nursing system and nursing machine adopting method
CN103985383A
Intelligent feeding method and device based on animal sound
CN117995200A