Pet language translation method and system based on audio learning

By building a standardized pet sound library and conducting audio learning, the problem of pet translators being unable to accurately match has been solved, and accurate translation and stable recognition of pet language have been achieved.

CN120780863AActive Publication Date: 2025-10-14SHENZHEN KOLAMAMA TECH CO LTD

Patent Information

Application Number
CN202511196288.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-10-14
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing pet translation machines cannot accurately match the vocalizations of pets of different types and ages, resulting in mistranslations or incorrect translations, and are unable to self-verify the correctness of the translation.

Method used

Build a standardized pet sound library, use audio learning methods to play target sound signals and trigger feedback operations, collect pet sounds in real time and match them with the standardized sound library, and output semantic translation results of pet needs.

Benefits of technology

It achieves accurate and reasonable pet language translation, reduces the probability of mistranslation or incorrect translation, and provides training logic that can be stably recognized and replicated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780863A_ABST
    Figure CN120780863A_ABST
Patent Text Reader

Abstract

The invention relates to a pet language translation method and system based on audio learning, and belongs to the technical field of animal training. The method comprises the following steps: constructing a pet standardized sound library; when the preset scene is triggered, playing the target sound signal; the target sound signal is a sound signal related to a preset scene in a pet standardized sound library; when it is detected that the first sound signal sent by the pet is matched with the sound signal in the pet standardized sound library, triggering a feedback operation corresponding to the matched sound signal; collecting pet sound in the current environment in real time, responding to the matching of the pet sound and the sound signal in the pet standardized sound library, and outputting a pet demand of the matched sound signal; the pet demand corresponds to a semantic translation result of the pet sound. By means of the mode, the training logic which can be stably recognized and can be copied and executed can be provided, accurate and reasonable pet language translation can be achieved, and the probability of mistranslation or wrong translation is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of animal training, and in particular to a pet language translation method and system based on audio learning. BACKGROUND

[0002] At present, pets have become indispensable members in many families, and the pet-keeping population continues to expand worldwide. With the upgrading of pet-keeping needs, people no longer satisfy with simple feeding, walking and other basic care, but yearn for deeper interaction with pets. People hope to understand the expression of pet needs and expect pets to respond accurately to their instructions. This demand for two-way communication between pets and owners has driven the continuous development of pet interaction and training related technologies.

[0003] Although some pet translation machine products have appeared on the current market, trying to realize the "translation" of pet language through sound collection and intelligent analysis, there are still many limitations in actual application. However, most of these products rely on simple recognition of pet natural sound, but it is difficult to establish a standardized sound corresponding logic. It is found in research that the core problem of existing pet translation machines is "inability to prove the correctness of translation", for example, different types and ages of pets, even if expressing the same needs (such as hunger), their sound tone and frequency may have significant differences, which leads to the fact that existing translation machines cannot accurately match and are prone to misinterpretation or mistranslation. SUMMARY

[0004] To solve the above technical problems, the present application provides a pet language translation method and system based on audio learning.

[0005] In a first aspect, the present application provides a pet language translation method based on audio learning, comprising: constructing a pet standardized sound library; the pet standardized sound library includes a plurality of sound signals, each sound signal being associated with a pet demand; when a preset scene is triggered, playing a target sound signal and triggering a feedback operation corresponding to the target sound signal; the target sound signal is a sound signal in the pet standardized sound library and related to the preset scene; in the learning stage, when a first sound signal emitted by a pet is detected to match a sound signal in the pet standardized sound library, triggering a feedback operation corresponding to the matching sound signal; real-time collection of pet sound in the current environment, in response to the pet sound matching a sound signal in the pet standardized sound library, outputting a pet demand corresponding to the matching sound signal; the pet demand corresponds to the semantic translation result of the pet sound.

[0006] Optionally, each sound signal in the pet standardized sound library has different acoustic characteristic parameters.

[0007] Optionally, the acoustic feature parameters include at least one of fundamental frequency, frequency bandwidth, resonance peak, time domain characteristics, amplitude, spectrum centroid, and Mel-frequency cepstral coefficients.

[0008] Optionally, the process of matching the pet sound with the sound signal in the pet standardized sound library includes: extracting the acoustic feature parameters of the pet sound; performing cosine similarity calculation on the acoustic feature parameters of the pet sound and the acoustic feature parameters of multiple sound signals in the pet standardized sound library; if there is a sound signal whose cosine similarity calculation result is greater than a set threshold, the match is successful.

[0009] Optionally, the set threshold is dynamically updated, and the dynamic update method includes: inputting the pet sound into a pre-trained classification matching model and outputting a target adjustment parameter; wherein the classification matching model is used to identify background interference in the pet sound; based on the target adjustment parameter, optimizing the basic threshold to obtain the set threshold.

[0010] Optionally, when a preset scene is triggered, playing the target sound signal includes: obtaining a work and rest mapping table; wherein the work and rest mapping table includes a mapping relationship between multiple time nodes and the pet's work and rest patterns; in response to the current time node matching the time node in the work and rest mapping table, extracting the corresponding target sound signal from the pet standardized sound library for playing; wherein the current time node matching the time node in the work and rest mapping table represents the triggering of the preset scene.

[0011] Optionally, when a preset scene is triggered, playing the target sound signal includes: obtaining the behavioral characteristics of the pet; predicting the pet's demand tendencies based on the pet's behavioral characteristics; wherein the predicted demand tendencies of the pet represent the triggering of the preset scene; and based on the predicted demand tendencies of the pet, extracting the corresponding target sound signal from the pet standardized sound library for playing.

[0012] Optionally, when a preset scene is triggered, playing the target sound signal includes: obtaining the position information of the pet and the position information of M identification points; M is a positive integer greater than 2; based on the position information of the pet and the position information of the M identification points, determining the relative position relationship between the pet and the M identification points; based on the relative position relationship between the pet and the M identification points, extracting the corresponding target sound signal from the pet standardized sound library for playing; wherein the relative position relationship between the pet and the M identification points is used to trigger the preset scene.

[0013] Optionally, the feedback operation comprises at least one of controlling the feeder to feed, controlling the water dispenser to dispense water, controlling the toy device to start the toy, and triggering the communication device of the user to prompt the user about the pet demand.

[0014] In a second aspect, the present application provides a pet language translation system based on audio learning, comprising: a construction module configured to construct a pet standardized sound library; the pet standardized sound library comprises a plurality of sound signals, each sound signal being associated with a pet demand; an audio learning module configured to play a target sound signal and trigger a feedback operation corresponding to the target sound signal when a preset scene is triggered; the target sound signal is a sound signal in the pet standardized sound library and related to the preset scene; a feedback reinforcement module configured to trigger a feedback operation corresponding to a matched sound signal in the pet standardized sound library when a first sound signal emitted by a pet is detected to match the sound signal in the pet standardized sound library during a learning stage; and a translation module configured to collect pet sounds in a current environment in real time, output a pet demand corresponding to a matched sound signal in the pet standardized sound library in response to the pet sound matching the sound signal, and the pet demand corresponding to a semantic translation result of the pet sound.

[0015] The present application provides a pet language translation method based on audio learning, which core is to construct a pet standardized sound library, which is equivalent to defining a language system for pets, and then in the environment where the pet is located, a sound signal in the pet standardized sound library is extracted and played in response to a preset scene trigger, and a control response is performed based on the pet demand corresponding to the sound signal, so that the pet can learn and know what physical feedback it can get in the current language environment, and then when a first sound signal emitted by the pet is detected to match a sound signal in the pet standardized sound library during a learning stage, a feedback operation corresponding to the matched sound signal is triggered to reinforce the feedback. Finally, in the application process, pet sounds in the current environment are collected in real time, and a pet demand corresponding to a matched sound signal in the pet standardized sound library is output in response to the pet sound matching the sound signal, and the pet demand corresponds to a semantic translation result of the pet sound. As can be seen, the above method can provide a stable and reproducible training logic, and can realize accurate and reasonable pet language translation and reduce the probability of misinterpretation or misinterpretation. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 A step flowchart of a pet language translation method based on audio learning provided by an embodiment of the present application; Figure 2 A step flowchart of another pet language translation method based on audio learning provided by an embodiment of the present application; Figure 3 A module block diagram of a pet language translation system based on audio learning provided by an embodiment of the present application is shown in FIG. 1. Figure 4 A module block diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 2. DETAILED DESCRIPTION

[0017] In the following description, for purposes of explanation and not limitation, specific details are set forth, such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, and

[0018] It is found in the research that the core problem existing in the existing pet translation machine is "inability to prove the correctness of translation", for example, different types and ages of pets, even if expressing the same demand (such as hunger), the tone and frequency of the sound may be significantly different, thereby causing the existing translation machine to be unable to accurately match, and easy to appear mis-translation or wrong translation.

[0019] In view of the above problems, the present application proposes the following embodiments to solve the above technical problems.

[0020] Please refer to Figure 1 The present application provides a pet language translation method based on audio learning, which specifically comprises steps 101-104.

[0021] Step 101: Construct a pet standardized sound library.

[0022] The pet standardized sound library includes a plurality of sound signals, and each sound signal is associated with a pet demand.

[0023] The pet standardized sound library described above can be constructed according to the collected existing pet sound data. According to the determined pet sound data, the pet demand associated with each sound signal is determined. At the same time, the sound signal is stored.

[0024] For example, in the pet standardized sound library, the following corresponding relationship is set: Sound signal A, its corresponding pet demand is hunger; Sound signal B, its corresponding pet demand is thirst; Sound signal C, its corresponding pet demand is toy; Sound signal D, its corresponding pet demand is to go out and play.

[0025] The above is only an example of the pet standardized sound library, and the specific sound signals and corresponding relationships included in the specific pet standardized sound library are not limited.

[0026] Step 102: When the preset scene is triggered, the target sound signal is played, and the feedback operation corresponding to the target sound signal is triggered.

[0027] The target sound signal is a sound signal in the pet standardized sound library and related to the preset scene.

[0028] The preset scene can be indicative of the pet eating, the pet playing, the pet drinking, and the like. When these scenes are triggered, the sound signal (i.e., the target sound signal) related to the current scene is extracted from the pet standardized sound library and played. Then, the feedback operation corresponding to the target sound signal is triggered, which corresponds to the preset scene.

[0029] For example, when the preset scene is the pet eating, the sound signal A (e.g., a sound signal indicative of hunger) related to the current scene of the pet eating is extracted and played, and then the feeder can be controlled to feed. Thus, the pet can understand in this environment that when the sound signal A appears, the pet can obtain food.

[0030] The above process is a continuous learning and perception process of the pet based on audio, and is also a learning stage of the pet.

[0031] Step 103: In the learning stage, when it is detected that the first sound signal emitted by the pet matches a sound signal in the pet standardized sound library, a feedback operation corresponding to the matched sound signal is triggered.

[0032] That is, in the continuous learning stage, if the pet itself triggers a sound signal that matches a sound signal in the pet standardized sound library, a feedback operation corresponding to the matched sound signal is triggered to achieve feedback reinforcement in the learning stage of the pet.

[0033] Step 104: Real-time collection of pet sounds in the current environment, in response to the pet sounds matching a sound signal in the pet standardized sound library, outputting the pet demand corresponding to the matched sound signal.

[0034] The pet demand corresponds to the semantic translation result of the pet sound.

[0035] That is, after the foregoing steps complete the playing and learning of the sound signals in the pet standardized sound library, in the actual translation stage, the pet sounds in the current environment are collected in real time, and when it is determined that the pet sounds match a sound signal in the pet standardized sound library, the pet demand corresponding to the matched sound signal is output.

[0036] For example, a text prompt can be sent to the pet owner's terminal device indicating that the pet needs, which represents the translation result of the current pet sound. Of course, the pet needs can also be directly played through the microphone, and the present application is not limited in this regard.

[0037] Considering that the language system is a carrier tool for communication, which maps the content that needs to be expressed (real existence), but the language itself is defined, and there is no such language system definition between the growth environment of pets and different types, which further leads to the fact that the existing translation machine cannot accurately match and is prone to misinterpretation or misinterpretation. Based on this, the pet language translation method based on audio learning provided by the embodiments of the present application is to build a set of pet standardized sound library, which is equivalent to defining a set of language system for pets. Then, in the environment where the pet is located, the sound signal in the pet standardized sound library is extracted and played in response to the preset scene trigger, and the control response is performed based on the pet needs corresponding to the sound signal, so that the pet can learn and know what kind of physical feedback can be obtained by emitting sound in the current language environment. Then, in the learning stage, when the first sound signal emitted by the pet is detected to match the sound signal in the pet standardized sound library, the feedback operation corresponding to the matched sound signal is triggered to strengthen the feedback. Finally, in the application process, the pet sound in the current environment is collected in real time, and the pet needs corresponding to the matched sound signal are output in response to the pet sound matching the sound signal in the pet standardized sound library; the pet needs correspond to the semantic translation result of the pet sound. As can be seen, the above method can provide a stable recognition and reproducible training logic, which can realize accurate and reasonable pet language translation and reduce the probability of misinterpretation or misinterpretation.

[0038] Optionally, when the preset scene is triggered, the target sound signal is played, and after playing the target sound signal for a preset time, the feedback operation corresponding to the target sound signal is triggered.

[0039] The above-mentioned preset time can be set to 1-3 seconds. That is, the reaction time of the pet to the sound is provided. In other words, the preset time can give the pet its own learning feedback time. For example, when the pet hears the target sound signal, it will react, and at this time, the corresponding physical feedback is provided so that the pet can effectively familiarize the association between the sound and the feedback response.

[0040] Optionally, each sound signal in the pet standardized sound library has different acoustic characteristic parameters.

[0041] That is, the process of matching the pet sound signal with the sound signal in the pet standardized sound library can be based on acoustic characteristic parameters.

[0042] Optionally, the acoustic feature parameters include at least one of fundamental frequency, frequency bandwidth, resonance peak, time domain characteristics, amplitude, spectrum centroid, and Mel-frequency cepstral coefficients.

[0043] It is understandable that the acoustic characteristic parameters of each sound signal may be one of the above examples, or a combination of at least two. For example, the acoustic characteristic parameters of a sound signal may include fundamental frequency and amplitude.

[0044] The following explains the acoustic characteristic parameters listed above. The fundamental frequency is the most basic frequency in a sound signal and determines the pitch of the sound. For example, a pet's whimpering sound typically has a lower fundamental frequency, while an excited bark has a higher fundamental frequency. Pets can quickly perceive the basic tonality of a sound based on this difference in fundamental frequency.

[0045] Frequency bandwidth refers to the frequency range of a sound signal, that is, the difference between the highest and lowest frequencies. For example, the frequency bandwidth of a pet's barking is wide, while the frequency bandwidth of a pet's gentle murmur is narrow.

[0046] Resonance peak is the frequency area where energy is concentrated due to resonance during the propagation of sound, and it is the key feature that determines the timbre of the sound.

[0047] Temporal features primarily describe the temporal characteristics of a sound, encompassing its duration, rhythm, and waveform. For example, a pet's "short warning sound" has a short duration and a rapid rhythm, while its "coquettish hum" has a longer duration and a smoother waveform. Pets can use these temporal characteristics to discern the emotion conveyed by a sound.

[0048] Amplitude reflects the energy of a sound signal and directly corresponds to its loudness. The larger the amplitude, the louder the sound. For example, a pet's excited yell has a large amplitude, while its whimpering sound has a small amplitude. Pets use amplitude to detect the intensity of a sound and, in turn, their urgency.

[0049] The spectral centroid is the center of gravity of a sound's spectral energy distribution and reflects the sound's brightness. A higher spectral centroid indicates a brighter and sharper sound; a lower spectral centroid indicates a duller sound. For example, a pet's "high-pitched bark" has a higher spectral centroid, while a "deep growl" has a lower spectral centroid, which influences the pet's perception of the sound.

[0050] Mel-frequency cepstral coefficients are feature parameters extracted based on the perceptual characteristics of sound. They effectively capture the spectral envelope of sound and have excellent noise immunity. In pet sound applications, they can better adapt to the auditory perception characteristics of pets and improve the accuracy of sound recognition and association.

[0051] Optionally, the feedback operation includes at least one of: controlling the feeder to feed, controlling the waterer to release water, controlling the toy device to start the toy, and triggering the user's communication device to prompt the user of the pet's needs.

[0052] Example 1: When the pet demand associated with the target sound signal represents "hunger", the feedback operation is to control the feeder to feed. Specifically, the feedback operation control instruction is sent to the feeder to open the discharge port and release the set amount of pet food (the set amount can be set according to the pet's weight), so that the pet gradually associates the sound signal with "eating".

[0053] Example 2: When the pet demand associated with the target sound signal represents "thirst", the feedback operation is to control the waterer to release water. Then, the feedback operation control instruction is sent to the waterer to open the water outlet and release the set volume of water, so that the pet gradually associates the sound signal with "drinking water".

[0054] In Example 3, when the pet's need associated with the target sound signal represents a "toy," the feedback operation is to control the toy device to activate the toy. Then, a feedback operation control instruction is sent to the toy device to activate the toy. Specifically, the toy device may be a cat toy that, when activated, can swing. Alternatively, the toy may be a ball-dispensing device that, when activated, ejects a ball. This feedback control allows the pet to gradually associate the sound signal with "playing with a toy."

[0055] In Example 4, when the pet's need associated with the target sound signal represents "going out to play," the feedback operation is to trigger the user's communication device to prompt the user of the pet's need. For example, a reminder can be sent to the user's communication device via text message or app to prompt the pet owner to take the pet out to play. This feedback control causes the pet to gradually associate the sound signal with "going out to play."

[0056] See also Figure 2 Optionally, the process of matching the pet sound with the sound signal in the pet standardized sound library may specifically include: steps 201 to 203.

[0057] Step 201: Extracting acoustic feature parameters of pet sounds.

[0058] Among them, the acoustic characteristic parameters can refer to the description in the above embodiments and will not be repeated here.

[0059] Step 202: Calculate cosine similarity between the acoustic feature parameters of the pet sound and the acoustic feature parameters of multiple sound signals in the pet standardized sound library.

[0060] Step 203: If there is a sound signal with a cosine similarity calculation result greater than a set threshold, the matching is successful.

[0061] The set threshold can be 0.85, 0.8, 0.7, etc. The specific value is not limited.

[0062] Specifically, in the learning phase, when the pet dog emits a first sound signal, the sound is collected by the microphone, and the acoustic feature parameter A is extracted. Then, the acoustic feature parameter A and the acoustic feature parameters of various sound signals in the standardized sound library are calculated by cosine similarity, and the cosine similarity between the acoustic feature parameter A and the acoustic feature parameter B is 0.88. Since 0.88>0.85 (set threshold), the comparison is successful, that is, the pet dog's current sound expression is recognized as a hunger demand, and the feedback behavior of the feeder is triggered to feed, thereby strengthening the learning feedback.

[0063] In practical application, the pet sound is collected by the microphone, and the acoustic feature parameter A is extracted. Then, the acoustic feature parameter A and the acoustic feature parameters of various sound signals in the standardized sound library are calculated by cosine similarity, and the cosine similarity between the acoustic feature parameter A and the acoustic feature parameter B is 0.88. Since 0.88>0.85 (set threshold), the comparison is successful, that is, the pet dog's current sound expression is recognized as a hunger demand, and the pet demand is output to represent the translation result of the pet sound.

[0064] It should be noted that the cosine similarity calculation can effectively measure the similarity of two acoustic feature parameters, improve the recognition accuracy of the pet's real demand, and complete the comparison of the pet sound and the standardized sound library in a short time, quickly determine the pet demand and trigger the corresponding feedback.

[0065] In addition, if there is a pet sound that has not been matched successfully, the unknown type of pet sound can be collected and uploaded to the server for subsequent in-depth analysis to identify new features.

[0066] Optionally, the set threshold adopts a dynamic updating mode, and the dynamic updating mode comprises: inputting the pet sound into a pre-trained classification matching model to output a target adjustment parameter; wherein the classification matching model is used to identify background interference in the pet sound; and based on the target adjustment parameter, the base threshold is optimized to obtain the set threshold.

[0067] In one embodiment, the classification matching model is used to identify the background noise level in pet sounds. For example, in a quiet environment, the background noise level is 0; in a relatively quiet environment, the background noise level is 1; in a moderately noisy environment, the background noise level is 2; and in a relatively noisy environment, the background noise level is 3. Different background noise levels correspond to different target adjustment parameters.

[0068] For example, the basic threshold can be set to 0.7. When the background interference level is 0, the corresponding target adjustment parameter can be 0, and accordingly, the threshold is set to 0.7+0=0.7. When the background interference level is 1, the corresponding target adjustment parameter can be 0.05, and accordingly, the threshold is set to 0.7+0.05=0.75. When the background interference level is 2, the corresponding target adjustment parameter can be 0.08, and accordingly, the threshold is set to 0.7+0.08=0.78. When the background interference level is 3, the corresponding target adjustment parameter can be 0.1, and accordingly, the threshold is set to 0.7+0.1=0.8. When the background interference level is 3, the corresponding target adjustment parameter can be 0.12, and accordingly, the threshold is set to 0.7+0.12=0.82.

[0069] It's important to note that strong interference can distort the acoustic characteristics of pet sounds, potentially overlapping with parameters like frequency and amplitude of pet calls, leading to deviations in the extracted pet sound feature vectors. Therefore, by increasing the threshold, we can filter out sounds that maintain a high degree of similarity even under strong interference (i.e., pet calls whose core features haven't been significantly compromised), thus reducing such misjudgments.

[0070] Of course, in another embodiment, the classification matching model can be used to identify the background interference type in pet sounds, such as identifying TV sound interference, traffic interference, thunderstorm interference, etc., and different target adjustment parameters are set for different background interference types.

[0071] In summary, the embodiment of the present application can improve the recognition and anti-interference ability in complex environments by adopting a dynamic update method for the set threshold, enhance the adaptability of the model to pet sound interference, and reduce the probability of missed judgment or false judgment.

[0072] The following describes the preset scenario triggering strategy provided in the embodiments of the present application.

[0073] In one embodiment, when a preset scene is triggered, playing the target sound signal may specifically include: obtaining a work and rest mapping table; wherein the work and rest mapping table includes a mapping relationship between multiple time nodes and the pet's work and rest patterns; in response to the current time node matching the time node in the work and rest mapping table, extracting the corresponding target sound signal from the pet standardized sound library for playing; wherein the current time node matching the time node in the work and rest mapping table represents the triggering of the preset scene.

[0074] For example, the above schedule mapping table can be constructed according to historical data. For example, according to the schedule of a pet dog, such as eating from 7:00 to 9:00 in the morning, playing with toys from 2:00 to 3:00 in the afternoon, and going out to play from 6:00 to 8:00 in the evening. The schedule mapping table can be constructed according to the schedule.

[0075] Specifically, when it is determined that the current time node is 7:30 in the morning, according to the schedule mapping table, the target sound signal representing the pet's demand as "hunger" can be extracted from the pet standardization sound library for playing.

[0076] When it is determined that the current time node is 2:00 in the afternoon, according to the schedule mapping table, the target sound signal representing the pet's demand as "toy" can be extracted from the pet standardization sound library for playing.

[0077] Since pets have a strong time perception instinct, such as expecting to eat at a fixed time, in the embodiments of the present application, the sound signal is bound to the schedule, so that the pet is in a "time-sound-demand" cycle, the memory of the sound signal at each time node is strengthened, the association between the sound and the demand is more profound, the training effect is more excellent, and the pet can make more focused learning and training on the sound in a specific time period.

[0078] In another embodiment, playing a target sound signal when a preset scene is triggered can specifically include: obtaining a behavior feature of the pet; predicting a demand tendency of the pet based on the behavior feature of the pet; wherein the predicted demand tendency of the pet represents triggering the preset scene; and extracting a corresponding target sound signal from the pet standardization sound library for playing based on the predicted demand tendency of the pet.

[0079] Generally, the behavior feature of the pet can show the current demand tendency of the pet. For example, if the pet scratches the edge of the food bowl with its claws, it can be determined that the pet has a "eating" demand. For example, if the pet scratches the door with its claws, it can be determined that the pet has a "going out to play" demand tendency.

[0080] Further, when it is identified that the pet scratches the edge of the food bowl with its claws, it can be determined that the pet has a "eating" tendency, and the target sound signal representing the pet's demand as "hunger" can be extracted from the pet standardization sound library for playing.

[0081] It should be noted that the above method can quickly respond when the pet actively shows a demand behavior, avoid missing the best training opportunity, make the training more suitable for the real-time state of the pet, and strengthen the immediate linkage memory of "behavior-sound-feedback". Moreover, the method can adapt to individual behavior differences of the pet, improve the training flexibility, and strengthen the pet's active participation, and improve the interactivity of the training.

[0082] In yet another embodiment, the playing of the target sound signal triggered by the preset scene can specifically include: obtaining the position information of the pet and the position information of the M markers; M is a positive integer greater than 2; determining the relative position relationship between the pet and the M markers based on the position information of the pet and the position information of the M markers; extracting the corresponding target sound signal from the pet standardized sound library based on the relative position relationship between the pet and the M markers to play; wherein the relative position relationship between the pet and the M markers is used to trigger the preset scene.

[0083] The above-mentioned markers can specifically refer to the point of the "feeder", the point of the "water feeder", the point of the "toy device", and the point of the "front door of the home". The above-mentioned relative position relationship can refer to the relative distance (such as the straight-line distance) between the pet and the M markers.

[0084] When it is detected that the straight-line distance between the pet position and the point of the "feeder" is less than 0.3 meters (i.e., the relative position relationship between the pet and the point of the "feeder"), the target sound signal associated with "hunger" is extracted from the pet standardized sound library to play. The above-mentioned distance value can be set according to the actual environment, such as 0.5 meters, 0.1 meters, etc.

[0085] Of course, in this embodiment, other playing trigger conditions can also be set, such as when it is detected that the straight-line distance between the pet position and the point of the "feeder" is less than 0.3 meters for 5 times, the target sound signal associated with "hunger" is extracted from the pet standardized sound library to play. The purpose of setting this playing trigger condition is to ensure that the pet currently indeed has the need to eat. If it is detected that the straight-line distance between the pet position and the point of the "feeder" is less than 0.3 meters for 5 times, it can represent that the pet repeatedly wanders around the "feeder".

[0086] As can be seen, the above-mentioned manner can be adapted to the scene, strengthen the pet's cognition and memory of the space, and strengthen the instant linkage memory of "space-sound-feedback".

[0087] Through the above-mentioned manner, the training of the pet language can be realized. In addition, it should be noted that the training time can be set as a fixed period, such as the pet can be continuously trained for one month or three months as a fixed period, and the sound signals in the pet standardized sound library are played at intervals within the fixed period, so that each sound signal reaches the set playing times. For example, the set playing times can be 50 times or 100 times to ensure that the pet can effectively learn and train.

[0088] In an embodiment, considering that different types of pets have large differences in sound, different pet standardized sound libraries can be constructed for different types of pets.

[0089] For example, different pet standardized sound libraries are constructed for pet cats and pet dogs. The pet standardized sound library for pet cats can be constructed according to the known sound data of the pet cats collected. The pet standardized sound library for pet dogs can be constructed according to the known sound data of the pet dogs collected.

[0090] In actual application, different pet standardized sound libraries can be obtained according to different types of pets selected.

[0091] It should be noted that constructing exclusive standardized sound libraries for different types of pets can fully adapt to the sound production characteristics of various pets, significantly improve the accuracy and acceptance of sound signals perceived by pets, and construct a sound environment suitable for different pets, so as to realize efficient and accurate training of pets of different types, and further improve the universality and adaptability of the training method.

[0092] Of course, in another embodiment, different pet standardized sound libraries can also be constructed for pets of different breeds of the same type.

[0093] For example, the breeds of pet dogs include huskies and bears, and the pet standardized sound library for pet dogs of the husky breed can be constructed according to the known sound data of the pet dog husky collected. The pet standardized sound library for pet dogs of the bear breed can be constructed according to the known sound data of the known pet dog bear collected.

[0094] In actual application, different pet standardized sound libraries can be obtained according to different breeds of pets selected.

[0095] It should be noted that constructing exclusive standardized sound libraries for pets of different breeds of the same type can accurately adapt to the sound differences between breeds, and can further enhance the accuracy of training and the stability of training results.

[0096] Please refer to Figure 3Based on the same inventive concept, the application provides a pet language translation system 300 based on audio learning, comprising: a construction module 301, configured to construct a pet standardized sound library; the pet standardized sound library comprises a plurality of sound signals, each sound signal being associated with a pet demand; an audio learning module 302, configured to play a target sound signal and trigger a feedback operation corresponding to the target sound signal when a preset scene is triggered; the target sound signal is a sound signal in the pet standardized sound library and related to the preset scene; a feedback reinforcement module 303, configured to, in a learning stage, trigger a feedback operation corresponding to a matched sound signal when detecting that a first sound signal emitted by a pet matches a sound signal in the pet standardized sound library; a translation module 304, configured to collect pet sounds in a current environment in real time, and output a pet demand corresponding to the matched sound signal in response to the pet sounds matching a sound signal in the pet standardized sound library; the pet demand corresponds to a semantic translation result of the pet sounds.

[0097] Please refer to Figure 4 Based on the same inventive concept, the application provides a module frame of an electronic device 400 applying the above method. The electronic device 400 comprises at least one processor 401 (only one is shown in the figure), a memory 402, a computer program 403 stored in the memory 402 and executable on the at least one processor 401, and the processor 401 implements the steps of the method in any of the preceding embodiments when executing the computer program 403. Figure 4

[0098] The electronic device 400 can be a device associated with a pet, such as a pet smart collar, a pet foot ring, a pet monitoring device, and a smart pet door. It can also be a household device, such as a household security monitoring device. The application does not limit this.

[0099] Those skilled in the art can understand that Figure 4 The electronic device 400 is only an example and does not limit the electronic device 400, which can include more or fewer components than shown, or combine certain components, or different components. For example, positioning sensors, cameras, speakers, microphones, and the like.

[0100] ​The processor 401 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0101] The memory 402 can be an internal storage unit of the electronic device 400, such as a hard disk or a memory of the electronic device 400 in some embodiments. The memory 402 can also be an external storage device of the electronic device 400, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 400 in other embodiments. Further, the memory 402 can include both the internal storage unit and the external storage device of the electronic device 400.

[0102] It should be noted that the above system, device, etc. are based on the same concept as the method embodiments of the present application, and the modules designed by the system and the steps performed by the device and the resulting technical effects can be referred to the method embodiments part, which will not be repeated here.

[0103] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the above described functions. Each functional unit or module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or software. In addition, the specific names of each functional unit or module are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the unit or module in the system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0104] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0105] An embodiment of the present application provides a computer program product. When the computer program product is run on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned various method embodiments when executing the computer program product.

[0106] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or system capable of carrying the computer program code to the camera system / electronic device, a recording medium, computer memory, read-only memory (ROM), random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium. Examples include a USB flash drive, a removable hard drive, a magnetic disk, or an optical disk.

[0107] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0108] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0109] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0110] In the embodiments provided by the present application, it should be understood that the disclosed system / device and method can be implemented in other manners. For example, the embodiments of the system / device described above are merely illustrative, and the division of the modules or units can be different from the above. For example, one or more units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be implemented by using some interfaces, and the indirect couplings or communication connections can be implemented in electronic, mechanical or other forms.

[0111] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0112] The above-described embodiments are merely used to illustrate the technical solutions of the present application, but not limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalent replacements; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A pet language translation method based on audio learning, characterized in that: include: Building a pet standardized sound library, wherein the pet standardized sound library includes a plurality of sound signals, each sound signal being associated with a pet need; When a preset scene is triggered, a target sound signal is played and a feedback operation corresponding to the target sound signal is triggered; the target sound signal is a sound signal in the pet standardized sound library and related to the preset scene; During the learning phase, when it is detected that a first sound signal emitted by a pet matches a sound signal in the pet standardized sound library, a feedback operation corresponding to the matched sound signal is triggered; collecting pet sounds in the current environment in real time, and in response to matching the pet sounds with sound signals in the pet standardized sound library, outputting the pet needs corresponding to the matched sound signals; The pet demand corresponds to the semantic translation result of the pet sound.

2. The pet language translation method based on audio learning according to claim 1, characterized in that: In the pet standardized sound library, each sound signal has different acoustic characteristic parameters.

3. The pet language translation method based on audio learning according to claim 2, characterized in that: The acoustic feature parameters include at least one of fundamental frequency, frequency bandwidth, resonance peak, time domain characteristics, amplitude, spectrum centroid, and Mel-frequency cepstral coefficient.

4. The pet language translation method based on audio learning according to claim 2, characterized in that: The process of matching the pet sound with the sound signal in the pet standardized sound library includes: extracting acoustic feature parameters of the pet sound; performing cosine similarity calculation on the acoustic feature parameters of the pet sound and the acoustic feature parameters of multiple sound signals in the pet standardized sound library; If there is a sound signal whose cosine similarity calculation result is greater than the set threshold, the match is successful.

5. The pet language translation method based on audio learning according to claim 4, characterized in that: The threshold is set in a dynamic update manner, and the dynamic update manner includes: Inputting the pet sound into a pre-trained classification matching model and outputting a target adjustment parameter; wherein the classification matching model is used to identify background interference in the pet sound; Based on the target adjustment parameter, the basic threshold is optimized to obtain the set threshold.

6. The pet language translation method based on audio learning according to claim 1, characterized in that: When the preset scene is triggered, playing the target sound signal includes: Obtaining a work and rest mapping table; wherein the work and rest mapping table includes a mapping relationship between multiple time nodes and the pet's work and rest patterns; In response to a current time node matching a time node in the work and rest mapping table, extracting a corresponding target sound signal from the pet standardized sound library and playing the signal; The matching of the current time node with the time node in the work and rest mapping table represents triggering of a preset scene.

7. The pet language translation method based on audio learning according to claim 1, characterized in that: When the preset scene is triggered, playing the target sound signal includes: Obtain behavioral characteristics of pets; Based on the behavioral characteristics of the pet, predict the demand tendency of the pet; wherein the predicted demand tendency of the pet represents the triggering of a preset scenario; Based on the predicted demand tendency of the pet, a corresponding target sound signal is extracted from the pet standardized sound library for playing.

8. The pet language translation method based on audio learning according to claim 1, characterized in that: When the preset scene is triggered, playing the target sound signal includes: Obtain the location information of the pet and the location information of M identification points; M is a positive integer greater than 2; Determine the relative position relationship between the pet and the M identification points based on the location information of the pet and the location information of the M identification points; Based on the relative positional relationship between the pet and the M identification points, extract a corresponding target sound signal from the pet standardized sound library and play it; The relative position relationship between the pet and the M identification points is used to trigger a preset scene.

9. The pet language translation method based on audio learning according to claim 1, characterized in that: The feedback operation includes at least one of: controlling the feeder to feed, controlling the waterer to release water, controlling the toy device to start the toy, and triggering the user's communication device to prompt the user of the pet's needs.

10. A pet language translation system based on audio learning, characterized in that: include: Building blocks for constructing a standardized pet sound library; The pet standardized sound library includes a variety of sound signals, each sound signal is associated with a pet need; An audio learning module, configured to play a target sound signal when a preset scene is triggered, and trigger a feedback operation corresponding to the target sound signal; the target sound signal is a sound signal in the pet standardized sound library and related to the preset scene; A feedback reinforcement module is configured to trigger a feedback operation corresponding to the matched sound signal when detecting that a first sound signal emitted by the pet matches a sound signal in the pet standardized sound library during the learning phase; a translation module for collecting pet sounds in a current environment in real time, and outputting pet needs corresponding to the matched sound signals in response to matching the pet sounds with sound signals in the pet standardized sound library; The pet demand corresponds to the semantic translation result of the pet sound.

Citation Information

Patent Citations

  • Infant or pet nursing method and nursing system and nursing machine adopting method

    CN103985383A

  • Automatic pet feeding method and device, computer storage medium and electronic equipment

    CN109729990A

  • Intelligent feeding method and device based on animal sound

    CN117995200A

  • Pet voice translation method and system, electronic equipment and storage medium

    CN119626263A

  • Pet voice interaction device and pet voice interaction method

    CN119943060A

Cited By

  • Pet voice recognition translation cloud training method and application

    CN121708907A