Intelligent feeding system and method based on bird voice recognition

Through the intelligent feeding system, using AI technology and speech recognition technology, the problem of traditional parrot training methods is solved, and the effect of precise feeding and improving learning efficiency is achieved.

CN119949263APending Publication Date: 2025-05-09BEIJING JIAOTONG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510127257.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-02
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Traditional parrot training methods take a long time and lack objectivity, making it difficult to accurately evaluate the learning progress and effectiveness of parrots.

Method used

An intelligent feeding system based on bird speech recognition is adopted, and automated voice recognition and feeding control is realized through AI voice generation module, sound acquisition module, imitation recognition module and reward module.

Benefits of technology

It provides objective and scientific evaluation standards, precise feeding strategies, improves parrot learning efficiency and interest, and reduces manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119949263A_ABST
    Figure CN119949263A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent feeding system based on bird voice recognition, and the system comprises an AI voice generation module which employs an AI technology, and automatically generates human languages with different age features; the birds listen to human languages and make simulated sounds, and the sound collection module collects the simulated sounds and the human languages; the imitation recognition module extracts feature information of the bird voice signals and human languages, constructs a similarity model by adopting an AI intelligent algorithm, calculates a similarity result, presets a threshold value in the similarity model, compares the threshold value with the similarity result, and outputs a comparison result; a reward mechanism is arranged in the reward module, and birds are fed according to a comparison result; through the voice recognition technology, intelligent feeding control is achieved, manual intervention is reduced, and the feeding efficiency is improved; meanwhile, the reward strategy can be dynamically adjusted according to the simulation condition of the birds, the individual requirements of different birds are met, and healthy growth of the birds is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sound recognition technology, and in particular to an intelligent feeding system and method based on bird voice recognition. Background Art

[0002] In the traditional parrot breeding model, people generally use manual feeding and oral teaching methods to train parrots to speak human language. Although this method can establish a preliminary communication bridge between parrots and humans to a certain extent, its inherent limitations and shortcomings are also obvious.

[0003] For example, breeders need to spend a lot of time and energy to repeatedly dictate and demonstrate, while parrots need to gradually master human language through constant attempts and imitation. This process is not only time-consuming, but also often unsatisfactory, because parrots have individual differences in learning ability and speed, and a unified teaching schedule is difficult to meet the needs of all parrots.

[0004] Moreover, in the traditional method, the learning effect of parrots is completely dependent on the subjective judgment of the breeder. The breeder observes the parrot's imitative behavior and uses personal experience and intuition to evaluate its learning progress and effectiveness. However, this evaluation method lacks objectivity and scientificity, and is easily affected by the breeder's personal emotions, experience level and other factors, resulting in inaccurate and inconsistent evaluation results.

[0005] Therefore, there is an urgent need for an intelligent feeding system based on bird voice recognition to solve the technical problems existing in the above-mentioned prior art. Summary of the invention

[0006] The present invention overcomes the deficiencies of the prior art and provides an intelligent feeding system based on bird voice recognition.

[0007] To achieve the above-mentioned purpose, the technical solution adopted by the present invention is: an intelligent feeding system and method based on bird voice recognition, comprising: an AI voice generation module, a sound collection module, an imitation recognition module and a reward module;

[0008] The AI ​​voice generation module uses AI technology to automatically generate human language with different age characteristics, and plays it in a loop through a sound amplifier;

[0009] The bird listens to the human language and makes simulated sounds, the sound collection module collects the simulated sounds and the human language, and converts the simulated sounds into bird voice signals and forwards them to the imitation recognition module;

[0010] The imitation recognition module extracts characteristic information of the bird voice signal and the human language, constructs a similarity model using an AI intelligent algorithm, calculates a similarity result, and presets a threshold in the similarity model for comparison with the similarity result, and outputs a comparison result;

[0011] The reward module is connected to the imitation recognition module through an electrical signal. A reward mechanism is arranged in the reward module. The reward mechanism adopts intelligent mechanical arm technology and feeds the birds according to the comparison result.

[0012] In a preferred embodiment of the present invention, the imitation recognition module uses deep learning technology to extract the spectral features, factor features and rhythmic features included in the bird voice signal, and recognizes the personalized pronunciation characteristics of different birds through learning.

[0013] In a preferred embodiment of the present invention, the imitation recognition module pre-processes the collected human language and bird voice signals, aligns the features of the human language and bird voice signals through dynamic time warping, and then uses PCA (principal component analysis) and LDA (linear discriminant analysis) methods to reduce the dimensionality of the features while retaining key information.

[0014] In a preferred embodiment of the present invention, the key information is used as data input of the similarity model, wherein the calculation formula of the similarity model is expressed as:

[0015]

[0016] In the formula, Similarity(X,Y) is used to express the familiarity result by the Ohm distance between the key information of human language and bird speech signals; x is used to represent the key information of human language; y is used to represent the key information of bird speech signals; i is used to represent the time order of key information; d is used to represent the total number of key information in the whole time period; (x i -y i ) is used to represent the difference between the key information of human language and the key information of bird speech signals in the same time period.

[0017] In a preferred embodiment of the present invention, the threshold is used to determine whether the bird successfully imitates the human language; when the threshold is less than the similarity result, the bird successfully imitates the human language; conversely, when the threshold is greater than the similarity result, the bird fails to imitate the human language; wherein the comprehensive expression of the threshold is:

[0018]

[0019] Where T is the preset similarity threshold; Output is the output of the model (1 indicates successful imitation, 0 indicates imitation failure).

[0020] In a preferred embodiment of the present invention, the sound collection module uses a microphone array and combines a noise reduction algorithm to capture the simulated sounds of birds.

[0021] In a preferred embodiment of the present invention, the intelligent feeding system is equipped with a remote monitoring and data analysis platform, and the administrator can view the birds' learning status, imitation effect and reward records in real time through a mobile phone APP or a web page.

[0022] In a preferred embodiment of the present invention, there is provided an intelligent feeding method based on bird speech recognition, which is applied to the above-mentioned intelligent feeding system and comprises the following steps:

[0023] S1. Automatically generate human language with different age characteristics through AI speech generation module, and use sound amplifier to play human language in a loop for birds to listen and imitate;

[0024] S2, the sound collection module captures and collects the imitation voices of birds and human language, and sends them to the imitation recognition module;

[0025] S3, the imitation recognition module uses deep learning technology to extract the spectral features, factor features and rhythmic features in the bird voice signal; at the same time, it extracts the corresponding features of human language, aligns the features of human language and bird voice signals through dynamic time warping, uses PCA (principal component analysis) and LDA (linear discriminant analysis) to reduce the dimension of the features, retains key information, uses the key information as the data input of the similarity model, calculates the similarity results between human language and bird voice signals, compares the similarity results with the preset threshold, and determines whether the bird successfully imitates the human language;

[0026] S4. Based on the comparison results of the imitation recognition module, the reward module feeds the birds through the intelligent robotic arm technology. If the similarity result is higher than the threshold, it means that the birds have successfully imitated, triggering the reward mechanism and feeding. If the similarity result is lower than the threshold, it means that the imitation has failed, and the reward is not triggered.

[0027] In a preferred embodiment of the present invention, the reward module dynamically adjusts the reward strategy according to the comparison results of the imitation recognition module; at the same time, the intelligent feeding system integrates a behavioral analysis function to monitor the learning status, points of interest and potential problems of the birds in real time, providing a basis for adjusting the reward mechanism and learning content.

[0028] The present invention solves the defects existing in the background technology and has the following beneficial effects:

[0029] (1) The present invention uses an imitation recognition module to accurately extract the spectral features, factor features and rhythmic features of the parrot's imitation of human language, and constructs a similarity model to calculate the similarity of the imitation, thereby providing an objective and scientific evaluation standard for the parrot's learning effect, allowing the breeder to more accurately understand the parrot's learning progress and effect.

[0030] (2) The present invention can dynamically adjust the feeding strategy through the reward module according to the imitation results of the parrot, and reward the parrot that has successfully imitated by feeding it, thereby achieving precise feeding, which not only helps the healthy growth of the parrot, but also stimulates the parrot's interest and enthusiasm in learning.

[0031] (3) The present invention uses voice recognition technology to achieve intelligent feeding control, reduce manual intervention, and improve feeding efficiency. It can also dynamically adjust the reward strategy according to the imitation of birds, meet the personalized needs of different birds, and promote the healthy growth of birds. At the same time, by looping human language and encouraging birds to imitate, it provides birds with opportunities to learn and communicate, which helps to improve the intelligence level and social ability of birds. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work.

[0033] Figure 1 It is a flow chart of the intelligent feeding system of the preferred embodiment of the present invention. DETAILED DESCRIPTION

[0034] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0035] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0036] Before that, it should be noted that there is an area in the brain of parrots that is specifically responsible for processing sound, and this area shows extremely high sensitivity when processing sound information. In addition, the brain of parrots contains special neurons that help them imitate the sounds made by humans. This brain structure allows parrots to accurately capture and copy the sounds they hear, including human language.

[0037] Proper training and guidance are also important factors in improving the parrot's ability to learn to talk. By repeatedly playing recordings and practicing conversations with parrots, you can effectively improve the parrot's language ability. In addition, individual differences, species, and living environment of parrots will also affect their ability to learn to talk.

[0038] However, the current process of training parrots to speak requires breeders to spend a lot of time and energy to repeatedly dictate and demonstrate, while parrots need to gradually master human language through constant attempts and imitation. This process is not only time-consuming, but also often unsatisfactory, because parrots have individual differences in learning ability and speed, and a unified teaching schedule is difficult to meet the needs of all parrots.

[0039] Moreover, in the traditional method, the learning effect of parrots is completely dependent on the subjective judgment of the breeder. The breeder observes the parrot's imitative behavior and uses personal experience and intuition to evaluate its learning progress and effectiveness. However, this evaluation method lacks objectivity and scientificity, and is easily affected by the breeder's personal emotions, experience level and other factors, resulting in inaccurate and inconsistent evaluation results.

[0040] Therefore, the present invention proposes an intelligent feeding system and method based on bird voice recognition to solve the technical problems existing in the prior art.

[0041] like Figure 1 As shown, an intelligent feeding system based on bird voice recognition includes: an AI voice generation module, a sound collection module, an imitation recognition module and a reward module;

[0042] The AI ​​speech generation module uses AI technology to automatically generate human language with different age characteristics and play it in a loop through a sound amplifier; the automatically generated human language includes simple words, phrases or sentences to adapt to the learning and imitation abilities of different birds.

[0043] The birds listen to the human language and make simulated sounds. The sound collection module collects the simulated sounds and the human language, and converts the simulated sounds into bird voice signals and forwards them to the imitation recognition module. The sound collection module uses a microphone array and combines it with a noise reduction algorithm to capture the simulated sounds of birds. The microphone array can enhance the sensitivity and directionality of sound collection, while the noise reduction algorithm can effectively reduce the interference of background noise and improve the clarity of sound collection.

[0044] In a preferred embodiment, considering the cost control of the sound collection module, a high-performance computer is placed on the server side, and a low-cost single-chip microcomputer is used at the terminal to perform data collection, transmission and mechanical control.

[0045] ESP32-WROOM-32 is a powerful general-purpose Wi-Fi+BT+BLE (Bluetooth Low Energy) MCU (microcontroller) module with a built-in dual-core 32-bit microprocessor and a clock frequency of up to 240MHz. It has 512KB (or 448KB of ROM) + 520KB of SRAM and 4MB of flash memory, which is enough to handle audio file streaming.

[0046] ESP32-WROOM-32 integrates a wealth of peripheral interfaces and sensors, including but not limited to SD card interface, Ethernet interface, high-speed SDIO / SPI, UART, I2S and I2C, etc. It has 2 UART interfaces, up to 25 PWM outputs, and 2 8-bit DAC channels, which greatly facilitates the control of motors and servos.

[0047] In such a typical IoT application scenario, we developed based on the ESP32-wroom32 module and used the GY-SPH0645LM4H sound sensor. The total cost can be controlled within 50 yuan, which has a market competitive advantage.

[0048] This solution is developed using the rich Arduino library of ESP32. The software architecture of the network microphone mainly includes the audio acquisition module, audio processing module, network communication module and main control module. The modules communicate with each other through function calls or message passing mechanisms to jointly realize the acquisition, processing and transmission of audio data. Among them, the audio acquisition module uses the I2S communication protocol to read audio data from the microphone sensor.

[0049] Establish a network connection and encode the processed audio data into a format suitable for network transmission (such as PCM, WAV, MP3, etc.). Use protocols such as TCP / UDP to send the audio data to a remote server or client.

[0050] The imitation recognition module extracts the characteristic information of bird voice signals and human language, uses AI intelligent algorithm to build a similarity model, calculates the similarity result, and presets a threshold in the similarity model for comparison with the similarity result, and outputs the comparison result;

[0051] The reward module is connected to the imitation recognition module through electrical signals. A reward mechanism is set up in the reward module. The reward mechanism adopts intelligent robotic arm technology to feed the birds according to the comparison results.

[0052] The imitation recognition module uses deep learning technology to extract the spectral features, factor features and rhythmic features included in the bird voice signals, and recognizes the personalized pronunciation characteristics of different birds through learning.

[0053] It is further explained here that the spectral feature reflects the distribution characteristics of the sound signal in the frequency domain and is one of the important features in speech recognition. For the imitation sound of the parrot, the spectral feature can reveal the frequency components and energy distribution of its sound signal, which helps to identify whether the parrot accurately imitates the tone and timbre of human language.

[0054] Factor features are language units such as phonemes or syllables in sound signals. In speech recognition, by extracting and analyzing factor features in sound signals, we can identify the specific phonemes or syllables that the parrot imitates, thereby judging the accuracy and completeness of its imitation.

[0055] Prosodic features include the pitch, rhythm, and stress of sounds, and are an important part of language expression. For parrots’ imitation sounds, prosodic features can reveal the intonation changes and rhythm of their sound signals, which helps to determine whether parrots accurately imitate the prosodic characteristics of human language.

[0056] The imitation recognition module pre-processes the collected human language and bird voice signals, aligns the features of human language and bird voice signals through dynamic time warping, and then uses PCA (principal component analysis) and LDA (linear discriminant analysis) methods to reduce the dimensionality of the features while retaining key information.

[0057] It is further explained that in the process of imitation recognition, since the speed and rhythm of parrot imitation are different from human language, it is necessary to align the parrot imitation sound signal with the human language signal in time through the dynamic time warping algorithm. The dynamic time warping feature reflects the change and matching degree of the sound signal during this alignment process, which is an important basis for judging the accuracy of parrot imitation.

[0058] In order to improve the computational efficiency and accuracy of the similarity model, it is necessary to reduce the dimension of the extracted features. The reduced-dimensional features retain the key information in the original features while reducing redundancy and noise, and are important input data for building a similarity model.

[0059] The key information is used as the data input of the similarity model, where the calculation formula of the similarity model is expressed as:

[0060]

[0061] In the formula, Similarity(X,Y) is used to express the familiarity result by the Ohm distance between the key information of human language and bird speech signals; x is used to represent the key information of human language; y is used to represent the key information of bird speech signals; i is used to represent the time order of key information; d is used to represent the total number of key information in the whole time period; (x i -y i ) is used to represent the difference between the key information of human language and the key information of bird speech signals in the same time period.

[0062] The threshold is used to determine whether the bird successfully imitates human language; when the threshold is less than the similarity result, the bird successfully imitates human language; conversely, when the threshold is greater than the similarity result, the bird fails to imitate human language; the comprehensive expression of the threshold is:

[0063]

[0064] Where T is the preset similarity threshold; Output is the output of the model (1 indicates successful imitation, 0 indicates imitation failure).

[0065] The intelligent feeding system is equipped with a remote monitoring and data analysis platform, and the administrator can view the birds’ learning status, imitation effects and reward records in real time through the mobile APP or website.

[0066] Provided is an intelligent feeding method based on bird speech recognition, which is applied to the above-mentioned intelligent feeding system and includes the following steps:

[0067] S1. Automatically generate human language with different age characteristics through AI speech generation module, and use sound amplifier to play human language in a loop for birds to listen and imitate;

[0068] S2, the sound collection module captures and collects the imitation voices of birds and human language, and sends them to the imitation recognition module;

[0069] S3, the imitation recognition module uses deep learning technology to extract the spectral features, factor features and rhythmic features in the bird voice signal; at the same time, it extracts the corresponding features of human language, aligns the features of human language and bird voice signals through dynamic time warping, uses PCA (principal component analysis) and LDA (linear discriminant analysis) to reduce the dimension of the features, retains key information, uses the key information as the data input of the similarity model, calculates the similarity results between human language and bird voice signals, compares the similarity results with the preset threshold, and determines whether the bird successfully imitates the human language;

[0070] S4. Based on the comparison results of the imitation recognition module, the reward module feeds the birds through the intelligent robotic arm technology. If the similarity result is higher than the threshold, it means that the birds have successfully imitated, triggering the reward mechanism and feeding. If the similarity result is lower than the threshold, it means that the imitation has failed, and the reward is not triggered.

[0071] The reward module dynamically adjusts the reward strategy based on the comparison results of the imitation recognition module; for birds with high imitation and fast learning progress, more or higher-quality rewards (such as more favorite food, longer interaction, etc.) are given. At the same time, the intelligent feeding system integrates behavioral analysis functions to monitor the learning status, interests and potential problems of birds in real time, providing a basis for adjusting the reward mechanism and learning content.

[0072] When the present invention is used, the AI ​​voice generation module automatically generates and plays human languages ​​with different age characteristics; the sound collection module captures and collects the imitation voices and human languages ​​of birds and sends them to the imitation recognition module; the imitation recognition module extracts features, aligns features, performs dimensionality reduction processing, and calculates similarity results, compares them with the preset threshold, and determines whether the birds have successfully imitated. According to the comparison results of the imitation recognition module, the reward module feeds the birds through intelligent mechanical arm technology; the administrator uses the remote monitoring and data analysis platform to view the learning status, imitation effect and reward records of the birds in real time, and adjusts the feeding strategy and learning content according to the analysis results.

[0073] The above is based on the ideal embodiment of the present invention. Through the above description, relevant personnel can make various changes and modifications without departing from the technical concept of the present invention. The technical scope of the present invention is not limited to the content in the specification, and the technical scope must be determined according to the scope of the claims.

Claims

1. An intelligent feeding system based on bird voice recognition, comprising: AI voice generation module, sound collection module, imitation recognition module and reward module; the features are: The AI ​​voice generation module uses AI technology to automatically generate human language with different age characteristics, and plays it in a loop through a sound amplifier; The bird listens to the human language and makes simulated sounds, the sound collection module collects the simulated sounds and the human language, and converts the simulated sounds into bird voice signals and forwards them to the imitation recognition module; The imitation recognition module extracts characteristic information of the bird voice signal and the human language, constructs a similarity model using an AI intelligent algorithm, calculates a similarity result, and presets a threshold in the similarity model for comparison with the similarity result, and outputs a comparison result; The reward module is connected to the imitation recognition module through an electrical signal. A reward mechanism is arranged in the reward module. The reward mechanism adopts intelligent mechanical arm technology and feeds the birds according to the comparison result.

2. The intelligent feeding system based on bird voice recognition according to claim 1, characterized in that: The imitation recognition module uses deep learning technology to extract the spectrum features, factor features and rhythm features included in the bird voice signal, and recognizes the personalized pronunciation characteristics of different birds through learning.

3. An intelligent feeding system based on bird voice recognition according to claim 1 or 2, characterized in that: The imitation recognition module pre-processes the collected human language and bird voice signals, aligns the features of the human language and bird voice signals through dynamic time warping, and then uses PCA (principal component analysis) and LDA (linear discriminant analysis) methods to reduce the dimensionality of the features while retaining key information.

4. The intelligent feeding system based on bird voice recognition according to claim 1, characterized in that: The key information is used as data input of the similarity model, wherein the calculation formula of the similarity model is expressed as: In the formula, Similarity(X,Y) is used to express the familiarity result by the Ohm distance between the key information of human language and bird speech signals; x is used to represent the key information of human language; y is used to represent the key information of bird speech signals; i is used to represent the time order of key information; d is used to represent the total number of key information in the whole time period; (x i -y i ) is used to represent the difference between the key information of human language and the key information of bird speech signals in the same time period.

5. The intelligent feeding system based on bird voice recognition according to claim 1 or 4, characterized in that: The threshold is used to determine whether the bird successfully imitates the human language; when the threshold is less than the similarity result, the bird successfully imitates the human language; On the contrary, when the threshold is greater than the similarity result, the bird fails to imitate the human language; wherein the comprehensive expression of the threshold is: Where T is the preset similarity threshold; Output is the output of the model (1 indicates successful imitation, 0 indicates imitation failure).

6. The intelligent feeding system based on bird voice recognition according to claim 1, characterized in that: The sound collection module uses a microphone array and combines a noise reduction algorithm to capture the simulated sounds of birds.

7. The intelligent feeding system based on bird voice recognition according to claim 1, characterized in that: The intelligent feeding system is equipped with a remote monitoring and data analysis platform, and the administrator can view the birds' learning status, imitation effects and reward records in real time through a mobile phone APP or a web page.

8. An intelligent feeding method based on bird speech recognition, applied to the intelligent feeding system according to any one of claims 1 to 8, characterized in that: The following steps are involved: S1. Automatically generate human language with different age characteristics through AI speech generation module, and use sound amplifier to play human language in a loop for birds to listen and imitate; S2, the sound collection module captures and collects the imitation voices of birds and human language, and sends them to the imitation recognition module; S3, the imitation recognition module uses deep learning technology to extract the spectral features, factor features and rhythmic features in the bird voice signal; at the same time, it extracts the corresponding features of human language, aligns the features of human language and bird voice signals through dynamic time warping, uses PCA (principal component analysis) and LDA (linear discriminant analysis) to reduce the dimension of the features, retains key information, uses the key information as the data input of the similarity model, calculates the similarity results between human language and bird voice signals, compares the similarity results with the preset threshold, and determines whether the bird successfully imitates the human language; S4. Based on the comparison results of the imitation recognition module, the reward module feeds the birds through the intelligent robotic arm technology. If the similarity result is higher than the threshold, it means that the birds have successfully imitated, triggering the reward mechanism and feeding. If the similarity result is lower than the threshold, it means that the imitation has failed, and the reward is not triggered.

9. The intelligent feeding method based on bird speech recognition according to claim 8, characterized in that: The reward module dynamically adjusts the reward strategy according to the comparison results of the imitation recognition module; at the same time, the intelligent feeding system integrates a behavioral analysis function to monitor the learning status, interest points and potential problems of the birds in real time, providing a basis for adjusting the reward mechanism and learning content.

Citation Information

Patent Citations

  • Animal training method and system

    CN105766687A

  • Speech similarity detection method and apparatus

    CN106935248A

  • Feeding device for training pet bird to speak

    CN108739510A

  • Bird training method and device

    CN109197674A

  • Human and animal situation language intelligent communication system

    CN117711369A