Bird active shooting method and system based on voiceprint recognition and automatic positioning

Through voiceprint recognition and multi-point cross-positioning technology, birds are automatically identified and actively photographed, which solves the problems of low efficiency and insufficient data accuracy in traditional bird monitoring methods, and achieves efficient and accurate bird monitoring.

CN120602781APending Publication Date: 2025-09-05XIAN QUELINGFEI INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510715927.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional bird monitoring methods are inefficient, have limited coverage, and are difficult to combine voiceprint characteristics and location information, resulting in insufficient real-time and accuracy of monitoring data.

Method used

Voiceprint recognition technology is used to process bird chirping through discrete Fourier transform, and bird species are identified using the improved ECAPA-TDNN model, and bird location is determined in combination with multi-point cross-positioning technology, and actively photographed through high-resolution cameras.

Benefits of technology

It realizes efficient and accurate bird monitoring, improving the real-time and accuracy of data collection without manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602781A_ABST
    Figure CN120602781A_ABST
Patent Text Reader

Abstract

The invention discloses an active bird shooting method and system based on voiceprint recognition and automatic localization, and belongs to the technical field of intelligent bird voiceprint collection and recognition, and the method comprises the steps: collecting bird sounds through a voiceprint collection device, and carrying out the discrete Fourier transform of the collected bird sounds; inputting the processed bird chirp data into a voiceprint recognition neural network model based on deep learning to obtain a high-dimensional feature vector; comparing the high-dimensional feature vector with each piece of bird voiceprint feature data in a bird voiceprint feature library to obtain bird species corresponding to the voiceprint feature data; according to the obtained closest bird species, based on a positioning sensor, determining the positions of the birds by using a multi-point cross positioning technology; according to the determined positions of the birds, the camera adjusts the shooting angle and the focal length to shoot the birds. According to the method, manual intervention is not needed, so that efficient and accurate bird monitoring is realized, and the monitoring efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of intelligent collection and identification of bird voiceprints, and in particular to a method and system for active bird photography based on voiceprint recognition and automatic positioning. Background Art

[0002] In recent years, with the growing demand for biodiversity conservation and ecological monitoring, bird research has become an important topic in fields such as environmental science, ecology, and animal behavior. Traditional bird monitoring methods mainly rely on manual observation or fixed camera shooting, but these technologies have significant limitations: manual observation is inefficient, has limited coverage, and is easily affected by subjective factors; static camera equipment has difficulty actively tracking targets, and is prone to missing shooting opportunities, especially in complex natural environments. In addition, bird activities are temporally and spatially random, and traditional passive monitoring methods cannot effectively combine the bird's voiceprint characteristics and location information, making it difficult to meet the real-time and accurate data collection requirements, resulting in insufficient accuracy and completeness of monitoring data.

[0003] Therefore, it is of great significance to develop an intelligent monitoring system or method that can automatically identify bird sound patterns, accurately locate and actively photograph them. Summary of the Invention

[0004] In response to the above-mentioned deficiencies in the prior art, the present application provides a method and system for active bird photography based on voiceprint recognition and automatic positioning, which solves the problems of the prior art in the lack of real-time and accuracy of monitoring data, as well as low collection efficiency.

[0005] In order to achieve the above-mentioned invention objectives, the technical solutions adopted in this application are: In a first aspect, the present application provides a method for actively photographing birds based on voiceprint recognition and automatic positioning, comprising: S1: Use the voiceprint collection device to collect bird calls and perform discrete Fourier transform processing on the collected bird calls; S2: Input the processed bird song data into a deep learning-based voiceprint recognition neural network model to obtain a high-dimensional feature vector; S3: Compare the high-dimensional feature vector with each bird voiceprint feature data in the bird voiceprint feature library to obtain the most similar bird species; S4: Based on the closest bird species obtained, the location of the bird is determined using multi-point cross-positioning technology based on the positioning sensor; S5: According to the determined position of the bird, the camera adjusts the shooting angle and focal length to shoot the bird.

[0006] Furthermore, the specific formula for the discrete Fourier transform processing in S1 includes:

[0007] in, is the frequency domain, is the discrete time domain signal of birds, , is the frequency index, is the time index, is the total number of sampling points, is the imaginary unit, is the base of natural logarithms.

[0008] Furthermore, the deep learning-based voiceprint recognition neural network model in S2 is an improvement of the ECAPA-TDNN model. The improvement process specifically includes: A1: Use dilated convolutions with different dilation rates to capture multi-scale context, where the capture formula is:

[0009] Where, The time scale is of is the dilated convolution with dilation rates of 1, 2, and 3, is the time scale index, The raw data or feature maps that need to be captured in multi-scale context; A2: Concatenate features of all scales and perform standard convolution processing:

[0010] in, For scale 1 is the dilated convolution with dilation rates of 1, 2, and 3, For scale 1 is the dilated convolution with dilation rates of 1, 2, and 3, For scale 1 The dilated convolutions are of dilation rates 1, 2, and 3; A3: Global average pooling and fully connected layers are used to calculate channel weights to obtain high-dimensional feature vectors. The channel weight calculation formula is:

[0011] in, The features generated by concatenating features of all scales and performing standard convolution processing are: is the activation function, is global average pooling.

[0012] Furthermore, the step S3 compares the high-dimensional feature vector with each bird voiceprint feature data in the bird voiceprint feature library to obtain the most similar bird species, specifically including: Compare the high-dimensional feature vector with each bird voiceprint feature data in the bird voiceprint feature library. If the bird voiceprint feature library contains corresponding feature data, take the bird species corresponding to the corresponding feature data as the closest bird species; otherwise, re-compare. If the re-comparison is successful, obtain the bird species corresponding to the corresponding feature data; otherwise, record an error message.

[0013] Furthermore, the comparing the high-dimensional feature vector with each bird voiceprint feature data in the bird voiceprint feature library specifically includes: The high-dimensional feature vector is sequentially compared with each bird voiceprint feature data in the bird voiceprint feature library by using the angle cosine algorithm to obtain the most similar bird species. The comparison formula is:

[0014] Where, A high-dimensional feature vector representing the current bird's voiceprint, Represents each feature vector in the bird voiceprint feature library, Represents the dot product of the current bird voiceprint feature vector and each vector in the bird feature library, Represents the Euclidean norm of the current bird voiceprint feature vector and each vector in the bird feature library.

[0015] Furthermore, in S4, based on the obtained closest bird species, the location of the bird is determined using a multi-point cross-positioning technology based on a positioning sensor, specifically including: S401: According to the obtained closest bird species, obtain information of a positioning sensor near the bird species, and calculate the distance between the positioning sensor and the bird species; S402: Obtain the position of the bird through geometric calculation based on the position coordinates of the positioning sensor and the calculated distance.

[0016] Second aspect: This application provides a bird active photography system based on voiceprint recognition and automatic positioning, including: A voiceprint recognition module, comprising a voiceprint collection device, a voiceprint processing unit, and a voiceprint recognition algorithm unit connected in sequence, for capturing bird calls in real time and identifying bird species; Positioning sensor module, including ultrasonic sensor, infrared sensor and GPS module, used to automatically locate the bird's position after voiceprint recognition; An active shooting module, including a high-resolution camera and an autofocus system, for actively adjusting the shooting angle and focal length according to the position coordinates provided by the positioning sensor module and shooting the birds; The data processing and storage module includes an image receiving module, an embedded processor, a data transmission device, and a storage device. It is used to process the received voiceprint, location, and image data, and store the data in a local server library and upload it to the cloud for backup; A display module is connected to the data processing and storage module and is used to display a data analysis report.

[0017] Furthermore, the voiceprint collection device adopts a high-sensitivity microphone array composed of multiple microphone units, and the microphone array has a wide frequency response range and a high signal-to-noise ratio; The voiceprint processing unit performs discrete Fourier transform processing on the collected bird call sound data; The voiceprint recognition algorithm unit uses a voiceprint recognition neural network model based on deep learning to identify the processed chirping sound data to obtain similar bird species.

[0018] Furthermore, the autofocus system utilizes phase detection focusing or contrast detection focusing technology to achieve autofocus.

[0019] Furthermore, the local server and the cloud have data analysis software for performing comprehensive analysis on the collected voiceprint, location and image data, and generating a data analysis report.

[0020] The beneficial effects of this application are: The present application provides a method and system for active bird photography based on voiceprint recognition and automatic positioning. Birds are identified through voiceprint equipment, the positioning sensor is used to automatically locate the position of the birds, and the camera is actively mobilized to photograph the located birds without human intervention, thereby achieving efficient and accurate bird monitoring and improving the efficiency and accuracy of monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0022] Figure 1 A schematic flow chart of a method for actively photographing birds based on voiceprint recognition and automatic positioning provided in an embodiment of the present application.

[0023] Figure 2 Schematic diagram of a bird active photography system based on voiceprint recognition and automatic positioning provided in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field based on this application are within the scope of protection of this application.

[0025] Example 1: The present application embodiment provides a method for actively photographing birds based on voiceprint recognition and automatic positioning. Figure 1 , Figure 1 The figure shows a flow chart of a method for actively photographing birds based on voiceprint recognition and automatic positioning provided by an embodiment of the present application, including: S1: Use a voiceprint collection device to collect bird calls, and perform discrete Fourier transform processing on the collected bird calls.

[0026] Furthermore, the specific formula for the discrete Fourier transform processing in S1 includes:

[0027] in, is the frequency domain, is the discrete time domain signal of birds, , is the frequency index, is the time index, is the total number of sampling points, is the imaginary unit, is the base of natural logarithms.

[0028] In one embodiment of the present application, a voiceprint collection device uses a highly sensitive microphone array consisting of multiple microphone units that can capture bird calls in real time. The microphone array has a wide frequency response range and a high signal-to-noise ratio, enabling clear pickup of bird calls of varying frequencies. For example, in one embodiment, the microphone array can include eight microphone units arranged in a ring, capturing sound signals in a 360-degree omnidirectional manner. The sampling rate can be set to 48kHz, ensuring accurate sampling of bird calls.

[0029] Among them, the collected audio data is subjected to discrete Fourier transform, that is, the bird sounds are converted from the time domain to the frequency domain to meet the input requirements of the voiceprint recognition neural network model.

[0030] S2: Input the processed bird singing data into the voiceprint recognition neural network model based on deep learning to obtain a high-dimensional feature vector.

[0031] Furthermore, the deep learning-based voiceprint recognition neural network model in S2 is an improvement of the ECAPA-TDNN model. The improvement process specifically includes: A1: Use dilated convolutions with different dilation rates to capture multi-scale context, where the capture formula is:

[0032] Where, The time scale is of is the dilated convolution with dilation rates of 1, 2, and 3, is the time scale index, The raw data or feature maps that need to be captured in multi-scale context; A2: Concatenate features of all scales and perform standard convolution processing:

[0033] in, For scale 1 is the dilated convolution with dilation rates of 1, 2, and 3, For scale 1 is the dilated convolution with dilation rates of 1, 2, and 3, For scale 1 The dilated convolutions are of dilation rates 1, 2, and 3; A3: Global average pooling and fully connected layers are used to calculate channel weights to obtain high-dimensional feature vectors. The channel weight calculation formula is:

[0034] in, The features generated by concatenating features of all scales for the above context and performing standard convolution processing are: is the activation function, is global average pooling.

[0035] S3: Compare the high-dimensional feature vector with each bird voiceprint feature data in the bird voiceprint feature library to obtain the most similar bird species.

[0036] Furthermore, the step S3 compares the high-dimensional feature vector with each bird voiceprint feature data in the bird voiceprint feature library to obtain the most similar bird species, specifically including: Compare the high-dimensional feature vector with each bird voiceprint feature data in the bird voiceprint feature library. If the bird voiceprint feature library contains corresponding feature data, take the bird species corresponding to the corresponding feature data as the closest bird species; otherwise, re-compare. If the re-comparison is successful, obtain the bird species corresponding to the corresponding feature data; otherwise, record an error message.

[0037] The comparing of the high-dimensional feature vector with each bird voiceprint feature data in the bird voiceprint feature library specifically includes: The high-dimensional feature vector is sequentially compared with each bird voiceprint feature data in the bird voiceprint feature library by using the angle cosine algorithm to obtain the most similar bird species. The comparison formula is:

[0038] Where, A high-dimensional feature vector representing the current bird's voiceprint, Represents each feature vector in the bird voiceprint feature library, Represents the dot product of the current bird voiceprint feature vector and each vector in the bird feature library, Represents the Euclidean norm of the current bird voiceprint feature vector and each vector in the bird feature library.

[0039] In one embodiment of the present application, a set of high-dimensional feature vectors is obtained by the output of the voiceprint recognition neural network model obtained through the above-mentioned improvement, and then each bird voiceprint feature data in the bird voiceprint feature library is compared in turn through the angle cosine algorithm. The range value of the angle cosine algorithm is [-1,1]. The closer to 1, the more similar the two are, and finally the most similar bird species is determined.

[0040] S4: Based on the closest bird species obtained, the location of the bird is determined using multi-point cross-positioning technology based on the positioning sensor.

[0041] Furthermore, in S4, based on the obtained closest bird species, the location of the bird is determined using a multi-point cross-positioning technology based on a positioning sensor, specifically including: S401: According to the obtained closest bird species, obtain information of a positioning sensor near the bird species, and calculate the distance between the positioning sensor and the bird species; S402: Obtain the position of the bird through geometric calculation based on the position coordinates of the positioning sensor and the calculated distance.

[0042] In one embodiment of the present application, the positioning sensor may be one or more of an ultrasonic sensor, an infrared sensor, and a GPS. These sensors may be used alone or in combination to improve the accuracy and reliability of positioning. For example, an ultrasonic sensor may calculate distance by measuring the reflection time of sound waves; an infrared sensor may determine a position by detecting infrared radiation from an object; and a GPS may provide precise position information worldwide.

[0043] The multi-point cross-localization technology used here specifically involves the coordinated operation of multiple positioning sensors, using methods such as triangulation and multilateral measurement to accurately calculate the bird's three-dimensional coordinates (X, Y, Z). For example, in one embodiment, three ultrasonic sensors can be deployed to measure the distance to the bird. The bird's precise location is then determined through geometric calculations based on these three distance values ​​and the sensor's position coordinates.

[0044] S5: According to the determined position of the bird, the camera adjusts the shooting angle and focal length to shoot the bird.

[0045] In one embodiment of the present application, a camera with high pixels, high frame rate, and large aperture is selected to capture clear and detailed images of birds. The camera can support optical zoom and autofocus functions, and can quickly adjust the shooting angle and focal length based on the coordinates provided by the positioning sensor to ensure that the birds are always within the clear range of the picture. For example, in one embodiment, a 20-megapixel high-definition camera with 5x optical zoom can be used, with a lens focal length range of 24-120mm and an aperture size of F2.8-F4.0, which can meet the shooting requirements of different scenarios.

[0046] Example 2: The embodiment of the present application provides a bird active photography system based on voiceprint recognition and automatic positioning, which can be seen in Figure 2 , Figure 2 The figure shows a schematic diagram of a bird active photography system based on voiceprint recognition and automatic positioning provided by an embodiment of the present application, including: The voiceprint recognition module includes a voiceprint collection device, a voiceprint processing unit and a voiceprint recognition algorithm unit connected in sequence, which is used to capture the singing sounds of birds in real time and identify the bird species.

[0047] The voiceprint collection device adopts a high-sensitivity microphone array composed of multiple microphone units, and the microphone array has a wide frequency response range and a high signal-to-noise ratio; The voiceprint processing unit performs discrete Fourier transform processing on the collected bird call sound data; The voiceprint recognition algorithm unit uses a voiceprint recognition neural network model based on deep learning to identify the processed chirping sound data to obtain similar bird species.

[0048] The positioning sensor module includes an ultrasonic sensor, an infrared sensor, and a GPS module, which is used to automatically locate the bird's position after voiceprint recognition.

[0049] The active shooting module includes a high-resolution camera and an autofocus system, which is used to actively adjust the shooting angle and focal length according to the position coordinates provided by the positioning sensor module and shoot the birds.

[0050] Furthermore, the autofocus system utilizes phase detection focusing or contrast detection focusing technology to achieve autofocus.

[0051] In one embodiment of the present application, fast and accurate autofocus is achieved using technologies such as phase detection autofocus (PDAF) or contrast detection autofocus (CDAF). Based on the bird's position information provided by the positioning sensor, the autofocus system can predict the bird's movement trends and adjust the focus position in advance to capture clear bird images. Furthermore, image recognition algorithms can be combined to perform real-time analysis of the captured image to determine its clarity and quality. If the image does not meet the requirements, the focus is automatically re-adjusted and the shot is taken.

[0052] The data processing and storage module includes an image receiving module, an embedded processor, a data transmission device and a storage device, which is used to process the received voiceprint, location and image data, and store the data in a local server library and upload it to the cloud for backup.

[0053] In one embodiment of the present application, an embedded processor serves as the core processing unit of the system, responsible for integrating and processing voiceprint, location, and image data. Embedded processors offer low power consumption and high performance, enabling them to run a variety of data processing algorithms and applications. For example, high-performance processors based on the ARM architecture, such as the NVIDIA Jetson series, can be selected. These processors offer powerful graphics processing capabilities and artificial intelligence computing capabilities, enabling them to meet complex data processing needs.

[0054] In one embodiment of the present application, a storage device is used to store voiceprints, location, and image data, as well as system configuration files and applications. The storage device can be a local SD card, solid-state drive (SSD), or network-attached storage (NAS), depending on the data volume and storage requirements. The system also supports data classification and labeling, organizing and managing data by bird species, shooting time, and shooting location, facilitating subsequent analysis and research. For example, the storage device can be divided into different partitions to store raw data, processed data, and analysis results, respectively, to improve data manageability and security.

[0055] In one embodiment of the present application, the data transmission device can select different data transmission methods, such as Wi-Fi, Bluetooth, 4G / 5G wireless communication modules or wired network interfaces, to transmit monitoring data to a local server or cloud storage platform according to the actual application scenario and needs. The data transmission device should be high-speed, stable, and reliable to ensure the integrity and timeliness of the data. For example, in a field monitoring scenario, a 4G / 5G wireless communication module can be used to transmit data to a cloud server via a mobile network to achieve remote monitoring and data sharing.

[0056] In one embodiment of the present application, the system employs a storage strategy that allows for the configuration of regular data backup tasks, backing up important data to multiple storage devices or cloud storage platforms to prevent data loss and corruption. Furthermore, storage devices can be managed hierarchically based on the frequency and importance of data, storing frequently accessed data on high-speed storage devices and less frequently accessed data on lower-speed storage devices, thereby improving the overall performance and cost-effectiveness of the storage system.

[0057] Furthermore, the local server and the cloud have data analysis software for performing comprehensive analysis on the collected voiceprint, location and image data, and generating a data analysis report.

[0058] In one embodiment of the present application, data analysis software or tools running on a local server or cloud platform are used to conduct in-depth analysis and mining of monitoring data. These analysis tools may include statistical analysis software, machine learning algorithm libraries, image recognition software, etc., which can comprehensively analyze voiceprint, location, and image data to extract valuable information and knowledge. For example, the Python programming language and its related data analysis libraries (such as NumPy, Pandas, Matplotlib, etc.) can be used for data cleaning, preprocessing, statistical analysis, and visualization; machine learning algorithms (such as cluster analysis, classification algorithms, regression analysis, etc.) can be used to mine and model bird behavior patterns and distribution patterns; and image recognition software can be used to analyze and study bird morphological characteristics and activity trajectories.

[0059] A display module is connected to the data processing and storage module and is used to display a data analysis report.

[0060] Based on the results of data analysis, a detailed monitoring report is generated, including information such as bird species, numbers, distribution, behavioral characteristics, and related multimedia materials such as charts, images, and videos. Reports can be presented in formats such as HTML, PDF, and Word, making them convenient for users to view and share. The system's display module also features a corresponding visual interface or mobile application, presenting monitoring data and analysis results to users in an intuitive manner. For example, maps can be used to mark the distribution of birds, timelines can be used to display bird activity patterns, and charts can be used to display bird population trends, thereby improving users' understanding and utilization of monitoring information.

[0061] The present application provides a method and system for active bird photography based on voiceprint recognition and automatic positioning. Birds are identified through voiceprint equipment, the positioning sensor is used to automatically locate the position of the birds, and the camera is actively mobilized to photograph the located birds without human intervention, thereby achieving efficient and accurate bird monitoring and improving the efficiency and accuracy of monitoring.

[0062] It should be noted that those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of this application, and it should be understood that the scope of protection of this application is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in this application without departing from the essence of this application, and such variations and combinations are still within the scope of protection of this application.

Claims

1. A bird active photography method based on voiceprint recognition and automatic positioning, characterized in that: include: S1: Use the voiceprint collection device to collect bird calls and perform discrete Fourier transform processing on the collected bird calls; S2: Input the processed bird song data into a deep learning-based voiceprint recognition neural network model to obtain a high-dimensional feature vector; S3: Compare the high-dimensional feature vector with each bird voiceprint feature data in the bird voiceprint feature library to obtain the most similar bird species; S4: Based on the closest bird species obtained, the location of the bird is determined using multi-point cross-positioning technology based on the positioning sensor; S5: According to the determined position of the bird, the camera adjusts the shooting angle and focal length to shoot the bird.

2. The bird active photography method based on voiceprint recognition and automatic positioning according to claim 1 is characterized in that: The specific formula for the discrete Fourier transform processing in S1 includes: in, is the frequency domain, is the discrete time domain signal of birds, , is the frequency index, is the time index, is the total number of sampling points, is the imaginary unit, is the base of natural logarithms.

3. The bird active photography method based on voiceprint recognition and automatic positioning according to claim 1 is characterized in that: The deep learning-based voiceprint recognition neural network model in S2 is an improvement on the ECAPA-TDNN model. The improvement process specifically includes: A1: Use dilated convolutions with different dilation rates to capture multi-scale context, where the capture formula is: Where, The time scale is of is the dilated convolution with dilation rates of 1, 2, and 3, is the time scale index, The raw data or feature maps that need to be captured in multi-scale context; A2: Concatenate features of all scales and perform standard convolution processing: in, For scale 1 is the dilated convolution with dilation rates of 1, 2, and 3, For scale 1 is the dilated convolution with dilation rates of 1, 2, and 3, For scale 1 The dilated convolutions are of dilation rates 1, 2, and 3; A3: Global average pooling and fully connected layers are used to calculate channel weights to obtain high-dimensional feature vectors. The channel weight calculation formula is: in, The features generated by concatenating features of all scales and performing standard convolution processing are: is the activation function, is global average pooling.

4. The bird active photography method based on voiceprint recognition and automatic positioning according to claim 1 is characterized in that: The step S3 compares the high-dimensional feature vector with each bird voiceprint feature data in the bird voiceprint feature library to obtain the most similar bird species, specifically including: Compare the high-dimensional feature vector with each bird voiceprint feature data in the bird voiceprint feature library. If the bird voiceprint feature library contains corresponding feature data, take the bird species corresponding to the corresponding feature data as the closest bird species; otherwise, re-compare. If the re-comparison is successful, obtain the bird species corresponding to the corresponding feature data; otherwise, record an error message.

5. The bird active photography method based on voiceprint recognition and automatic positioning according to claim 1 or 4, characterized in that: The comparing of the high-dimensional feature vector with each bird voiceprint feature data in the bird voiceprint feature library specifically includes: The high-dimensional feature vector is sequentially compared with each bird voiceprint feature data in the bird voiceprint feature library by using the angle cosine algorithm to obtain the most similar bird species. The comparison formula is: Where, A high-dimensional feature vector representing the current bird's voiceprint, Represents each feature vector in the bird voiceprint feature library, Represents the dot product of the current bird voiceprint feature vector and each vector in the bird feature library, Represents the Euclidean norm of the current bird voiceprint feature vector and each vector in the bird feature library.

6. The bird active photography method based on voiceprint recognition and automatic positioning according to claim 1 is characterized in that: In S4, based on the obtained closest bird species, the location of the bird is determined using a multi-point cross-positioning technology based on a positioning sensor, specifically including: S401: According to the obtained closest bird species, obtain information of a positioning sensor near the bird species, and calculate the distance between the positioning sensor and the bird species; S402: Obtain the position of the bird through geometric calculation based on the position coordinates of the positioning sensor and the calculated distance.

7. A bird active photography system based on voiceprint recognition and automatic positioning, characterized in that: include: A voiceprint recognition module, comprising a voiceprint collection device, a voiceprint processing unit, and a voiceprint recognition algorithm unit connected in sequence, for capturing bird calls in real time and identifying bird species; Positioning sensor module, including ultrasonic sensor, infrared sensor and GPS module, used to automatically locate the bird's position after voiceprint recognition; An active shooting module, including a high-resolution camera and an autofocus system, for actively adjusting the shooting angle and focal length according to the position coordinates provided by the positioning sensor module and shooting the birds; The data processing and storage module includes an image receiving module, an embedded processor, a data transmission device, and a storage device. It is used to process the received voiceprint, location, and image data, and store the data in a local server library and upload it to the cloud for backup; A display module is connected to the data processing and storage module and is used to display a data analysis report.

8. The bird active photography system based on voiceprint recognition and automatic positioning according to claim 7 is characterized in that: The voiceprint collection device uses a high-sensitivity microphone array composed of multiple microphone units, and the microphone array has a wide frequency response range and a high signal-to-noise ratio; The voiceprint processing unit performs discrete Fourier transform processing on the collected bird call sound data; The voiceprint recognition algorithm unit uses a voiceprint recognition neural network model based on deep learning to identify the processed chirping sound data to obtain similar bird species.

9. The bird active photography system based on voiceprint recognition and automatic positioning according to claim 7, characterized in that: The automatic focusing system uses phase detection focusing or contrast detection focusing technology to achieve automatic focusing.

10. The bird active photography system based on voiceprint recognition and automatic positioning according to claim 7, characterized in that: The local server and the cloud have data analysis software for performing comprehensive analysis on the collected voiceprint, location and image data and generating a data analysis report.

Citation Information

Patent Citations

  • Bird identification method and device, terminal equipment and computer readable storage medium

    CN110033777A

  • Bird online monitoring system and method combining image and acoustic recognition technology

    CN110730331A

  • Bird feature recognition-based habitat environment adjustment method and system

    CN117809662A

  • Animal voiceprint monitoring method and device, medium and product

    CN119091890A

Cited By

  • Intelligent bird identification and observation system based on edge calculation

    CN121527814A

  • Ecological intelligent capturing method and system based on acoustic-infrared multi-mode cooperation

    CN122063694A