Bird sound directional acquisition method and device based on target identification and positioning, and storage medium
By using the camera and YOLOv8 network model for bird target recognition and positioning, and combining the gimbal recording equipment to realize directional collection of bird sounds, it solves the problems of high energy consumption and high waste of storage equipment resources in the existing technology, and realizes high-quality bird audio acquisition and automatic tracking functions.
Patent Information
- Application Number
- CN202510556212.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, bird sound acquisition strategies have high energy consumption and high waste of storage equipment resources, and bird audio acquisition cannot track the bird's orientation at any time, resulting in low audio quality.
Bird video data is collected through the camera, and after preprocessing, the YOLOv8 network model is used to identify and locate bird targets, and the directional collection of bird sounds is achieved with the gimbal recording equipment.
It realizes automatic collection of bird sound data all-weather, saves power and storage resources, improves audio acquisition quality, and can automatically track bird orientation.
Smart Images

Figure CN120070511A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of target recognition and tracking, and particularly to a bird sound directional acquisition method, device, and storage medium based on target recognition and positioning. Background Art
[0002] The sounds of birds are not only an important part of their biological characteristics but also provide important data support for ecological research. Through the collection and analysis of bird sounds, researchers can monitor information such as the distribution of bird populations, behavioral patterns, and changes in habitat environments. In field ecological research, bird sounds are widely used for species identification, habitat monitoring, and behavioral research.
[0003] However, current bird sound collection mainly relies on portable recording devices and directional microphones. Most traditional recording devices are single microphones or microphone arrays, which can collect sound signals in the environment but cannot achieve precise collection of sounds from specific sources. To achieve directional acquisition of bird sounds, researchers often use directional microphone arrays or highly customized collection devices. Directional microphones can suppress noise from other directions by focusing on the propagation direction of sound waves and improve the signal quality of the target sound. Since birds may make sounds during movement, the above collection methods cannot track the orientation of birds at any time for collection, resulting in low-quality collected audio, and the collection strategy has high energy consumption and high waste of storage device resources. Therefore, enabling the recording device to follow the orientation of birds for collection is an important means to ensure the quality of the collected audio. Summary of the Invention
[0004] The purpose of this application is to provide a bird sound directional acquisition method, device, and storage medium based on target recognition and positioning to solve the problems in the existing bird sound collection strategy, such as high energy consumption, high waste of storage device resources, and the inability to track the orientation of birds at any time for bird audio collection, resulting in low-quality collected audio.
[0005] To achieve the above purpose, an embodiment of this application provides a bird sound directional acquisition method based on target recognition and positioning, including the following steps: collecting bird videos through a camera to obtain video data;
[0006] preprocessing the collected video data;
[0007] performing bird target recognition and positioning on the preprocessed video data to obtain recognition and positioning results;
[0008] tracking the bird target according to the recognition and positioning results to obtain tracking results;
[0009] performing directional acquisition of bird sounds according to the tracking results in combination with a recording device mounted on a pan-tilt head.
[0010] Optionally, the preprocessing of the collected video data specifically includes:
[0011] Decode the video stream into individual frames, extract a certain number of frames per second, and resize each frame image to a uniform size;
[0012] Convert each frame of image into a grayscale image;
[0013] Applying Gaussian filtering to the grayscale image to remove noise in the image, using median filtering to remove noise points, adjusting the brightness and contrast of the image, and performing contrast enhancement in local areas to improve image quality;
[0014] The Canny edge detection algorithm is used to extract the edge information of the image, and morphological operations are used to further improve the structure of the image. By first corroding and then expanding, small noise points are removed, the target boundary is smoothed, and the target features are highlighted.
[0015] Optionally, the bird target identification and positioning may include:
[0016] The preprocessed video data is input into the YOLOv8 network model, and a set of high-dimensional feature maps is obtained after passing through the backbone network;
[0017] After the high-dimensional feature map is subjected to feature fusion of the neck and convolution processing of the head, a plurality of pre-selection boxes are obtained;
[0018] Non-maximum suppression is used for the multiple pre-selected boxes to obtain the IoU of all candidate boxes. For each target category, the box with the highest confidence is retained, and other boxes that overlap with it by more than a set threshold are deleted to obtain and output the target detection result.
[0019] Optionally, the YOLOv8 network model is a bird recognition model obtained through pre-training, and the specific steps of obtaining the bird recognition model include:
[0020] Obtain a bird identification dataset, annotate each image in the bird identification dataset, and divide the dataset into a training set, a validation set, and a test set;
[0021] Data augmentation for bird identification dataset using rotation, flipping, cropping and scaling, and noise addition methods;
[0022] Select the YOLOv8 architecture and hyperparameter configuration suitable for the bird target detection task;
[0023] Use the annotated bird recognition dataset to train the YOLOv8 network model;
[0024] Use the test set to evaluate the final performance of the YOLOv8 network model;
[0025] Export the trained YOLOv8 network model as a PyTorch model, convert it to the ONNX format, and optimize the model with TensorRT to finally obtain the bird recognition model.
[0026] Optionally, the completion of the tracking of the bird target specifically includes:
[0027] When a bird target is first detected, initialize the direction of the pan-tilt head so that the camera points to the target. At this time, calculate the moving direction of the bird target as the diagonal direction between the current angle of the camera and the target position;
[0028] In two consecutive frames of images, estimate the moving direction and speed of the bird target by calculating the position change of the center point of the target bounding box;
[0029] Use the historical trajectory and speed information of the bird target, and adopt a Kalman filter to predict the possible position of the bird target in the next frame;
[0030] According to the change of the bird target relative to the current field of view, control the pitch angle and yaw angle of the pan-tilt head;
[0031] Control the pan-tilt head through a PID control algorithm;
[0032] If the position of the bird target changes, update the position and direction of the bird target.
[0033] Optionally, the combination with the recording device mounted on the pan-tilt head to complete the directional acquisition of bird sounds specifically includes:
[0034] First, determine whether there is a bird in the camera image through bird target recognition and positioning. If there is no bird, maintain visual capture; if there is a bird, issue a prompt and turn on the recording device to complete the directional acquisition of bird audio data in combination with the pan-tilt head.
[0035] Optionally, the use of the labeled bird recognition dataset to train the YOLOv8 network model specifically includes:
[0036] During the training process, use the labeled bird recognition dataset for model training and save the model weights after each epoch;
[0037] After the training is completed, use the test set to evaluate the model performance. The evaluation metrics include mAP, precision, recall rate, etc., and further analyze the effect of the model and make optimizations.
[0038] To achieve the above object, the present application also provides a bird sound directional acquisition device based on target recognition and positioning, including:
[0039] A camera, which is used to collect bird videos;
[0040] A recording device for directionally collecting bird sounds;
[0041] A pan-tilt unit for enabling the camera and the recording device to face the direction of the bird target when directionally collecting bird sounds;
[0042] A memory storing a computer program;
[0043] A processor for using the computer program stored in the memory to complete the method for directionally collecting bird sounds based on target recognition and positioning described in any one of the foregoing;
[0044] A power supply unit for powering all devices.
[0045] To achieve the above object, the present application also provides a computer storage medium having a computer program stored thereon, wherein when the computer program is executed by a machine, the steps of the method described above are implemented.
[0046] The embodiments of the present application have the following advantages:
[0047] Through the above method, 1. it is possible to automatically collect the sound data of the calls of special birds all day long, solve the problem of limited manual processing ability, improve work efficiency, and save human resources;
[0048] 2. it is able to automatically grasp the timing of turning on the recording device, and achieve that the recording collection is only turned on when a bird appears, saving power and storage resources;
[0049] 3. The method for tracking bird targets based on a deep learning network model can automatically complete the tracking of the bird's orientation, and the collection device follows the direction of the bird, improving the quality of audio collection. Description of the Drawings
[0050] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.
[0051] Figure 1 A flowchart of a method for directionally collecting bird sounds based on target recognition and positioning provided for at least one embodiment of the present application;
[0052] Figure 2 A schematic diagram of a device for directionally collecting bird sounds based on target recognition and positioning provided for at least one embodiment of the present application;
[0053] Figure 3 A simplified flowchart of a bird sound directional acquisition method based on target recognition and positioning provided by at least one embodiment of the present application;
[0054] Figure 4 A schematic diagram of the specific operation implementation process of a bird sound directional acquisition method based on target recognition and positioning provided by at least one embodiment of the present application;
[0055] Figure 5 A structural diagram of the YOLOv8 network model, a bird recognition model used in a bird sound directional acquisition method based on target recognition and positioning provided by at least one embodiment of the present application;
[0056] Figure 6 Schematic diagrams of the overall detailed processes of a bird sound directional acquisition method based on target recognition and positioning provided by at least one embodiment of the present application. Detailed implementation manners
[0057] The following specific embodiments illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0058] It should be noted that in the claims and the description of the present application, the steps can be basically executed in parallel or in the reverse order under appropriate circumstances, depending on the functions involved.
[0059] In addition, the technical features involved in different implementation manners of the present application described below can be combined with each other as long as they do not conflict with each other.
[0060] With the development of computer vision and deep learning technologies, target detection and tracking technologies have shown great potential in bird sound monitoring. Target detection technology is usually used to identify and locate the targets of interest in images or videos, while target tracking technology is used to maintain the tracking of the targets in consecutive video frames. How to combine target detection models to complete the directional acquisition of bird sounds is a problem worthy of exploration.
[0061] The embodiments of the present application take the YOLOv8 network model algorithm as the main body, and combine a two-degree-of-freedom pan-tilt and a recording device to complete the directional acquisition of bird sounds, which can provide optional technical support and reference for the collection of bird soundprints in the wild.
[0062] The design concept of this application is to use a high-definition camera and a recording device mounted on a two-degree-of-freedom pan-tilt to complete the directional acquisition of bird sounds, including the following steps:
[0063] Step 1: Use the high-definition camera mounted on the two-degree-of-freedom pan-tilt to capture the moving images of birds in real time, and use the deep learning object detection algorithm to identify the specific position coordinates of the birds and determine their moving directions within the field of view;
[0064] Step 2: Input the bird target position data captured by the high-definition camera into the motion tracking algorithm, calculate the bird motion trajectory and speed, and judge the real-time relative position and direction of the target;
[0065] Step 3: According to the bird position and direction data output by the motion tracking algorithm, control the motor actions on the two-degree-of-freedom pan-tilt, and adjust the pitch angle and horizontal rotation angle of the pan-tilt to keep the pan-tilt always aligned with the bird target;
[0066] Step 4: After the pan-tilt completes the direction adjustment, ensure that the sound pickup direction of the recording device is aligned with the bird target, and use the high-sensitivity recording device to collect the voiceprint data emitted by the birds;
[0067] Step 5: Store the image data captured by the high-definition camera and the voiceprint data collected by the recording device simultaneously, and perform feature extraction and classification on the collected voiceprints through the voiceprint recognition algorithm to establish a bird voiceprint sample library;
[0068] Step 6: According to the changes in the moving speed and direction of the bird target, adjust the motion tracking and pan-tilt control parameters in real time, optimize the tracking accuracy and reaction speed of the system, and ensure the stability and accuracy of the acquisition process.
[0069] The above are the steps of the design concept of this application. For the specific method of completing the acquisition of bird voiceprint data, refer to the following embodiments.
[0070] An embodiment of this application provides a method for directional acquisition of bird sounds based on target recognition and positioning, refer to Figures 1 to 3 , Figure 1 and Figure 3 are the flowcharts of a method for directional acquisition of bird sounds based on target recognition and positioning provided in at least one embodiment of this application. It should be understood that the method may also include additional boxes not shown and / or may omit the boxes shown. The scope of this application is not limited in this regard.
[0071] At step 101, collect the bird video through the camera to obtain video data.
[0072] Specifically, this acquisition method uses a high-definition camera to perform image acquisition at a fixed frequency.
[0073] Specifically, in order to collect images on a pan-tilt equipped with a high-definition camera, the acquisition strategy will be adjusted according to the dynamic characteristics of the target and the hardware resources. Since birds move at a relatively fast speed, in this embodiment, the acquisition frequency is set between 30 FPS according to the requirements of real-time detection, and the image resolution is selected as 1920x1080 high definition, which is suitable for most detection and tracking tasks. This can ensure the effectiveness of monitoring the acquisition area while reducing the resource waste caused by continuous acquisition.
[0074] At step 102, preprocess the acquired video data.
[0075] Specifically, through multiple image preprocessing techniques, gradually remove noise, enhance the image quality, and extract useful edge information, making the contour of the target clearer and improving the recognition accuracy of subsequent steps.
[0076] In some embodiments, the preprocessing of the acquired video data specifically includes:
[0077] First, decode the video stream into individual frames, and extract a certain number of frames per second to ensure data continuity. Then, adjust each frame of the image to a unified size to avoid the impact of inconsistent image sizes on subsequent processing. Next, convert each frame of the image into a grayscale image to reduce the computational complexity. This can maintain the brightness information of the image while removing the color information and reducing the amount of calculation. Then apply Gaussian filtering to remove the noise in the image, then use median filtering to remove the noise points, and then adjust the brightness and contrast of the image to make the grayscale distribution of the image more uniform. Finally, improve the image quality through local area contrast enhancement. Then use the Canny edge detection algorithm to extract the edge information of the image, and then use morphological operations to further improve the structure of the image. By eroding first and then dilating, remove small noise points, smooth the target boundary, and highlight the target features.
[0078] Specifically, for the above operations, among them, Gaussian filtering is used to reduce the high-frequency noise in the image. Its core idea is to calculate the weighted average of the neighborhood pixel values, and its calculation formula is:
[0079]
[0080] where G(x,y) represents the Gaussian kernel weight value at position (x,y), are the coordinates of the filter kernel, is the standard deviation, which controls the width of the filter. It first applies the Gaussian kernel to each pixel and its neighborhood and calculates the weighted average value.
[0081] Specifically, median filtering is used to remove salt and pepper noise. The implementation operation is as follows: for each pixel, consider a window (such as a 3x3 or 5x5 window), sort the pixel values within the window, and replace the current pixel value with the sorted median. This operation can effectively remove the isolated noise points in the above image acquisition and processing process.
[0082] Specifically, the brightness and contrast of the image are also adjusted. Specifically, by increasing or decreasing the brightness value of the image, the overall image brightness is brought to an ideal level. By linearly transforming the gray levels of the image, the contrast of the image is enhanced, making the bright and dark areas of the image more distinct.
[0083] Specifically, the specific formula for contrast adjustment is:
[0084]
[0085] where I old represents the gray value or color channel value of the corresponding pixel in the original image, and I new represents the pixel value after contrast adjustment. α is the contrast factor, used to control the enhancement or weakening of the image contrast, and β is the brightness offset, used to adjust the overall brightness level of the image. Increasing α can enhance the contrast, and adjusting β can change the overall brightness of the image.
[0086] Specifically, the Canny edge detection is used to detect the edges in the image. The steps of the Canny edge detection algorithm for extracting the significant change regions in the image include:
[0087] Gaussian filtering: Perform Gaussian filtering on the image to reduce noise.
[0088] Gradient calculation: Calculate the gradient of the image, usually using the Sobel operator.
[0089] Non-maximum suppression: Suppress non-edge points in the gradient direction and retain local maxima.
[0090] Double-threshold processing: Set two thresholds, high and low, to distinguish strong edges and weak edges.
[0091] Edge connection: Connect weak edges to strong edges through edge connection.
[0092] The formula for gradient calculation in Canny edge detection is:
[0093]
[0094] where M represents the magnitude of the gradient of the image pixel points, used to measure the severity of the gray level change in the local area of the image, and are the gradients of the image in the horizontal and vertical directions, respectively.
[0095] Specifically, the morphological operation is finally used to improve the image structure, remove noise, and smooth the target boundary. The process is to first apply the erosion operation to remove small noise points. Then apply the dilation operation to smooth the target boundary and highlight the target features.
[0096] Among them, the operation formulas of corrosion and expansion are:
[0097]
[0098]
[0099] in, is the input binary image, represents the structural element centered at position z, represents the intersection operation of sets, Represents the dilation operation of the structural element. E(A) represents the result image after the erosion processing of image A, which is mainly used to reduce the foreground area; D(A) represents the result image after the dilation processing of image A, which is mainly used to expand the foreground area.
[0100] In step 103, bird target recognition and positioning are completed for the pre-processed video data to obtain recognition and positioning results.
[0101] Specifically, in this step, the image data preprocessed in step 102 is input into the target recognition network YOLOv8, and the recognition result is obtained after being processed by the YOLOv8 network model. The recognition result includes whether the image data contains bird information and the type of bird.
[0102] In some embodiments, the detailed steps for completing bird target recognition and positioning are as follows:
[0103] The video data preprocessed in step 102 is input into the YOLOv8 network model. After passing through the backbone network, a set of high-dimensional feature maps are obtained. Then, after the high-dimensional feature maps are subjected to feature fusion of the neck and convolution processing of the head, multiple pre-selected boxes are obtained. Then, the multiple pre-selected boxes are subjected to non-maximum suppression (NMS) to calculate the IoU (Intersection over Union) of all candidate boxes. Then, for each target category, the box with the highest confidence is retained, and other boxes that overlap with it by more than the set threshold are deleted to obtain the target detection result, and finally the target detection result is output.
[0104] Among them, NMS is a very important step to obtain the correct result. NMS is a post - processing method for screening object detection results, aiming to remove redundant detection boxes. Multiple pre - selected boxes may cover the same object. NMS calculates the IoU (Intersection over Union) value between candidate boxes to decide which box to retain.
[0105] Specifically, the intersection - over - union (IoU) is an important indicator to measure the overlapping degree of two bounding boxes, and its calculation formula is:
[0106]
[0107] Among them, |J∩H| represents the intersection area of box J and box H, and |J∪H| represents the union area of box J and box H.
[0108] The processing process of NMS is specifically as follows: For each pair of candidate boxes, calculate their intersection - over - union (IoU); for each object category, select the box with the highest confidence (the probability predicted by the model); for boxes with an overlapping area greater than the set threshold (usually 0.5), delete these boxes. Set an IoU threshold θ, usually θ = 0.5; if the IoU of two boxes is greater than the threshold, delete one of the boxes, usually the one with lower confidence.
[0109] At step 104, according to the recognition and positioning results, complete the tracking of the bird target to obtain the tracking result.
[0110] Specifically, it means that when video data is captured by a high - definition camera and the target bird is detected by YOLOv8, through the steps of the target tracking algorithm, the pan - tilt head moves, so that the camera and the recording device face the bird.
[0111] In some embodiments, the detailed processing steps for completing the tracking of the bird target are as follows: When the bird target is first detected, initialize the pitch angle and yaw angle of the pan - tilt head by calculating the diagonal direction between the target position and the initial position of the camera, so that the camera points to the target. In subsequent consecutive frames, calculate the position change of the center point of the target bounding box, and estimate the moving direction and speed of the target. Use the historical trajectory and speed information of the target, and adopt a Kalman filter to predict the possible position of the target in the next frame to reduce the risk of target loss. According to the position change of the target relative to the current field of view, adjust the pitch angle and yaw angle of the pan - tilt head in real time. Precisely control the adjustment of the pan - tilt head through the PID control algorithm to avoid excessive oscillation and position error, and ensure that the pan - tilt head tracks the target smoothly and precisely. If the target position changes, update the target position and direction, and repeat this process to achieve stable target tracking.
[0112] Among them, V is the velocity vector. Assuming the time interval between each frame is Δt, the center point positions of the target in the current frame and the previous frame are (x 1 , y 1 ) and (x 2 , y 2 ). The calculation of the velocity change is as follows:
[0113]
[0114] The moving direction of the target can be estimated by the direction of the line connecting the target position in the current frame and the target position in the previous frame. Specifically, by calculating the angle of the connecting line, which is the moving direction of the target (assuming in a two-dimensional plane), then the angle θ direction is:
[0115]
[0116] Specifically, the PID control algorithm is also used in step 104 to precisely control the pan-tilt head. To make it face the target and avoid excessive oscillation and position error, by adjusting the pitch angle and yaw angle of the pan-tilt head, ensure the smooth movement of the pan-tilt head and minimize the target tracking error. The formula of the PID controller is:
[0117]
[0118] Among them, u(t) represents the controller output at time t, which is used to guide the adjustment direction and amplitude of the pan-tilt head. e(t) is the current error, that is, the deviation between the target position and the current position of the pan-tilt head. K p , K i , K d are the proportional gain, integral gain, and derivative gain coefficients respectively. e(τ) represents the error signal at any historical moment τ (where 0 ≤ τ ≤ t). The integral term is used to accumulate historical errors to correct the system deviation. τ is the integral variable, representing the time point elapsed during the integration process.
[0119] In step 105, according to the tracking result, combined with the recording device mounted on the pan-tilt head, the directional acquisition of bird sounds is completed.
[0120] Specifically, after using step 104 to complete the tracking effect, the recording device can be turned on to complete the acquisition.
[0121] In some embodiments, the specific steps of the directional acquisition of bird sounds combined with the recording device mounted on the pan-tilt head are:
[0122] First, judge whether there are birds in the camera screen through bird target recognition and positioning. If there are no birds, maintain visual capture; if there are birds, give a prompt and turn on the recording device to complete the acquisition of directional bird audio data in combination with the two-degree-of-freedom pan-tilt head.
[0123] As Figure 4 shown is the specific operation process, which shows the specific process of the device from startup to the completion of the acquisition work. The detailed process explanation is as follows:
[0124] First, power on and start the device. After all devices are turned on normally and initialized, the high-definition camera will collect environmental images at the pre-set acquisition frequency. During the acquisition process, at the same time, the image data will also be input into the pre-trained YOLOv8 network model after preprocessing operations for feature extraction, bird recognition, and positioning. Only when a bird is recognized will bird tracking be carried out, a prompt will be issued, and the recording device will be turned on for recording.
[0125] In some embodiments, the YOLOv8 network model used is a bird recognition model obtained by pre-training through a large number of data sets. The network model of YOLOv8 is as Figure 5 shown. The specific steps to obtain this bird recognition model are as follows:
[0126] First, prepare a large enough data set, including various bird images, and annotate each picture. The annotation content includes the bird category and the position of the bounding box.
[0127] Then divide the data set into a training set, a validation set, and a test set.
[0128] Next, perform data augmentation. By methods such as rotation, flipping, cropping, scaling, and noise addition, increase the diversity of the data and improve the generalization ability of the model.
[0129] Then, select the YOLOv8 architecture suitable for the bird target detection task and set hyperparameters such as the learning rate, batch size, and number of training epochs, etc., so that the model can handle bird detection tasks under different scales and complex backgrounds.
[0130] During the training process, use the annotated data set to train the model and save the model weights after each epoch.
[0131] After training is completed, use the test set to evaluate the model performance. Commonly used evaluation metrics include mAP, precision, and recall, etc., to further analyze the model effect and make optimizations.
[0132] If the model effect meets the standard, the trained YOLOv8 model can be exported as the PyTorch format and converted to the ONNX format for cross-platform deployment.
[0133] The model can also be optimized through TensorRT to accelerate the inference speed and ensure that the model can run efficiently on edge devices or mobile terminals, ultimately achieving the efficient deployment of the bird recognition task.
[0134] As shown Figure 6 in the figure is the process of all the above operations and the flow of the bird sound directional acquisition method based on target recognition and positioning of this application.
[0135] The corresponding bird sound directional acquisition device based on target recognition and positioning includes a memory and a processor. The memory is used to store computer programs, and the computer programs include all program codes and instructions. The processor is used to execute all steps in the above bird sound directional acquisition method based on target recognition and positioning. On the acquisition device, there are: a high-definition camera for shooting high-definition video data, a recording device for collecting bird sound data, a two-degree-of-freedom servo pan-tilt head: when used for bird sound directional acquisition, the high-definition camera and the recording device can face the direction of the bird, and a power supply unit for supplying power to all devices.
[0136] This application can be a method, a device, a system, and / or a computer program product. The computer program product can include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of this application are loaded.
[0137] The computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the above. The computer-readable storage medium used here is not construed as an instantaneous signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses through an optical fiber cable), or electrical signals transmitted through wires.
[0138] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0139] The computer program instructions for performing the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages - such as Smalltalk, C++, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of this application.
[0140] Aspects of the present application are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0141] These computer-readable program instructions may be provided to a processing unit of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0142] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0143] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0144] Note that, unless otherwise directly stated, all features disclosed in this specification (including any appended claims, abstract, and drawings) may be replaced by alternative features that serve the same, equivalent, or similar purposes. Therefore, unless otherwise expressly stated, each feature disclosed is only an example of a set of equivalent or similar features. Where used, "furthermore", "preferably", "moreover", and "even more preferably" are simple introductions for elaborating another embodiment based on the foregoing embodiments. The content following the "furthermore", "preferably", "moreover", or "even more preferably" is combined with the foregoing embodiments to form a complete composition of another embodiment. The components that can be arbitrarily combined among several "furthermore", "preferably", "moreover", or "even more preferably" settings following the same embodiment form yet another embodiment.
[0145] Although the present application has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it based on the present application, which will be obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present application fall within the scope of protection required by the present application.
Claims
1. A method for directional collection of bird sounds based on target recognition and positioning, characterized in that: The following steps are involved: The bird videos are collected by the camera to obtain video data; Preprocessing the collected video data; For the pre-processed video data, bird target recognition and positioning are completed to obtain recognition and positioning results; According to the identification and positioning results, the bird target is tracked to obtain a tracking result; According to the tracking results, the directional collection of bird sounds is completed in combination with the recording equipment mounted on the gimbal.
2. The method for directional collection of bird sounds based on target recognition and positioning according to claim 1 is characterized in that: The preprocessing of the collected video data specifically includes: Decode the video stream into individual frames, extract a certain number of frames per second, and resize each frame image to a uniform size; Convert each frame of image into a grayscale image; Applying Gaussian filtering to the grayscale image to remove noise in the image, using median filtering to remove noise points, adjusting the brightness and contrast of the image, and performing contrast enhancement in local areas to improve image quality; The Canny edge detection algorithm is used to extract the edge information of the image, and morphological operations are used to further improve the structure of the image. By first corroding and then expanding, small noise points are removed, the target boundary is smoothed, and the target features are highlighted.
3. The method for directional collection of bird sounds based on target recognition and positioning according to claim 1, characterized in that: The bird target identification and positioning is completed, specifically including: The preprocessed video data is input into the YOLOv8 network model, and a set of high-dimensional feature maps is obtained after passing through the backbone network; After the high-dimensional feature map is subjected to feature fusion of the neck and convolution processing of the head, a plurality of pre-selection boxes are obtained; Non-maximum suppression is used for the multiple pre-selected boxes to obtain the IoU of all candidate boxes. For each target category, the box with the highest confidence is retained, and other boxes that overlap with it by more than a set threshold are deleted to obtain and output the target detection result.
4. The method for directional collection of bird sounds based on target recognition and positioning according to claim 3 is characterized in that: The YOLOv8 network model is a bird recognition model obtained through pre-training. The specific steps of obtaining the bird recognition model include: Obtain a bird identification dataset, annotate each image in the bird identification dataset, and divide the dataset into a training set, a validation set, and a test set; Data augmentation for bird identification dataset using rotation, flipping, cropping and scaling, and noise addition methods; Select the YOLOv8 architecture and hyperparameter configuration suitable for the bird target detection task; Use the annotated bird recognition dataset to train the YOLOv8 network model; Use the test set to evaluate the final performance of the YOLOv8 network model; The trained YOLOv8 network model is exported as a PyTorch model, converted into ONNX format, and optimized using TensorRT to finally obtain the bird recognition model.
5. The method for directional collection of bird sounds based on target recognition and positioning according to claim 1, characterized in that: The tracking of the bird target specifically includes: When a bird target is detected for the first time, the direction of the gimbal is initialized so that the camera points to the target. At this time, the moving direction of the bird target is calculated as the diagonal direction of the current camera angle and the target position; In two consecutive frames, the moving direction and speed of the bird target are estimated by calculating the position change of the center point of the target bounding box; Using the historical trajectory and velocity information of the bird target, a Kalman filter is used to predict the possible position of the bird target in the next frame; Control the pitch and yaw angles of the gimbal according to the changes of the bird target relative to the current field of view; The PTZ is controlled by PID control algorithm; If the bird target position changes, update the bird target position and direction.
6. The method for directional collection of bird sounds based on target recognition and positioning according to claim 1, characterized in that: The directional collection of bird sounds by combining the recording device mounted on the gimbal specifically includes: First, bird target recognition and positioning are used to determine whether there are birds in the camera image. If no birds are seen, visual capture is maintained. If birds are seen, a prompt is given and the recording device is turned on in conjunction with the gimbal to complete the collection of directional bird audio data.
7. The method for directional collection of bird sounds based on target recognition and positioning according to claim 4, characterized in that: The YOLOv8 network model is trained using the labeled bird recognition data set, specifically including: During the training process, the labeled bird recognition dataset is used for model training, and the model weights are saved after each epoch. After training is completed, the test set is used to evaluate the model performance. Evaluation indicators include mAP, precision, and recall, etc., to further analyze the effect of the model and make adjustments.
8. A bird sound directional collection device based on target recognition and positioning, characterized in that: include: A camera, wherein the camera is used to collect bird videos; A recording device, wherein the recording device is used for directionally collecting bird sounds; A pan / tilt, which is used for directional bird sound collection so that the camera and recording equipment can be oriented in the direction of the bird target; a memory storing a computer program; A processor, the processor being used to use a computer program stored in a memory to implement the bird sound directional collection method based on target recognition and positioning as described in any one of claims 1 to 7; A power supply unit, wherein the power supply unit is used to supply power to all devices.
9. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a machine, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Transmission tower bird repelling method, device and system based on bird species identification
CN117612087A
Intelligent directional tracking bird repelling method, device and system
CN118104635A
Bird monitoring system with combination of dual-light camera carried by unmanned aerial vehicle and deep learning
CN118196660A
Natural environment bird monitoring method based on multi-modal fusion deep learning and computer device
CN119027775A
Cited By
Bird observation method, device and equipment based on camera group and readable medium
CN120635943A
Image feature recognition test method and device
CN122090143A