Intelligent detection method and system for bird population in Tibetan wetland, electronic equipment and storage medium

Through multi-sensor fusion and intelligent identification models, the problem of inefficiency of traditional bird monitoring methods is solved, efficient and accurate bird population monitoring is achieved, cost reduction and automatic monitoring is provided all-weather.

CN120070992APending Publication Date: 2025-05-30TIBET MUSEUM OF NATURAL SCIENCE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510150771.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional bird population monitoring methods are inefficient, are limited by manpower and time, and are difficult to cover a wide area, and data accuracy and reliability are disturbed by human factors. Modern technologies such as radio telemetry and satellite tracking, although improve monitoring range and accuracy, are costly, complex in operation and disturbing bird life.

Method used

Using multi-sensor fusion, data processing and compensation, intelligent identification model construction and spatial statistical analysis methods, multi-dimensional data is obtained through high-definition cameras, audio collectors and environmental sensors, denoising and compensation processing is carried out, and recognition model combining convolutional neural networks and recurrent neural networks is constructed to realize intelligent detection of bird populations.

Benefits of technology

It improves the efficiency and accuracy of bird population monitoring, reduces dependence on human resources, reduces costs, and realizes all-weather and fully automatic monitoring, providing strong technical support for ecological protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070992A_ABST
    Figure CN120070992A_ABST
Patent Text Reader

Abstract

The invention discloses a Tibetan wetland bird population intelligent detection method and system, electronic equipment and a storage medium, and the method comprises the steps: S1, arranging a plurality of high-definition cameras, an audio collector and a plurality of environment sensors around a wetland, and obtaining a high-definition image, bird voiceprint information and environment information of the wetland; s2, denoising the high-definition image and the bird voiceprint information respectively, and compensating the denoised image and the denoised voiceprint based on the environment information to obtain a compensated image and compensated voiceprint information; s3, constructing a bird image recognition model and a bird audio recognition model, respectively recognizing the compensated image and the compensated voiceprint information by using the two models, and performing mutual verification and complementation on recognition results to obtain a final recognition result; and S4, based on the final identification result, combining with the set spatial distribution of the equipment, and adopting spatial statistics to analyze the bird population number, the activity range and the migration rule to obtain a population detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ecological monitoring, and particularly to an intelligent detection method, system, electronic device and storage medium for bird populations in wetlands in Tibetan areas. Background Art

[0002] Wetlands in Tibetan areas, as the treasures of the Qinghai-Tibet Plateau, are not only hotspots of global biodiversity but also paradises for many rare birds. The wetland ecosystem here is fragile and complex, and changes in bird populations often reflect the health status of the entire ecosystem. Therefore, accurate monitoring of bird populations in wetlands in Tibetan areas has important scientific value and practical significance for maintaining ecological balance and protecting biodiversity.

[0003] Traditional bird population monitoring methods, like ancient explorations, rely on the eyes and ears of ornithologists. They observe at the edges of wetlands or lurk in the grass, using telescopes to observe and notebooks to record. Although this method is direct, it is inefficient, limited by manpower and time, difficult to cover vast areas, and even more difficult to achieve long-term continuous monitoring. In addition, interference from human factors, such as differences in observers' experience and fatigue levels, will affect the accuracy and reliability of data.

[0004] With the progress of technology, some modern technologies have begun to be applied to bird monitoring, such as radio telemetry and satellite tracking. Although these technologies have increased the monitoring range and accuracy, they are costly, complex to operate, and cause certain interference to the lives of birds. Therefore, these methods have not fully solved the problems existing in traditional monitoring methods.

[0005] In recent years, the rapid development of artificial intelligence technology, especially the breakthrough of deep learning in the field of image recognition, has brought new hope for bird population monitoring. Intelligent detection methods can process a large amount of image and video data, automatically identify bird species, and count population numbers, thus greatly improving the efficiency and accuracy of monitoring. This method not only reduces the dependence on human resources, lowers costs, but also can achieve all-weather and fully automatic monitoring, providing strong technical support for ecological protection.

[0006] However, applying artificial intelligence technology to bird population monitoring in wetlands in Tibetan areas still faces many challenges. For example, the particularity of the plateau environment, the diversity of bird species, and the changes in light and weather conditions all pose higher requirements for the robustness and accuracy of intelligent detection algorithms. Therefore, developing an intelligent detection method that adapts to the special environment of wetlands in Tibetan areas and can accurately identify and count bird populations has become an important research direction in the current field of ecological monitoring technology. Summary of the Invention

[0007] To solve the above technical problems, the present invention provides

[0008] An intelligent detection method for wetland bird populations in Tibetan areas, the method comprising:

[0009] S1. Set a number of high-definition cameras, audio collectors, and a number of environmental sensors around the wetland to obtain high-definition images of the wetland, bird soundprint information, and environmental information;

[0010] S2. Denoise the high-definition image and the bird soundprint information respectively, and compensate the denoised image and the denoised soundprint based on the environmental information to obtain a compensated image and compensated soundprint information;

[0011] S3. Construct a bird image recognition model and a bird audio recognition model, and use the two models to recognize the compensated image and the compensated soundprint information respectively, and mutually verify and supplement the recognition results to obtain a final recognition result;

[0012] S4. Based on the final recognition result, combined with the spatial distribution of the set devices, use spatial statistics to analyze the bird population quantity, activity range, and migration pattern to obtain a population detection result.

[0013] Preferably, the S2 includes:

[0014] Perform grayscale processing on the high-definition image to obtain a grayscale image, and use median filtering to denoise the grayscale image to obtain the denoised image;

[0015] Perform sampling rate conversion, frame segmentation, and windowing on the bird soundprint information to obtain preliminarily processed soundprint information, and then use spectral subtraction to denoise the preliminarily processed soundprint information to obtain the denoised soundprint;

[0016] Compensate the denoised image and the denoised soundprint based on the environmental information to obtain the compensated image and the compensated soundprint information.

[0017] Preferably, the S3 includes:

[0018] Select a convolutional neural network to construct an initial image recognition model, obtain historical bird population images and perform preprocessing, and train the initial image recognition model based on the preprocessed historical bird population images to obtain the bird image recognition model;

[0019] Select a combination of a convolutional neural network and a recurrent neural network to construct an initial audio recognition model, obtain historical bird audio information and perform preprocessing, and train the initial audio recognition model based on the preprocessed historical bird audio information to obtain the bird audio recognition model;

[0020] Use the bird image recognition model and the bird audio recognition model to respectively recognize the compensated image and the compensated voiceprint information, and obtain the bird population image recognition result and the bird population audio recognition result;

[0021] Perform weighted fusion on the bird population image recognition result and the bird population audio recognition result, and complement each other to obtain the final recognition result.

[0022] The present invention also provides an intelligent detection system for bird populations in wetland areas of Tibetan regions. The system is used to implement the method described in any one of the above, and includes: a data acquisition module, a data processing module, an identification module, and an analysis module;

[0023] The data acquisition module is used to set a number of high-definition cameras, audio collectors, and a number of environmental sensors around the wetland to obtain high-definition images of the wetland, bird voiceprint information, and environmental information;

[0024] The data processing module respectively denoises the high-definition image and the bird voiceprint information, and compensates the denoised image and the denoised voiceprint based on the environmental information to obtain the compensated image and the compensated voiceprint information;

[0025] The identification module is used to construct a bird image recognition model and a bird audio recognition model, and use the two models to respectively recognize the compensated image and the compensated voiceprint information, and mutually verify and supplement the recognition results to obtain the final recognition result;

[0026] The analysis module is based on the final recognition result, combined with the spatial distribution of the set devices, and uses spatial statistics to analyze the bird population quantity, activity range, and migration law to obtain the population detection result.

[0027] Preferably, the data processing module includes: an image processing unit, an audio processing unit, and a compensation unit;

[0028] The image processing unit is used to obtain a grayscale image by grayscaling the high-definition image, and perform image denoising on the grayscale image using median filtering to obtain the denoised image;

[0029] The audio processing unit is used to perform sampling rate conversion, frame division, and windowing on the bird voiceprint information to obtain the preliminarily processed voiceprint information, and then use spectral subtraction to denoise the preliminarily processed voiceprint information to obtain the denoised voiceprint;

[0030] The compensation unit compensates the denoised image and the denoised voiceprint based on the environmental information to obtain the compensated image and the compensated voiceprint information.

[0031] Preferably, the recognition module includes: an image recognition model construction unit, an audio recognition model construction unit, a recognition unit, and an integration unit;

[0032] The image recognition model construction unit is used to select a convolutional neural network to construct an initial image recognition model, obtain historical bird population images and perform preprocessing, and train the initial image recognition model based on the preprocessed historical bird population images to obtain the bird image recognition model;

[0033] The audio recognition model construction unit is used to select a combination of a convolutional neural network and a recurrent neural network to construct an initial audio recognition model, obtain historical bird audio information and perform preprocessing, and train the initial audio recognition model based on the preprocessed historical bird audio information to obtain the bird audio recognition model;

[0034] The recognition unit uses the bird image recognition model and the bird audio recognition model to respectively recognize the compensated image and the compensated voiceprint information, and obtains a bird population image recognition result and a bird population audio recognition result;

[0035] The integration unit is used to perform weighted fusion on the bird population image recognition result and the bird population audio recognition result, and complement each other to obtain the final recognition result.

[0036] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the intelligent detection method for bird populations in Tibetan wetland areas described above is implemented.

[0037] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed, the intelligent detection method for bird populations in Tibetan wetland areas described above is implemented.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] Through means such as multi-sensor fusion, data processing and compensation, intelligent recognition model construction, and spatial statistical analysis, the present invention realizes the intelligent detection of wetland bird populations in Tibetan areas, and has significant beneficial effects compared with the prior art. First, through multi-dimensional data collection and fusion, high-definition cameras, audio collectors, and environmental sensors are used to obtain images, voiceprints, and environmental information of the wetland, and denoising and compensation processing are performed to improve the accuracy and reliability of the data. Secondly, an image recognition model based on a convolutional neural network and an audio recognition model combining a convolutional neural network and a recurrent neural network are constructed. Through weighted fusion and mutual supplementation, the accuracy of bird recognition is significantly improved. In addition, spatial statistical methods are used to analyze the population quantity, activity range, and migration pattern of birds, providing a scientific basis for wetland protection and management. The system realizes the automation of data collection, processing, recognition, and analysis, reduces manual intervention, and improves the monitoring efficiency and continuity. Finally, this technology provides rich data support for scientific researchers, promotes wetland ecological protection and sustainable development, and has important scientific research and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention;

[0042] Figure 2 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention.

[0043] Description of the reference numerals:

[0044] 1010, processor; 1020, memory; 1030, input / output interface; 1040, communication interface; 1050, bus. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative position relationships, and when the absolute position of the object being described changes, the relative position relationship may also change accordingly.

[0047] Embodiment 1

[0048] In this embodiment, as Figure 1 shown, an intelligent detection method for wetland bird populations in Tibetan areas, the method includes:

[0049] S1. Set a number of high-definition cameras, audio collectors and a number of environmental sensors around the wetland to obtain high-definition images of the wetland, bird vocalization information and environmental information. The environmental information among them includes: light intensity, humidity and temperature.

[0050] In this embodiment, S1 includes:

[0051] S1.1. Equipment site selection and installation: S1.1.1. Conduct a detailed environmental survey around the wetland to determine the wetland's topography, vegetation distribution, bird activity hotspots and other information. Select appropriate equipment installation points through field surveys and historical data to ensure that the equipment can cover the main areas of the wetland. S1.1.2. Select a suitable installation location based on the wetland's topography and bird activity patterns. When installing a high-definition camera, consider the camera's viewing angle, height and stability. The camera should be installed in an area where birds are active, and the height should be higher than the wetland vegetation to avoid vegetation blocking the view. During installation, use a tripod or fixed bracket to ensure the stability of the camera, and adjust the camera's viewing angle to cover the largest range of wetland areas. S1.1.3. Audio collector installation: The audio collector should be installed near the bird activity area, and try to avoid interference from wind noise and man-made noise. During installation, use a windshield and shock-absorbing bracket to reduce the impact of wind noise and vibration on audio collection. The installation height of the audio collector should be moderate, usually consistent with the bird activity height, to ensure that clear bird soundprint information is collected. S1.1.4. Environmental sensor installation: Environmental sensors, including light sensors, humidity sensors, and temperature sensors, should be installed in locations that can accurately reflect the environmental conditions of the wetland. For example, light sensors should be installed in open areas to avoid being blocked by vegetation; humidity sensors and temperature sensors should be installed near the wetland surface to accurately measure the humidity and temperature changes of the wetland.

[0052] S1.2. Equipment debugging and calibration: S1.2.1. After installation, debug the HD camera to ensure that it can work properly. During the debugging process, check the resolution, frame rate and exposure settings of the camera to ensure that the image is clear and without obvious noise. Use the test image for calibration and adjust the focal length and viewing angle of the camera to enable it to accurately capture the image information of the wetland. S1.2.2. Audio collector debugging: Debug the audio collector to ensure that it can normally collect audio signals. During the debugging process, check the sampling rate, bit depth and gain settings of the audio collector to ensure that the audio signal is clear and without obvious noise. Use the test audio for calibration and adjust the gain and filter settings of the audio collector to enable it to accurately collect the soundprint information of birds. S1.2.3. Environmental sensor debugging: Debug the environmental sensor to ensure that it can accurately measure the environmental information of the wetland. During the debugging process, check the measurement range and accuracy of the sensor to ensure that it can accurately reflect the light intensity, humidity and temperature changes of the wetland. Use standard equipment for calibration to ensure that the measurement results of the sensor are accurate and reliable.

[0053] S1.3. Data Acquisition and Synchronization: S1.3.1. Start all devices and begin to collect high-definition images of the wetland, bird vocalization information, and environmental information. The high-definition camera collects image data at the set frame rate and resolution, the audio collector collects audio data at the set sampling rate and bit depth, and the environmental sensor collects environmental data at the set time interval. S1.3.2. To ensure data consistency and comparability, it is necessary to synchronize the collected data. Use timestamps to mark all data to ensure that there is a corresponding relationship between the image data, audio data, and environmental data at the same time point. For example, GPS time synchronization or the Network Time Protocol (NTP) can be used for time synchronization to ensure that the timestamps of all devices are accurate. S1.3.3. Store the collected data in a local storage device or a cloud server. Use a database management system to classify and store the data to ensure data security and accessibility. For example, a relational database can be used to store the meta-information of image data and audio data, and a distributed file system can be used to store the actual image and audio files.

[0054] S2. Denoise the high-definition images and bird vocalization information respectively, and compensate the denoised images and denoised vocalizations based on the environmental information to obtain the compensated images and compensated vocalization information.

[0055] S2 includes: performing grayscale conversion on the high-definition images to obtain grayscale images, and using median filtering to denoise the grayscale images to obtain denoised images; performing sampling rate conversion, framing, and windowing on the bird vocalization information to obtain preliminarily processed vocalization information, and then using spectral subtraction to denoise the preliminarily processed vocalization information to obtain denoised vocalizations; compensating the denoised images and denoised vocalizations based on the environmental information to obtain the compensated images and compensated vocalization information.

[0056] In this embodiment, S2 includes:

[0057] S2.1. Perform grayscale conversion on the high-definition images to obtain grayscale images, and use median filtering to denoise the grayscale images to obtain denoised images:

[0058] S2.1.1. First, preprocess the collected high-definition images, including operations such as grayscale conversion and binarization of the images, to reduce the dimension and complexity of the images while retaining the main features of the images. The grayscale conversion formula can be used to convert a color image into a grayscale image:

[0059] I gray = 0.299×I red + 0.587×I green + 0.114×I blue

[0060] where, I grayRepresents the pixel value of a grayscale image, I red 、I green and I blue respectively represent the pixel values of the red, green, and blue channels of a color image. S2.1.2. Select a suitable image denoising algorithm, such as median filtering, Gaussian filtering, wavelet transform, etc. In this embodiment, median filtering can effectively remove salt-and-pepper noise. Its basic principle is to replace the value of each pixel point in the image with the median of the neighboring pixel values of that pixel point. After filtering, the denoised image is obtained. The specific operation steps are as follows: (1) Select a window of odd size, such as 3×3, 5×5, etc.; (2) Slide the window over the image and sort the pixel values within the window; (3) Replace the value of the pixel point at the center of the window with the median of the pixel values within the window. S2.1.3. Use the peak signal-to-noise ratio (PSNR) index to evaluate the denoised image. The calculation formula of PSNR is:

[0061]

[0062] where MAX represents the maximum pixel value of the image, and MSE represents the mean square error.

[0063] S2.2. Perform sampling rate conversion, framing, and windowing on the avian voiceprint information to obtain the preliminarily processed voiceprint information, and then use spectral subtraction to denoise the preliminarily processed voiceprint information to obtain the denoised voiceprint:

[0064] S2.2.1. Audio preprocessing: Preprocess the collected audio signal, uniformly convert the sampling rate of the audio signal to 16 kHz, then divide the audio signal into several frames, with each frame having a length of 25 ms and a frame shift of 10 ms. The windowing operation can use a Hamming window, and its formula is:

[0065]

[0066] where w(n) represents the value of the window function, n represents the index of the window function, and N represents the length of the window function. S2.2.2. Select a suitable audio denoising algorithm, such as spectral subtraction, Wiener filtering, deep learning, etc.; In this embodiment, spectral subtraction is selected. Its basic principle is to first estimate the power spectrum of the noise, then subtract the noise power spectrum from the power spectrum of the noisy speech to obtain the power spectrum of the clean speech, and finally obtain the clean speech signal through inverse Fourier transform. The specific operation steps are as follows: (1) Perform Fourier transform on the noisy speech signal to obtain its spectrum; (2) Estimate the power spectrum of the noise, which can be estimated using the first few frames of the noisy speech signal; (3) Subtract the noise power spectrum from the power spectrum of the noisy speech to obtain the power spectrum of the clean speech; (4) Perform inverse Fourier transform on the power spectrum of the clean speech to obtain the denoised voiceprint. S2.2.3. Use the signal-to-noise ratio (SNR) index to evaluate the denoised audio. The calculation formula of SNR is:

[0067]

[0068] Among them, P signal represents the power of the signal, and P noise represents the power of the noise.

[0069] S2.3. Compensate the denoised image and the denoised voiceprint based on the environmental information to obtain the compensated image and the compensated voiceprint information:

[0070] S2.3.1. Compensate the denoised image according to the environmental information: If the illumination intensity is low, the image can be processed for brightness enhancement, and the Laplace operator can be used for contrast enhancement; if the humidity is high, the image can be processed for contrast enhancement. Histogram equalization method can be used for brightness enhancement, and its basic principle is to redistribute the pixel values of the image so that the distribution of pixel values is more uniform.

[0071] S2.3.2. Compensate the denoised audio according to the environmental information: If the humidity is high, the audio signal can be downsampled to reduce the influence of humidity on the audio signal; if the temperature is high, the audio signal can be high-pass filtered to remove low-frequency noise. Simple downsampling method can be used for downsampling, and Butterworth filter can be used for high-pass filtering, and its formula is:

[0072]

[0073] Among them, H(jω) represents the frequency response of the filter, ω represents the frequency, and ω c represents the medium frequency, and b represents the order of the filter.

[0074] S3. Construct a bird image recognition model and a bird audio recognition model, and use the two models to respectively recognize the compensated image and the compensated voiceprint information, and mutually verify and supplement the recognition results to obtain the final recognition result.

[0075] S3 includes: selecting a convolutional neural network to construct an initial image recognition model, obtaining historical bird population images and preprocessing them, training the initial image recognition model based on the preprocessed historical bird population images to obtain a bird image recognition model; selecting a combination of a convolutional neural network and a recurrent neural network to construct an initial audio recognition model, obtaining historical bird audio information and preprocessing it, training the initial audio recognition model based on the preprocessed historical bird audio information to obtain a bird audio recognition model; using the bird image recognition model and the bird audio recognition model to respectively recognize the compensated image and the compensated voiceprint information to obtain a bird population image recognition result and a bird population audio recognition result; performing weighted fusion on the bird population image recognition result and the bird population audio recognition result and complementing each other to obtain a final recognition result.

[0076] In this embodiment, S3 includes:

[0077] S3.1. Select a convolutional neural network to construct an initial image recognition model, obtain historical bird population images and preprocess them, and train the initial image recognition model based on the preprocessed historical bird population images to obtain a bird image recognition model:

[0078] S3.1.1 Obtain historical bird population images and preprocess the historical bird population images, including operations such as image normalization and data augmentation. Normalization can scale the pixel values of the image to the range [0, 1], and data augmentation can include operations such as rotation, flipping, and cropping to increase the generalization ability of the model. S3.1.2. Select a convolutional neural network (CNN) as the bird image recognition model. CNN has strong feature extraction capabilities and is suitable for processing image data. S3.1.3. Construct a CNN model containing multiple convolutional layers, pooling layers, and fully connected layers. In this embodiment, ResNet-50 is used as the basic model, and its structure includes: ① Convolutional layer: Extract local features of the image; ② Pooling layer: Perform downsampling on the feature map to reduce the size of the feature map; ③ Fully connected layer: Flatten the feature map and perform classification. S3.1.4. Use the preprocessed historical bird population images to train the model, and the optimization objective function is the cross-entropy loss function. S3.1.5. Use the validation set to evaluate the model, and the evaluation metrics include accuracy, recall, and F1 score, etc.

[0079] S3.2. Select a combination of a convolutional neural network and a recurrent neural network to construct an initial audio recognition model, obtain historical bird audio information and preprocess it, and train the initial audio recognition model based on the preprocessed historical bird audio information to obtain a bird audio recognition model:

[0080] S3.2.1. Obtain historical bird audio information and preprocess the historical bird audio information, including operations such as audio framing, windowing, and Fourier transform. For example, the audio signal can be divided into several frames, each frame having a length of 25 ms and a frame shift of 10 ms. Then, windowing is performed on each frame, and finally, Fourier transform is performed to obtain a spectrogram. S3.2.2. Select a combined model of a convolutional neural network (CNN) and a recurrent neural network (RNN) as the bird audio recognition model. The CNN is used to extract local features of the audio, and the RNN is used to capture the time series features of the audio. S3.2.3. Construct a CNN-RNN model containing multiple convolutional layers, pooling layers, and recurrent layers. For example, the following structure can be used: ① Convolutional layer: Extract local features of the audio; ② Pooling layer: Downsample the feature map to reduce the size of the feature map; ③ Recurrent layer: Capture the time series features of the audio. S3.2.4. Use the preprocessed historical bird audio information to train the model, and the optimization objective function is the cross-entropy loss function. S3.2.5. Use the validation set to evaluate the model, and the evaluation metrics include accuracy, recall, and F1 score, etc.

[0081] S3.3. Use the bird image recognition model and the bird audio recognition model to respectively recognize the compensated image and the compensated voiceprint information, and obtain the bird population image recognition result and the bird population audio recognition result.

[0082] S3.4. Weightedly fuse the bird population image recognition result and the bird population audio recognition result, and complement each other to obtain the final recognition result:

[0083] S3.4.1. Fuse the recognition results of the bird image recognition model and the bird audio recognition model. The weighted voting method can be used to perform weighted fusion according to the confidence levels of the two models. S3.4.2. Verify the fused result. If the prediction results of the two models are consistent, the recognition result is considered reliable; if the prediction results are inconsistent, further analysis is required, and it can be judged by combining environmental information and expert knowledge. S3.4.3. If the recognition result of a certain model is inaccurate, the result of the other model can be used for supplementation. For example, if the recognition result of the image model is inaccurate, the result of the audio model can be used for supplementation, and vice versa.

[0084] S4. Based on the final recognition result, combined with the spatial distribution of the set devices, use spatial statistics to analyze the bird population quantity, activity range, and migration pattern, and obtain the population detection result.

[0085] In this embodiment, S4 includes:

[0086] S4.1. Population quantity statistics:

[0087] S4.1.1. Based on the final recognition result, obtain the number of bird individuals detected by each device. Let the number of birds detected by the i-th device be N i , and the total number of devices be M. S4.1.2. Consider the influence of the spatial distribution of devices on the detection result. Since devices at different locations may have different detection efficiencies due to environmental differences, correction is required. Let the correction coefficient of the i-th device be C i , then the corrected number of birds is:

[0088] N′ i = N i × C i

[0089] where C i can be determined by historical data or expert experience. S4.1.3. Add up the corrected numbers of all devices to obtain the estimated total population:

[0090]

[0091] where N total represents the estimated total population.

[0092] S4.2. Activity range analysis:

[0093] S4.2.1. Obtain the spatial coordinates (x i , y i ) of each device and the position information of the detected bird individuals. Assume the position of the birds detected by each device is (x ie , y ie ), where e represents the e-th individual. S4.2.2. Use the convex hull algorithm in spatial statistics to determine the boundary of the bird activity range. The convex hull algorithm can extract the outermost boundary of all detection points to form a polygon. Let the vertices of the convex hull be (x hull , y hull ). S4.2.3. Calculate the area A of the convex hull polygon, that is, the activity range area. The formula is:

[0094]

[0095] where k represents a natural number and L represents the number of vertices of the convex hull. S4.2.4. Use the kernel density estimation method to identify the activity hot spot areas. The kernel density estimation formula is:

[0096]

[0097] where K represents the kernel function, h represents the bandwidth vegetable, t represents the number of detection points. By calculating the density value at each position, the hot spot areas with higher density can be identified.

[0098] S4.3. Migration pattern analysis:

[0099] S4.3.1. Combine the time information t e and spatial information (x ie , y ie ) of the detected bird individuals to construct time series data. Let the time series of the e-th individual be (t e , x ie , y ie ). S4.3.2. Use the trajectory smoothing algorithm to process the time series data and infer the migration path of the birds. In this embodiment, the moving average method is used to smooth the trajectory points:

[0100]

[0101] where v represents the radius of the smoothing window. S4.3.3. Use time series analysis methods (such as Fourier transform) to analyze the time pattern of bird migration. By analyzing the spectrogram, the main frequency components of the migration activities can be identified, thereby inferring the time pattern of migration.

[0102] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0103] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, it should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. The actions or steps recited in the claims can be executed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0104] Embodiment 2

[0105] An intelligent detection system for wetland bird populations in Tibetan areas, comprising: a data acquisition module, a data processing module, an identification module, and an analysis module.

[0106] The data acquisition module is used to set up a number of high-definition cameras, audio collectors and a number of environmental sensors around the wetland to obtain high-definition images, bird vocal print information and environmental information of the wetland.

[0107] The data processing module denoises the high-definition images and bird vocal print information respectively, and compensates the denoised images and denoised vocal prints based on the environmental information to obtain the compensated images and compensated vocal print information.

[0108] The data processing module includes: an image processing unit, an audio processing unit and a compensation unit. The image processing unit is used to obtain a grayscale image from the high-definition image, and perform image denoising on the grayscale image using median filtering to obtain the denoised image; the audio processing unit is used to perform sampling rate conversion, frame segmentation and windowing on the bird vocal print information to obtain the preliminarily processed vocal print information, and then use spectral subtraction to denoise the preliminarily processed vocal print information to obtain the denoised vocal print; the compensation unit compensates the denoised image and denoised vocal print based on the environmental information to obtain the compensated image and compensated vocal print information.

[0109] The recognition module is used to construct a bird image recognition model and a bird audio recognition model, and use the two models to recognize the compensated image and compensated vocal print information respectively, and mutually verify and supplement the recognition results to obtain the final recognition result.

[0110] The recognition module includes: an image recognition model construction unit, an audio recognition model construction unit, a recognition unit and an integration unit. The image recognition model construction unit is used to select a convolutional neural network to construct an initial image recognition model, obtain historical bird population images and perform preprocessing, and train the initial image recognition model based on the preprocessed historical bird population images to obtain a bird image recognition model; the audio recognition model construction unit is used to select a combination of a convolutional neural network and a recurrent neural network to construct an initial audio recognition model, obtain historical bird audio information and perform preprocessing, and train the initial audio recognition model based on the preprocessed historical bird audio information to obtain a bird audio recognition model; the recognition unit uses the bird image recognition model and the bird audio recognition model to recognize the compensated image and compensated vocal print information respectively to obtain a bird population image recognition result and a bird population audio recognition result; the integration unit is used to perform weighted fusion on the bird population image recognition result and the bird population audio recognition result and mutually supplement them to obtain the final recognition result.

[0111] Based on the final recognition result and combined with the spatial distribution of the set devices, the analysis module uses spatial statistics to analyze the bird population quantity, activity range and migration pattern to obtain a population detection result.

[0112] The system of the above embodiments is used to implement the corresponding intelligent detection method for the wetland bird population in Tibetan areas in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0113] It should be noted that the above intelligent detection system for the wetland bird population in Tibetan areas is embodied in the form of functional units. The term "module" here can be implemented in the form of software and / or hardware, and no specific limitation is made thereto.

[0114] For example, the "module" can be a software program, a hardware circuit, or a combination of both to implement the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group of processors, etc.) for executing one or more software or firmware programs, and a memory, a merged logic circuit, and / or other suitable components to support the described functions.

[0115] Embodiment III

[0116] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the intelligent detection method for the wetland bird population in Tibetan areas described in any of the above embodiments.

[0117] Figure 2 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0118] The processor 1010 can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0119] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0120] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0121] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module can implement communication in a wired manner (such as USB (Universal Serial Bus), network cable, etc.) or can implement communication in a wireless manner (such as a mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).

[0122] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0123] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary for implementing the solution of the embodiments of this specification and does not necessarily include all the components shown in the figure.

[0124] The system of the above embodiment is used to implement the corresponding intelligent detection method for Tibetan wetland bird populations in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0125] Embodiment 4

[0126] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present invention further provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the intelligent detection method for bird populations in wetlands in Tibetan areas described in any of the above embodiments.

[0127] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0128] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the intelligent detection method for bird populations in wetlands in Tibetan areas described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0129] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity.

[0130] In addition, for simplicity of explanation and discussion, and so as not to make the embodiments of the present disclosure difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that details regarding the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be entirely within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be regarded as illustrative rather than restrictive.

[0131] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0132] Therefore, the units of the examples described in the embodiments of this application can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0133] The embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. An intelligent detection method for bird populations in Tibetan wetlands, characterized in that: The method comprises: S1. Set up a number of high-definition cameras, audio collectors and environmental sensors around the wetland to obtain high-definition images of the wetland, bird soundprint information and environmental information; S2. Denoising the high-definition image and the bird voiceprint information respectively, and compensating the denoised image and the denoised voiceprint based on the environmental information to obtain a compensated image and a compensated voiceprint information; S3. Constructing a bird image recognition model and a bird audio recognition model, and using the two models to recognize the compensated image and the compensated voiceprint information respectively, and verifying and supplementing the recognition results to obtain the final recognition result; S4. Based on the final identification result, combined with the spatial distribution of the installed equipment, spatial statistics are used to analyze the bird population, activity range and migration patterns to obtain population detection results.

2. According to claim 1, a method for intelligent detection of bird populations in Tibetan wetlands is characterized by: The S2 includes: Gray-scaling the high-definition image to obtain a gray-scale image, and performing image denoising on the gray-scale image using a median filter to obtain the denoised image; Performing sampling rate conversion, framing and windowing on the bird voiceprint information to obtain initially processed voiceprint information, and then denoising the initially processed voiceprint information by using spectral subtraction to obtain the denoised voiceprint; The denoised image and the denoised voiceprint are compensated based on the environmental information to obtain the compensated image and the compensated voiceprint information.

3. According to claim 1, a method for intelligent detection of bird populations in Tibetan wetlands is characterized by: The S3 includes: Selecting a convolutional neural network to construct an initial image recognition model, obtaining and preprocessing historical bird population images, and training the initial image recognition model based on the preprocessed historical bird population images to obtain the bird image recognition model; Select a convolutional neural network and a recurrent neural network to combine and construct an initial audio recognition model, obtain historical bird audio information and perform preprocessing, and train the initial audio recognition model based on the preprocessed historical bird audio information to obtain the bird audio recognition model; Using the bird image recognition model and the bird audio recognition model to respectively recognize the compensated image and the compensated voiceprint information, to obtain a bird population image recognition result and a bird population audio recognition result; The bird population image recognition result and the bird population audio recognition result are weighted and fused, and complement each other to obtain the final recognition result.

4. An intelligent detection system for bird populations in Tibetan wetlands, the system being used to implement the method described in any one of claims 1 to 3, characterized in that: include: Data acquisition module, data processing module, identification module and analysis module; The data acquisition module is used to set up a number of high-definition cameras, audio collectors and a number of environmental sensors around the wetland to obtain high-definition images of the wetland, bird soundprint information and environmental information; The data processing module denoises the high-definition image and the bird voiceprint information respectively, and compensates the denoised image and the denoised voiceprint based on the environmental information to obtain a compensated image and a compensated voiceprint information; The recognition module is used to construct a bird image recognition model and a bird audio recognition model, and use the two models to respectively recognize the compensated image and the compensated voiceprint information, and verify and supplement the recognition results to obtain a final recognition result; The analysis module analyzes the bird population quantity, activity range and migration pattern using spatial statistics based on the final recognition result and in combination with the spatial distribution of the set equipment to obtain a population detection result.

5. According to claim 3, a method for intelligent detection of bird populations in Tibetan wetlands is characterized in that: The data processing module includes: an image processing unit, an audio processing unit and a compensation unit; The image processing unit is used to grayscale the high-definition image to obtain a grayscale image, and perform image denoising on the grayscale image using a median filter to obtain the denoised image; The audio processing unit is used to perform sampling rate conversion, frame division and windowing on the bird voiceprint information to obtain the initially processed voiceprint information, and then use spectral subtraction to denoise the initially processed voiceprint information to obtain the denoised voiceprint; The compensation unit compensates the denoised image and the denoised voiceprint based on the environmental information to obtain the compensated image and the compensated voiceprint information.

6. According to claim 3, a method for intelligent detection of bird populations in Tibetan wetlands is characterized by: The recognition module includes: an image recognition model building unit, an audio recognition model building unit, a recognition unit and an integration unit; The image recognition model construction unit is used to select a convolutional neural network to construct an initial image recognition model, obtain historical bird population images and perform preprocessing, and train the initial image recognition model based on the preprocessed historical bird population images to obtain the bird image recognition model; The audio recognition model construction unit is used to select a convolutional neural network and a recurrent neural network to combine to construct an initial audio recognition model, obtain historical bird audio information and perform preprocessing, and train the initial audio recognition model based on the preprocessed historical bird audio information to obtain the bird audio recognition model; The recognition unit recognizes the compensated image and the compensated voiceprint information respectively by using the bird image recognition model and the bird audio recognition model to obtain a bird population image recognition result and a bird population audio recognition result; The integration unit is used to perform weighted fusion on the bird population image recognition result and the bird population audio recognition result, and complement each other to obtain the final recognition result.

7. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for intelligent detection of bird populations in wetlands in the Tibetan area as claimed in any one of claims 1 to 3 is implemented.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the intelligent detection method for bird populations in Tibetan wetlands as described in any one of claims 1 to 3 is implemented.

Citation Information

Cited By

  • Bird voiceprint and vision fusion real-time identification method for oil exploitation operation area

    CN121861392A

  • Green low-carbon coastal wetland intelligent bird watching method and system

    CN122087556A