A wild bird monitoring system and method based on cloud-edge collaboration

By using bird singing detection model to screen effective audio information in the cloud-edge wild bird monitoring system, combined with the bird species identification model of the cloud platform, using environmental and historical data, the problems of large energy consumption and low recognition accuracy of bird monitoring equipment in the existing technology are solved, and longer equipment service life and higher recognition accuracy are achieved.

CN116386649BActive Publication Date: 2025-06-06BEIJING FORESTRY UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310367657.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-07
Publication Date
2025-06-06
Estimated Expiration
2043-04-07

AI Technical Summary

Technical Problem

There are two main problems in the existing bird monitoring method based on passive acoustic technology: First, the audio data obtained by the field deployed equipment contains a large amount of non-singing data, which leads to large energy consumption and affects the service life of the equipment; Second, the recognition accuracy of species recognition methods based on the acoustic characteristics of sound is low.

Method used

The field bird monitoring system based on cloud edge collaboration is adopted to obtain audio information and environmental data through the audio acquisition module and the environmental data acquisition module. The main controller uses the bird singing detection model to filter the audio information containing the bird singing and transmit it to the cloud platform. The cloud platform uses bird species identification models, combining bird singing audio information, environmental data and historical ecological data to identify bird species.

Benefits of technology

By screening effective bird singing audio information, the data transmission volume is reduced and the service life of the monitoring equipment is extended. At the same time, the accuracy of bird species identification is improved by combining environmental and historical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386649B_ABST
    Figure CN116386649B_ABST
Patent Text Reader

Abstract

The present application provides a wild bird monitoring system and method based on cloud-edge collaboration, the system comprising: an audio acquisition module for collecting audio information of the monitoring area where the target birds are located; an environmental data acquisition module for collecting current environmental data of the monitoring area; a main controller module, connected to the audio acquisition module and the environmental data acquisition module respectively, for determining whether the audio information contains bird calls through a bird call detection model, and if so, determining that the audio information is bird call information; a cloud platform for receiving the bird call information and current environmental data sent by the main controller module, and processing the bird call information, current environmental data and pre-stored ecological historical data through a bird species recognition model to obtain the species information of the target birds. The present application realizes wild bird monitoring, realizes the screening of effective audio information through a bird call detection model, and improves the accuracy of bird species recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of bird monitoring technology, and in particular to a wild bird monitoring system and method based on cloud-edge collaboration. Background Art

[0002] Bird communities are an important part of ecosystems and are indicator species for evaluating the health of ecosystems. The survey and monitoring of bird species is an important means to understand the composition of ecosystem biological communities and their health status.

[0003] Traditional bird monitoring methods mainly rely on long-term field work by ecological conservation workers, which is not only time-consuming and laborious, but also the bird information collected is very limited. In recent years, with the development of passive acoustic technology, the deployment of acoustic collection equipment in the wild to record bird call information, and the use of deep learning methods to automatically identify bird species based on the information characteristics contained in the calls, thereby achieving bird monitoring, has begun to receive more and more attention.

[0004] There are currently two problems with the bird monitoring method based on passive acoustic technology: first, the audio data obtained by deploying acoustic collection equipment in the wild contains a large amount of non-sound data, which consumes a lot of energy during transmission and affects the service life of the equipment in the wild; second, the current species identification method based on bird calls is only based on the acoustic characteristics contained in the calls, which has a bottleneck in recognition accuracy, resulting in a low accuracy rate in bird species identification. Summary of the invention

[0005] The purpose of the embodiment of the present application is to provide a wild bird monitoring system and method based on cloud-edge collaboration to solve the problem of low accuracy in bird species identification. The specific technical solution is as follows:

[0006] In a first aspect, a wild bird monitoring system based on cloud-edge collaboration is provided, the system comprising:

[0007] An audio collection module, used to collect audio information of a monitoring area where a target bird is located, wherein the target bird is a bird to be identified as a bird species;

[0008] An environmental data collection module, used to collect current environmental data of the monitoring area;

[0009] a main controller module, connected to the audio acquisition module and the environmental data acquisition module respectively, and used to determine whether the audio information contains bird sounds through a bird sound detection model, and if so, determine that the audio information is bird sound information;

[0010] The cloud platform is used to receive the bird singing information and the current environmental data sent by the main controller module, and process the bird singing information, the current environmental data and the pre-stored ecological history data through a bird species recognition model to obtain the species information of the target bird.

[0011] In a second aspect, a method for monitoring wild birds based on cloud-edge collaboration is provided, which is applied to a cloud platform, and the method includes:

[0012] The main controller obtains bird singing audio information and current environmental data of the monitoring area where the target bird is located, wherein the main controller determines whether the received audio information contains bird singing through a bird singing detection model, and if so, determines that the audio information is bird singing information, and the target bird is a bird to be identified as a bird species;

[0013] The pre-stored historical ecological data, the bird song audio information and the current environmental data are input into a bird species recognition model to obtain the species information of the target bird output by the bird species recognition model.

[0014] Optionally, the bird species recognition model includes a bird song recognition network and an ecological niche information network, and the target bird includes at least one bird; the pre-stored historical ecological data, the bird song audio information and the current environmental data are input into the bird species recognition model, and the target bird species information output by the bird species recognition model is obtained, including:

[0015] Processing the bird call audio information through the bird call recognition network to obtain a bird species prediction vector for each bird;

[0016] According to the historical ecological data and the current environmental data, obtaining a bird species existence prior vector of each bird through the ecological niche information network, wherein the bird species existence prior vector indicates the suitability of the bird to survive in the monitoring area;

[0017] Multiply the bird species prediction vector corresponding to each bird species and the bird species existence prior vector to obtain the candidate species prediction value corresponding to each bird species;

[0018] A target species prediction value greater than a prediction value threshold is selected from the candidate species prediction values, and the bird species corresponding to the target species prediction value is used as the species information of the target bird.

[0019] Optionally, the processing of the bird call audio information by the bird call recognition network to obtain a bird species prediction vector of each bird includes:

[0020] Generate a bird song spectrum diagram according to the bird song audio information;

[0021] The bird song spectrum graph is input into the bird song recognition network to obtain the bird species prediction vector of each bird output by the bird song recognition network.

[0022] Optionally, inputting the bird song spectrum graph into the bird song recognition network to obtain the bird species prediction vector of each bird output by the bird song recognition network comprises:

[0023] Input the bird song spectrum graph into a 1*1 convolutional layer and a 3*3 depth-separable convolutional layer in sequence to obtain a feature graph vector;

[0024] Inputting the feature graph vector into a plurality of convolutional layers of different sizes respectively, obtaining the first bird song region feature output by each convolutional layer, wherein the duration and frequency range of the bird song in the first bird song region features output by different convolutional layers are different;

[0025] Inputting the first bird song regional features into the deconvolution layers corresponding to the convolution layers respectively, to obtain the second bird song regional features output by each deconvolution layer, wherein the deconvolution layer is used to remodel the global information of the bird song spectrum graph;

[0026] After vector addition of each of the second bird song region features, the added vectors are scaled through an activation function layer to obtain an attention vector;

[0027] The attention vector and the feature map vector are multiplied and then passed through a 1*1 convolution layer to obtain a bird species prediction vector.

[0028] Optionally, according to the historical ecological data and the current environmental data, obtaining the bird species existence prior vector of each bird through the ecological niche information network includes:

[0029] Inputting the current environmental data and the historical ecological data into a species distribution prediction model to obtain a priori probabilities of different birds existing in the monitoring area output by the species distribution prediction model;

[0030] The prior probabilities of different birds are input into the ecological niche information network to obtain the bird species existence prior vector of each bird output by the ecological niche information network.

[0031] In a third aspect, a method for monitoring wild birds based on cloud-edge collaboration is provided, which is applied to a main controller, and the method includes:

[0032] Acquire audio information in the monitoring area, and input the audio information into a bird song detection model, wherein the bird song detection model includes a low-level feature extraction block, a high-level feature extraction block and a classifier;

[0033] Processing the audio information using the low-level feature extraction block to extract low-level features from the audio information;

[0034] Processing the low-level features using the high-level feature extraction block to extract high-level features from the audio information;

[0035] Inputting the high-level features into the classifier to obtain a probability value that the audio information contains bird calls;

[0036] If the probability value is greater than a preset probability threshold, it is determined that the audio information contains bird singing, and the audio information containing the bird singing is used as the bird singing audio information of the target bird.

[0037] Optionally, the low-level feature extraction block includes two one-dimensional convolutional layers connected in sequence, and the one-dimensional convolutional layer is used to extract low-level features in the audio information.

[0038] Optionally, the high-level feature extraction block includes:

[0039] The first 1*1 convolutional layer is used to increase the number of input channels and improve the feature dimension;

[0040] A 3*3 depth-separable convolutional layer is connected to the output end of the 1*1 first convolutional layer to decouple the spatial dimension and channel dimension of the input features;

[0041] An attention module, connected to the output end of the 3*3 depth-separable convolutional layer, is used to reduce information loss caused by dimensionality reduction;

[0042] A 1*1 second convolutional layer is connected to the output end of the attention module, and the 1*1 second convolutional layer is used to reduce the number of channels.

[0043] In a fourth aspect, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, any of the steps of the method for identifying bird species is implemented.

[0044] Beneficial effects of the embodiments of the present application:

[0045] The embodiment of the present application provides a wild bird monitoring system based on cloud-edge collaboration. The present application uses a bird song detection model to screen bird song audio information, and uses audio information containing bird songs as valid audio information. This can reduce the amount of data transmitted to the cloud platform, avoid energy loss caused by invalid data transmission, and extend the service life of monitoring equipment in the wild. Bird species are identified through a bird species recognition model. Bird species are identified not only using bird song audio information, but also using current environmental data and historical ecological data. Adding current environmental data and historical ecological data can determine the probability of birds surviving in the monitoring area. The present application can improve the accuracy of bird species recognition by adding current environmental data and historical ecological data.

[0046] Of course, implementing any product or method of the present application does not necessarily require achieving all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0048] Figure 1 A flow chart of a method for monitoring wild birds based on cloud-edge collaboration provided in an embodiment of the present application;

[0049] Figure 2 A flow chart of a method for obtaining bird song audio information provided in an embodiment of the present application;

[0050] Figure 3 A schematic diagram of the structure of a bird sound detection model provided in an embodiment of the present application;

[0051] Figure 4 A schematic diagram of the processing flow of the bird species identification model provided in an embodiment of the present application;

[0052] Figure 5 An enlarged schematic diagram of the inverted residual block provided in an embodiment of the present application;

[0053] Figure 6 A schematic diagram of the structure of a bird species identification system provided in an embodiment of the present application;

[0054] Figure 7 A schematic diagram of a process for determining bird song audio information provided in an embodiment of the present application;

[0055] Figure 8 A schematic diagram of the structure of a wild bird monitoring device based on cloud-edge collaboration provided in an embodiment of the present application;

[0056] Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0058] In the subsequent description, the suffixes such as "module", "component" or "unit" used to represent elements are only used to facilitate the description of the present application and have no specific meaning. Therefore, "module" and "component" can be used interchangeably.

[0059] In order to solve the problems mentioned in the background technology, according to one aspect of an embodiment of the present application, an embodiment of a wild bird monitoring method based on cloud-edge collaboration is provided, which can apply a server and a main controller to improve the accuracy of bird species identification.

[0060] The following will describe in detail a method for monitoring wild birds based on cloud-edge collaboration provided by an embodiment of the present application in combination with a specific implementation method. Figure 1 As shown, the specific steps are as follows:

[0061] Step 101: Obtain bird call audio information and current environmental data of the monitoring area where the target bird is located through the main controller.

[0062] The main controller determines whether the received audio information contains bird sounds through a bird sound detection model, and if so, determines that the audio information is bird sound information. The target bird is a bird to be identified as a bird species.

[0063] In an embodiment of the present application, in a wild environment where birds live, an area is divided as a monitoring area, and the birds to be identified as bird species in the monitoring area are target birds. The main controller obtains audio information of the monitoring area through an audio acquisition device, and then determines whether the received audio information contains bird sounds through a bird sound detection model. If it does, it determines that the audio information is bird sound audio information. The main controller also collects current environmental data of the monitoring area through an environmental information acquisition device, and then sends the bird sound audio information and current environmental data to the server. The server also pre-stores historical ecological data of the monitoring area.

[0064] Among them, the current environmental data include the current temperature, humidity, longitude and latitude, and light intensity of the monitoring area; the historical ecological data include the historical annual average temperature, historical annual precipitation, historical daily temperature difference, historical annual temperature difference, and historical temperature seasonal variation coefficient of the monitoring area.

[0065] Step 102: inputting the pre-stored historical ecological data, bird song audio information and current environmental data into a bird species recognition model to obtain the species information of the target bird output by the bird species recognition model.

[0066] The server inputs the bird song audio information, current environmental data and historical ecological data into a bird species recognition model, and the bird species recognition model outputs the species information of the target bird.

[0067] The present application uses a bird song detection model to screen bird song audio information, and uses audio information containing bird songs as valid audio information. This can reduce the amount of data transmitted to the cloud platform, avoid energy loss caused by invalid data transmission, and extend the service life of monitoring equipment in the wild; bird species are identified through a bird species recognition model. The identification of bird species not only uses bird song audio information, but also uses current environmental data and historical ecological data. Adding current environmental data and historical ecological data can determine the probability of birds surviving in the monitoring area. The present application can improve the accuracy of bird species recognition by adding current environmental data and historical ecological data.

[0068] As an optional implementation, Figure 2 As shown, the main controller obtains the bird call audio information of the monitoring area where the target bird is located, including:

[0069] Step 201: Acquire audio information in the monitoring area and input the audio information into a bird sound detection model.

[0070] An audio acquisition device is set up in the monitoring area to obtain audio information of the monitoring area. The audio information includes bird singing data and environmental sound data. The main controller inputs the audio information into the bird singing detection model. The audio information can be in .wav format. Before the bird singing detection model inputs the audio information, it only needs to perform a simple segmentation operation, without the need for pre-processing operations such as framing and windowing of the audio information, thereby improving the detection efficiency of the bird singing audio information.

[0071] Step 202: Use a low-level feature extraction block to process the audio information to extract low-level features in the audio information.

[0072] The bird song detection model includes a low-level feature extraction block, a high-level feature extraction block and a classifier. Figure 3 Schematic diagram of the bird song detection model.

[0073] The low-level feature extraction block includes two one-dimensional convolutional layers connected in sequence, a Maxpooling layer and a transposition layer, wherein each one-dimensional convolutional layer is connected to a BatchNormalization layer (batch normalization BN layer) and a ReLU activation function layer, and the one-dimensional convolutional layer is used to extract low-level features of the audio. Among them, the step size of the one-dimensional convolutional layer can be 2 steps or 3 steps, and this application does not impose specific restrictions on the step size. The Maxpooling layer is used to reduce the dimension of the features and remove redundant information; the transposition layer is used to perform a transpose operation on the feature vector to obtain the low-level features of the audio.

[0074] The one-dimensional convolutional layer can adaptively extract better discriminative features, avoiding the limitations of using only MFCC (Mel-Frequency Cepstral Coefficients, a feature widely used in automatic speech and speaker recognition) or Logmel. It also reduces the complexity of software design. There is no need to write MFCC or Logmel calculation programs according to different devices, which facilitates model deployment on different hardware device platforms and facilitates platform porting.

[0075] The BatchNormalization layer is used to obtain the mean and variance of the data, thereby standardizing the input data to accelerate convergence during training and prevent overfitting. The ReLU activation function layer is used to introduce nonlinear representation and enhance the representation ability of the model.

[0076] Step 203: Use a high-level feature extraction block to process the low-level features to extract high-level features from the audio information.

[0077] The main controller inputs the extracted low-level features into the high-level feature extraction block. The high-level feature extraction block includes multiple 3*3 depth separable convolutional layers, where each 3*3 depth separable convolutional layer is connected to a 1*1 convolutional layer to increase the number of input channels and improve the feature dimension. The 3*3 depth separable convolutional layer is used to decouple the spatial dimension and channel dimension of the input feature, and reduce the number of parameters required for calculation while extracting feature information, thereby improving calculation efficiency. An attention module is added after each 3*3 depth separable convolutional layer to reduce information loss caused by dimensionality reduction. A 1*1 convolutional layer and a ReLU6 activation function layer are connected after the attention module. The 1*1 convolutional layer is used to reduce the number of channels and further reduce the amount of calculation. Optionally, a residual connection is added between the two 1*1 convolutional layers to prevent gradient disappearance during training.

[0078] Exemplarily, the attention module may be an ESE (Effective Squeeze and Extraction) attention module, and the present application does not impose any specific limitation on the attention module.

[0079] Step 204: input the high-level features into the classifier to obtain a probability value of the audio information containing bird calls.

[0080] The classifier consists of a fully connected layer and a softmax layer. The fully connected layer is used to integrate low-level features and high-level features, map the feature information to the category space, and achieve classification; the softmax layer is used to map the output of the fully connected layer to between (0, 1) and generate a probability value for the audio information containing bird song data.

[0081] Step 205: If the probability value is greater than a preset probability threshold, it is determined that the audio information contains bird calls.

[0082] The main controller compares the probability value with the preset probability threshold. If the probability value is less than the preset probability threshold, it is determined that the audio information does not contain bird singing data, the collected audio information is invalid, and the audio information is deleted; if the probability value is greater than the preset probability threshold, it is determined that the audio information contains bird singing data, the collected audio information is valid, and the audio information is saved.

[0083] Step 206: The audio information containing the bird's singing is used as the bird's singing audio information of the target bird.

[0084] The main controller uses the audio information containing the bird's singing as the bird's singing audio information of the target bird.

[0085] This application determines that the collected audio information is valid only after it determines that the audio information contains bird singing audio information, and regards the audio information containing bird singing as valid audio information. This can reduce the amount of data transmitted to the cloud platform, avoid energy loss caused by invalid data transmission, extend the service life of monitoring equipment in the wild, and facilitate long-term monitoring of birds in the wild.

[0086] As an optional implementation, the bird species identification model includes a bird song identification network and an ecological niche information network, and the target bird includes at least one bird species; the cloud platform obtains the species information of the target bird in the following manner:

[0087] The bird song audio information, current environmental data and historical ecological data are input into the bird species recognition model to obtain the species information of the target bird output by the bird species recognition model, including: processing the bird song audio information through a bird song recognition network to obtain a bird species prediction vector for each bird; obtaining a bird species existence prior vector for each bird through an ecological niche information network based on historical ecological data and current environmental data, wherein the bird species existence prior vector indicates the suitability of the bird to survive in the monitored area; multiplying the bird species prediction vector corresponding to each bird with the bird species existence prior vector to obtain a candidate species prediction value corresponding to each bird; selecting a target species prediction value greater than a prediction value threshold from the candidate species prediction values, and using the bird species corresponding to the target species prediction value as the species information of the target bird.

[0088] Figure 4 Schematic diagram of the processing flow of the bird species identification model. Figure 4 The following steps can be identified.

[0089] First, the server generates a bird song spectrum graph based on the bird song audio information. The bird song spectrum graph can be implemented through the existing Python code. Then, the bird song spectrum graph is input into the bird song recognition network to obtain the bird species prediction vector of each bird output by the bird song recognition network.

[0090] Optionally, inputting the bird song spectrum graph into the bird song recognition network to obtain the bird species prediction vector of each bird output by the bird song recognition network includes: inputting the bird song spectrum graph into a 1*1 convolution layer and a 3*3 depth-separable convolution layer in sequence to obtain a feature graph vector; inputting the feature graph vector into a plurality of convolution layers of different sizes respectively to obtain a first bird song region feature output by each convolution layer, wherein the duration and frequency range of the bird song are different in the first bird song region features output by different convolution layers; inputting the first bird song region feature into the deconvolution layer corresponding to each convolution layer respectively to obtain a second bird song region feature output by each deconvolution layer, wherein the deconvolution layer is used to remodel the global information of the feature graph; after vector addition of each second bird song region feature, scaling the added vector through an activation function layer to obtain an attention vector; multiplying the attention vector and the feature graph vector, and then passing through a 1*1 convolution layer to obtain a bird species prediction vector.

[0091] Figure 5The enlarged schematic diagram of the inverted residual block. The bird song recognition network consists of multiple sequentially connected inverted residual blocks, each of which includes two 1*1 convolutional layers, one 3*3 depth-separable convolutional layer, and one multi-head convolutional attention block. The multi-head convolutional attention module is used to focus on the difference feature information of the song contained in the spectrum graph, including three convolutional layers, three deconvolutional layers corresponding to the convolutional layers, and one sigmoid activation function layer.

[0092] In the embodiment of the present application, the bird song spectrum graph is sequentially input into a 1*1 convolutional layer and a 3*3 depth-separable convolutional layer to obtain a feature graph vector. The input of the multi-head convolutional attention module is a feature graph vector. The three attention heads in the multi-head convolutional attention module are a 1*1 convolutional layer, a 3*3 convolutional layer, a 5*5 convolutional layer, and a corresponding deconvolutional layer (the number of convolution kernels is the quotient of the number of feature graph channels and the reduction rate, where the reduction rate is used to characterize the degree of loss of cross-channel information during the convolution process. The smaller the reduction rate, the less cross-channel information is lost during the convolution process, but more convolution kernels are needed to extract cross-channel information, which can easily cause overfitting). Exemplarily, the reduction rate can be 0.5. Convolutional layers of different sizes have different receptive fields, which are conducive to extracting the regional features of bird songs with different durations and frequency ranges in the bird song spectrum. The deconvolution layer is used to upsample the output of the convolution layer to achieve remodeling of the global information of the feature map. The dimension of the deconvolution layer output is consistent with the dimension of the feature map. Then the vectors obtained by all attention heads are added and passed through the sigmoid activation function layer to scale the vector element values ​​to between (0, 1) to obtain the final attention vector. By multiplying the attention vector with the feature map vector, the bird song area in the spectrum feature map is focused on, thereby achieving attention to the difference information of different categories of songs and distinguishing the differences in songs of different birds.

[0093] At the same time, the server inputs the current environmental data and historical ecological data into the species distribution prediction model to obtain the prior probability of different birds existing in the monitoring area output by the species distribution prediction model, and then inputs the prior probability of different birds into the niche information network to obtain the bird species existence prior vector of each bird output by the niche information network. Among them, the prior probability of different birds existing in the monitoring area refers to the probability that different birds may exist in the monitoring area.

[0094] The species distribution prediction model may be a MaxEnt model, or may be a Bioclim or Domain model, etc. This application does not impose any specific restrictions on the type of species distribution prediction model.

[0095] Secondly, the server multiplies the bird species prediction vector corresponding to each bird species and the bird species existence prior vector to obtain the candidate species prediction value corresponding to each bird species, and then selects the target species prediction value greater than the prediction value threshold from the candidate species prediction values, and uses the bird species corresponding to the target species prediction value as the species information of the target bird.

[0096] Finally, the server compares each candidate species prediction value with the prediction value threshold. If the candidate species prediction value is greater than or equal to the prediction value threshold, indicating that the prediction is accurate, the bird species corresponding to the candidate species prediction value is used as the identified species information. Preferably, if the number of identified species information is greater than the set number threshold, the top n species information is selected in descending order of number as the identified target bird species information, where n is a positive integer greater than 1.

[0097] If the predicted value of the candidate species is less than the predicted value threshold, it is considered that a rare bird may appear, and an alarm is issued to prompt staff to verify.

[0098] In this application, by adding a priori vectors of bird species existence, the impact of current environmental data and historical ecological data on the survival probability of birds can be added, which can avoid the problem of identification errors that easily occur when the calls of two birds are similar. This application combines bird species predictions and the survival suitability of bird species in the monitoring area to improve the accuracy of bird species identification.

[0099] Exemplarily, the prediction value of the candidate species is obtained by vector dot product. For example, the bird species prediction vectors corresponding to bird species A, B, and C obtained by the bird song recognition network are [0.4, 0.4, 0.2] respectively. At this time, the network cannot determine whether it is bird A or bird B because the songs of the two birds A and B are similar. The bird species existence prior vectors corresponding to bird species A, B, and C obtained by the niche network are [0.8, 0.3, 0.1] respectively. Among them, the monitoring area is most suitable for bird A, and the candidate species prediction values ​​obtained are [0.32, 0.12, 0.02] respectively. Finally, it is considered that the bird to be identified is bird A.

[0100] Optionally, the embodiment of the present application also provides a schematic diagram of a system for bird species identification, such as Figure 6As shown in the figure, the system consists of edge devices and cloud platforms. Among them, the edge devices include: audio acquisition module, environmental information acquisition module, main controller module, data transmission module and power module. The audio acquisition module and the environmental information acquisition module can realize the long-term automatic acquisition of bird calls and current environmental data (including temperature, humidity, light intensity, longitude and latitude, etc.) in the monitoring area; the main controller module is connected to the audio acquisition module and the environmental data acquisition module respectively. After the data acquisition is completed, the bird call detection model in the main controller module is used to screen the bird call segments, and then the bird call audio information and current environmental data of the bird call segments are uploaded to the cloud platform through the data transmission module; the cloud platform pre-stores historical ecological data, and the cloud platform uses a bird species recognition model that integrates acoustic information and ecological niche information to identify bird species according to bird call audio information, current environmental data and historical ecological data. After the species recognition is completed, the recognition results and environmental information are visualized, and a bird situation database is established for easy query by staff.

[0101] Specifically, the audio collection module includes multiple electret microphones, which are used to collect bird calls and environmental sound data in the monitoring target area from multiple directions.

[0102] The environmental information acquisition module includes a temperature sensor, a humidity sensor, a longitude and latitude sensor, and a light intensity sensor, which are used to obtain current environmental data such as temperature, humidity, longitude and latitude, and light intensity in the monitoring area.

[0103] The main controller module uses a microcontroller with a Cortex-M7 core to control edge devices and other modules. It is also used to process the data obtained by the audio acquisition module and the environmental information acquisition module, and use the bird song detection model to detect whether the audio information contains bird song fragments.

[0104] The data transmission module uses 4G or 5G communication according to the actual situation of the monitoring area, and uploads the bird song audio information including the singing clips and current environmental data such as temperature, humidity, longitude and latitude, and light intensity to the cloud platform.

[0105] The power module adopts solar energy + lithium battery power supply mode. When the weather is fine, solar energy is used to directly power the device and charge the lithium battery at the same time; in rainy weather, large-capacity lithium batteries are used to power the device.

[0106] The cloud platform uses cloud servers to receive and store bird song audio information and current environmental data uploaded by edge devices, and uses a bird species recognition model that integrates acoustic information and ecological niche information to identify bird species.

[0107] This application uses low-cost devices for hardware implementation, which facilitates large-scale field deployment and reduces monitoring costs.

[0108] The technical solution of the present invention is further specifically described below through embodiments and in conjunction with the accompanying drawings.

[0109] A bird species identification system, which consists of edge devices and a cloud platform. The edge devices include: an audio acquisition module, an environmental information acquisition module, a main controller module, a data transmission module, and a power supply module. The overall structure of the system is as follows: Figure 6 The specific description is as follows:

[0110] After the edge device is installed, the power module is turned on and the monitoring device is powered by solar energy + lithium battery. To ensure the normal operation of the device, solar energy is used to directly power the device and charge the lithium battery when the weather is fine; large-capacity lithium batteries are used to power the device in rainy weather. To ensure a smooth charging process, the CN3791 chip is selected for solar charging management in actual applications.

[0111] The audio acquisition module collects the bird sounds and environmental sounds in the monitoring area according to the pre-set sampling interval and sampling frequency. In actual applications, the sampling interval can be 1min sampling 30s, and the sampling frequency is 44100Hz. Considering that the position of birds in the monitoring area is difficult to determine in advance, the audio acquisition module uses multiple electret microphones to collect bird sounds and environmental sounds from multiple directions. While the audio is sampling, the environmental information acquisition module is started to collect environmental information such as temperature, humidity, longitude and latitude, and light intensity in the monitoring area. Considering the overall power consumption of the monitoring equipment, in actual applications, the temperature sensor uses the DS18B20 sensor, the humidity sensor uses the DHT11 sensor, the longitude and latitude sensor uses the WT-NEO6M sensor, and the light intensity sensor uses the BH1750FVI sensor. Among them, this application only exemplifies the sensors and does not make specific restrictions.

[0112] After one sampling is completed, the main controller module starts to process the audio information collected by the audio acquisition module and the current environmental data collected by the environmental information acquisition module, and uses the bird singing detection model to detect whether the audio data contains bird singing fragments. Preferably, a lightweight bird singing detection model is used. The lightweight detection model has a small amount of computation and is more suitable for edge devices. Figure 7 A flow chart for determining the audio information of bird calls, such as Figure 7 As shown, it mainly includes the following steps:

[0113] S1, segment the audio data and input it into a lightweight bird song detection model;

[0114] S2, the low-level feature extraction block of the lightweight bird sound detection model extracts low-level features of the audio information, generates low-level features that replace audio spectrum features, and then inputs them into the high-level feature extraction block to extract high-level features;

[0115] S3, the high-level feature extraction block of the lightweight bird song detection model extracts high-level features of the audio information, generates audio embeddings representing high-level features of the audio, and then inputs them into the classifier for classification;

[0116] S4, the classifier outputs a probability value that the audio information contains a bird song segment;

[0117] S5. Compare the probability value with a preset probability threshold. If the probability value is less than the preset probability threshold, it is determined that the audio information does not contain bird sound data, the collected audio information is invalid, and the audio information is deleted. If the probability value is greater than the preset probability threshold, it is determined that the audio information contains bird sound data, the collected audio information is valid, and the audio information is saved.

[0118] After the detection is completed, if the collected audio information is valid, the data transmission module will be turned on to upload the bird singing information and current environmental data (temperature, humidity, longitude and latitude, light intensity, etc.) to the cloud platform via 4G or 5G.

[0119] The cloud platform uses a cloud server to receive and store audio information and current environmental data uploaded by edge devices. It combines pre-stored historical ecological data and uses a bird species recognition model that integrates acoustic information and ecological niche information to identify bird species. Figure 7 Processing flow chart for the bird species identification model, such as Figure 7 As shown, it mainly includes the following steps:

[0120] S1. Input the current environmental data of the monitoring area uploaded by the edge device and the historical ecological data pre-stored in the cloud platform into the MaxEnt software to obtain the prior probability of the existence of different birds in the monitoring area;

[0121] S2, inputting the prior probability of the existence of different bird species in the monitoring area calculated by MaxEnt software into the ecological niche information network to obtain the bird species existence prior vector of different bird species based on the ecological niche information;

[0122] S3, inputting the bird call audio information uploaded by the edge device into the bird call recognition network to obtain bird species prediction vectors of different birds based on acoustic information;

[0123] S4. Multiply the bird species prediction vector by the bird species existence prior vector to obtain the candidate species prediction value of each bird, compare the maximum species prediction value with the prediction value threshold, and if the maximum species prediction value is greater than or equal to the prediction value threshold, the bird species corresponding to the maximum species prediction value is used as the identified species information; if the maximum species prediction value is less than the prediction value threshold, it is considered that a rare bird has appeared, and an alarm is issued to prompt the staff to verify.

[0124] in, Figure 7 Step S3 and step S1 in the embodiment can be parallel steps.

[0125] After species identification is completed, the identification results and environmental information will be visualized, and a bird information database will be established to facilitate staff inquiries.

[0126] Based on the same technical concept, the embodiment of the present application also provides a wild bird monitoring device based on cloud-edge collaboration, which is applied to a cloud platform, such as Figure 8 As shown, the device comprises:

[0127] The acquisition module 801 is used to acquire the bird singing audio information and current environment data of the monitoring area where the target bird is located through the main controller, wherein the main controller determines whether the received audio information contains bird singing through the bird singing detection model, and if so, determines that the audio information is bird singing information, and the target bird is a bird to be identified as a bird species;

[0128] The input-output module 802 is used to input the pre-stored historical ecological data, bird song audio information and current environmental data into the bird species recognition model to obtain the target bird species information output by the bird species recognition model.

[0129] Optionally, the bird species recognition model includes a bird song recognition network and an ecological niche information network, and the target bird includes at least one bird; the input and output module 802 includes:

[0130] A first processing unit is used to process the bird call audio information through a bird call recognition network to obtain a bird species prediction vector of each bird;

[0131] A second processing unit is used to obtain a bird species existence prior vector of each bird through an ecological niche information network according to historical ecological data and current environmental data, wherein the bird species existence prior vector indicates the fitness of the bird to survive in the monitoring area;

[0132] A calculation unit, used for multiplying the bird species prediction vector corresponding to each bird species and the bird species existence prior vector to obtain a candidate species prediction value corresponding to each bird species;

[0133] The selection unit is used to select a target species prediction value greater than a prediction value threshold from the candidate species prediction values, and use the bird species corresponding to the target species prediction value as the species information of the target bird.

[0134] Optionally, the first processing unit is used for:

[0135] Generate a bird song spectrum graph based on the bird song audio information;

[0136] The bird song spectrum graph is input into the bird song recognition network to obtain the bird species prediction vector of each bird output by the bird song recognition network.

[0137] Optionally, the first processing unit is used for:

[0138] The bird song spectrum graph is input into the 1*1 convolution layer and the 3*3 depth-separable convolution layer in sequence to obtain the feature graph vector;

[0139] Inputting the feature graph vectors into a plurality of convolutional layers of different sizes respectively, obtaining the first bird song region features output by each convolutional layer, wherein the duration and frequency range of the bird song in the first bird song region features output by different convolutional layers are different;

[0140] The first bird song regional features are input into the deconvolution layers corresponding to each convolution layer, respectively, to obtain the second bird song regional features output by each deconvolution layer, wherein the deconvolution layer is used to remodel the global information of the bird song spectrum graph;

[0141] After adding the vectors of the features of each second bird song region, the added vectors are scaled through the activation function layer to obtain the attention vector.

[0142] The attention vector and the feature map vector are multiplied, and then passed through a 1*1 convolution layer to obtain the bird species prediction vector.

[0143] Optionally, the second processing unit is used for:

[0144] Input current environmental data and historical ecological data into the species distribution prediction model to obtain the prior probability of different bird species existing in the monitoring area output by the species distribution prediction model;

[0145] The prior probabilities of different bird species are input into the niche information network, and the bird species existence prior vectors of each bird species output by the niche information network are obtained.

[0146] A wild bird monitoring device based on cloud-edge collaboration is applied to a main controller and is used to:

[0147] Acquire audio information in the monitoring area, and input the audio information into a bird song detection model, wherein the bird song detection model includes a low-level feature extraction block, a high-level feature extraction block, and a classifier;

[0148] The audio information is processed using a low-level feature extraction block to extract low-level features from the audio information;

[0149] The high-level feature extraction block is used to process the low-level features and extract the high-level features in the audio information;

[0150] The high-level features are input into the classifier to obtain the probability value of the audio information containing bird calls;

[0151] If the probability value is greater than a preset probability threshold, it is determined that the audio information contains bird singing, and the audio information containing the bird singing is used as the bird singing audio information of the target bird.

[0152] Optionally, the low-level feature extraction block includes two sequentially connected one-dimensional convolutional layers, and the one-dimensional convolutional layers are used to extract low-level features in the audio information.

[0153] Optionally, the advanced feature extraction block includes:

[0154] The first 1*1 convolutional layer is used to increase the number of input channels and improve the feature dimension;

[0155] A 3*3 depth-wise separable convolutional layer is connected to the output of the first 1*1 convolutional layer to decouple the spatial and channel dimensions of the input features.

[0156] The attention module is connected to the output of the 3*3 depth-separable convolutional layer to reduce the information loss caused by dimensionality reduction;

[0157] The 1*1 second convolutional layer is connected to the output of the attention module, and the 1*1 second convolutional layer is used to reduce the number of channels.

[0158] According to another aspect of the embodiments of the present application, the present application provides an electronic device, such as Fig. 9 As shown, it includes a memory 903, a processor 901, a communication interface 902 and a communication bus 904. The memory 903 stores a computer program that can be run on the processor 901. The memory 903 and the processor 901 communicate through the communication interface 902 and the communication bus 904. When the processor 901 executes the computer program, the steps of the above method are implemented.

[0159] The memory and processor in the above electronic device communicate through a communication bus and a communication interface. The communication bus can be a peripheral component interconnect standard (PCI) bus, a serial peripheral interface (SPI) bus or an integrated circuit bus (IIC) bus. The communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0160] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Alternatively, the memory may also be at least one storage device located away from the aforementioned processor.

[0161] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a microcontroller unit (MCU), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0162] According to another aspect of the embodiments of the present application, a computer-readable medium having a non-volatile program code executable by a processor is provided.

[0163] Optionally, in an embodiment of the present application, a computer-readable medium is configured to store program code for the processor to execute the above method.

[0164] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and this embodiment will not be described in detail here.

[0165] When the embodiments of the present application are specifically implemented, reference may be made to the above-mentioned embodiments, which have corresponding technical effects.

[0166] It is understood that the embodiments described herein can be implemented by hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPDevice, DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in the present application or a combination thereof.

[0167] For software implementation, the technology described herein can be implemented by a unit that performs the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0168] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0169] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0170] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0171] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0172] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0173] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk. It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such a process, method, article or device. Without more constraints, an element defined by the phrase "comprising a..." does not exclude the existence of other identical elements in the process, method, article or apparatus comprising the element.

[0174] The above description is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest range consistent with the principles and novel features applied for herein.

Claims

1. A wild bird monitoring system based on cloud-edge collaboration, It is characterized in that The system comprises: An audio collection module, used to collect audio information of a monitoring area where a target bird is located, wherein the target bird is a bird to be identified as a bird species; An environmental data collection module, used to collect current environmental data of the monitoring area; a main controller module, connected to the audio acquisition module and the environmental data acquisition module respectively, and used to determine whether the audio information contains bird sounds through a bird sound detection model, and if so, determine that the audio information is bird sound information; The cloud platform is used to receive the bird singing information and the current environmental data sent by the main controller module, and process the bird singing information, the current environmental data and the pre-stored ecological history data through a bird species recognition model to obtain the species information of the target bird.

2. A method for monitoring wild birds based on cloud-edge collaboration. It is characterized in that Applied to a cloud platform, the method includes: The main controller obtains bird singing audio information and current environmental data of the monitoring area where the target bird is located, wherein the main controller determines whether the received audio information contains bird singing through a bird singing detection model, and if so, determines that the audio information is bird singing information, and the target bird is a bird to be identified as a bird species; The pre-stored historical ecological data, the bird song audio information and the current environmental data are input into a bird species recognition model to obtain the species information of the target bird output by the bird species recognition model.

3. The method according to claim 2, It is characterized in that The bird species recognition model includes a bird song recognition network and an ecological niche information network, and the target bird includes at least one bird; the pre-stored historical ecological data, the bird song audio information and the current environmental data are input into the bird species recognition model, and the species information of the target bird obtained by the bird species recognition model output includes: Processing the bird call audio information through the bird call recognition network to obtain a bird species prediction vector for each bird; According to the historical ecological data and the current environmental data, obtaining a bird species existence prior vector of each bird through the ecological niche information network, wherein the bird species existence prior vector indicates the suitability of the bird to survive in the monitoring area; Multiply the bird species prediction vector corresponding to each bird species and the bird species existence prior vector to obtain the candidate species prediction value corresponding to each bird species; A target species prediction value greater than a prediction value threshold is selected from the candidate species prediction values, and the bird species corresponding to the target species prediction value is used as the species information of the target bird.

4. The method according to claim 3, It is characterized in that The bird call audio information is processed by the bird call recognition network to obtain a bird species prediction vector of each bird, including: Generate a bird song spectrum diagram according to the bird song audio information; The bird song spectrum graph is input into the bird song recognition network to obtain the bird species prediction vector of each bird output by the bird song recognition network.

5. The method according to claim 4, It is characterized in that Inputting the bird song spectrum graph into the bird song recognition network to obtain the bird species prediction vector of each bird output by the bird song recognition network includes: Input the bird song spectrum graph into a 1*1 convolutional layer and a 3*3 depth-separable convolutional layer in sequence to obtain a feature graph vector; Inputting the feature graph vector into a plurality of convolutional layers of different sizes respectively, obtaining the first bird song region feature output by each convolutional layer, wherein the duration and frequency range of the bird song in the first bird song region features output by different convolutional layers are different; Inputting the first bird song regional features into the deconvolution layers corresponding to the convolution layers respectively, to obtain the second bird song regional features output by each deconvolution layer, wherein the deconvolution layer is used to remodel the global information of the bird song spectrum graph; After vector addition of each of the second bird song region features, the added vectors are scaled through an activation function layer to obtain an attention vector; The attention vector and the feature map vector are multiplied and then passed through a 1*1 convolution layer to obtain a bird species prediction vector.

6. The method according to claim 3, It is characterized in that According to the historical ecological data and the current environmental data, the bird species existence prior vector of each bird species obtained through the ecological niche information network includes: Inputting the current environmental data and the historical ecological data into a species distribution prediction model to obtain a priori probabilities of different birds existing in the monitoring area output by the species distribution prediction model; The prior probabilities of different birds are input into the ecological niche information network to obtain the bird species existence prior vector of each bird output by the ecological niche information network.

7. A method for monitoring wild birds based on cloud-edge collaboration. It is characterized in that Applied to a main controller, the method comprises: Acquire audio information in the monitoring area, and input the audio information into a bird song detection model, wherein the bird song detection model includes a low-level feature extraction block, a high-level feature extraction block and a classifier; Processing the audio information using the low-level feature extraction block to extract low-level features from the audio information; Processing the low-level features using the high-level feature extraction block to extract high-level features from the audio information; Inputting the high-level features into the classifier to obtain a probability value that the audio information contains bird calls; If the probability value is greater than a preset probability threshold, it is determined that the audio information contains bird calls, and the audio information containing the bird calls is used as the bird call audio information of the target bird; The high-level feature extraction block includes: The first 1*1 convolutional layer is used to increase the number of input channels and improve the feature dimension; A 3*3 depth-wise separable convolutional layer connected to the output end of the 1*1 first convolutional layer, for decoupling the spatial dimension and channel dimension of the input features; An attention module, connected to the output end of the 3*3 depth-separable convolutional layer, is used to reduce information loss caused by dimensionality reduction; A 1*1 second convolutional layer is connected to the output end of the attention module, and the 1*1 second convolutional layer is used to reduce the number of channels.

8. The method according to claim 7, It is characterized in that The low-level feature extraction block includes two one-dimensional convolutional layers connected in sequence, and the one-dimensional convolutional layer is used to extract low-level features in the audio information.

9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 2-6 or 7-8 is implemented.

Citation Information

Patent Citations

  • Bird sound recognition method based on mixed feature selection and GWO-KELM model

    CN113066481A

  • Bird buzzing automatic identification system in real environment

    CN115294994A