A method for monitoring bird activity distribution in urban streets based on bird call intensity

By using multi-channel audio and video collaborative processing, combined with a visually guided spatial sound source masking model and beamforming technology, the problem of data distortion in bird call monitoring under high background noise in urban streets has been solved, achieving highly accurate and reliable monitoring of bird activity distribution.

CN122090876BActive Publication Date: 2026-07-17SICHUAN AGRI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN AGRI UNIV
Filing Date
2026-04-23
Publication Date
2026-07-17

Smart Images

  • Figure CN122090876B_ABST
    Figure CN122090876B_ABST
Patent Text Reader

Abstract

This invention relates to the fields of ecological environment monitoring and bioacoustics, specifically disclosing a method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls. This method simultaneously collects multi-channel audio and video data from street areas, uses visual information to identify non-bird dynamic sound sources and constructs a spatial sound source mask model, and combines beamforming technology to directionally focus the audio signal and attenuate interference frequencies. Subsequently, it detects bird call events and calculates their intensity, correlates them with spatial positioning information to generate a spatiotemporal distribution map of bird activity intensity, and verifies the validity of the call events through video verification. This invention improves the signal-to-noise ratio and monitoring accuracy through audio-visual fusion and a closed-loop verification mechanism, achieving highly reliable and high spatial resolution monitoring of urban bird activity distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of ecological environment monitoring and bioacoustics technology, specifically relating to a method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls. Background Technology

[0002] With the acceleration of global urbanization, biodiversity monitoring of urban ecosystems has become an important means of assessing environmental quality and building green and livable cities. Utilizing acoustic sensing technology to capture and analyze bird calls provides non-invasive data support for studying species distribution and activity rhythms in urban street ecological networks. In complex urban soundscapes, acquiring accurate real-time data on bird activity intensity is of scientific value for urban planning and the formulation of ecological balance protection policies, which places high demands on the system's signal acquisition accuracy, anti-interference capabilities, and spatial positioning efficiency.

[0003] Bird activity distribution monitoring technology based on acoustic signal intensity aims to acquire raw acoustic parameters through distributed acquisition devices to invert the dynamic distribution of birds in street space. This type of technology typically involves frequency domain preprocessing of acoustic signals, feature vector extraction, and spatial density mapping. The core lies in accurately identifying the effective signals of target birds from complex combinations of sound sources and combining sound pressure level data to achieve quantitative analysis of activity scale.

[0004] Existing technologies perform poorly in handling the high background noise interference unique to urban streets. This makes it difficult for traditional filtering algorithms to accurately isolate high-decibel non-steady-state noise such as vehicle horns and construction machinery, resulting in severely inflated or distorted monitoring data. Single acoustic monitoring methods lack multi-modal verification mechanisms when faced with overlapping calls from multiple species and complex reflected sound fields, making it difficult to distinguish interference sources in similar frequency bands through linear spectrum analysis, leading to monitoring blind spots. Traditional monitoring models, lacking dynamic perception of physical space, cannot identify and remove interference sources in real time during signal processing, resulting in decreased accuracy in species identification under low signal-to-noise ratio environments. Therefore, a monitoring method for bird activity distribution in urban streets based on bird call intensity is needed. Summary of the Invention

[0005] The purpose of this invention is to provide a method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls, which can solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: a method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls, comprising the following specific steps:

[0007] Step 1: Simultaneously acquire ambient acoustic signals using multi-channel audio acquisition devices deployed in urban street areas, and simultaneously capture dynamic visual information within the corresponding spatiotemporal range using video acquisition devices;

[0008] Step 2: Perform target detection processing on the dynamic visual information to identify non-bird dynamic sound source entities in the street scene, including moving vehicles, pedestrians and construction machinery, and extract their spatial position coordinates and motion state parameters;

[0009] Step 3: Based on the spatial location coordinates and motion state parameters, construct a visually guided spatial sound source mask model. This model is used to identify the location and coverage of high-decibel non-bird sound sources in three-dimensional space.

[0010] Step 4: The spatial sound source mask model is fused with the acoustic signal acquired by the multi-channel audio acquisition device. Beamforming technology is used to focus the sound field directionally, and the frequency components corresponding to non-bird sound sources are attenuated or eliminated in the frequency domain according to the mask model.

[0011] Step 5: Detect bird calls on the noise-reduced acoustic signal, extract the acoustic segments that match the characteristics of bird calls, and calculate the intensity value of each call event by combining the sound pressure level data.

[0012] Step 6: Associate and map the intensity value of the bird call event with its spatial location information at the time of occurrence to generate a spatiotemporal distribution map of bird activity intensity within the street area;

[0013] Step 7: Use the dynamic visual information to perform reverse verification on the located call events, and determine whether there is a visual target with bird morphological characteristics in the corresponding spatiotemporal coordinates. If there is, retain the call event as valid data; otherwise, discard it.

[0014] Preferably, in step 1, the multi-channel audio acquisition device adopts a ring array layout with no less than 8 channels, a sampling frequency of no less than 44.1 kHz, and each channel has a time synchronization mechanism to ensure the accuracy of sound field reconstruction; the video acquisition device is a visible light camera or an infrared thermal imager with a frame rate of no less than 25 frames per second and a resolution of no less than 1280×720 pixels, and its field of view covers the effective sound pickup area of ​​the audio acquisition device.

[0015] Preferably, the target detection processing in step 2 adopts a target recognition algorithm based on deep learning. This algorithm is optimized using a labeled image dataset containing typical interference sources of urban streets during the training phase. It can output the bounding boxes and category labels of various non-bird dynamic entities in real time, and smoothly track the target trajectory in continuous frames through Kalman filtering.

[0016] Preferably, in step 3, the spatial sound source mask model is constructed based on a spherical coordinate system. The non-bird sound source entities obtained by visual recognition are projected onto the sound field space to form a cone-shaped or fan-shaped mask area centered on the sound source and covering its sound emission direction and propagation path. This mask area is dynamically updated with the target's motion state.

[0017] Preferably, in step 4, the beamforming technology uses a delayed summation method to weight and synthesize multi-channel audio signals. The weighting coefficients are dynamically adjusted according to the direction of the sound source and the geometric relationship between the microphone array and the source. An attenuation factor is applied to a specific frequency band of the synthesized signal within the masked area. The attenuation level is adaptively adjusted according to the predicted sound intensity level of the non-bird sound source.

[0018] Preferably, in step 5, the sound event detection extracts the time-frequency features of the acoustic signal through short-time Fourier transform and performs matching analysis in combination with a preset bird call template library. Signal segments with a matching degree exceeding a preset threshold are determined as valid sound events, and their intensity value is determined by the root mean square value of the sound pressure level in the corresponding time period.

[0019] Preferably, in step 6, the spatiotemporal distribution map is based on the street geographic information system as the grid, mapping the intensity value of each valid call event to the grid cell corresponding to its sound source location coordinates, and performing cumulative statistics according to time windows to form a heat map output reflecting the density of bird activity.

[0020] Preferably, in step 7, the reverse verification process compares the video frame at the moment the call event occurs with the sound source location coordinates, and searches for visual targets with bird outlines, body proportions and movement patterns within a preset spatial tolerance range. If a target that matches the biological characteristics of birds is detected, the call event is confirmed as real and valid data.

[0021] Compared with the prior art, the present invention has the following beneficial effects:

[0022] By introducing a collaborative processing mechanism of computer vision information and acoustic signals, the problem of distortion in the measurement of bird call intensity in urban street high background noise environments is effectively overcome; a visually guided spatial mask model is used to accurately suppress non-avian high-decibel interference sources, improving the signal-to-noise ratio of audio signals; through cross-validation of acoustic localization and visual targets, misjudgments caused by overlapping calls and invalid sound fields are eliminated, improving the purity of bird activity distribution data and the accuracy of species identification; the overall method constructs a closed-loop verification system from "hearing" to "seeing", enabling urban street bird ecological monitoring to have higher reliability and spatial resolution capabilities. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the overall technical solution architecture of the present invention;

[0024] Figure 2 This is a schematic diagram of the core principle framework of the visually guided spatial sound source mask model construction and multi-channel audio signal fusion processing in this invention;

[0025] Figure 3 This is a flowchart illustrating the logical process of target detection and spatial orientation parameter extraction for non-avian dynamic sound source entities in this invention.

[0026] Figure 4 This is a schematic diagram of the reverse verification interaction relationship and data flow between the location information of the sounding event and the biometric features of the visual target in this invention;

[0027] Figure 5 This is a logical flowchart of the generation of the effective sound event intensity correlation mapping and the spatiotemporal distribution map of street bird activity in this invention. Detailed Implementation

[0028] Example 1: Please refer to the appendix Figure 1 To be continued Figure 5 To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments.

[0029] In the method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls, step 1 is first performed: environmental acoustic signals are acquired synchronously by multi-channel audio acquisition devices deployed in urban street areas, and dynamic visual information in the corresponding spatiotemporal range is captured synchronously by video acquisition devices.

[0030] In step 1, the multi-channel audio acquisition device employs a highly symmetrical circular array layout with at least eight channels, aiming to achieve complete recording of the three-dimensional sound field through spatial sampling. Each channel includes a high-sensitivity omnidirectional condenser microphone with a frequency response range covering 20 Hz to 20,000 Hz, ensuring complete capture of all details from low-frequency ambient background noise to high-frequency bird calls. The sampling frequency of the audio acquisition device is set to at least 44.1 kHz, and a 24-bit sampling depth is used during analog-to-digital conversion to obtain an extremely high dynamic range, enabling the resolution of weak acoustic signals even in noisy environments.

[0031] To achieve time synchronization between channels, a high-precision clock distribution module is integrated into the system. The synchronization clock signal generated by this module is fed to the sample-and-hold circuit of each audio channel via differential transmission cables. Its synchronization error is controlled within 1 microsecond, a precision crucial for subsequent sound field reconstruction based on the time difference of arrival.

[0032] The video acquisition device is configured to be mounted coaxially with or at a known spatial offset from the audio acquisition device. The video acquisition device employs a visible light camera or an infrared thermal imager with a frame rate of at least 25 frames per second to ensure the continuity of dynamic target trajectories. The camera provides high-resolution images of at least 1280×720 pixels, ensuring clear identification of distant vehicles, pedestrians, and construction machinery in complex street backgrounds. The field of view of the video acquisition device is optically designed to fully cover the effective pickup area of ​​the audio acquisition device, typically set to a wide angle of 120 degrees horizontally and 90 degrees vertically. The video and audio streams are timestamped at the data encapsulation layer; each frame is associated with a precise system timestamp, which is aligned with the sequence number of the audio sampling frame at the millisecond level, establishing a unified spatiotemporal coordinate reference system.

[0033] After obtaining the raw data in step 1 above, step 2 is executed: target detection processing is performed on the dynamic visual information to identify non-bird dynamic sound source entities in the street scene, including vehicles, pedestrians and construction machinery, and their spatial position coordinates and motion state parameters are extracted.

[0034] Specifically, in step 2, the system employs a target recognition algorithm based on a deep convolutional neural network. The backbone network of this algorithm adopts a cross-stage local network structure and achieves robust recognition of targets of different sizes in the image through a multi-scale feature fusion strategy. During the pre-training phase, the algorithm loads a large-scale dataset containing hundreds of thousands of annotated street traffic scene images, covering typical urban interference sources such as vehicles, pedestrians, motorcycles, and excavators under various lighting conditions, occlusion ratios, and weather backgrounds.

[0035] When the video stream is input to the detection module, the algorithm generates a series of candidate regions on each frame and calculates the category probability distribution and bounding box regression bias for each region. The detection results are output in the format of the center pixel coordinates, width, height, and confidence score of various non-avian dynamic entities. To transform pixel coordinates into three-dimensional spatial coordinates, the system establishes a projective transformation model from the two-dimensional image plane to the three-dimensional world coordinate system based on the camera's intrinsic and extrinsic parameters. For objects with known geometric dimensions (such as standard-sized cars), the radial length of its distance sensor is inferred from the pixel span.

[0036] To address situations where the target is temporarily occluded or lost during detection, step 2 introduces a Kalman filter algorithm based on a state-space model. This algorithm encapsulates the target's current position, velocity, and acceleration in a state vector, calculates the expected position of the target in the next frame through a prediction step, and updates it using actual observations. This process enables smooth tracking of the trajectory of non-avian entities and outputs continuous motion state parameters, including instantaneous displacement vectors and angular velocity changes. These parameters not only describe the target's spatial orientation at the current moment but also predict its displacement trend over extremely short time intervals, providing dynamic input for subsequent spatial mask construction.

[0037] Based on the spatial location coordinates and motion state parameters, a visually guided spatial sound source mask model is constructed. This model is used to identify the location and coverage of high-decibel non-bird sound sources in three-dimensional space.

[0038] Specifically, in step 3, the spatial sound source mask model is constructed based on a spherical coordinate system, with its origin set at the array geometric center of the multi-channel audio acquisition device. The system converts the three-dimensional spatial position coordinates extracted in step 2 into azimuth, pitch, and radial distance in a spherical coordinate system.

[0039] The spatial sound source masking model defines each identified non-avian entity as a dynamic noise contribution region. Based on the entity's geometric center and physical envelope dimensions, a mask core centered on the sound source is generated in spherical coordinate space. Considering the diffusion characteristics of sound waves in atmospheric propagation and the entity's directional features, this mask core extends radially outward, forming a cone-shaped or fan-shaped mask region. The angular size of this region is dynamically determined by the ratio of the non-avian entity's physical width to its distance.

[0040] The masking model internally defines a weighting function. At the center of the masked region, the weight is set to 1, representing an extremely high noise energy distribution in that direction. As the angle of deviation from the center increases, the weight decays according to a Gaussian distribution until it reaches 0 at the edge of the region. This weight distribution simulates the energy spillover effect of a real sound source in space. Since non-avian entities are in continuous motion, the spatial sound source masking model is updated in real time based on the predicted state of the Kalman filter output. The update frequency is consistent with the frame rate of the video capture device, ensuring that the masked region always closely follows the dynamic noise source.

[0041] After constructing the spatial sound source mask model, step 4 is executed: the spatial sound source mask model is fused with the acoustic signal acquired by the multi-channel audio acquisition device, beamforming technology is used to directionally focus the sound field, and the frequency components corresponding to non-bird sound sources are attenuated or eliminated in the frequency domain according to the mask model.

[0042] Specifically, in step 4, the system utilizes the time difference information of the multi-channel signals to achieve beamforming. Beamforming technology employs a delay-summing method to weighted synthesize the multi-channel audio signals. The system calculates the propagation delay from a specific target direction to each microphone unit. This delay is equal to the geometric distance between the microphone and the target point divided by the speed of sound. The speed of sound is calculated and compensated based on the real-time temperature feedback from the ambient temperature sensor.

[0043] To suppress interference within the masked region during synthesis, the system dynamically adjusts the gain coefficient of each beam direction based on the weight distribution of the spatial sound source mask model. The gain coefficient remains at its maximum when the main lobe of the beam points to the non-masked region (i.e., the possible bird activity area); while when the side lobes or secondary main lobes of the beam cover the non-bird sound source orientation defined in the masked region, the system automatically applies a depth attenuation factor.

[0044] Furthermore, at the spectral domain processing level, the system performs a short-time Fourier transform on each channel signal, converting the time-domain signal into a time-frequency distribution map. For strong noise sources located within the masked region (such as a moving heavy truck), the system predicts their main spectral distribution characteristics. For example, vehicle engine noise is typically concentrated in the low-frequency band, while braking sounds contain high-frequency components. The mask model generates an attenuation mask array within the corresponding frequency window, where each element is a gain adjustment value between 0 and 1. By multiplying the time-frequency map element-wise with this attenuation mask array, the system can accurately "erase" noise interference in specific directions and frequencies, while retaining undisturbed signal components in other directions. This dual spatial and frequency domain filtering mechanism improves the signal-to-noise ratio of the target audio segment, laying the foundation for subsequent accurate detection.

[0045] Then, step 5 is performed: the noise-reduced acoustic signal is subjected to sound event detection, the acoustic segments that match the characteristics of bird calls are extracted, and the intensity value of each sound event is calculated by combining the sound pressure level data.

[0046] The denoised mono synthesized signal is fed into a buzzing event detector. This detector first extracts the energy envelope of the signal using a sliding window technique. The sliding window length is set to 10 to 50 milliseconds, with an overlap rate of 50%, to balance temporal and frequency resolution.

[0047] To accurately separate bird calls from background noise, the system uses a pre-defined bird call template library for matching analysis. This library contains typical time-frequency patterns of calls from various common urban birds, including frequency modulation characteristics, duration, and harmonic structure. The system extracts the Mel-Cepstral Coefficient features of the segment to be detected and calculates its cosine similarity or Euclidean distance with each feature vector in the template library. When the matching score exceeds a preset confidence threshold (e.g., 0.75), the signal segment is determined to be a valid call event, and its start and end times are recorded.

[0048] For each valid call event, the system further calculates its acoustic intensity value. The intensity value is calculated based on sound pressure level (SPL) data. Specifically, within the duration corresponding to the call event, the system extracts the amplitude sequence of digital audio samples, calculates the average of the squared values ​​of each sample point in the sequence, and then takes the square root of this average to obtain the root mean square (RMS) value. Subsequently, combining the microphone array's sensitivity parameters and the preamplifier's gain coefficient, this RMS value is converted into a SPL physical quantity in decibels. To improve calculation accuracy, the system also subtracts the energy contribution of residual environmental background noise estimated through a spatial masking model within that time period, obtaining the call intensity purely generated by bird activity.

[0049] Next, step 6 is executed: the intensity value of the sound event is associated with its spatial location information at the time of occurrence to generate a spatiotemporal distribution map of bird activity intensity in the street area.

[0050] Specifically, in step 6, the system integrates the spatial orientation data output by the sound source localization module with the intensity data calculated in step 5. Sound source localization calculates the cross-correlation function of the multi-channel signals, searches for the arrival delay difference corresponding to the cross-correlation peak, and inversely determines the azimuth and elevation angles of the sound source in the spherical coordinate system.

[0051] The spatiotemporal distribution map is generated using a street geographic information system (GIS) as the base map. This base map is divided into finely divided two-dimensional or three-dimensional grid cells, with the size of each cell set to 1 meter × 1 meter or 5 meters × 5 meters according to the monitoring accuracy requirements. The system uses the intensity value of each valid sound event as a weight and maps it to the corresponding grid cell based on its location coordinates.

[0052] The mapping process employs a kernel density estimation method. For each occurrence of a bird call, its energy is not only contributed to the central grid but also smoothly distributed to surrounding grids according to a preset diffusion function, simulating the spatial distribution probability of the sound source. In the time dimension, the system sets adjustable observation windows (e.g., 15 minutes, 1 hour, or 24 hours). Within the same time window, the intensity values ​​mapped to each grid cell are accumulated or averaged. Finally, these statistical data are converted into a heatmap output using a color mapping algorithm, where the hue or brightness of the colors represents the intensity density of bird activity. This map visually displays the spatial distribution preferences of birds within a street area at different times, such as in green belts, treetops, or the edges of specific buildings.

[0053] Finally, step 7 is executed: the dynamic visual information is used to perform reverse verification on the located call events to determine whether there are visual targets with bird morphological characteristics in the corresponding spatiotemporal coordinates. If there are, the call event is retained as valid data; otherwise, it is discarded.

[0054] Specifically, in step 7, the system establishes a closed-loop logic for audio-visual feedback. Whenever a high-confidence beeping event is detected in step 5 and its spatial location is locked in step 6, the system immediately backtracks the video cache data at the corresponding moment.

[0055] The reverse verification process unfolds within a preset spatial tolerance range. Due to the inherent physical errors in sound source localization, the system defines a search window with a certain pixel radius centered on the sound source mapping point in the visual image. Within this window, the system invokes a recognition model specifically optimized for bird biometrics. This model focuses on detecting visual targets that possess bird outlines, specific body proportions, and high-frequency vibration patterns (such as wing flapping or throat movements during calls).

[0056] If the recognition model successfully locates a visual target matching bird characteristics within the search window, and the target's trajectory matches the displacement trend of the sound source localization, the system determines the sound event as a "genuine event with both visual and auditory verification," retains it, and marks it with the highest reliability level. Conversely, if the direction pointed to by the sound source visually only contains reflective glass, swaying leaves, or objects that, while producing sound, are not birds, the system considers the sound event to be invalid interference caused by multiple reflections in a complex sound field or algorithmic misjudgment, and removes it from the spatiotemporal distribution map. This cross-validation mechanism filters out "ghost sound sources" in the urban environment, ensuring extremely high purity of the monitoring results.

[0057] Example 2: This example further refines the technical implementation scheme of the system in handling multi-target overlapping calls, based on Example 1. When multiple individual birds are simultaneously making calls from different directions in a street environment, the system's resolution and data processing logic are as follows:

[0058] In step 4, the beamforming module not only generates the main beam but also generates multiple directional beams in parallel based on the suspected bird target locations detected visually in step 2. Each beam independently performs spatial filtering, and its weighting coefficient matrix is ​​dynamically adjusted according to the instantaneous azimuth of each target. For any target beam, the system treats all other known non-bird dynamic sound sources (from the mask model in step 3) and other detected bird sound sources as interference terms and sets null points in the gain map of the current beam.

[0059] During the detection of the call event in step 5, the system performs time-frequency overlap analysis on the multiple beam signals output in parallel. If two beams from different directions detect calls at the same time, the system calculates their spectral correlation. If the correlation is extremely low, it is determined that the calls are from independent bird individuals; if the correlation is high, the system compares the energy attenuation gradients of the two signals to determine whether acoustic crosstalk exists. In this case, the system uses the independent component analysis logic in the blind source separation algorithm to decompose the mixed signal into several statistically independent sub-components, accurately reconstructing the call intensity value of each individual bird.

[0060] In step 6, when generating the spatiotemporal distribution map, a multi-agent tracking framework is introduced for multi-target scenarios. Each confirmed bird call source is assigned a unique identifier, and its intensity contribution is recorded over time. The generation of the heatmap is no longer a simple energy superposition, but a superposition based on the probability distribution of individuals, avoiding inflated intensity data caused by the aggregation of multiple birds.

[0061] For nighttime scenes with extremely poor street lighting conditions, in this embodiment, the video acquisition device in step 1 automatically switches to infrared thermal imaging mode. The infrared image is processed using pseudo-color to extract the temperature distribution features of the target. The target detection algorithm in step 2 switches to a deep learning model specifically designed for infrared wavelengths, locating pedestrians or vehicles by capturing the temperature difference contours between them and their background. Since birds are warm-blooded animals, they appear as high-temperature spots in the infrared image. The reverse verification process in step 7 calculates the area, shape eccentricity, and correspondence between the high-temperature spots and their call intensity based on this, maintaining monitoring accuracy even in completely dark environments.

[0062] In terms of signal transmission and storage, this embodiment employs an edge computing architecture. The computational tasks in steps 1 to 5 are directly completed in an embedded processing unit deployed on-site in the street. This unit is equipped with a high-performance graphics processing chip and a digital signal processor. Only the extracted valid sound event parameters, location information, and verified statistical data are uploaded to the cloud server via a low-power wide area network or a 5G network. This approach reduces the bandwidth pressure on data transmission and protects the visual privacy of pedestrians on the street.

[0063] Example 3: This example focuses on describing the technical details of how the present invention maintains monitoring accuracy under adverse weather conditions, such as heavy rainfall or strong winds.

[0064] In step 1, when the environmental sensor detects that the rainfall exceeds a preset threshold or the wind speed exceeds level 5, the system triggers the anti-wind noise mode. Each microphone unit of the multi-channel audio acquisition device is wrapped with an acoustically transparent hydrophobic windproof cover, which aims to physically reduce the low-frequency turbulence noise generated by the direct impact of airflow on the diaphragm.

[0065] In the fusion process of step 4, the system introduces a dynamic compensation model based on environmental meteorological parameters. Wind causes a shift in the propagation path of sound waves in the air, resulting in a sound ray bending effect. Based on the wind direction and speed vector provided by the wind speed sensor, the system uses ray acoustics theory to correct the time delay parameter in beamforming. The adjustment amount of the delay time is equal to the projection of the microphone spacing onto the wind direction vector divided by the component of the wind speed in the direction of sound propagation. This correction ensures that the beam can still accurately target visually identified bird targets or noise sources even in strong winds.

[0066] To address the broadband random noise generated by rainfall, the spectral attenuation operation in step 4 employs a non-stationary noise suppression algorithm. The system estimates the noise floor generated by raindrops hitting the ground or vegetation in real time and extracts noise feature vectors from the multi-channel signal. When performing spectral attenuation, the system does not simply use a visually guided mask, but rather combines spectral subtraction logic to adaptively adjust the attenuation coefficient according to the ratio of signal energy to estimated noise energy. For frequency bands with extremely low signal-to-noise ratios, the attenuation factor approaches 0; for frequency bands with high signal-to-noise ratios, the attenuation factor remains at 1.

[0067] In the reverse verification of step 7, considering that rain may cause visual image blurring, the system activates image dehazing and enhancement algorithms. By extracting the dark primary color prior features of the image, the outline details of birds affected by occlusion are restored. If visual verification fails due to low visibility (e.g., visibility less than 10 meters), the system automatically reduces the weight of visual verification and instead relies on the statistical distribution characteristics of acoustic features, such as analyzing the coherence function of the call in different channels to help determine the validity of the sound source.

[0068] When performing step 5 to calculate the sound intensity value, the system introduces an energy attenuation correction coefficient based on the intensity level of the rain sound. Since raindrops scatter and absorb high-frequency sound waves, the system calculates the additional loss of sound waves along the propagation path based on the rainfall intensity and compensates it back into the calculated intensity value, ensuring that the bird activity intensity measured under different weather conditions is comparable and consistent.

[0069] Example 4: This example describes the application of the present invention in a specific urban construction scenario, such as bird monitoring near a construction area. In this scenario, noise generated by non-bird sound source entities (such as excavators and impact drills) has extremely high instantaneous intensity and complex pulse characteristics.

[0070] In step 2, the target detection algorithm is specifically designed to identify the working posture of construction machinery. When the impact drill is detected to be in operation, the system extracts its impact frequency and pulse width parameters.

[0071] In step 3, when constructing the spatial sound source masking model, the masking region is given higher temporal resolution for this type of high-decibel impulse noise. The mask weights are no longer static, but are pulse-modulated according to the rhythm of the impact drill's operation. At the moment of impact, the mask attenuation value is adjusted to the maximum; during the interval between impacts, the attenuation value is rapidly restored to capture faint bird calls that may have been masked.

[0072] In step 4, to address the strong ground vibration interference generated by the construction machinery, the system acquires the vibration signal of the array support using an accelerometer. An adaptive cancellation filter is then used, with the vibration signal as a reference input, to subtract the non-acoustic spurious signal components caused by physical vibration from the audio channel. This process further purifies the acoustic signal and prevents sound level estimation errors due to vibration.

[0073] In step 6, when generating the spatiotemporal distribution map, the system incorporates baseline comparison logic. By analyzing the changes in bird activity intensity distribution in the street area before and after construction, the system can quantify the specific impact of construction noise on bird habitat preferences. The spatiotemporal distribution map supports multi-layer overlay display, allowing users to simultaneously view the overlap between noise heatmaps and bird activity heatmaps, intuitively observing birds' avoidance behavior towards high-decibel areas.

[0074] In the reverse verification of step 7, for frequently moving objects within the construction area, the system utilizes multi-camera collaborative positioning technology to improve the accuracy of spatial coordinates. Stereoscopic visual information of the target is acquired through cameras installed at different angles, and the target's depth information is calculated. The sound source localization result must simultaneously satisfy the projection geometric constraints of multiple cameras to be considered valid. This multi-dimensional cross-comparison enables the system to maintain accurate bird call recognition even under extremely complex industrial background noise.

[0075] At the data management level, this embodiment employs distributed ledger technology to store monitoring data. Each generated valid sound event record, including its time, location, intensity value, and corresponding visual verification feature summary, is encrypted and stored in a local database. This approach ensures the authenticity and immutability of the monitoring data, providing legally valid original evidence for urban ecological assessment.

[0076] Example 5: This example details the collaborative working mechanism of the present invention in a large-scale distributed deployment of a street network. Multiple monitoring nodes (including audio acquisition devices and video acquisition devices) are installed on continuous street light poles spaced 50 to 100 meters apart.

[0077] In step 1, all monitoring nodes achieve network-wide microsecond-level synchronization via Precise Time Protocol (PTP). Each node has an independent Global Navigation Satellite System (GNSS) receiver to obtain precise geographic coordinates and absolute time synchronization.

[0078] When generating the spatiotemporal distribution map in step 6, the system is no longer limited to the field of view of a single node. When a bird flies from the sound pickup range of one monitoring node to the range of another, the system utilizes cross-camera target tracking technology and a sound source trajectory continuity algorithm to achieve seamless tracking of the individual bird. The intensity values ​​acquired by two adjacent nodes are merged using a weighted fusion algorithm. The weight allocation depends on the geometric distance of the bird from the center of each node array. The closer the node, the greater its contribution weight. This multi-node collaborative mechanism eliminates monitoring blind spots and improves the positioning accuracy of fast-moving targets.

[0079] In the fusion process of step 4, the distributed deployment allows the system to perform multi-array coherent processing using audio signals from neighboring nodes. By calculating the spatial cross-spectral density matrix, the system is able to suppress background noise from traffic flow at the far end of the street. This spatial diversity processing capability enhances the suppression effect on low-frequency background noise.

[0080] To address the display of spatiotemporal distribution maps under large datasets, step 6 introduces hierarchical detail rendering technology. In a panoramic street view, the map displays coarse-grained activity trends; when the user zooms in near a specific building or tree, the system automatically retrieves high-precision grid data to display activity distributions down to the decimeter level. Simultaneously, the system establishes an AI-based trend prediction model. By learning from historical monitoring data, the model can predict the probability distribution of bird activity during specific future time periods (such as early morning or evening), providing predictive guidance for street greening maintenance and ecological protection decisions.

[0081] In the reverse verification stage of step 7, the distributed system implements multi-view verification. If a node's visual field is obstructed by trees, the system automatically schedules the visual sensors of adjacent nodes to compensate for the deficiencies of single-point observations using redundant information from multiple angles. This collaborative verification method reduces the false negative rate caused by environmental occlusion.

[0082] The system in this embodiment also possesses self-healing capabilities. If the video acquisition device at a certain node malfunctions, the system automatically increases the weight of audio location information from surrounding nodes and corrects it by combining it with historical activity probability maps. If the audio channel is damaged, sound level estimation is compensated through video analysis. This highly redundant and complementary architecture ensures the long-term stable operation of the city-scale monitoring network.

[0083] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls, characterized in that, Includes the following steps: Step 1: Simultaneously acquire ambient acoustic signals using multi-channel audio acquisition devices deployed in urban street areas, and simultaneously capture dynamic visual information within the corresponding spatiotemporal range using video acquisition devices; Step 2: Perform target detection processing on the dynamic visual information to identify non-bird dynamic sound source entities in the street scene, and extract the spatial position coordinates and motion state parameters of the dynamic sound source entities; Step 3: Based on the spatial location coordinates and motion state parameters, construct a visually guided spatial sound source mask model. The spatial sound source mask model is used to identify the location and coverage of non-bird sound sources in three-dimensional space. Step 4: The spatial sound source mask model is fused with the environmental acoustic signal acquired by the multi-channel audio acquisition device. Beamforming technology is used to focus the sound field directionally, and the frequency components corresponding to non-bird sound sources are attenuated or eliminated in the frequency domain according to the spatial sound source mask model. Step 5: Detect bird calls on the noise-reduced acoustic signal, extract acoustic segments that match the characteristics of bird calls, and calculate the intensity value of each call event by combining the sound pressure level data. Step 6: Associate and map the intensity value of the bird call event with its spatial location information at the time of occurrence to generate a spatiotemporal distribution map of bird activity intensity within the street area; Step 7: Use the dynamic visual information to perform reverse verification on the located call events, and determine whether there is a visual target with bird morphological characteristics in the corresponding spatiotemporal coordinates. If there is, retain the call event as valid data; otherwise, discard it. In the data encapsulation layer, timestamps are injected into the video and audio streams, each frame of image is associated with a system time stamp, and the system time stamp is aligned with the sequence number of the audio sampling frame with millisecond-level precision, thus constructing a unified spatiotemporal coordinate reference system.

2. The method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls according to claim 1, characterized in that, Step one specifically includes: The multi-channel audio acquisition device adopts a highly symmetrical ring array layout, each channel contains an omnidirectional condenser microphone, and each channel has an independent sample-and-hold circuit. A high-precision clock distribution module is used to generate a synchronous clock signal, and the synchronous clock signal is fed to the sample and hold circuit of each channel through a differential transmission cable to ensure that the time synchronization error between each channel is within a preset microsecond level accuracy range. The video acquisition device is coaxially mounted with the multi-channel audio acquisition device or maintains a preset spatial offset, and its field of view completely covers the sound pickup area of ​​the multi-channel audio acquisition device.

3. The method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls according to claim 2, characterized in that, Step two specifically includes: A deep convolutional neural network is used to analyze each frame of a video stream. A multi-scale feature fusion strategy is used to identify non-bird dynamic sound source entities of different sizes in the images. The center pixel coordinates, bounding box size, and category confidence of the identified non-bird dynamic sound source entities are output. Based on the intrinsic and extrinsic parameters of the video acquisition device, a projective transformation model from a two-dimensional image plane to a three-dimensional world coordinate system is established, and the coordinates of the center pixel are transformed into three-dimensional spatial position coordinates. Kalman filtering based on a state-space model is introduced to encapsulate the current position, velocity, and acceleration of the non-bird dynamic sound source entity in a state vector. The expected position of the target in the next frame is calculated through a prediction step, and the actual observation values ​​are used for correction and update to output continuous motion state parameters.

4. The method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls according to claim 3, characterized in that, Step three specifically includes: A spherical coordinate system is constructed with the array geometric center of the multi-channel audio acquisition device as the origin, and the three-dimensional spatial position coordinates are transformed into azimuth, pitch and radial distance in the spherical coordinate system. Each identified non-bird dynamic sound source entity is defined as a dynamic noise contribution region. A mask core is generated in spherical coordinate space based on the geometric center and physical envelope size of the entity. The mask core is then extended radially outward to form a cone-shaped or fan-shaped mask region covering the sound emission direction and propagation path. A weight function is defined inside the cone-shaped or fan-shaped mask region, such that the weight value is maximized at the center of the region and decreases according to a Gaussian distribution as the angle of deviation from the center increases, until it is reduced to the minimum value at the edge of the region. The cone or fan-shaped mask region is updated in real time based on the predicted state output by the Kalman filter, so that the update frequency of the mask region is consistent with the frame rate of the video acquisition device.

5. The method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls according to claim 4, characterized in that, Step four specifically includes: Using the time difference information of multi-channel audio signals, the signals are weighted and synthesized by delay summation. The propagation delay from the target direction to each microphone unit is calculated based on the geometric distance between the microphone and the target point and the speed of sound after real-time ambient temperature compensation. The gain coefficient of each beam direction is dynamically adjusted according to the weight distribution of the spatial sound source mask model. When the side lobe or secondary main lobe of the beam covers the azimuth defined by the mask area, a depth attenuation factor is applied. In the spectral domain processing, a short-time Fourier transform is performed on each channel signal to obtain a time-frequency distribution map. The spectral distribution characteristics of non-bird dynamic sound sources located within the mask area are predicted, and an attenuation mask array composed of gain adjustment values ​​is generated in the corresponding frequency window. By multiplying the time-frequency distribution map element-wise with the attenuation mask array, noise interference from the target direction that falls within a specific frequency range is eliminated. The specific frequency range is a frequency window dynamically determined based on the spectral distribution characteristics of the non-bird dynamic sound source.

6. The method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls according to claim 5, characterized in that, Step five specifically includes: The energy envelope of the denoised signal is extracted using the sliding window technique, and the time resolution and frequency resolution are balanced by adjusting the length and overlap of the sliding window. Extract the Mel-Cepstral coefficient features of the segment to be detected, and calculate the cosine similarity or Euclidean distance between it and each feature vector in the preset bird call template library. When the cosine similarity or Euclidean distance exceeds the preset confidence threshold, it is determined to be a valid call event and its start time and end time are recorded. Within the duration corresponding to the effective sound event, the amplitude sequence of digital audio samples is extracted and its root mean square value is calculated. Combined with the sensitivity parameters of the microphone array and the gain coefficient of the preamplifier, the root mean square value is converted into a sound pressure level physical quantity. During the calculation, the energy contribution of residual environmental background noise estimated by the spatial sound source mask model within this duration is deducted.

7. The method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls according to claim 6, characterized in that, Step six specifically includes: Using the street geographic information system as the base map, the monitoring area is divided into grid cells of preset size, and the cross-correlation function of multi-channel signals is used to determine the azimuth and elevation angles of the sound source in the spherical coordinate system; The intensity values ​​of effective sound events are mapped to the grid cells using the kernel density estimation method, so that the energy of each sound event is not only contributed to the central grid, but also smoothly distributed to the surrounding grids according to the preset diffusion function, in order to simulate the spatial distribution probability of the sound source. Set an adjustable time observation window to accumulate or calculate the average intensity value mapped to each grid cell within the same time observation window; The statistical data is converted into a heatmap output using a color mapping algorithm, and the intensity and density of bird activity are represented by the hue or brightness of the colors.

8. The method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls according to claim 7, characterized in that, Step seven specifically includes: A closed-loop audio-visual feedback logic is established. After detecting a call event and locking its spatial location, the video cache data at the corresponding moment is retrieved. A search window is defined within a preset spatial tolerance range. Within the search window, a recognition model optimized for bird biological characteristics is called. The recognition model is used to detect visual targets with bird outlines, body proportions, and high-frequency vibration patterns. If the recognition model locks onto a visual target within the search window, and the movement trajectory of the visual target matches the displacement trend of the sound source localization, then the sound event is determined to be real and valid data and its reliability level is marked. If no target matching the biological characteristics is detected in the search window, the sound event is determined to be interference caused by reflection in a complex sound field or algorithm misjudgment, and it is removed from the spatiotemporal distribution map.

9. The method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls according to claim 8, characterized in that, The method also involves processing and adapting to the environment in multi-target overlapping scenarios: when multiple bird individuals are simultaneously making calls from different directions in a street environment, in step four, multiple directional beams are generated in parallel based on the visually detected locations of suspected bird targets. Each directional beam independently performs spatial filtering, and null points are set for non-target sound sources in the gain diagram. In step five, time-frequency overlap analysis is performed on the parallel output signals of multiple directional beams. The presence of acoustic crosstalk is determined by calculating the spectral correlation, and the mixed signal is decomposed into several statistically independent sub-components using independent component analysis logic, thereby restoring the call intensity value of each individual. For nighttime monitoring scenarios, the video acquisition device automatically switches to infrared thermal imaging mode, identifies high-temperature spots in the infrared image by extracting the temperature distribution characteristics of the target, and performs target detection and reverse verification by combining the area and shape eccentricity of the high-temperature spots.

10. The method for monitoring the distribution of bird activity in urban streets based on the intensity of bird calls according to claim 9, characterized in that, The method also involves compensation and multi-node distributed collaborative processing under severe weather conditions: in strong wind or rainy weather, based on the wind direction and wind speed vectors provided by environmental sensors, the time delay parameter in beamforming is corrected using ray acoustics theory to counteract the sound ray bending effect caused by wind. To address random noise generated by rainfall, the spectral attenuation coefficient is adaptively adjusted based on the ratio of signal energy to estimated noise energy using spectral subtraction logic. The energy loss of the sound wave along its propagation path is calculated according to the rainfall intensity, and the calculated intensity value is compensated. When multiple monitoring nodes are deployed in a distributed manner, each node achieves network-wide synchronization through a precise time protocol. Cross-camera target tracking technology is used to track individual birds that cross the sound pickup range of different monitoring nodes. The intensity values ​​acquired by multiple adjacent nodes are weighted and fused based on the geometric distance of the birds from the center of each node array, and multi-array coherent processing is performed using a spatial cross-spectral density matrix to suppress background noise from traffic flow at distant points on the street.