Forest ecology monitoring method and system based on sound analysis

Through the forest ecological monitoring method based on sound analysis, combined with environmental sensor data and drone image data, real-time integration of biological behavior, environmental parameters and spatial details in forest monitoring is achieved, and the simple problem of decision support systems in the existing technology is solved, and the real-time and accuracy of monitoring is improved.

CN120014813APending Publication Date: 2025-05-16日照朝力信息科技有限公司 +1

Patent Information

Application Number
CN202510472233.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to integrate biological behavior, environmental parameters and spatial details in real time in forest monitoring, resulting in a simple decision support system and difficulty in achieving closed-loop feedback and dynamic adjustment.

Method used

The forest ecological monitoring method based on sound analysis is adopted, and through continuous recording, energy detection, spectrum conversion, noise reduction, MFCC feature extraction and deep learning models, sound classification and abnormal detection are realized, and multi-source data fusion is combined with environmental sensor data and drone image data to generate comprehensive ecological health indicators.

Benefits of technology

It realizes high-frequency and real-time monitoring of biological activities, improves the real-time and accuracy of monitoring, can capture biological behavior dynamics more comprehensively, and provides more accurate ecological health assessment and recovery decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014813A_ABST
    Figure CN120014813A_ABST
Patent Text Reader

Abstract

The invention discloses a forest ecology monitoring method and system based on sound analysis, and relates to the technical field of sound processing, and the method comprises the steps: obtaining the real-time continuous recording of a forest monitoring area, embedding a timestamp and geographic position information in the recording, and guaranteeing that all recording data have precise space-time identifiers; eliminating a silent part in the recording data by using short-time energy calculation, and dynamically setting an energy threshold value and retaining an effective signal to obtain effective audio data; the effective audio data are converted into a time-frequency matrix through STFT, filter parameters are adjusted in combination with real-time meteorological data, dynamic noise reduction processing is carried out, and a frequency spectrum after noise reduction is obtained; mFCC features are extracted through Mel filter group transformation and DCT, and standardized feature vectors are generated; and performing deep learning classification on the MFCC features by using a pre-trained CNN and LSTM, setting a statistical threshold based on historical data, detecting abnormality, and finally generating output data including a classification result, activity intensity and abnormality early warning information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sound analysis and processing, and in particular to a forest ecological monitoring method and system based on sound analysis. Background Art

[0002] At present, forest ecological monitoring mainly uses satellite remote sensing and UAV images to collect large-scale vegetation coverage, land use and disease detection information, but is subject to resolution limitations, cloud cover and untimely spatiotemporal updates. Fixed environmental sensors are used to monitor temperature, humidity, soil pH, light and other parameters in real time, and provide local environmental status information, but their data usually only reflect the physical environment status and cannot fully capture biological activity information. Although multimodal data such as the fusion of environmental data and remote sensing images are also used for environmental monitoring, they often fail to integrate biological behavior, environmental parameters and spatial details in real time, and the decision support system is relatively simple, and it is difficult to achieve closed-loop feedback and dynamic adjustment. Summary of the invention

[0003] This application provides a forest ecological monitoring method and system based on sound analysis, which solves the technical problems in the prior art that forest monitoring often fails to integrate biological behavior, environmental parameters and spatial details in real time, the decision support system is relatively simple, and it is difficult to achieve closed-loop feedback and dynamic adjustment.

[0004] In view of the above problems, the present application provides a forest ecology monitoring method and system based on sound analysis.

[0005] In a first aspect, the present application provides a forest ecological monitoring method based on sound analysis, wherein the method is applied to a forest ecological monitoring system based on sound analysis, and the method comprises: Obtain real-time continuous recordings of forest monitoring areas and embed timestamps and geographic location information in the recordings to ensure that all recording data has accurate time and space identification; Use short-time energy calculation to remove the silent part in the recording data, retain only the time period containing obvious acoustic activities, and dynamically set the energy threshold to retain the valid signal to obtain valid audio data. The dynamic energy threshold can be automatically adjusted according to the ambient noise level to prevent missing weak signals and ensure that weak signals will not be missed; The effective audio data is converted into a time-frequency matrix through STFT to analyze the energy distribution of different frequency bands, and the filter parameters are adjusted in combination with real-time meteorological data (such as wind speed and rainfall). Dynamic noise reduction is performed on specific frequency bands (such as low-frequency wind sound), and the noise spectrum is estimated by detecting silent segments. The estimated noise spectrum is then subtracted from the original spectrum to obtain the denoised spectrum. Mel filter bank transform and DCT are used to extract MFCC features from the denoised spectrum, and standardized feature vectors are generated to replace the original audio data for uploading, thereby significantly reducing the amount of data; Use pre-trained CNN and LSTM to perform deep learning classification on MFCC features, set statistical thresholds based on historical data, detect anomalies, and ultimately generate output data containing classification results, activity intensity, and anomaly warning information; That is, the pre-trained convolutional neural network (CNN) is used to process the MFCC feature map and extract spatiotemporal features; Use long short-term memory networks (LSTM or bidirectional LSTM) to capture the temporal dependencies in audio sequences for classification and recognition (such as distinguishing between "birdsong" and "insect calls"); Setting statistical thresholds based on historical data: For example, when the average frequency or total energy of animal calls in a certain area suddenly drops, or abnormal noise (which may represent mechanical noise, illegal activities, etc.) suddenly increases, an abnormal warning is triggered; Generates data including classification results, activity intensity indicators, anomaly detection tags, GPS locations of abnormal areas, and warning levels.

[0006] Obtain environmental sensor data from forest monitoring areas and ensure that the data has clear time and space identification; Perform preliminary verification, noise smoothing and format unification on the collected raw data to ensure data accuracy and consistency; Use historical data and high-precision reference sensors to build regression models, automatically calibrate sensor outputs, eliminate errors caused by long-term drift, and calculate real-time averages, variances and other statistics through sliding window technology to generate instant environmental status indicators (such as soil moisture index and temperature fluctuations); When the sound monitoring module detects an abnormal warning, according to the time and GPS information when the sound monitoring module detects the abnormal warning, the statistical indicators in the corresponding time window and area are extracted from the environmental sensor data, and the environmental data in the statistical indicators are used to determine whether the abnormality is related to the actual environmental conditions (such as drought and high temperature), thereby providing auxiliary information to verify the reliability of the abnormal warning detected by the sound monitoring module, and to determine whether the environmental indicators deviate significantly from the normal range during the warning period, to prevent misjudgment due to local acoustic fluctuations, thereby verifying the reliability of the sound warning, and generating a joint score through weighted synthesis to ensure that only truly abnormal areas trigger subsequent drone collection tasks.

[0007] Without environmental sensor data, false alarms may be triggered or real abnormal areas may be missed due to sound information alone, affecting the accuracy of the drone mission.

[0008] Furthermore, the method for judging whether the environmental indicators significantly deviate from the normal range during the warning period is as follows: Compare the sound activity information corresponding to the warning time period with the historical data to determine whether the current sound activity is significantly lower than the normal level or whether there is abnormal noise; and give a sound abnormality warning score: Utilize long-term monitoring data and obtain the normal range and average level of each environmental parameter, and generate a standard value range for each parameter; When the real-time collected environmental data deviates significantly from the historical normal range, the environmental parameters are considered abnormal. The system will convert these deviations into an environmental anomaly score to show the degree of environmental anomaly; The weights of the sound anomaly warning score and the environmental anomaly score are obtained respectively, and a weighted combination calculation is performed according to the sound anomaly warning score, the environmental anomaly score and their respective weights to obtain the comprehensive environmental score. The comprehensive environmental score is compared with a preset threshold. If it exceeds the threshold, it is confirmed that there is a real ecological anomaly in the area, and the drone image acquisition task is triggered.

[0009] Based on the sound warning information (such as a sudden drop in animal activity or abnormal noise marks in a certain area) and environmental sensor data (showing drought, temperature anomalies, etc.), and by integrating the spatiotemporal heat map generated by the sound data and environmental data, the specific GPS range of the abnormal hotspot area is identified, which is the target area for drone image acquisition; The UAV task scheduling system automatically plans the flight path, sets the flight altitude, speed and coverage, and refers to meteorological data to ensure flight safety; Obtain image information collected by drones; The collected images are geometrically corrected, orthorectified and stitched to ensure that the images are aligned with the geographic coordinate system; vegetation indexes (such as NDVI), texture features and target detection information (such as diseased areas) are further extracted.

[0010] Data from sound monitoring, environmental sensors and drone imagery are unified in format, time stamped and georeferenced to ensure that the data is on the same basis for comparison and fusion.

[0011] Construct a multimodal feature vector for each abnormal area, including sound data (animal activity intensity, abnormal warning score, MFCC statistical indicators), environmental data (temperature and humidity, soil indicators and other statistical data) and image data (vegetation coverage, NDVI, disease proportion, etc.); For time series data (such as temperature, humidity, and sound activity index), Kalman filtering is used for spatiotemporal smoothing and prediction to reduce random noise interference.

[0012] Model the uncertainty of different data sources and obtain the dependencies between the data through conditional probability calculation to generate more stable ecological health assessment indicators; For example, when integrating sound monitoring and ground environmental data, the Bayesian network can be used for modeling, and the dependency relationship between sound monitoring and the ground environment can be obtained through conditional probability calculation, thereby generating more stable ecological health assessment indicators.

[0013] Construct a fusion model, take each modal data as input, capture the spatiotemporal relationship through convolutional layer, fully connected layer and LSTM layer, realize nonlinear fusion, and output a comprehensive ecological health index; The output results of each model are integrated into the final fusion index using a weighted average or voting mechanism.

[0014] Perform anomaly detection on the integrated ecological health indicators (e.g., using isolation forests or autoencoders) and generate risk warnings when the indicators exceed the preset normal range; Use time series models such as ARIMA or LSTM to predict future ecological trends and provide reference for long-term ecological monitoring and restoration strategy formulation; Combining the expert rule base and decision support model, the comprehensive indicators, anomaly detection and trend prediction results are converted into specific restoration suggestions (such as increasing inspections, adjusting irrigation, replanting local species, etc.) and on-site operation instructions, and heat maps, trend curves and risk assessment reports of abnormal areas are generated.

[0015] In a second aspect, the present application also provides a forest ecological monitoring system based on sound analysis, wherein the system comprises: Sound monitoring module: realizes sound classification and anomaly detection through continuous recording, energy detection, spectrum conversion, noise reduction, MFCC feature extraction and deep learning model, and generates abnormal warning data; Environmental sensor data acquisition and processing module: Through real-time environmental data collection, automatic calibration, statistical analysis and context verification, it provides background information for abnormal sound warning and ensures the reliability of warning; UAV image acquisition module: Determine the target area based on sound and environmental warnings, automatically plan flight missions to collect high-resolution images, and perform image correction, stitching, and ecological index extraction; Multi-source data fusion and intelligent decision support module: standardize and align various data sources in time and space to construct a multimodal feature vector, realize data fusion through Kalman filtering, Bayesian reasoning and multi-input neural network, generate comprehensive ecological health indicators, and output ecological restoration suggestions and operation instructions through anomaly detection and time series prediction.

[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application uses sound monitoring to achieve 24-hour continuous recording and provide millisecond-level time resolution, enabling it to capture the dynamic changes of animal activities and sudden anomalies in the ecosystem in real time. The early warning based on sound data can detect ecological anomalies (such as a sudden drop in animal activity or the appearance of abnormal noise) in a very short time, thereby triggering subsequent drone missions in a timely manner, greatly improving the real-time nature of monitoring.

[0017] This application improves sound monitoring and can directly capture animal sound information, provide intuitive feedback on the biodiversity and ecological vitality in the area, and through the classification and statistics of animal calls, it can analyze the activity status of different species and then judge the health status of the ecosystem. It is difficult to capture such biological behavior dynamics by relying solely on environmental sensors or remote sensing images. That is, this application can achieve a more comprehensive forest monitoring effect.

[0018] This application mainly monitors the forest environment based on sound monitoring. It has a low dependence on weather conditions. Unlike remote sensing images, which are easily affected by clouds, rain and snow, it can operate stably under various weather conditions.

[0019] Moreover, this method can further reduce the risk of misjudgment caused by local noise fluctuations and improve the accuracy of overall early warning by combining environmental sensor data to contextually verify sound data. Compared with relying solely on a combination of remote sensing and environmental data, this method provides direct feedback on biological activities through sound, making monitoring more comprehensive.

[0020] Finally, this method improves the fusion of sound data with environmental sensors and drone image data to obtain comprehensive ecological information from biological behavior to environmental status to spatial details. This multi-source data fusion can generate more accurate comprehensive ecological health indicators and provide scientific decision-making support for ecological anomaly detection and ecological restoration.

[0021] By combining anomaly detection, time series prediction and expert rule systems, this method can automatically generate restoration recommendations and on-site operation instructions to help managers adjust ecological restoration strategies in a timely manner, which is often difficult to achieve in solutions that rely solely on remote sensing or environmental sensing. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A schematic diagram of a forest ecological monitoring method based on sound analysis for this application; Figure 2 This is a flow chart of step S100 in a forest ecological monitoring method based on sound analysis of the present application; Figure 3 This is a flow chart of step S400 in a forest ecological monitoring method based on sound analysis of the present application; Figure 4This is a structural schematic diagram of a forest ecological monitoring system based on sound analysis in this application.

[0023] Explanation of the reference numerals: sound monitoring module 10 , environmental sensor data acquisition and processing module 20 , drone image acquisition module 30 , multi-source data fusion and intelligent decision support module 40 . DETAILED DESCRIPTION

[0024] This application provides a forest ecological monitoring method and system based on sound analysis. It solves the technical problems of incomplete forest monitoring and poor real-time performance in the prior art. It uses sound monitoring to achieve high-frequency and real-time biological activity monitoring, and then combines environmental sensors and drone image data to form multimodal data fusion. It can not only make up for the shortcomings of remote sensing and environmental sensor solutions in terms of time resolution and biological information capture, but also provide more comprehensive, accurate and intelligent ecological health assessment and restoration decision support. These advantages make this method have obvious technical advantages in the field of forest ecological monitoring and restoration.

[0025] Embodiment 1 Please see attached Figure 1 The present application provides a forest ecological monitoring method based on sound analysis, wherein the method is applied to a forest ecological monitoring system based on sound analysis, and the method specifically comprises the following steps: S100: Generates abnormal warning data through continuous recording, energy detection, spectrum conversion, noise reduction, MFCC feature extraction, sound classification and anomaly detection.

[0026] S110 obtains real-time continuous recordings of the forest monitoring area and embeds timestamps and geographic location information in the recordings to ensure that all recording data has accurate time and space identification; Specifically, it deploys a highly sensitive and weather-resistant microphone array in the target forest area to ensure coverage of the entire monitoring area, and regularly calibrates the equipment to ensure stable recording quality. It implements 24-hour continuous recording to capture all acoustic activities in the area, including animal calls, environmental noise, and other possible abnormal sound sources. During the recording process, each audio data segment is automatically bound to a precise timestamp and GPS location information to ensure that the temporal and spatial context of data collection can be accurately restored during subsequent analysis. S120 uses short-time energy calculation to remove the silent part in the recording data, retaining only the time period containing obvious acoustic activities, and dynamically sets the energy threshold to retain the valid signal to obtain valid audio data. Its dynamic energy threshold can be automatically adjusted according to the ambient noise level to prevent missing weak signals and ensure that weak signals will not be missed; Specifically: it processes the continuous recording signal in frames (for example, 20 milliseconds per frame) and calculates the energy value of each frame to quantify the activity level of each audio segment. It then automatically removes low-energy, basically silent frames based on the preset energy threshold, retaining only fragments containing obvious acoustic activity, thereby reducing the amount of subsequent processing data. It also dynamically adjusts the energy threshold based on the real-time environmental noise level to ensure that weak signals can be captured even in a noisy environment, preventing valuable information from being missed due to an excessively high fixed threshold.

[0027] S130 converts valid audio data into a time-frequency matrix through short-time Fourier transform (STFT) to analyze the energy distribution of different frequency bands, adjusts filter parameters in combination with real-time meteorological data (such as wind speed and rainfall), and performs dynamic noise reduction on specific frequency bands (such as low-frequency wind sound). It then estimates the noise spectrum by detecting silent segments, and then subtracts the estimated noise spectrum from the original spectrum to obtain the noise-reduced spectrum. Specifically: This step performs STFT conversion on the effective audio data filtered by S120 to generate a time-frequency matrix, revealing the energy distribution of each frequency band at each time point, and then uses real-time meteorological data (such as wind speed and rainfall conditions) as an auxiliary to automatically adjust the filter parameters according to the current environmental conditions, implement dynamic attenuation on specific frequency bands (such as low-frequency wind sound), reduce environmental noise interference, and estimate the spectral characteristics of background noise by detecting silent segments, and then subtract the estimated noise spectrum from the original spectrum to obtain clearer denoised spectrum data, laying a good foundation for subsequent feature extraction; S140 uses Mel filter bank transform and DCT to extract MFCC features from the denoised spectrum, generating standardized feature vectors to replace the original audio data for uploading, thereby significantly reducing the amount of data; Specifically: This step uses a set of Mel filters to convert the denoised spectrum data obtained in S130 to the Mel scale, and applies discrete cosine transform (DCT) to the Mel-filtered data to extract the Mel-frequency cepstrum coefficients to form a fixed-dimensional MFCC feature vector. The generated MFCC feature vector is normalized to reduce the variation caused by recording conditions and equipment differences, thereby ensuring data consistency and significantly reducing the data volume, which is convenient for subsequent transmission and processing.

[0028] S150 uses pre-trained CNN and LSTM to perform deep learning classification on MFCC features, and sets statistical thresholds based on historical data to detect anomalies, and finally generates output data containing classification results, activity intensity, and anomaly warning information; Specifically: This step inputs the standardized MFCC feature vector into the pre-trained convolutional neural network (CNN) to extract spatiotemporal features, and then combines the long short-term memory network (LSTM or bidirectional LSTM) to capture the time dependency in the audio sequence, classify and identify the sound (for example, distinguish between "birdsong", "insect sounds", etc.), establish a statistical baseline based on historical data, and set the animal activity level and acoustic index range under normal conditions. When the average frequency or total energy of animal calls in a certain area suddenly decreases, or abnormal noise (such as mechanical sounds, illegal activity sounds) suddenly increases, an abnormal warning is triggered and warning data is generated; Sound classification results (such as animal species identification, illegal human activity identification, such as illegal logging, fire warning, etc.); Activity intensity index (reflecting sound energy or frequency changes); Anomaly detection tagging (recording the time of anomaly occurrence, regional GPS location and warning level); These data will serve as an important basis for subsequent data fusion, drone mission triggering and ecological anomaly verification.

[0029] S200 provides environmental context for sound data through real-time data collection, automatic calibration, statistical analysis and contextual verification to ensure the accuracy of abnormal warnings. The specific steps are as follows: S210 obtains environmental sensor data from the forest monitoring area and ensures that the data has clear time and space identification; Specifically: In the forest monitoring area, select a variety of environmental sensors according to monitoring needs, such as temperature sensors, humidity sensors, soil pH sensors, light sensors and CO2 sensors, so that the sensors should be evenly distributed in the target area. At the same time, the deployment density can be increased in key areas to ensure sufficient data coverage, and each sensor node is equipped with a GPS module and clock so that the collected data automatically carries accurate geographic location information and time stamps, which is convenient for subsequent spatiotemporal alignment and analysis.

[0030] S220 performs preliminary verification, noise smoothing (such as using sliding average or exponential smoothing) and format unification on the collected raw environmental data to ensure data accuracy; Specifically: it first performs a rationality check on the environmental data sampled each time to ensure that the values ​​are within the expected range (such as temperature, humidity, pH value, etc.); for readings that are obviously beyond the normal range or have sudden changes, they are marked, filtered or temporarily stored for backup to avoid affecting the overall data quality; continuous sampling data is processed using algorithms such as sliding average or exponential smoothing to eliminate short-term fluctuations and random noise, making the data curve smoother and more stable; this step helps reduce spike data caused by instantaneous errors in equipment or external interference, ensuring that the data reflects the true environmental status; finally, all sensor output data is converted into a unified data format (such as JSON or CSV format), and each data record is ensured to contain unified fields, such as timestamp, sensor ID, GPS coordinates, and the values ​​of various environmental parameters.

[0031] The S230 uses historical data and high-precision reference sensors to build a regression model, automatically calibrate sensor outputs, eliminate errors caused by long-term drift, and calculate real-time averages, variances and other statistics through sliding window technology to generate instant environmental status indicators (such as soil moisture index and temperature fluctuations); Specifically: This step uses long-term monitoring data to establish a historical baseline for each environmental parameter, determine the average level and fluctuation range under normal conditions, and identify possible drift trends or systematic deviations of the equipment by comparing the historical records of each sensor. High-precision reference sensors are deployed in some areas as standard data sources. Using the differences between these data and data collected by other sensors, a regression model is established to automatically calculate the correction coefficient so that it can eliminate errors caused by long-term drift or equipment aging, ensuring the accuracy of real-time data. The sliding window technology is used to perform real-time statistics on the calibrated data, and calculate statistical quantities such as the average, standard deviation, maximum and minimum values. These statistical indicators can reflect the environmental status, such as soil moisture index, temperature fluctuations, etc., to form an instant environmental status report, providing a basis for contextual verification of sound data.

[0032] When the S240 sound monitoring module detects an abnormal warning, according to the time and GPS information when the sound monitoring module detects the abnormal warning, the statistical indicators in the corresponding time window and area are extracted from the environmental sensor data, and the environmental data in the statistical indicators are used to determine whether the abnormality is related to the actual environmental conditions (such as drought, high temperature), thereby providing auxiliary information to verify the reliability of the abnormal warning detected by the sound monitoring module, to determine whether the environmental indicators significantly deviate from the normal range during the warning period, to prevent misjudgment due to local acoustic fluctuations, thereby verifying the reliability of the sound warning, and generating a joint score through weighted synthesis to ensure that only truly abnormal areas trigger subsequent drone collection tasks; The method for judging whether the environmental indicators significantly deviate from the normal range during the warning period is as follows: Compare the sound activity information corresponding to the warning time period with the historical data to determine whether the current sound activity is significantly lower than the normal level or whether there is abnormal noise; and give a sound abnormality warning score: Utilize long-term monitoring data and obtain the normal range and average level of each environmental parameter, and generate a standard value range for each parameter; When the real-time collected environmental data deviates significantly from the historical normal range, the environmental parameters are considered abnormal. The system will convert these deviations into an environmental anomaly score to show the degree of environmental anomaly; The weights of the sound anomaly warning score and the environmental anomaly score are obtained respectively, and a weighted combination calculation is performed according to the sound anomaly warning score, the environmental anomaly score and their respective weights to obtain the comprehensive environmental score. The comprehensive environmental score is compared with a preset threshold. If it exceeds the threshold, it is confirmed that there is a real ecological anomaly in the area, and the drone image acquisition task is triggered.

[0033] Specifically: this step extracts the specific time period when the warning occurs based on the abnormal warning information generated by the sound monitoring module, and determines the corresponding GPS coordinate range to ensure that the statistical indicators subsequently extracted from the environmental data match the sound warning in time and space. During the warning period, all environmental parameter records of the area are extracted from the environmental sensor data, such as temperature, humidity, soil pH, light intensity, etc., and the data within the time period are statistically analyzed using the sliding window technology to calculate the average value, standard deviation, maximum value and minimum value of each parameter, so that the normal range and average level of each environmental parameter can be established using long-term monitoring data. Then, the real-time collected data is compared with the historical baseline. If the comparison result shows that one or more environmental parameters deviate significantly from the historical normal level (such as a sharp drop in humidity or an abnormal increase in temperature) during the warning period, it is considered that the environmental parameter is abnormal. According to the degree of deviation of each parameter, its abnormality is determined. For example, the greater the deviation from the normal range, the higher the degree of abnormality. Through expert experience and historical data, the abnormal classification rules of each parameter are defined, such as slight, moderate or severe abnormalities, and the degree of abnormality of each parameter is normalized to form an overall environmental abnormality score to reflect the abnormal situation of the current regional environmental status. At the same time, the system will review the abnormal warning score output by the sound monitoring module (the score reflects the sudden drop in animal activity or the intensity of abnormal noise) to ensure that the sound data has been preprocessed and classified to form a preliminary abnormality score. Then, based on historical data and field verification results, the relative weights of the sound abnormality warning score and the environmental abnormality score are preliminarily set. For example, equal weights can be used initially, and then adjusted according to the actual prediction effect through regression analysis and cross-validation, so that the contribution of the two data sources to the comprehensive score is more reasonable. Using the preset weights, the sound abnormality warning score and the environmental abnormality score are weighted and combined to form a comprehensive score to determine whether there is a real ecological abnormality in the current area. If it exceeds the preset threshold, the abnormality is confirmed and the subsequent drone image acquisition task is triggered; The joint score, statistical indicators of various environmental parameters and sound warning scores are recorded in the database for long-term analysis and model optimization. Based on subsequent drone image acquisition and on-site verification results, the warning threshold and weight distribution are continuously adjusted to ensure that the system can adaptively improve the accuracy of abnormal warnings.

[0034] S300 determines the target area based on the warning information, plans the flight mission to collect high-resolution images, and performs image preprocessing and ecological index extraction. The steps are as follows: S310 uses sound warning information (such as a sudden drop in animal activity or abnormal noise markers in a certain area) and environmental sensor data (indicating drought, temperature anomalies, etc.), and integrates the spatiotemporal heat map generated by sound data and environmental data to identify the specific GPS range of abnormal hot spots, which are the target areas for drone image acquisition; Specifically, this step is based on the statistical indicators of abnormal warning information detected by the sound monitoring module (such as a sudden drop in animal activity or abnormal noise marks in a certain area) and environmental sensor data (such as drought, high temperature or humidity abnormalities). The system uses the existing spatiotemporal heat map to overlay these data on the map to form an image that comprehensively displays the acoustic activity and environmental status in the area, and identifies the specific distribution and range of abnormal hot spots on the superimposed heat map through data visualization methods. The system automatically extracts the GPS coordinates of these areas and marks the target areas that need to focus on collecting images. If there is a suspected false alarm, the system can perform secondary verification by comparing historical data; at the same time, environmental data is used as context information to verify the sound anomaly, ensuring that the target area is determined with high accuracy to prevent false alarms caused by single noise fluctuations.

[0035] The S320 automatically plans the flight path through the drone mission scheduling system, sets the flight altitude, speed and coverage, and refers to meteorological data to ensure flight safety; Specifically, in this step, according to the GPS range of the target area obtained in S310, the system uses the automatic task scheduling module to formulate a UAV task plan and determine the flight altitude, flight speed and image acquisition frequency.

[0036] It can obtain higher image resolution at lower flight altitudes, but the coverage area is smaller; In flight path planning, it is also necessary to ensure that there is sufficient overlap between images (usually 60% to 80%) to facilitate subsequent image stitching; By using GIS data and planning algorithms (such as the shortest path algorithm or genetic algorithm) to generate the best flight route, it is ensured that the UAV can efficiently cover the target abnormal area and avoid terrain obstacles and high-risk areas as much as possible. In mission planning, the system also refers to public meteorological data to ensure that weather conditions are suitable during the flight (avoid severe weather such as strong winds and heavy rains), thereby ensuring flight safety and the quality of collected images.

[0037] S330 obtains image information collected by drones. Its drones can collect RGB images and multispectral images at the same time, providing data for subsequent ecological indicator extraction; Specifically, its drone automatically takes off according to the planned mission and collects images in the target area according to the predetermined flight route. The drone is equipped with a high-resolution RGB camera and a multispectral camera, which can collect different types of image data at the same time to meet the needs of subsequent vegetation index and other ecological indicator extraction. Each collected image is embedded with a timestamp and GPS location information to ensure that the image data can be aligned with the sound and environmental data in time and space, and the drone transmits the image data in real time or in batches to the ground control station or central platform to ensure that the data is available for subsequent processing in a short time. During the collection process, the drone stores the image data in a local storage device and combines it with the real-time transmission data to ensure data security and integrity and prevent data loss due to transmission interruptions.

[0038] S340 performs geometric correction, orthorectification and image stitching on the collected images to ensure that the images are aligned with the geographic coordinate system; it further extracts vegetation indexes (such as NDVI), texture features and target detection information (such as diseased areas).

[0039] Specifically: This step performs geometric correction and orthorectification on the collected original image data, uses the GPS / IMU data and ground control points (GCP) of the drone to align the image with the actual geographic coordinate system, eliminates geometric distortion caused by flight altitude and terrain undulations, and forms a continuous regional orthophoto by splicing multiple images covering the same area, ensuring seamless spatial information and providing a complete image basis for subsequent vegetation and ecological analysis; Use multispectral image data to calculate indicators such as NDVI (Normalized Difference Vegetation Index) to assess vegetation coverage and growth in the area; Use image processing and deep learning methods (such as CNN models) to analyze image textures and identify abnormal features such as diseased areas, dead vegetation, or signs of illegal logging; The extracted ecological indicator data (such as NDVI value, vegetation coverage, area of ​​abnormal target region, etc.) will serve as important input for subsequent data fusion.

[0040] S400 standardizes sound, environment and image data, constructs multimodal feature vectors, and uses fusion algorithms to generate comprehensive ecological health indicators. It outputs ecological restoration suggestions and operation instructions through anomaly detection and trend prediction. The specific steps are as follows: The S410 unifies the data from sound monitoring, environmental sensors and drone images into a unified format, timestamp and geographic coordinates, ensuring that the data is on the same basis for comparison and fusion.

[0041] Specifically: This step parses the data from sound monitoring, environmental sensors, and drone images into a unified format (such as JSON or CSV format), ensures that each record contains standard fields (such as timestamp, GPS coordinates, data type, value, etc.), and performs time correction on each data source to ensure that all data is recorded according to a unified time standard (such as UTC), uses interpolation methods to align data with different sampling frequencies, and then uses the GPS information of the sensor and the geographic alignment of the drone image to map all data to the same geographic coordinate system to ensure that the data can be accurately compared in space.

[0042] S420 constructs a multimodal feature vector for each abnormal area, including sound data (animal activity intensity, abnormal warning score, MFCC statistical indicators), environmental data (temperature and humidity, soil indicators and other statistical data) and image data (vegetation coverage, NDVI, disease proportion and other spatial details); For each abnormal area, the above three types of data are aligned in time and space to construct a multimodal feature vector, so that the vector can comprehensively describe the ecological status of the area and provide input for subsequent fusion.

[0043] S430 uses Kalman filtering to perform spatiotemporal smoothing and prediction on time series data (such as temperature, humidity, and sound activity index) to reduce random noise interference.

[0044] Model the uncertainty of different data sources and obtain the dependencies between the data through conditional probability calculation to generate more stable ecological health assessment indicators; For example, when integrating sound monitoring and ground environmental data, the Bayesian network can be used for modeling, and the dependency relationship between sound monitoring and the ground environment can be obtained through conditional probability calculation, thereby generating more stable ecological health assessment indicators.

[0045] Construct a fusion model, take each modal data as input, capture the spatiotemporal relationship through convolutional layer, fully connected layer and LSTM layer, realize nonlinear fusion, and output a comprehensive ecological health index; The output results of each model are integrated into the final fusion index using a weighted average or voting mechanism.

[0046] Specifically, this step uses the Kalman filter method to smooth and predict various time series data (such as temperature, humidity, sound activity indicators, etc.) to reduce random noise interference, make the data more stable and reliable, and use Bayesian reasoning to model the uncertainty of different data sources, and obtain the dependency relationship between the data through conditional probability calculation, thereby generating more stable ecological health assessment indicators; By constructing a multi-input neural network model, the feature vectors of sound, environment and image data are taken as input, and the convolutional layer, fully connected layer and LSTM layer are used to capture the nonlinear spatiotemporal relationship between the data, and the comprehensive ecological health index is output. The model can be pre-trained on a public data set and then fine-tuned using local data to improve adaptability to specific areas. Finally, the output results of the Kalman filter, Bayesian reasoning and deep neural network model are weighted averaged or integrated using a voting mechanism to generate the final comprehensive ecological health index.

[0047] S440 performs anomaly detection on the fused ecological health indicators (e.g., using isolation forests or autoencoders) and generates risk warnings when the indicators exceed the preset normal range; Use time series models such as ARIMA or LSTM to predict future ecological trends and generate trend forecast reports to provide reference for long-term ecological monitoring and restoration strategy formulation; Specifically, this step uses anomaly detection algorithms (such as isolated forests, autoencoders, etc.) to monitor whether the integrated ecological health indicators exceed the normal fluctuation range in real time. When the indicators exceed the set threshold, risk warnings are automatically generated, indicating that there are potential ecological anomalies in the area; Use the ARIMA model or LSTM model to perform time series modeling on the fused indicators, predict the changing trend of ecological health indicators in the future, generate a trend forecast report, provide a reference for subsequent ecological restoration decisions, and help determine whether the ecological status is likely to further deteriorate or gradually recover.

[0048] S450 combines an expert rule base and decision support model to convert comprehensive indicators, anomaly detection and trend prediction results into specific restoration suggestions (such as increasing inspections, adjusting irrigation, replanting native species, etc.) and on-site operation instructions; The results are finally presented through intuitive dashboards and reports, including heat maps of abnormal areas, trend curves and risk assessment reports.

[0049] Specifically: This step combines the fusion indicators, anomaly detection results, and trend prediction results with the preset expert rule base to formulate a specific ecological restoration strategy. For example: If the composite index continues to be below the normal range and the forecast shows further deterioration, it is recommended to increase inspections, adjust irrigation measures or carry out replanting.

[0050] The decision model can use a rule engine or a decision tree based on machine learning to automatically generate operational instructions, so that it can ultimately generate intuitive dashboards and reports, including heat maps of abnormal areas, trend prediction curves, risk assessment reports, and specific recovery recommendations and operational instructions for on-site managers to refer to for decision-making.

[0051] Embodiment 2 Based on the same inventive concept as the forest ecology monitoring method based on sound analysis in the aforementioned embodiment, the present invention also provides a forest ecology monitoring system based on sound analysis, the system comprising: Sound monitoring module 10: realizes sound classification and anomaly detection through continuous recording, energy detection, spectrum conversion, noise reduction, MFCC feature extraction and deep learning model, and generates abnormal warning data; The sound monitoring module 10 comprises: The real-time recording and data labeling unit 11 is used to automatically embed accurate timestamps and GPS location information during the recording process to ensure that the data has time and space identification.

[0052] A silence elimination and energy detection unit 12 is used to calculate the short-time energy of the recording data and divide the audio signal into short-time frames; Low-energy frames are removed based on a set dynamic energy threshold (automatically adjusted according to the ambient noise level) to obtain valid audio segments and energy statistics for each segment.

[0053] The spectrum conversion and noise reduction processing unit 13 is used to obtain a time-frequency matrix by applying short-time Fourier transform (STFT) to the effective audio, and dynamically adjust the filter parameters using real-time meteorological data (such as wind speed and rainfall) to attenuate low-frequency wind sounds, etc., estimate the noise spectrum for detecting silent segments, and then subtract the noise spectrum from the original spectrum to obtain the denoised spectrum data, so as to obtain the denoised and spectrally converted audio data, which is convenient for subsequent feature extraction.

[0054] The MFCC feature extraction module 14 is used to apply the Mel filter bank transform to convert the spectrum to the Mel scale, calculate the Mel frequency cepstral coefficient (MFCC) through discrete cosine transform (DCT), and then normalize the MFCC vector to eliminate the influence caused by the recording conditions and equipment differences, so as to obtain a standardized MFCC feature vector sequence for subsequent sound classification and anomaly detection, while greatly reducing the amount of data.

[0055] A sound classification and anomaly detection unit 15, for extracting audio features using a pre-trained convolutional neural network (CNN) and combining a long short-term memory network (LSTM or bidirectional LSTM) to capture time dependencies in audio sequences for classification and recognition (e.g., "birdsong", "insect calls", etc.); Based on historical data and statistical thresholds, when a significant decrease in sound activity or a sudden increase in abnormal noise is detected, an abnormal warning is triggered to obtain data including sound classification results, activity intensity indicators, anomaly detection marks, GPS location of the abnormal area, and warning level.

[0056] Environmental sensor data acquisition and processing module 20: provides background information for abnormal sound warning through real-time environmental data acquisition, automatic calibration, statistical analysis and context verification to ensure the reliability of warning; The environmental sensor data acquisition and processing module 20 includes: Environmental data collection unit 21, used to collect data from sensors such as temperature, humidity, pH, light and CO2 deployed in the forest, and add timestamp and GPS information to each sampled data; The preliminary check and noise smoothing unit 22 is used to check the rationality of the data, filter abnormal or erroneous data, and smooth the data fluctuations using a sliding average or exponential smoothing method to convert the data into a unified format and ensure that all data have standard fields to obtain pre-processed environmental data records.

[0057] The automatic calibration and statistical analysis unit 23 is used to establish a historical baseline, use a regression model and a high-precision reference sensor for automatic calibration, and use a sliding window to calculate statistical indicators such as the mean value and standard deviation of each parameter, generate real-time status data such as soil moisture index and temperature fluctuation, so as to obtain calibrated environmental data and real-time statistical indicators.

[0058] The joint warning score generation module 24 is used to extract statistical indicators of the corresponding time window in the environmental data according to the time and area of ​​the sound warning and compare the real-time environmental indicators with the historical baseline to determine whether there is a significant deviation. If a significant deviation occurs, the degree of deviation is converted into an environmental anomaly score, and the sound anomaly warning score and the environmental anomaly score are weighted and combined according to the preset weights or the weights obtained by historical analysis to generate a joint warning score for verifying the reliability of the sound anomaly warning and guiding the triggering of the UAV mission. If the warning score exceeds the preset threshold, it is confirmed that there is a real ecological anomaly in the area, thereby triggering the subsequent UAV image acquisition task.

[0059] UAV image acquisition module 30: determines the target area based on sound and environmental warnings, automatically plans flight missions to acquire high-resolution images, and performs image correction, stitching, and ecological index extraction; The drone image acquisition module 30 includes: The target area determination and hotspot positioning unit 31 is used to integrate the sound warning data and the environmental sensor data, generate an abnormal hotspot map, and automatically extract the GPS information of the abnormal area to determine the target area where the image needs to be collected to obtain the geographic coordinates and range of the abnormal hotspot area.

[0060] The UAV mission planning and flight path generation unit 32 is used to utilize GIS data and GPS information of the target area to plan the optimal flight path through the shortest path algorithm or genetic algorithm, and set the flight altitude and overlap rate according to the image resolution requirements, while referring to the weather forecast data to ensure flight safety, so as to obtain detailed flight mission instructions, including the flight path, parameter settings and expected coverage area.

[0061] The image acquisition and data transmission unit 33 is used to obtain the RGB image and multispectral image information collected when the UAV automatically performs the flight mission according to the mission instruction, and embed the timestamp and GPS location information for each image, and upload the collected data to the central platform through real-time transmission or offline storage to obtain the collected original image data file and its metadata.

[0062] Image preprocessing and ecological index extraction 34 is used to perform geometric correction and orthorectification on the collected images to ensure that the images are aligned with the geographic coordinate system, and use image stitching algorithms to generate regional seamless orthophotos. Multispectral analysis is then used to calculate NDVI, extract texture features and detect targets (such as diseased areas), and generate ecological indicators such as vegetation coverage and disease proportion.

[0063] Multi-source data fusion and intelligent decision support module 40: standardize and align all data sources in time and space to construct a multimodal feature vector, realize data fusion through Kalman filtering, Bayesian reasoning and multi-input neural network, generate comprehensive ecological health indicators, and output ecological restoration suggestions and operation instructions through anomaly detection and time series prediction.

[0064] The multi-source data fusion and intelligent decision support module 40 includes: The data standardization and spatiotemporal alignment unit 41 is used to parse the data obtained by each module, convert it into a unified data format, and perform time interpolation on different data sampling frequencies. Finally, GPS and GIS tools are used to map all data to a unified coordinate system to obtain a standardized and spatiotemporally aligned multi-source data set.

[0065] The multimodal feature vector construction unit 42 is used to construct a feature vector reflecting the ecological status of each abnormal area, integrating three types of data: sound, environment and image, wherein: Sound data: extract animal activity intensity, abnormal warning score, and MFCC statistical indicators; Environmental data: using temperature, humidity, soil indicators and other statistical data; Image data: extract vegetation coverage, NDVI, texture features, disease detection results, etc.

[0066] After aligning these data in time and space, they are integrated into a multimodal feature vector.

[0067] The data fusion and model building unit 43 is used to smooth and predict time series data (such as temperature, humidity and sound activity) using Kalman filtering, and to model the uncertainty of each data source using Bayesian reasoning, and to calculate the dependency between each data through conditional probability, thereby generating a more stable ecological health assessment indicator.

[0068] A multi-input neural network model is constructed, and the feature vectors of each modality are taken as input. The nonlinear spatiotemporal relationship is captured through convolutional layers, fully connected layers and LSTM, and a comprehensive ecological health index is output. The output results of each model are integrated using a weighted average or voting mechanism to ensure the stability of the fusion index, so as to obtain a comprehensive ecological health index for evaluating the regional ecological status.

[0069] The anomaly detection and trend prediction unit 44 is used to detect anomalies of the integrated ecological health indicators after fusion, monitor the comprehensive indicators in real time by using anomaly detection algorithms (such as isolation forests or autoencoders), determine whether they exceed the normal fluctuation range, and apply ARIMA or LSTM models to perform time series analysis and prediction on the comprehensive indicators to generate a trend prediction report.

[0070] Decision support and restoration recommendation generation is used to combine comprehensive ecological health indicators, anomaly detection and trend prediction results with expert rule bases and decision support models to analyze comprehensive indicator data and prediction results, and generate restoration recommendations based on risk assessment (such as increasing inspections, adjusting irrigation, replanting local species, etc.) to obtain specific ecological restoration recommendations and operational instructions, as well as charts, heat maps and risk assessment reports.

[0071] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0072] This specification and the drawings are merely exemplary illustrations of the present application. If modifications and variations of the present invention fall within the scope of the present invention and its equivalent technology, the present invention is intended to include these modifications and variations.

Claims

1. A forest ecological monitoring method based on sound analysis, characterized in that: Obtain real-time continuous recordings of forest monitoring areas and embed timestamps and geographic location information in the recordings to ensure that all recording data has accurate time and space identification; Use short-time energy calculation to remove the silent part in the recorded data, and dynamically set the energy threshold to retain the valid signal to obtain valid audio data; The effective audio data is converted into a time-frequency matrix through STFT, and the filter parameters are adjusted in combination with real-time meteorological data to perform dynamic noise reduction processing to obtain the noise-reduced spectrum; Mel filter bank transform and DCT are used to extract MFCC features and generate standardized feature vectors; The pre-trained CNN and LSTM are used to perform deep learning classification on MFCC features, and statistical thresholds are set based on historical data to detect anomalies, ultimately generating output data containing classification results, activity intensity, and anomaly warning information.

2. A forest ecology monitoring method based on sound analysis according to claim 1, characterized in that: Obtain environmental sensor data from forest monitoring areas and ensure that the data has clear time and space identification; Perform preliminary verification, noise smoothing and format unification on the collected raw data to ensure data accuracy and consistency; Use historical data and high-precision reference sensors to build a regression model, automatically calibrate sensor outputs, and generate real-time environmental status indicators through sliding window statistics; When the sound monitoring module detects an abnormal warning, according to the time and GPS information when the sound monitoring module detects the abnormal warning, the statistical indicators in the corresponding time window and area are extracted from the environmental sensor data. The environmental data in the statistical indicators are used to determine whether the abnormality is related to the actual environmental conditions. This is used to determine whether the environmental indicators deviate significantly from the normal range during the warning period, thereby verifying the reliability of the sound warning and generating a joint score through weighted synthesis to ensure that only truly abnormal areas trigger subsequent drone collection tasks.

3. The forest ecology monitoring method based on sound analysis according to claim 2 is characterized in that: The method for judging whether the environmental indicators significantly deviate from the normal range during the warning period is as follows: Compare the sound activity information corresponding to the warning time period with the historical data to determine whether the current sound activity is significantly lower than the normal level or whether there is abnormal noise; and give a sound abnormality warning score: Utilize long-term monitoring data and obtain the normal range and average level of each environmental parameter, and generate a standard value range for each parameter; When the real-time collected environmental data deviates significantly from the historical normal range, the environmental parameters are considered abnormal, and the degree of deviation is converted into an environmental anomaly score to show the degree of environmental anomaly; The weights of the sound anomaly warning score and the environmental anomaly score are obtained respectively, and a weighted combination calculation is performed according to the sound anomaly warning score, the environmental anomaly score and their respective weights to obtain the comprehensive environmental score. The comprehensive environmental score is compared with a preset threshold. If it exceeds the threshold, it is confirmed that there is a real ecological anomaly in the area, and the drone image acquisition task is triggered.

4. The forest ecology monitoring method based on sound analysis according to claim 1, characterized in that: Based on the sound warning information and environmental sensor data, and by integrating the spatiotemporal heat map generated by the sound data and environmental data, the specific GPS range of the abnormal hotspot area is identified, which is the target area for drone image acquisition; The UAV task scheduling system automatically plans the flight path, sets the flight altitude, speed and coverage, and refers to meteorological data to ensure flight safety; Obtain image information collected by drones; The collected images are geometrically corrected, orthorectified and stitched to ensure that the images are aligned with the geographic coordinate system; vegetation index, texture features and target detection information are further extracted.

5. The forest ecology monitoring method based on sound analysis according to claim 1, characterized in that: Unify data from sound monitoring, environmental sensors and drone images in a unified format, time stamp and geo-coordinate to ensure that the data is on the same basis for comparison and fusion; Construct a multimodal feature vector, including sound data, environmental data, and image data; For time series data, Kalman filtering is used for spatiotemporal smoothing and prediction to reduce random noise interference; Model the uncertainty of different data sources and obtain the dependencies between the data through conditional probability calculation to generate ecological health assessment indicators; Construct a fusion model, take each modal data as input, capture the spatiotemporal relationship through convolutional layer, fully connected layer and LSTM layer, realize nonlinear fusion, and output a comprehensive ecological health index; The output results of each model are integrated into the final fusion index using a weighted average or voting mechanism.

6. The forest ecology monitoring method based on sound analysis according to claim 1, characterized in that: Perform anomaly detection on the integrated ecological health indicators and generate risk warnings when the indicators exceed the preset normal range; Use time series models such as ARIMA or LSTM to predict future ecological trends and provide reference for long-term ecological monitoring and restoration strategy formulation; Combining the expert rule base and decision support model, the comprehensive indicators, anomaly detection and trend prediction results are converted into specific recovery suggestions and on-site operation instructions, and heat maps, trend curves and risk assessment reports of abnormal areas are generated.

7. A forest ecological monitoring system based on sound analysis, characterized in that: The system comprises: Sound monitoring module: realizes sound classification and anomaly detection through continuous recording, energy detection, spectrum conversion, noise reduction, MFCC feature extraction and deep learning model, and generates abnormal warning data; Environmental sensor data acquisition and processing module: Through real-time environmental data collection, automatic calibration, statistical analysis and context verification, it provides background information for abnormal sound warning and ensures the reliability of warning; UAV image acquisition module: Determine the target area based on sound and environmental warnings, automatically plan flight missions to collect high-resolution images, and perform image correction, stitching, and ecological index extraction; Multi-source data fusion and intelligent decision support module: standardize and align various data sources in time and space to construct a multimodal feature vector, realize data fusion through Kalman filtering, Bayesian reasoning and multi-input neural network, generate comprehensive ecological health indicators, and output ecological restoration suggestions and operation instructions through anomaly detection and time series prediction.

Citation Information

Patent Citations

  • Forest monitoring system based on deep learning and sound recognition

    CN116844570A

  • Forest patrol method and system

    CN117636169A

  • Forest fire prevention monitoring and alarming method and device

    CN118803433A

  • Forestry ecological environment monitoring system and method

    CN118840656A

  • Forest grassland fire danger monitoring, early warning and forecasting system and method

    CN118840817A

Cited By

  • Ammunition searching method and system based on unmanned aerial vehicle

    CN120434363A

  • Ammunition search method and system based on drone

    CN120434363B

  • Wild animal and plant species identification method

    CN120470544A

  • Noise distribution visual monitoring method and system based on GIS

    CN120611006A

  • Banking outlet safety inspection data evidence storage statistics platform based on distributed storage

    CN120849511A