Underground grub pest detection system and method based on deep learning multi-mode

Through the deep learning multimodal detection system, the activity of grubs is enhanced by lighting, vibration and humidity adjustment, and combined with sound and carbon dioxide sensors for accurate analysis, solving the problems of traditional detection efficiency and environmental pollution, and achieving efficient and accurate underground grubs monitoring.

CN120597026APending Publication Date: 2025-09-05QINGDAO AGRI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510641443.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional underground grub detection methods are inefficient, labor-intensive and material-intensive, and are prone to damage the ecological environment. Chemical control methods can easily lead to increased pest resistance and environmental pollution.

Method used

A multimodal detection system based on deep learning is adopted to enhance the activity of grubs through lighting, vibration and humidity adjustment, combine sound and carbon dioxide sensor to collect data, and use deep learning models for accurate analysis to output the existence probability, density level and location of grubs.

Benefits of technology

Real-time dynamic monitoring of underground grubs has been achieved, the accuracy and reliability of detection are improved, the risk of crop damage is reduced, and it is in line with the concept of green agriculture development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597026A_ABST
    Figure CN120597026A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of agricultural pest detection, and particularly relates to an underground grub pest detection system and method based on deep learning multi-mode. Firstly, in an area suspected to have underground grub pests, the activity of underground grubs is enhanced in a targeted manner by arranging a specific stimulation module. After the grub activity is enhanced, sound signals and carbon dioxide concentration changes generated by the grub activity are captured by using a data acquisition module, preprocessing and feature extraction are performed, and analysis and processing are performed in combination with a deep learning model, so that the existence condition, the number and the underground rough distribution position of underground grubs are accurately judged. According to the scheme, by actively enhancing the activity of grubs and combining the multi-mode detection and deep learning technology, the accuracy and reliability of detection are effectively improved, insect pests can be warned in advance, an efficient and accurate solution is provided for underground pest prevention and control in agricultural production, the prevention and control efficiency is improved, the prevention and control cost is reduced, and the negative influence on the environment is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of agricultural pest detection, and specifically relates to a subterranean grub pest detection system and method based on deep learning multimodality. Background Art

[0002] As a common and serious agricultural pest, ground grubs cause extensive damage to the roots and stems of crops, severely impacting their growth, development, and ultimately yield. Because they live underground and are highly elusive, traditional detection methods are difficult to effectively monitor and accurately locate.

[0003] Currently, underground pest detection relies primarily on manual excavation or simple soil sampling and analysis. The former is extremely inefficient, consumes significant manpower and material resources, and significantly disturbs the soil ecosystem. The latter is susceptible to sampling point limitations, making it difficult to fully and accurately reflect the distribution and extent of grub damage across an area. Furthermore, while chemical control methods can control pest populations to a certain extent, excessive use of chemical pesticides not only leads to increased pest resistance but also pollutes soil, water sources, and other ecological environments, disrupting the ecological balance. Summary of the Invention

[0004] In response to the shortcomings of the existing technology, the present invention proposes a system and method for detecting underground grub pests based on deep learning multimodality. The system collects the sound of underground grub activity and carbon dioxide concentration signals, optimizes them through a data processing module, and then performs precise analysis through a deep learning model. It can accurately detect the presence, quantity and location of underground grubs, and transmit the detection results to users in a timely manner, providing powerful support for pest and disease monitoring for agricultural production.

[0005] The present invention is implemented by adopting the following technical solution: a subterranean grub pest detection system based on deep learning multimodality, comprising:

[0006] a stimulation module, comprising a light emitting device, a vibration generating device, a temperature regulating device and a humidity regulating device, for emitting weak light of a specific wavelength, generating vibrations and regulating soil temperature and humidity to enhance grub activity;

[0007] A data acquisition module, including a sound sensor and a carbon dioxide concentration sensor, for collecting sound signals and carbon dioxide concentration data in the soil environment in real time;

[0008] The edge computing module includes a data processing module and a deep learning model. The data processing module is used to preprocess and extract features from the collected sound signals and carbon dioxide concentration data; the deep learning model is used to identify the extracted features and then determine the probability of existence and density level of underground grubs.

[0009] The result output module outputs a detection report containing the probability of existence, density level and location of grubs based on the recognition results of the deep learning model, and sends the report to the designated terminal or management system for prevention and control in combination with preset strategies.

[0010] Furthermore, the deep learning model includes an acoustic subnetwork and a carbon dioxide subnetwork, and the acoustic subnetwork and the carbon dioxide subnetwork are connected via a cross-modal attention fusion layer;

[0011] The acoustic subnetwork embeds a low-frequency enhancement convolution block (LEB) in EfficientNet-B4. The shallow network of the low-frequency enhancement convolution block introduces deformable convolution (Conv) to adaptively capture the irregular time-frequency distribution of low-frequency vibrations. The channel attention mechanism (SE Block) of the low-frequency enhancement convolution block dynamically weights key frequency bands and outputs a frequency-time domain feature vector.

[0012] The carbon dioxide sub-network is based on the Transformer-XL architecture, combined with the advantages of bidirectional LSTM, and adds a time series difference feature input layer: input original concentration sequence, first-order difference ΔC, second-order difference Δ 2 C, captures concentration mutation patterns; combined with the position encoding extension module PEM to enhance the modeling ability of long-term dependencies and output a time series feature vector;

[0013] The cross-modal attention fusion layer designs a dynamic weight allocation mechanism to achieve bimodal feature fusion of frequency-time domain feature vectors and time series feature vectors, and ultimately outputs the probability of existence and density level of white grub pests.

[0014] Furthermore, when the cross-modal attention fusion layer performs feature fusion, the following method is specifically adopted:

[0015] (1) Feature splicing and attention calculation:

[0016] The frequency-time domain feature vector and the time series feature vector are concatenated, and the feature weight α is generated through a two-layer fully connected network to dynamically allocate the contribution of the two types of feature vectors;

[0017] (2) Gated nonlinear fusion:

[0018] Introducing the gating weight g=σ(W g Attention) to achieve adaptive fusion, and input the fused feature vector into the fully connected layer for classification, and finally output the probability of pest existence and density level.

[0019] Furthermore, the data processing module processes data in the following manner:

[0020] (1) For sound signals:

[0021] Noise reduction: Adaptive wavelet packet decomposition is used to filter out noise above 5kHz, and a three-layer wavelet noise reduction algorithm is then used to filter out low-frequency environmental noise. A dynamic noise template is generated based on the soil noise library, and spectral subtraction + time-frequency masking technology is used to achieve time-frequency domain noise removal.

[0022] Feature engineering processing: Pre-emphasize the high-frequency components through a high-pass filter, then frame the de-noised sound signal. Fast Fourier transform, Mel filter bank processing, and discrete cosine transform are performed on the framed signal to extract MFCC features and frequency domain features. Time domain features are then calculated and fused to generate a multi-domain feature vector. The location of the sound source is then calculated using a positioning algorithm.

[0023] (2) For the carbon dioxide concentration signal: First, perform baseline calibration, based on the soil respiration dynamic model, eliminate the interference of plant root respiration, then use the Kalman filter algorithm to remove the random noise of the concentration data, generate a 24-hour concentration fluctuation curve, and finally use the second-order difference method Δ 2 C. Kalman filtering and first-order difference method were used to analyze the concentration change trend to quantify the metabolic activity intensity of the grub population and calculate the basal concentration value, peak concentration, concentration change rate, and diurnal fluctuation amplitude time series characteristics.

[0024] The present invention also proposes a method for detecting underground grub pests based on deep learning multimodality, comprising the following steps:

[0025] Step A: Underground environment regulation and grub activity enhancement: The stimulation module is used to stimulate the environment and actively guide the grubs to gather towards the data collection module;

[0026] Step B: Synchronous acquisition and processing of multimodal data:

[0027] Step B1: using a sound sensor and a carbon dioxide concentration sensor to collect sound signals and carbon dioxide concentration data in the soil environment in real time;

[0028] Step B2: performing noise reduction and filtering preprocessing operations on the collected sound signals and carbon dioxide concentration data to extract the time domain and frequency domain characteristics of the signals and characteristic parameters of carbon dioxide concentration changes;

[0029] Step C: Multimodal deep learning model construction and training: Combine the acoustic sub-network and the carbon dioxide sub-network, and fuse them through a cross-modal attention mechanism to achieve dynamic weighting of multi-source features. The fully connected layer outputs the probability of pest presence and density level.

[0030] Step D, Adaptive Feedback Control: Implement a dynamic detection strategy based on the output of the multimodal deep learning model. Determine whether to use conventional detection mode or weak signal enhancement / high-precision mode based on the detection confidence level, and optimize the output data.

[0031] Step E: Detection result processing and decision support: Based on the deep learning model in step D, a detection report containing the probability of existence, density level and location of white grubs is output, and the report is sent to the designated terminal or management system for prevention and control in combination with the preset strategy.

[0032] Furthermore, the step C is specifically implemented in the following manner:

[0033] (1) Frequency domain feature extraction through acoustic sub-network: using an improved version of EfficientNet-B4, embedding low-frequency enhancement convolution blocks: shallow networks introduce deformable convolution Conv to adaptively capture the irregular time-frequency distribution of low-frequency vibrations; the channel attention mechanism dynamically weights key frequency bands and outputs frequency-time domain feature vectors;

[0034] (2) Temporal feature modeling through the carbon dioxide sub-network:

[0035] Based on the Transformer-XL architecture, combined with bidirectional LSTM, a time series difference feature input layer is added, and the original concentration sequence, first-order difference ΔC, and second-order difference Δ 2 C, captures concentration mutation patterns; the position encoding extension module enhances the modeling ability of long-term dependencies and outputs a time series feature vector;

[0036] (3) Combined with the cross-modal attention fusion layer, a dynamic weight allocation mechanism is designed to achieve bimodal feature fusion:

[0037] Feature splicing and attention calculation:

[0038] The frequency-time domain feature vector is concatenated with the time series feature vector, and the feature weight α is generated through a two-layer fully connected network to dynamically allocate the contribution of the two types of features:

[0039] α=Softmax(W2ReLU(W1[h audio ;h CO2 ]))

[0040] Among them, h audio is the frequency-time domain feature vector, is the time series feature vector, W1 and W2 are weight matrices, and Softmax is used to generate a 0-1 feature weight α to achieve dynamic weighted fusion;

[0041] Gated nonlinear fusion:

[0042] Introducing the gating weight g=σ(W g Attention) to achieve adaptive fusion. The specific calculation process is as follows:

[0043]

[0044]

[0045] Among them, h audio is the frequency-time domain feature vector output by the acoustic sub-network, h CO2 is the time series feature vector, is the feature concatenation operation, W q , W k , W v is a learnable weight matrix used to map concatenated features to query Q, key K, value V, d k is the dimension of the key K, which is used to normalize the denominator of the scaling dot product attention. Attention is the attention output matrix. The dependency weights between features are calculated through softmax. g is the dynamic gating weight. The sigmoid function is used to generate a scalar between 0 and 1 to control the fusion ratio of acoustic features. Finally, the probability of pest existence and density level are output.

[0046] Furthermore, in step B, the collected data is preprocessed and feature extracted, specifically in the following manner:

[0047] For sound signals, adaptive wavelet packet decomposition + 3-layer wavelet denoising + spectral subtraction are used, and fast Fourier transform is used to convert time domain sound signals into frequency domain signals to extract the frequency and amplitude characteristic parameters of the sound;

[0048] For the carbon dioxide concentration data, a time series smoothing algorithm is used to extract the concentration change rate and peak time series characteristic parameters.

[0049] Furthermore, the step A is specifically implemented in the following manner:

[0050] Start the light emitting device and use weak light of a specific wavelength to illuminate the soil surface. Periodically irradiate in a pulsed mode, using the negative phototaxis of the grubs to encourage them to migrate to the data acquisition module area.

[0051] Activate the vibration generator, deploy it at different depths in the soil, apply low-frequency vibration of appropriate frequency, dynamically adjust the parameters in real time through sensors, and use the grubs' tendency to vibrations of specific frequencies to guide them to gather in the data acquisition module area.

[0052] Furthermore, the step D implements dynamic detection through adaptive feedback control, wherein:

[0053] Conventional detection mode: runs with preset parameters. When the detection confidence and concentration change rate meet the trigger conditions, it switches to high-precision sampling mode to balance energy consumption and detection accuracy.

[0054] Weak signal enhancement / high-precision mode: If the continuous detection confidence level does not reach the threshold, the environmental stimulation parameters are automatically adjusted until the signal quality reaches the standard, ensuring effective detection in complex environments;

[0055] The trigger conditions for the weak signal enhancement mode include: when the detection confidence level is less than 60% for three consecutive times, the vibration amplitude is increased in 10% steps; when the confidence level is less than 55% for five consecutive times, the UV light intensity is increased and the vibration frequency range is expanded to improve the weak signal signal-to-noise ratio until the signal confidence level is ≥60% or reaches the device threshold, ensuring effective collection of deep weak signals;

[0056] The triggering conditions for high-precision mode include: when the confidence level is greater than 90% and the CO2 change rate is greater than 20% of the historical average, high-precision sampling is triggered, and a three-dimensional pest heat map is generated in combination with sound source positioning; when the confidence level is greater than 95% and the CO2 concentration mutation rate exceeds the historical peak by 30%, redundant sensor array cross-validation is activated.

[0057] Compared with the prior art, the advantages and positive effects of the present invention are:

[0058] This solution uses lighting, vibration, temperature, and humidity adjustments to enhance the activity of underground grubs. It also uses deep learning to detect underground grub pests using CO2 concentrations combined with sound recognition. By collecting and transmitting sound and CO2 signals in real time, this system enables real-time dynamic monitoring of underground grubs, allowing for the timely detection of signs of grub activity, buying valuable time for prevention and control efforts and reducing the risk of crop damage.

[0059] In addition, the creative use of deep learning's pattern recognition capabilities and multimodal analysis technology can accurately identify the existence, quantity and location of underground white grubs, significantly improving the accuracy and reliability of detection and providing a basis for precise prevention and control. The entire detection process does not use chemical reagents, has no pollution to the soil environment, does not destroy the soil ecological balance, conforms to the concept of green agricultural development, and has wide promotion and use value. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 This is a flow chart of the detection method according to an embodiment of the present invention;

[0061] Figure 2 This is a schematic diagram of the structure of a multimodal deep learning model according to an embodiment of the present invention;

[0062] Figure 3 This is a schematic diagram of the underground grub detection principle according to an embodiment of the present invention;

[0063] Figure 4 Schematic diagram of a practical application scenario of the detection system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0064] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described below with reference to the accompanying drawings and embodiments. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can also be implemented in other ways than those described herein. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0065] Example 1: This example proposes a subterranean grub pest detection system based on deep learning multimodality, comprising:

[0066] The stimulation module is used to emit weak light of a specific wavelength, generate vibrations, and adjust soil temperature and humidity to enhance the activity of grubs. The stimulation module includes a light emitting device, a vibration generating device, a temperature regulating device, and a humidity regulating device. Each device works together to enhance the activity of grubs. By using an integrated light emitting device (650nm wavelength infrared LED) and a vibration generating device (dual-frequency piezoelectric vibrator) to dynamically adjust parameters, combined with temperature and humidity regulation, a suitable microenvironment is constructed.

[0067] The data acquisition module includes a sound sensor and a carbon dioxide concentration sensor, which are used to collect sound signals and carbon dioxide concentration data in the soil environment in real time. Specifically, a 6-channel microphone array (supporting TDOA positioning) and a micro NDIRCO2 sensor (integrated with a conductivity probe) can be used to achieve high-precision acquisition of deep signals.

[0068] The edge computing module includes a data processing module and a deep learning model. The data processing module is used to pre-process and extract features from the collected data. The deep learning model module stores a pre-trained deep learning model for identifying the extracted features and determining the probability of existence and density level of underground grubs.

[0069] The result output module outputs a detection report containing the probability of grub presence, density level, and location based on the recognition results of the deep learning model. It also has a data transmission function and can send the report to a designated terminal or system. Specifically, the detection report is transmitted via LoRa / 4G, supporting real-time interaction between PC and mobile terminals, and providing grub location distribution, population estimation, and damage assessment information.

[0070] Among them, such as Figure 2 As shown, the deep learning model includes an acoustic subnetwork and a carbon dioxide subnetwork, and the acoustic subnetwork and the carbon dioxide subnetwork are connected through a cross-modal attention fusion layer;

[0071] The acoustic subnetwork embeds a low-frequency enhancement convolution block (LEB) in EfficientNet-B4. The shallow network of the low-frequency enhancement convolution block introduces deformable convolution (Conv) to adaptively capture the irregular time-frequency distribution of low-frequency vibrations. The channel attention mechanism (SE Block) of the low-frequency enhancement convolution block dynamically weights key frequency bands and outputs a frequency-time domain feature vector.

[0072] The carbon dioxide sub-network is based on the Transformer-XL architecture, combined with the advantages of bidirectional LSTM, and adds a time series difference feature input layer: input original concentration sequence, first-order difference ΔC, second-order difference Δ 2 C, captures concentration mutation patterns; combined with the position encoding extension module PEM to enhance the modeling ability of long-term dependencies and output a time series feature vector;

[0073] The cross-modal attention fusion layer designs a dynamic weight allocation mechanism to achieve bimodal feature fusion of frequency-time domain feature vectors and time series feature vectors, and ultimately outputs the probability of presence and density level of grub pests. When the cross-modal attention fusion layer performs feature fusion, it specifically adopts the following method:

[0074] (1) Feature splicing and attention calculation:

[0075] The frequency-time domain feature vector and the time series feature vector are concatenated, and the feature weight α is generated through a two-layer fully connected network to dynamically allocate the contribution of the two types of feature vectors:

[0076]

[0077] Among them, h audio is the frequency-time domain feature vector, is the time series feature vector; W1 and W2 are weight matrices, and Softmax is used to generate the feature weight α of 0-1 to achieve dynamic weighted fusion;

[0078] (2) Gated nonlinear fusion:

[0079] Introducing dynamic gating weight g=σ(W g Attention) to achieve adaptive fusion. The specific calculation process is as follows:

[0080]

[0081] Among them, h audio represents the frequency-time domain feature vector, h CO2 represents the time series feature vector, W q , W k , W v is a learnable weight matrix used to map concatenated features to query Q, key K, value V, d kis the dimension of the key K, used to normalize the denominator of the scaled dot product attention, W g is a learnable weight vector, Attention is the attention output matrix, and the dependency weights between features are calculated through softmax. g is the dynamic gating weight, and a sigmoid function is used to generate a scalar between 0 and 1 to control the fusion ratio of acoustic features. ⊙ is the element-by-element multiplication, which is used to fuse the two modal features according to the weight. fusion The fused feature vector is input into the fully connected layer for classification, and the final output is the probability of pest existence and density level.

[0082] When determining the location, a positioning algorithm (such as a positioning algorithm based on time difference of arrival) can be used to locate the sound source using the collected sound signals. For example, the location of the sound source can be calculated by analyzing the propagation time differences of the sound signals between different sensors. However, this method may have certain deviations. In this embodiment, the CO concentration data and the sound signal are combined to determine the location.

[0083] This example uses a deep learning model running on the Jetson AGX Xavier platform. The acoustic subnetwork uses an improved version of EfficientNet-B4 (with embedded low-frequency enhancement convolution blocks), and the CO2 subnetwork is based on the Transformer-XL architecture combined with bidirectional LSTM, supporting real-time inference (single-sample processing time < 200ms). Of course, specific modules can also be implemented using other networks, for example, replacing EfficientNet-B4 with 1D ResNet+TCNK to achieve the same feature extraction goal.

[0084] The data processing module and deep learning model module are integrated into a processor with computing power to achieve efficient data processing and model reasoning. The result output module supports wired and wireless network transmission and can transmit the test report to a designated terminal or system.

[0085] Example 2: The detection method of the underground grub pest detection system based on deep learning multimodality proposed in Example 1 is as follows: Figure 1 As shown, the following steps are included:

[0086] Step A: underground environment regulation and grub activity enhancement;

[0087] The stimulation module provides environmental stimulation and actively guides grubs to gather towards the data collection module, solving the problem of weak signals of underground pest activity and difficulty in data collection, creating favorable conditions for subsequent data collection. Specifically:

[0088] Stimulation modules are deployed in the target soil area to create a suitable microenvironment through multi-dimensional stimulation to enhance grub activity:

[0089] Start the light emitting device, use weak light of a specific wavelength to illuminate the soil surface, and irradiate periodically in a pulse mode, using the negative phototaxis of the grubs to encourage them to migrate to the data acquisition sensor area. This embodiment uses a LED cold light source with a wavelength of 650nm (light intensity 100lux) to irradiate the soil surface (0-10cm) in a pulse mode. Each irradiation lasts for 30 seconds, and the cycle is repeated every 3 minutes. The grubs' phototaxis is used to encourage them to migrate to the data acquisition module area. The grubs' compound eyes are sensitive to 620-650nm red light (phototaxis threshold <50lux). Pulsed illumination (30s / 3min) simulates the plant maturity signal and triggers its migration behavior to the surface. Based on the grub light sensitivity curve (EC50=65lux), the effectiveness of the stimulation is ensured while avoiding excessive disturbance.

[0090] Activate the vibration generator, deploy the vibration generator at different depths in the soil, apply low-frequency vibration of appropriate frequency, collect the soil vibration signal intensity in real time through the sound sensor, dynamically adjust the parameters of the vibration generator (frequency, amplitude, duration, etc.), dynamically adjust the parameters in real time through the sensor, and use the white grubs' tendency to vibrations of specific frequencies to guide them to gather in the sensor deployment area, thereby improving the efficiency of activity signal collection.

[0091] This embodiment uses an embedded piezoelectric vibrator to apply a vibration signal with a frequency of 50Hz and an amplitude of 1mm. The vibration lasts for 30 minutes, and the vibration parameters are dynamically adjusted (for example, the signal amplitude is detected every 5 minutes, and if it is below the threshold, the amplitude is increased by 5%) to form a soil disturbance gradient, stimulate the stress-induced migration behavior of the grubs, and increase the activity signal intensity by 40%. The natural vibration frequency of the soil is 10-30Hz. The 50Hz vibration generates unnatural mechanical waves through the piezoelectric ceramic piece (resonance frequency 45Hz), stimulating the grubs' cutaneous mechanical receptors. A vibration amplitude of 1mm can cause an increase in the frequency of body wall peristalsis, prompting the insect body to leave its original habitat.

[0092] Step B: synchronous acquisition and processing of multimodal data;

[0093] An acoustic sensor array is deployed at the target depth to collect low-frequency vibration signals at a specific sampling rate. Acoustic event fragments and characteristic parameters are extracted through noise reduction processing, and the sound source is located in combination with a positioning algorithm.

[0094] Deploy a carbon dioxide concentration sensor network to collect deep concentration data at a specific frequency, combine time window analysis with environmental parameter compensation model to eliminate background interference, and extract the dynamic change characteristics of concentration.

[0095] Combine Figure 3 As shown, the specific steps include:

[0096] Step B1: Realize real-time collection of underground multi-dimensional data through a data acquisition module;

[0097] For sound signal acquisition: A sound sensor array collects low-frequency vibration signals from the soil at a specific sampling rate. A hexagonal waterproof MEMS microphone array (6 channels, 24kHz sampling rate, 10kHz anti-aliasing filter cutoff frequency) is deployed at a depth of 20-30cm, evenly spaced at 15cm intervals. The system continuously collects signals for 5 minutes, generating a time-domain signal sequence and storing it in PCM format, covering a detection area of ​​0-50cm underground.

[0098] For CO2 concentration and environmental parameter collection: A network of CO2 sensors is used to collect CO2 concentration data in deep soil layers at a specific frequency. Specifically, an SBA-5 CO2 sensor (range 0-10,000 ppm, (0-3,000 ppm) accuracy ±40 ppm) is used to synchronously collect deep soil CO2 concentrations, recording every 10 seconds. This creates a multidimensional dataset that includes sound signals, CO2 concentrations, and environmental parameters.

[0099] In this embodiment, the sound sensor uses a microphone array, combined with the coordinated deployment of a CO2 sensor network (carbon dioxide concentration sensor), to achieve the spatiotemporal synchronous acquisition of multimodal data of the underground environment, solving the one-sidedness of single-modal detection (for example, the acoustic signal alone cannot distinguish between soil vibration and mechanical interference, and the CO2 concentration alone cannot locate the specific location of insect pests).

[0100] Step B2: Perform noise reduction and feature engineering on the collected data:

[0101] The low-frequency vibration signal (10-200 Hz) generated by grub activity is easily interfered with by soil particle collision (high-frequency noise) and root friction (low-frequency noise). Traditional noise reduction methods are difficult to simultaneously suppress high- and low-frequency noise. This embodiment specifically adopts the following method:

[0102] (1) For sound signals (20-dimensional multi-domain feature generation):

[0103] Noise Reduction: Adaptive wavelet packet decomposition (5-layer decomposition) is used to filter out noise above 5kHz, and then a 3-layer wavelet noise reduction algorithm (soft threshold processing) is combined to filter out low-frequency environmental noise, such as root friction and water flow. A dynamic noise template is generated based on a 100-hour soil noise library, and spectral subtraction + time-frequency masking technology is used to achieve time-frequency spectrum domain noise removal. By innovatively introducing time-frequency masking technology (which improves the signal-to-noise ratio by 30%), a noise spectrum template is generated based on a 100-hour soil background noise library training, achieving accurate removal of noise components in the time-frequency spectrum domain and improving the signal-to-noise ratio.

[0104] Feature engineering processing:

[0105] Pre-emphasis: A high-pass filter (α = 0.97) is used to enhance high-frequency components. The formula is y(n) = x(n) - 0.97x(n-1), which improves the discernibility of high-frequency details in the sound signal.

[0106] Framing and windowing: Framing the signal after noise reduction, dividing the signal into 25ms frame length and 10ms frame shift, and applying Hamming window Reduce spectrum leakage and form short-term stable signal segments.

[0107] MFCC feature extraction: The framed signal is subjected to fast Fourier transform (FFT), Mel filter bank processing, and discrete cosine transform (DCT) to extract 13th-order MFCC features and 5-dimensional spectral entropy parameters (frequency domain features). 2-dimensional time domain features such as zero-crossing rate and short-time energy are calculated to reflect the time domain variation characteristics of the signal. The frequency domain and time domain features are integrated to generate a 20-dimensional multi-domain feature vector (18-dimensional frequency domain + 2-dimensional time domain).

[0108] Model input preparation: Time-frequency conversion generates a 128×128 time-frequency graph, which is input into the acoustic subnetwork;

[0109] (2) For carbon dioxide concentration signal (10-dimensional time series feature generation):

[0110] Baseline calibration: Based on the soil respiration dynamic model, it eliminates the interference of plant root respiration.

[0111] Considering that soil respiration (plant root metabolism) causes the CO2 concentration baseline to drift, the traditional fixed threshold method cannot adapt to diurnal / seasonal changes. When determining the carbon dioxide concentration, this embodiment constructs a soil respiration baseline dynamic model:

[0112]

[0113] Where T represents temperature, EC represents electrical conductivity, and α, β, and γ are environmental compensation coefficients. The actual CO2 concentration is calculated by fitting historical data using the least squares method:

[0114] C pest =C raw -C baseline

[0115] Time series smoothing: The Kalman filter algorithm is used to remove random noise in the concentration data and generate a 24-hour concentration fluctuation curve to highlight abnormal concentration changes caused by the metabolic activity of grubs.

[0116] Eliminate interference and extract statistical features: Comprehensive application of the second-order difference method Δ 2C (to identify mutation points), Kalman filtering and the first-order difference method ΔC with a 5-minute time window were used to analyze the concentration change trend, eliminate the interference of plant root respiration, and extract 10-dimensional time series features such as the mean, peak, change rate (ΔCO2 / Δt) and diurnal fluctuation of CO2 concentration to quantify the metabolic activity intensity of the grub population. The basal concentration value, peak concentration, concentration change rate (ΔCO2 / Δt), diurnal fluctuation of CO2 and other 10-dimensional time series features were calculated to quantify the metabolic activity intensity of the grub population.

[0117] By combining time window analysis with an environmental parameter compensation model, background interference caused by factors such as soil respiration was eliminated. A time series smoothing algorithm was used to process concentration data, extracting statistical and time-dependent features such as mean concentration, peak value, rate of change, and diurnal fluctuations. This generated a 10-dimensional time series feature vector, effectively quantifying the metabolic activity intensity of the grub population.

[0118] Output: 10-dimensional time series feature vector (original concentration + ΔC + Δ 2 C + 6 statistics)

[0119] Model input preparation: Input the CO2 sub-network (Transformer-XL + bidirectional LSTM).

[0120] Step C: Multimodal deep learning model construction and training:

[0121] The training process of the deep learning model is as follows:

[0122] Data collection: Collect a large amount of sample data including the sounds produced by grub activity and the corresponding changes in carbon dioxide concentration. At the same time, collect background sample data in the absence of grubs. Each sample data must record the environmental information at the time of collection.

[0123] Data labeling: Label the sample data to clarify which data is characteristic of grub activity and which is background data. For characteristic data of grub activity, further label relevant information about the grubs. The labeling process uses a combination of manual and automated labeling.

[0124] Data partitioning: The labeled sample data is divided into training set, validation set and test set according to a certain ratio. The stratified sampling method is used to ensure that the proportion of sample data types in each set is consistent with the overall data type.

[0125] Model selection and initialization: Select an appropriate deep learning network structure, determine network hyperparameters based on problem complexity and data characteristics, and use random initialization methods to initialize network weights and biases;

[0126] Model training: The training set data is fed into the network for training. Mini-batch stochastic gradient descent is used for optimization, with appropriate learning rates, batch sizes, and number of training epochs. During training, the cross-entropy loss function is used to measure the difference between the model predictions and the ground truth. The network parameters are adjusted using the backpropagation algorithm. After each training epoch, the model is evaluated using the validation set data, and hyperparameters are adjusted based on the evaluation results.

[0127] Model evaluation: Use the test set data to evaluate the performance of the trained model and calculate the model's accuracy, recall rate, F1 value and other indicators on the test set. If the model performance does not meet the requirements, return to the data collection step, increase sample data or adjust the model structure, and retrain.

[0128] In combination with Example 1, it can be seen that the model architecture of this scheme adopts the fusion of the acoustic sub-network and the carbon dioxide sub-network through the cross-modal attention mechanism to achieve dynamic weighting of multi-source features, and output the probability of pest existence and density level through the fully connected layer.

[0129] In this embodiment, a cross-modal fusion neural network (DSAF-Net) is designed to address the heterogeneous characteristics of sound and CO2 data. Figure 3 As shown, it contains three core modules.

[0130] 1. Acoustic sub-network (frequency domain feature extraction):

[0131] The low-frequency vibration signals propagating in the soil show irregular distribution in the time-frequency domain due to the influence of scattering and absorption by soil particles. Traditional convolutional neural networks find it difficult to effectively capture these features.

[0132] This example uses an improved version of EfficientNet-B4, embedding a low-frequency enhanced convolutional block (LEB):

[0133] The shallow network introduces deformable convolution (Deformable Conv) to adaptively capture the irregular time-frequency distribution of low-frequency vibrations of 10-50Hz; the channel attention mechanism (SE Block) dynamically weights key frequency bands (such as the main frequency band of white grub activity of 10-200Hz) and outputs 256-dimensional frequency-time domain feature vectors to enhance the characterization ability of weak vibration patterns.

[0134] 2. CO2 sub-network (temporal feature modeling)

[0135] Changes in soil CO2 concentration are influenced by a combination of factors, including grub activity, plant root respiration, ambient temperature, and humidity. These interactions make the temporal variation of CO2 concentration very complex. Traditional temporal models struggle to accurately capture this complex pattern of variation.

[0136] Based on the Transformer-XL architecture, combined with the advantages of bidirectional LSTM, a time series difference feature input layer is added: input original concentration sequence, first-order difference (ΔC), second-order difference (Δ 2 C) to capture concentration mutation patterns; the Position Encoding Extension Module (PEM) enhances the modeling capability of long-term dependencies such as diurnal cycles and seasonal changes, outputs a 128-dimensional time series feature vector, and addresses the problem of soil CO2 concentration baseline drift.

[0137] 3. Cross-modal attention fusion layer

[0138] Design a dynamic weight allocation mechanism to achieve bimodal feature fusion:

[0139] Sound features and CO2 features belong to different modalities and have different physical meanings and data distributions. How to effectively fuse these two heterogeneous features is a challenge, so a dynamic weight allocation mechanism is used.

[0140] Training process: Collect and annotate synchronously acquired multimodal data, divide it into training set, validation set and test set, use optimization algorithm and loss function to train the model, and ensure real-time inference efficiency through model compression technology.

[0141] like Figure 3 As shown in the figure, underground grub detection and positioning are achieved through cross-modal fusion model:

[0142] Feature splicing and attention calculation:

[0143] The 256-dimensional frequency-time domain feature vector and the 128-dimensional CO2 time series feature vector are concatenated into a 384-dimensional vector. The feature weight α (0-1) is generated through a two-layer fully connected network to dynamically allocate the contribution of the two types of features:

[0144]

[0145] Among them, h audio is a 256-dimensional acoustic feature, The 128-dimensional CO2 feature vector is concatenated into a 384-dimensional vector; W1 is a 384×256-dimensional matrix used to concatenate the feature vector [h audio ;h CO2 ] to perform linear transformation, that is, W1[h audio ;h CO2 The calculation process of ] is matrix multiplication, which maps the 384-dimensional input vector to the m-dimensional space. W2 is a 256×1-dimensional matrix, which is used to linearly transform the result after being processed by the activation function σ again, and finally obtain a scalar value. The 0-1 feature weight α is generated by Softmax to achieve dynamic weighted fusion.

[0146] Gated nonlinear fusion:

[0147] Traditional feature splicing methods easily introduce a large amount of redundant information into the model, which increases the model complexity, lengthens the training time, and may reduce the generalization ability of the model. Therefore, gated nonlinear fusion is introduced.

[0148] Introducing the gating weight g=σ(W g Attention) to achieve adaptive fusion. The specific calculation process is as follows:

[0149]

[0150] Among them, h audio h is the 256-dimensional frequency-time domain feature vector output by the acoustic subnetwork, encoding the low-frequency vibration mode of grub activity. CO2 It is the 128-dimensional time series feature vector output by the CO2 sub-network. audio ;h CO2 ] is a feature concatenation operation (“;” indicates concatenation), which combines the 256-dimensional acoustic features and the 128-dimensional CO2 features into a 384-dimensional vector. q ,W k ,W v is a learnable weight matrix with a dimension of 384×384 (matching the concatenated feature dimension), which is used to map the concatenated features into query (Q), key (K), and value (V). k is the dimension of the key (K), where d k =384 (consistent with the dimension of K), used to normalize the denominator of the dot product attention. Attention is the attention output matrix with a dimension of 384×384, and the dependency weights between features are calculated through softmax. W g is a learnable weight vector with a dimension of 384×1, which is used to map the attention output to the gating weight. g is the gating weight, which generates a scalar between 0 and 1 through the sigmoid function to control the fusion ratio of the acoustic features. ⊙ is the element-wise multiplication (Hadamard product), which is used to fuse the two modal features according to the weight. fusion The 384-dimensional feature vector after cross-modal fusion is input into the fully connected layer for classification to solve the feature redundancy problem of traditional splicing and fusion.

[0151] Step D: Adaptive feedback control

[0152] Implement dynamic detection strategies based on model output: Through confidence-driven closed-loop control (such as increasing vibration amplitude and increasing sampling rate), the cycle problem of "weak signal → inaccurate detection → feedback enhancement" is solved. The conventional detection mode or weak signal enhancement / high-precision mode is determined based on whether the detection confidence or concentration change rate meets the standards. This ensures the output data quality (such as more accurate sound source positioning and more obvious CO2 gradient).

[0153] Conventional detection mode: runs with preset parameters. When the detection confidence and concentration change rate meet the trigger conditions, it switches to high-precision sampling mode to balance energy consumption and detection accuracy.

[0154] Weak signal enhancement / high-precision mode: If the continuous detection confidence level does not reach the threshold, the environmental stimulation parameters are automatically adjusted until the signal quality meets the standard, ensuring effective detection in complex environments.

[0155] Specifically, the triggering conditions of the weak signal enhancement mode include: when the detection confidence is less than 60% for three consecutive times, the vibration amplitude is increased in steps of 10%; when the confidence is less than 55% for five consecutive times, the ultraviolet light intensity is enhanced and the vibration frequency range is expanded to improve the weak signal-to-noise ratio until the signal confidence is ≥60% or reaches the device threshold, ensuring the effective collection of deep weak signals.

[0156] The triggering conditions for high-precision mode include: when the confidence level is greater than 90% and the CO2 change rate is greater than 20% of the historical average, high-precision sampling is triggered, and a three-dimensional pest heat map is generated in combination with sound source positioning; when the confidence level is greater than 95% and the CO2 concentration mutation rate exceeds the historical peak by 30%, redundant sensor array cross-validation is activated.

[0157] Step E: Test result processing and decision support

[0158] like Figure 3 As shown, based on the high-precision sampling data optimized in step D (such as 96kHz acoustic signal, 2HzCO2 refresh rate), a detection report including the probability of existence, density level and location of grubs is generated.

[0159] When determining the location, this embodiment can use a positioning algorithm (such as a positioning algorithm based on time difference of arrival) to locate the sound source using the collected sound signals. For example, the location of the sound source can be calculated by analyzing the propagation time differences of the sound signals between different sensors. However, this method may have certain deviations. This embodiment uses a method of combining CO concentration data and sound signals to determine the location. Specifically:

[0160] 1. Spatiotemporal alignment and grid division to ensure data synchronization and spatial resolution;

[0161] Time alignment: Using a 1-second time window, the sound source location point P and the CO concentration data are aligned by timestamp to generate a synchronized dataset.

[0162] Spatial gridding: Divide the detection area into a 10cm×10cm×10cm three-dimensional grid (z=20-30cm), with a volume of 0.001m 3 , mapping the sound source location points and CO concentration data to the corresponding grid.

[0163] 2. Core algorithm for heat map generation: Sound source density calculation:

[0164] The number of sound source points in the grid is counted and the signal energy is weighted. The continuous sound source density field is generated by the Kriging interpolation method. The CO gradient fusion adopts dynamic weighting to fuse the sound source density and CO concentration gradient. When ΔC / Δt>50ppm / min, λ=0.5, otherwise λ=0.7 to generate the comprehensive density D fusion .

[0165] Based on the comprehensive density, the location of pests can be accurately determined by setting thresholds to screen high-probability grids, combining deep learning model predictions, analyzing spatial neighborhood relationships, dynamically tracking changing trends, or fusing multimodal data.

[0166] 3. Realize intelligent linkage: Generate pest distribution maps based on positioning results and concentration distribution, mark high-density areas, and generate detection reports including location, existence probability, and density level. Combined with crop root distribution data (such as corn root density), calculate the hazard index and grade it (low hazard blue, medium hazard yellow, high hazard red), providing a decision-making basis for precise prevention and control. The report is transmitted to the user terminal or management system through the communication module, and prevention and control is carried out in accordance with the preset strategy.

[0167] The above description is merely a preferred embodiment of the present invention and does not constitute any other form of limitation to the present invention. Any person skilled in the art may utilize the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes for application in other fields. However, any simple modification, equivalent change, and modification of the above embodiments made in accordance with the technical essence of the present invention without departing from the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A multimodal underground grub detection system based on deep learning, characterized by: include: a stimulation module, comprising a light emitting device, a vibration generating device, a temperature regulating device and a humidity regulating device, for emitting weak light of a specific wavelength, generating vibrations and regulating soil temperature and humidity to enhance grub activity; A data acquisition module, including a sound sensor and a carbon dioxide concentration sensor, for collecting sound signals and carbon dioxide concentration data in the soil environment in real time; The edge computing module includes a data processing module and a deep learning model. The data processing module is used to preprocess and extract features from the collected sound signals and carbon dioxide concentration data. The deep learning model is used to identify the extracted features and then determine the presence probability and density level of underground grub pests. The result output module outputs a detection report containing the probability of existence, density level and location of grubs based on the recognition results of the deep learning model, and sends the report to the designated terminal or management system for prevention and control in combination with preset strategies.

2. The underground grub pest detection system based on deep learning multimodality according to claim 1, characterized in that: The deep learning model includes an acoustic subnetwork and a carbon dioxide subnetwork, and the acoustic subnetwork and the carbon dioxide subnetwork are connected through a cross-modal attention fusion layer; The acoustic subnetwork embeds a low-frequency enhancement convolution block (LEB) in EfficientNet-B4. The shallow network of the low-frequency enhancement convolution block introduces deformable convolution (Conv) to adaptively capture the irregular time-frequency distribution of low-frequency vibrations. The channel attention mechanism (SE Block) of the low-frequency enhancement convolution block dynamically weights key frequency bands and outputs a frequency-time domain feature vector. The CO2 sub-network is based on the Transformer-XL architecture, combined with a bidirectional LSTM and an additional temporal difference feature input layer: it inputs the original concentration sequence, first-order difference ΔC, and second-order difference Δ2C to capture concentration mutation patterns. It also combines the position encoding extension module PEM to enhance the modeling capability of long-term dependencies and output a temporal feature vector. The cross-modal attention fusion layer designs a dynamic weight allocation mechanism to achieve bimodal feature fusion of frequency-time domain feature vectors and time series feature vectors, and ultimately outputs the probability of existence and density level of white grub pests.

3. The underground grub pest detection system based on deep learning multimodality according to claim 2, characterized in that: When the cross-modal attention fusion layer performs feature fusion, the following method is specifically adopted: (1) Feature splicing and attention calculation: The frequency-time domain feature vector and the time series feature vector are concatenated, and the feature weight α is generated through a two-layer fully connected network to dynamically allocate the contribution of the two types of feature vectors; (2) Gated nonlinear fusion: Introducing dynamic gating weight g=σ(W g Attention) to achieve adaptive fusion, W g is a learnable weight vector, Attention is the attention output matrix, and the fused feature vector is input into the fully connected layer for classification, and finally the probability of existence and density level of grub pests are output.

4. The underground grub pest detection system based on deep learning multimodality according to claim 1, characterized in that: The data processing module performs data processing in the following manner: For sound signals, adaptive wavelet packet decomposition + 3-layer wavelet denoising + spectral subtraction are used, and fast Fourier transform is used to convert time domain sound signals into frequency domain signals to extract the frequency and amplitude characteristic parameters of the sound; For the carbon dioxide concentration data, a time series smoothing algorithm is used to extract the concentration change rate and peak time series characteristic parameters.

5. The method of the underground grub pest detection system based on deep learning multimodality according to claim 2, characterized in that: The following steps are involved: Step A: Underground environment regulation and grub activity enhancement: The stimulation module is used to stimulate the environment and actively guide the grubs to gather towards the data collection module; Step B: Synchronous acquisition and processing of multimodal data: Step B1: using a sound sensor and a carbon dioxide concentration sensor to collect sound signals and carbon dioxide concentration data in the soil environment in real time; Step B2: performing noise reduction and filtering preprocessing operations on the collected sound signals and carbon dioxide concentration data to extract the time domain and frequency domain characteristics of the signals and characteristic parameters of carbon dioxide concentration changes; Step C: Multimodal deep learning model construction and training: Combining the acoustic sub-network and the carbon dioxide sub-network, through the cross-modal attention mechanism, dynamic weighting of multi-source features is achieved, and the probability of existence and density of grub pests are output through the fully connected layer; Step D, Adaptive Feedback Control: Implement a dynamic detection strategy based on the output of the multimodal deep learning model. Determine whether to use conventional detection mode or weak signal enhancement / high-precision mode based on the detection confidence level, and optimize the output data. Step E: Detection result processing and decision support: Based on the deep learning model in step D, a detection report containing the probability of existence, density level, and location of white grubs is output, and the report is sent to a designated terminal or management system for prevention and control in combination with preset strategies.

6. The method of the underground grub pest detection system based on deep learning multimodality according to claim 5, characterized in that: The step C is specifically implemented in the following manner: (1) Frequency domain feature extraction through acoustic sub-network: using the improved EfficientNet-B4 architecture, embedding low-frequency enhancement convolution blocks: shallow networks introduce deformable convolution Conv to adaptively capture the irregular time-frequency distribution of low-frequency vibrations; the channel attention mechanism dynamically weights key frequency bands and outputs frequency-time domain feature vectors; (2) Temporal feature modeling through the carbon dioxide sub-network: Based on the Transformer-XL architecture, combined with bidirectional LSTM, a temporal difference feature input layer is added to input the original concentration sequence, first-order difference ΔC, and second-order difference Δ2C to capture concentration mutation patterns. The position encoding extension module enhances the modeling capability of long-term dependencies and outputs a temporal feature vector. (3) Combined with the cross-modal attention fusion layer, a dynamic weight allocation mechanism is designed to achieve bimodal feature fusion: Feature splicing and attention calculation: The frequency-time domain feature vector is concatenated with the time series feature vector, and the feature weight α is generated through a two-layer fully connected network. The contribution of the two types of features is dynamically allocated to achieve dynamic weighted fusion. Gated nonlinear fusion: Introducing dynamic gating weight g=σ(W g Attention) to achieve adaptive fusion, W g is a learnable weight vector, Attention is the attention output matrix, and the fused feature vector is input into the fully connected layer for classification, and finally the probability of existence and density level of grub pests are output.

7. The method of the underground grub pest detection system based on deep learning multimodality according to claim 5, characterized in that: In step B, the collected data is preprocessed and feature extracted, specifically in the following manner: For sound signals, adaptive wavelet packet decomposition + 3-layer wavelet denoising + spectral subtraction are used, and fast Fourier transform is used to convert time domain sound signals into frequency domain signals to extract the frequency and amplitude characteristic parameters of the sound; For the carbon dioxide concentration data, a time series smoothing algorithm is used to extract the concentration change rate and peak time series characteristic parameters.

8. The method of the underground grub pest detection system based on deep learning multimodality according to claim 5, characterized in that: The step A is specifically implemented in the following manner: Start the light emitting device and use weak light of a specific wavelength to illuminate the soil surface. Periodically irradiate in a pulsed mode, using the negative phototaxis of the grubs to encourage them to migrate to the data acquisition module area. Activate the vibration generator, deploy it at different depths in the soil, apply low-frequency vibration of appropriate frequency, dynamically adjust the parameters in real time through sensors, and use the grubs' tendency to vibrations of specific frequencies to guide them to gather in the data acquisition module area.

9. The method of the underground grub pest detection system based on deep learning multimodality according to claim 5, characterized in that: In step E, when determining the location of the white grub pest, the location is determined by combining the CO concentration data and the sound signal. Specifically: (1) Spatiotemporal alignment and grid division; Using a 1-second time window, the sound source location points and CO2 concentration data were aligned by timestamp to generate a synchronized dataset. The detection area was then divided into m three-dimensional grids, and the sound source location points and CO2 concentration data were mapped to the corresponding grids. (2) The core algorithm for heat map generation is sound source density calculation: the number of sound source points in the grid is counted and the signal energy is weighted, and a continuous sound source density field is generated through the Kriging interpolation method; CO2 concentration gradient fusion, using dynamic weights to fuse the sound source density and CO2 concentration gradient, when ΔC / Δt>50ppm / min, λ=0.5, otherwise λ=0.7, to generate a comprehensive density, and then determine the location of the white grub pests based on the comprehensive density.

10. The method of the underground grub pest detection system based on deep learning multimodality according to claim 5, characterized in that: The step D implements dynamic detection through adaptive feedback control, wherein: Conventional detection mode: runs with preset parameters. When the detection confidence and concentration change rate meet the trigger conditions, it switches to high-precision sampling mode to balance energy consumption and detection accuracy. Weak signal enhancement / high-precision mode: If the continuous detection confidence level does not reach the threshold, the environmental stimulation parameters are automatically adjusted until the signal quality meets the standard, ensuring effective detection in complex environments.

Citation Information

Cited By

  • Garden underground pest monitoring method, electronic equipment and storage medium

    CN121687088A