Intelligent monitoring method and device based on multi-modal information fusion

By employing a multimodal information fusion-based intelligent monitoring method, combined with microwave radar and video image subsystems, high-precision identification and location of grain storage pests have been achieved. This solves the problem of inaccurate detection in existing technologies, adapts to complex grain storage environments, and meets the needs of modern grain storage management.

CN121365299APending Publication Date: 2026-01-20BEIJING AEROSPACE TIMES TECH DEV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511578881.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing grain storage pest monitoring technologies suffer from limitations such as manual sampling, insufficient performance of single-modal sensors, and poor environmental adaptability of multimodal fusion, resulting in inaccurate detection and high false positive rates. These technologies fail to meet the high-precision, full-coverage, and robust requirements of modern grain storage management.

Method used

A multimodal information fusion-based intelligent monitoring method is adopted, which simultaneously collects data through microwave radar and video image subsystems. Combined with adaptive preprocessing, multi-domain feature extraction and weighted fusion classification output, it can realize real-time identification, location and counting of pests from the surface of grain piles to a depth of 0.5m.

Benefits of technology

In environments with dust ≤20mg/m³ and humidity 10%~95%, the recognition accuracy is ≥93%, the counting error is ≤10%, and the positioning error is ≤0.1m. It is suitable for the intelligent transformation of existing grain warehouses and meets the high precision, full coverage, and strong robustness requirements of modern grain storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121365299A_ABST
    Figure CN121365299A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent monitoring method and device based on multi-modal information fusion, and the method comprises the steps: (1) collecting multi-modal data; (2) carrying out self-adaptive preprocessing on the microwave signal; (3) video image self-adaptive preprocessing; (4) multi-domain feature extraction; and (5) carrying out weighted fusion classification output. The method and the system are suitable for monitoring main grain storage pests such as maize weevil and grain beetles in the scenes of national grain depots and civil granaries, can realize real-time identification, positioning and counting of pests from the surface layer of a grain pile to the depth of 0.5 m, and solve the problems of insufficient performance of a single-mode sensor (vision, radar and acoustics) and poor adaptability of a multi-mode fusion environment in the prior art. The method adapts to intelligent transformation of the existing granary, and meets the requirements of high precision, full coverage and strong robustness of modern grain storage.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of grain storage safety monitoring, in particular to an intelligent monitoring method and device based on multi-modal information fusion. BACKGROUND

[0002] As a large grain producing and storing country, China's annual loss accounts for about 8%. Traditional monitoring technology cannot meet the core needs of "high precision, full coverage, strong robustness" of modern grain storage management.

[0003] There are three important bottlenecks in existing grain storage pest monitoring technology: 1. Limitations of manual sampling monitoring: It depends on the experience of monitoring personnel, and 2 to 3 people are needed for a day to detect a single warehouse (conventional small and medium-sized grain warehouses with a volume of 500-1000m³). It can only reflect local surface pests and cannot detect pests under the surface of the grain pile, which easily leads to "missed detection and misjudgment".

[0004] 2. Defects of single modal sensor technology: (1) Visual detection technology: It can only cover the surface of the grain pile, and is affected by dust (concentration often reaches 10-20mg / m³) and light fluctuation (50-1000lux) in the grain warehouse. The image noise is significant, the pest morphological features are blurred, and the recognition accuracy is less than 85%; (2) Traditional radar detection technology: Although it can penetrate the grain pile, it is greatly disturbed by the reflection of the metal warehouse wall (reflection coefficient >-10dB), it is difficult to locate, and it can only output a binary result of "presence or absence of pests", which cannot distinguish pest species, and the recognition accuracy is less than 70%; (3) Acoustic detection technology: There are significant technical defects. First, complex environmental noise interference (background noise such as grain pile stirring and mechanical ventilation) easily overwhelms the weak sound signal of pests, resulting in a very low signal-to-noise ratio and a high false positive rate. Second, there is a "detection blind area" that cannot identify early-stage larvae with weak acoustic activity, and the detection range is limited.

[0005] 3. Defects of existing multi-modal fusion technology: In existing technology, there is no public radar-visual joint detection method. The existing acoustic-visual multi-modal fusion technology has poor environmental adaptability and has not established a dynamic adaptation mechanism of "environmental parameters (temperature, humidity, dust) - algorithm parameters". In complex grain warehouse environment (humidity >80%, dust >15mg / m³), the performance of the sensor fluctuates significantly, the fusion accuracy decreases significantly, and it cannot meet the detection requirements stably.

[0006] In view of the above problems, it is urgent to provide a monitoring method and device to break through the bottlenecks of traditional technology and realize accurate monitoring of grain storage pests. SUMMARY

[0007] The purpose of the present application is to overcome the above technical deficiencies, provide a multi-modal information fusion based intelligent monitoring method and device, to solve the technical problems of insufficient performance of single modal sensor (vision, radar, acoustic) and poor environmental adaptability of multi-modal fusion in the related art.

[0008] To achieve the above technical purpose, the present application adopts the following technical scheme: According to one aspect of the present application, a multi-modal information fusion based intelligent monitoring method is provided, comprising: (1) Multi-modal data acquisition: collect microwave reflection signals of pests in the surface layer to the internal pests of the grain pile through a microwave radar subsystem, collect pest images of the surface layer of the grain pile through a video image subsystem, synchronously collect environmental parameters of temperature, humidity, dust concentration and light intensity in the grain depot, and realize timestamp binding of radar and image data; (2) Microwave signal adaptive preprocessing: sequentially perform temperature and humidity correction, dynamic adjustment of forgetting factor and RLS iterative filtering to realize microwave signal purification; (3) Video image adaptive preprocessing: sequentially perform multi-scale Retinex enhancement, dust noise judgment, adaptive threshold BM3D filtering and weighted reconstruction to improve image clarity; (4) Multi-domain feature extraction: extract radar signal domain features of preprocessed radar signals using a one-dimensional CNN network, extract radar frequency domain features by performing sliding window processing on the preprocessed radar signals using one-dimensional FFT, and extract visual image domain features of the denoised image using a Yolo series or MobileNet series lightweight convolutional neural network; (5) Weighted fusion classification output: calculate radar confidence based on radar echo signal-to-noise ratio, perform weighted embedding vector splicing on normalized radar signal domain features, radar frequency domain features and visual image domain features according to weight proportions of 15% to 25% for radar signal domain features, 15% to 25% for radar frequency domain features and 50% to 70% for visual image domain features, then use a multi-head Transformer model to strengthen the correlation between modal features, and finally output classification probability of each type of pest, number of pests in a single area and average density through a Softmax activation function.

[0009] Further, in step (1): the microwave radar subsystem uses a K-band radar with a transmission power of 0-30dBm, a bandwidth of 200-400MHz, a sampling rate of ≥10kHz, transmits signals through a gain of ≥8dBi antenna, and converts the amplified and filtered frequency-shifted signals into three-dimensional signals of time, amplitude and frequency through a 12-bit or higher ADC; the video image subsystem uses a camera with a pixel of ≥8 million, equipped with a light supplement module, and outputs images with a frame rate of ≥10fps; the environmental parameters are collected by corresponding sensors and transmitted to the front-end processing module through a RS485 standard bus.

[0010] Further, the specific process of the microwave signal adaptive preprocessing in step (2) is that the temperature and humidity correction dynamically adjusts the correction coefficient according to the temperature and humidity of the environment where the grain pile is located, and when the temperature or humidity deviates from the reference state, the correction coefficient is adjusted accordingly to compensate for the environmental impact; in the dynamic adjustment of the forgetting factor, the value of the forgetting factor is dynamically set according to the grain pile density, and the higher the grain pile density, the larger the value of the forgetting factor.

[0011] Further, in step (2), the RLS iterative filtering calibrates the covariance matrix based on the initial noise variance of the grain pile, and realizes clutter suppression through the iterative process of error calculation, gain update, coefficient and covariance update.

[0012] Further, the specific process of the video image adaptive preprocessing in step (3) is that the multi-scale Retinex enhancement decomposes the image into illumination component and reflection component, and smoothes the illumination component by using three groups of different Gaussian kernels with sizes of 20-40 pixels, 40-80 pixels and 80-160 pixels to remove background interference, extracts the corresponding reflection component, and weights and fuses the reflection components according to the weight proportions of 0.3, 0.5 and 0.2; the dust noise judgment divides the image into multiple square regions according to the side length of 16-64 pixels, calculates the gray variance of each region, and presets a gray variance threshold, and the gray variance exceeding the threshold is determined as a dust pollution area, and the one not exceeding is a normal area.

[0013] Further, in step (3), the adaptive threshold BM3D filtering is performed on the dust pollution area, a reference block with a side length of 16-64 pixels is taken, similar blocks are screened in a search window with a size of 3-5 times the reference block size and sorted from large to small according to the similarity to take the first 8-32, then 3D-DCT transformation is performed, and the adaptive threshold is calculated by using the local statistical method based on the gray mean and variance within the block, and the transformed coefficients are threshold processed; the weighted reconstruction weights and sums the image blocks after inverse transformation according to the similarity weight, and outputs the denoised image.

[0014] Further, the specific process of the multi-domain feature extraction in step (4) is that the radar signal domain feature extraction adopts a one-dimensional CNN network, which includes an input layer, at least 8 convolution layers, corresponding pooling layers and a full connection layer, and finally outputs 64-256 dimensional radar signal domain features; the radar frequency domain feature extraction adopts one-dimensional FFT to perform sliding window processing on the preprocessed radar signal, and the window size is 1s-10s, and the output dimension is 64-256 dimensional radar frequency domain features; the visual image domain feature extraction adopts Yolo series or MobileNet series lightweight convolutional neural network to process the denoised image, and outputs the pest contour coordinates, pest type and recognition probability as the visual image domain features.

[0015] Further, in step (5): when weighted fusion, the radar confidence is calculated according to SNR, SNR≥20dB takes 1.0, and SNR<0dB takes 0, the feature weights of the radar signal domain, the frequency domain and the visual domain are 15%-25%, 15%-25% and 50%-70% respectively, and the fusion sequence is formed by splicing; the multi-head Transformer model is used for classification, and after layer normalization, the pest classification probability, the count value and the average density are output through Softmax.

[0016] The application further provides an intelligent monitoring device based on multi-modal information fusion for implementing the above method, which comprises a perception module, a front-end processing module and a central processing module, the perception module comprises a microwave radar subsystem, a video image subsystem with a resolution of more than 8 million pixels, and an environment subsystem comprising temperature, humidity, dust concentration and light intensity sensors; the front-end processing module comprises a low-power processor with a single-precision floating-point operation capacity of more than 300 GFLOPS, a memory of more than 4G, a solid state disk of more than 256G and an Ethernet communication device of more than 100Mbps, which are used for performing adaptive preprocessing of microwave signals and video images; the central processing module comprises a high-performance processor with a full-core integer operation capacity of more than 400 GIPS, a GPU processor with a single-precision floating-point operation capacity of more than 40 TFLOPS, a memory of more than 64G, a hard disk of more than 8T and a network communication device of more than 1000Mbps, which are used for performing multi-domain feature extraction and weighted fusion classification output.

[0017] The application provides an intelligent monitoring method and device based on multi-modal information fusion, which is suitable for monitoring main grain storage pests such as Sitophilus zeamais, Rhyzopertha dominica, Tribolium castaneum and Sitophilus oryzae in state-owned grain depots, private grain stores and the like, and can realize real-time identification, positioning and counting of pests in a depth of 0.5 m from the surface of a grain pile. In view of the problems of insufficient performance of a single-mode sensor (vision, radar and acoustics) and poor adaptability of multi-modal fusion in the prior art, the application adopts a whole-link technical solution of'multi-modal perception-adaptive preprocessing-multi-domain feature extraction-weighted fusion classification output', and combines a three-level system architecture of 'perception module-front-end processing module-center processing module', so as to synchronously collect microwave signals of pests, surface images and environmental parameters such as temperature and humidity, dust, illumination and the like through a microwave radar subsystem and a video image subsystem; the front-end processing module performs temperature and humidity correction, dynamic adjustment of a forgetting factor and RLS iterative filtering on the microwave signals, and performs multi-scale Retinex enhancement and adaptive threshold BM3D filtering on the images; the center processing module extracts radar signal domain and frequency domain features through one-dimensional CNN and FFT, extracts visual image domain features through a deep neural network, and realizes fusion classification and counting through dynamic weighting and multi-head Transformer. The application has a wide detection range, and in an environment with dust of ≤20 mg / m³ and humidity of 10% to 95%, the identification accuracy is ≥93%, the counting error is ≤10%, and the positioning error is ≤0.1 m, so that the application is suitable for intelligent transformation of an existing grain store and meets the requirements of modern grain storage in high precision, full coverage and strong robustness. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 A flowchart of an intelligent monitoring method based on multi-modal information fusion is provided for the embodiments of the application. DETAILED DESCRIPTION

[0019] In order for those skilled in the art to better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.

[0020] According to the embodiment of the present application, a multi-modal information fusion-based intelligent monitoring method and system are provided to solve the problems of the prior art in grain storage pest monitoring. The technical solution completely covers all the technical features of the claims, and through the whole-process cooperation of "multi-modal data acquisition, microwave signal adaptive preprocessing, video image adaptive preprocessing, multi-domain feature extraction, and weighted fusion classification output", combined with the three-level system architecture of "perception module, front-end processing module, and center processing module", high-precision grain storage pest monitoring is realized. In the following, taking a small flat-roofed granary as an application scenario (granary size 20m x 5m x 4m, stored grain variety is wheat, grain pile height is 2m, and storage capacity is about 150t) as an example, the above-mentioned method and system provided by the embodiment of the present application are exemplarily and in detail introduced.

[0021] In the embodiment of the present application, the multi-modal information fusion-based intelligent monitoring method is executed in the whole-process cooperation of "multi-modal data acquisition, microwave signal adaptive preprocessing, video image adaptive preprocessing, multi-domain feature extraction, and weighted fusion classification output", and specifically includes the following steps: Step 1) Multi-modal data acquisition The core of this step is to realize the synchronous acquisition of radar signals, grain pile surface layer images, and granary environment parameters, and the data correlation is ensured through timestamp binding to provide a consistent data basis for subsequent fusion processing.

[0022] Substep (1): Time synchronization control An NTP (lightweight network time protocol) time synchronization mechanism is adopted to take the subsequent center processing module as a time reference node, and the clock synchronization is performed with the front-end processing module every 15s. The maximum measured time synchronization error is 45ms, which meets the technical requirement of "time synchronization error ≤100ms", and effectively avoids the misalignment problem caused by time deviation of radar and image data.

[0023] Substep (2): Microwave radar subsystem data acquisition A K-band microwave radar (center frequency 24GHz, adjustable transmission power 0-30dBm, adjustable bandwidth 200-400MHz, sampling rate 10kHz) is selected, and an antenna array with a gain of 8dBi is configured. The antenna is inclined downward by 25° to ensure that the microwave signal can cover a depth range of 0.5m after penetrating the surface layer of the grain pile.

[0024] The frequency shift signal caused by the pests in the grain pile is conditioned by a low-noise amplifier (gain 20dB) and a band-pass filter (adjustable bandwidth 200-400MHz), and then converted into a "time-amplitude-frequency" three-dimensional data frame (each frame time length 1-10s can be set) by a 12-bit ADC. The frame is transmitted in real time to the front-end processing module through Ethernet.

[0025] Sub-step (3): Video image subsystem data acquisition It uses an 8-megapixel camera (sensitivity ≥1 lux@F1.8) and is equipped with a high color rendering index fill light module (color rendering index 90, power 5W). The fill light brightness is automatically adjusted according to the light sensor data.

[0026] The camera outputs a frame rate of 10fps, with a single frame image size of 3840×2160 pixels. When transmitted to the front-end processing module, the measured image signal-to-noise ratio is 31dB.

[0027] Sub-step (4): Environmental parameter acquisition Environmental data is collected using a temperature and humidity sensor (measurement range -40~80℃ / 0~100% RH, accuracy ±0.5℃ / ±2% RH, sampling rate 1Hz), a dust concentration sensor (range 0~20mg / m³, accuracy ±1mg / m³, sampling rate 1Hz), and a light intensity sensor (range 0~2000lux, accuracy ±10lux, sampling rate 1Hz). All parameters are transmitted to the front-end processing module via a standardized RS485 bus for dynamic adaptation of subsequent preprocessing algorithm parameters.

[0028] Step 2) Adaptive preprocessing of microwave signals To address the relatively small fluctuations in the grain storage environment, a three-stage process—"temperature and humidity correction—dynamic adjustment of forgetting factor—RLS iterative filtering"—is implemented to purify the microwave signal. The specific operation is as follows: Sub-step (1): Temperature and humidity correction Using "20℃, 60%RH" as the baseline (corresponding to a correction factor of 1.0), the correction factor is dynamically adjusted based on the measured environmental parameters: Temperature correction: If the actual measured temperature of the grain silo is 28℃ (8℃ higher than the benchmark), according to the rule that "for every 10℃ change in temperature, the correction coefficient adjustment range is ≥5%", the temperature correction coefficient = 1.0 + (8 / 10) × 5% = 1.04. Humidity correction: If the measured humidity of the grain silo is 72% RH (not exceeding the 75% RH threshold), the humidity correction factor remains at 1.0; The comprehensive correction factor is 1.04 × 1.0 = 1.04. The radar signal amplitude value is multiplied by this factor to compensate for the influence of the change of the dielectric constant of the grain pile with temperature and humidity on the echo signal.

[0029] Sub-step (2): Dynamic adjustment of forgetting factor The forgetting factor is limited to a range of 0.95 to 0.99 and is dynamically set according to the grain pile density. When the wheat grain pile density is 700 to 800 kg / m³, the forgetting factor is set to 0.97 to 0.98 to enhance the suppression effect of clutter on the stable background of the grain pile.

[0030] Sub-step (3): RLS iterative filtering The initial parameters include the covariance matrix inverse matrix (based on the initial noise variance calibration of the grain pile) and the filter coefficients (the initial value is 0); through the iterative process of “error calculation→gain update→coefficient and covariance update”, the clutter suppression is realized, and the final clutter suppression is 58dB higher than the actual measurement, and the pest signal retention rate is >96%.

[0031] Step 3) Video image adaptive preprocessing In view of the characteristics of small grain storehouse, low dust concentration (5-15mg / m³) and small light fluctuation, a four-level process of “multi-scale Retinex enhancement—dust noise judgment—adaptive threshold BM3D filtering—weighted reconstruction” is performed to improve image clarity, and the specific operation is as follows: Sub-step (1): Multi-scale Retinex enhancement After converting the 3840×2160 pixel image into a grayscale image, it is decomposed into an illumination component representing the background and a reflection component representing the details of the pests; three groups of Gaussian kernels of different sizes (20 pixels, 40 pixels, and 100 pixels) are selected to smooth the illumination component; the three groups of smoothed reflection components are weighted and fused according to the weight ratio of 0.3, 0.5, and 0.2, and the edge contrast of the processed pests is improved by 28%.

[0032] Sub-step (2): Dust noise judgment The enhanced image is divided into 120×90=14400 square regions according to 24×24 pixels, and the gray variance of each region is calculated; when the dust concentration is measured to be 12mg / m³, according to the rule of “dust concentration dynamic adjustment threshold”, the preset gray variance threshold is 50+(12-10)×2=54 (the threshold increases by 2 for every 1mg / m³ increase in dust concentration); the regions with gray variance>54 (about 12% of the total image area) are determined as dust contaminated areas, and the rest are normal areas.

[0033] Sub-step (3): Adaptive threshold BM3D filtering Only the dust contaminated areas are filtered: a 24×24 pixel reference block is taken, similar blocks are screened in a 72×72 pixel search window, and the top 12 are sorted according to the similarity from large to small; 3D-DCT transformation is performed on the 3D matrix composed of 12 similar blocks; the adaptive threshold is calculated based on the gray mean and variance within the block, and the noise components below the threshold in the transformed coefficients are set to zero.

[0034] Sub-step (4): Weighted reconstruction The image blocks after inverse 3D-DCT transformation are weighted and summed according to similarity weights (the highest block weight 0.14, the lowest block weight 0.03), and the denoised image is output; the processing result: the actual measurement of image noise intensity is 8dB.

[0035] Step 4) Multi-domain feature extraction Differential features are extracted from the preprocessed radar signals and video images respectively to provide a basis for subsequent fusion classification, and the specific operations are as follows: Substep (1): Radar signal domain feature extraction (one-dimensional CNN) The network structure includes an input layer, 8 convolutional layers, corresponding pooling layers, and a fully connected layer: the input layer receives 128-dimensional preprocessed radar signals, the convolutional layer uses a 3x1 convolutional kernel, a ReLU activation function, and a batch normalization after each layer, the pooling layer uses a 2x1 maximum pooling kernel (step size 2), the fully connected layer has a neuron number decreasing from 256 to 128, and finally outputs 128-dimensional radar signal domain features (in the range of 64-256 dimensions).

[0036] The network is trained with 50,000 labeled radar signals for 50 training rounds, and the loss function converges to 0.045; during inference, 128-dimensional features are output every 1s, and the single inference time is about 0.2s, meeting the real-time requirement.

[0037] Substep (2): Radar frequency domain feature extraction (one-dimensional FFT) The preprocessed radar signal is processed using a 1s sliding window; one-dimensional FFT is performed on each sliding window segment, and the amplitude values of the first 128 frequency points are taken; 128-dimensional radar frequency domain features are output.

[0038] Substep (3): Visual image domain feature extraction (lightweight CNN) The Yolo V11 lightweight convolutional neural network is selected, and the denoised image is normalized to the range [0,1] before being input into the network; the network outputs pest outline coordinates (such as the corn weevil bounding box coordinates (x1=181, y1=189, x2=233, y2=257)), pest type (rice weevil), and recognition probability (94%), which constitute 256-dimensional visual image domain features.

[0039] Step 5) Weighted fusion classification output Pest classification, counting, and positioning are achieved through dynamic weighting and cross-modal feature association, and the specific operations are as follows: Substep (1): Weighted fusion Radar confidence calculation: according to the rule of "calculating confidence based on radar echo signal-to-noise ratio (SNR)", the measured radar echo SNR is 24dB (≥20dB), so the radar confidence is 1.0; Weight distribution: radar signal domain feature weight 20%, radar frequency domain feature weight 20%, visual image domain feature weight 60%; Feature splicing: the normalized 128-dimensional radar signal feature, 128-dimensional radar frequency feature and 256-dimensional visual feature are weighted according to the above weight and spliced into a 512-dimensional fusion feature sequence.

[0040] Substep (2): classification output (multi-head Transformer) Using a multi-head Transformer model to strengthen the correlation between modalities: the number of attention heads is set to 6 (4 cross-modal attention heads, accounting for 66.7%; 2 self-attention heads, accounting for 33.3%, the cross-modal proportion is within the range of 60%-80%), the number of hidden layer neurons in the feedforward network is 1024; the correlation between radar features and visual features is calculated through cross-modal attention to strengthen the complementarity of cross-modal information; after 3-layer Transformer encoding, the output is obtained through the Softmax activation function: • Classification probability: 96% for Sitophilus zeamais and 94% for Rhyzopertha dominica; • Counting: 4 heads of Sitophilus zeamais and 9 heads of Rhyzopertha dominica are measured in a 1m2 detection area, the model prediction value is 4 heads and 10 heads, the counting error is 7.7%, and the density is 28 heads / m3; • Positioning result: the center coordinates of Sitophilus zeamais are (1.57m, 1.83m), and the positioning error is 0.07m; the center coordinates of Rhyzopertha dominica are (2.31, 2.67), and the positioning error is 0.06m.

[0041] The embodiment of the application also provides an intelligent monitoring device based on multi-modal information fusion for implementing the above method.

[0042] The device is deployed according to a three-level architecture of "perception module-front end processing module-center processing module", the hardware selection and installation mode of each module are adapted to the space layout of the warehouse, and specifically includes the following parts: 1) Perception module selection and deployment The perception module is used for collecting multi-modal data and environmental parameters, and is composed of a microwave radar subsystem, a video image subsystem and an environment subsystem: Microwave radar subsystem: 24GHz frequency band radar module (transmit power 0-30dBm adjustable, bandwidth 200-400MHz adjustable, sampling rate 10kHz, 12-bit ADC), matched with an antenna array with a gain of 8dBi; Video image subsystem: 800 million pixel camera (sensitivity ≥ 1 lux @ F1.8), supporting color rendering index ≥ 90, power 5W light module; Environmental subsystem: temperature and humidity sensor (measurement range -40~80℃ / 0~100% RH, accuracy ±0.5℃ / ±2% RH), dust concentration sensor (range 0~20mg / m³, accuracy ±1mg / m³), light intensity sensor (0~2000lux, accuracy ±10lux), supporting RS485 bus transmission module.

[0043] 2) Front-end processing module selection and deployment The front-end processing module is used to perform adaptive preprocessing of microwave signals and video images, and the hardware selection and deployment meets the requirements of "low power consumption and high adaptation": A low-power processor (single-precision floating-point operation capability 320GFLOPS) is selected, supporting 4G memory, 512G solid state disk, 1000Mbps Ethernet communication device; support RS485 and Ethernet dual interface, can simultaneously receive multi-modal data and environmental parameters of sensing module.

[0044] 3) Center processing module selection and deployment The center processing module is used to perform multi-domain feature extraction and weighted fusion classification, and the hardware selection and deployment meets the requirements of "high performance and scalability": A high-performance processor (full-core integer operation capability 420GIPS) is selected, supporting GPU processor (single-precision floating-point operation capability 48TFLOPS), 64G memory, 8T hard disk, 1000Mbps Ethernet communication device; supporting display and database software, used for real-time display of monitoring results and storage of historical data.

[0045] The above only describes the preferred embodiments of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, can make several improvements and refinements, these improvements and refinements should also be considered as the protection scope of the present application.

Claims

1. A multi-modal information fusion-based intelligent monitoring method, characterized in that, Comprise: (1) Multi-modal data acquisition: collect microwave reflection signals of pests from the surface to the interior of the grain pile through the microwave radar subsystem, collect pest images on the surface of the grain pile through the video image subsystem, synchronously collect environmental parameters such as temperature, humidity, dust concentration and light intensity in the grain depot, and realize timestamp binding of radar and image data; (2) Microwave signal adaptive preprocessing: sequentially perform temperature and humidity correction, forgetting factor dynamic adjustment and RLS iterative filtering to realize microwave signal purification; (3) Video image adaptive preprocessing: sequentially perform multi-scale Retinex enhancement, dust noise judgment, adaptive threshold BM3D filtering and weighted reconstruction to improve image clarity; (4) Multi-domain feature extraction: extract radar signal domain features of preprocessed radar signals using one-dimensional CNN network, extract radar frequency domain features by performing sliding window processing on preprocessed radar signals using one-dimensional FFT, and extract visual image domain features of denoised images using Yolo series or MobileNet series lightweight convolutional neural network; (5) Weighted fusion classification output: calculate radar confidence based on radar echo signal-to-noise ratio, perform weighted embedding vector splicing on normalized radar signal domain features, radar frequency domain features and visual image domain features according to weight proportions of 15%-25% for radar signal domain features, 15%-25% for radar frequency domain features and 50%-70% for visual image domain features, then use a multi-head Transformer model to strengthen the correlation between modal features, and finally output the classification probability of each type of pest, the number of pests in a single area and the average density through a Softmax activation function.

2. The method of claim 1, wherein, In step (1), the microwave radar subsystem uses a K-band radar with a transmission power of 0-30 dBm, a bandwidth of 200-400 MHz, and a sampling rate of ≥10 kHz. The signal is transmitted through a gain ≥8dBi antenna, and the frequency-shifted signal is amplified and filtered before being converted into time, amplitude and frequency three-dimensional signals by a 12-bit or higher ADC. The video image subsystem uses a camera with a pixel count of ≥8 million, equipped with a light supplement module, and outputs images with a frame rate of ≥10fps. Environmental parameters are collected by corresponding sensors and transmitted to the front-end processing module through an RS485 standard bus.

3. The method of claim 1, wherein, The specific process of the microwave signal adaptive preprocessing in step (2) is as follows: the temperature and humidity correction adjusts the correction coefficient according to the temperature and humidity of the environment where the grain pile is located. When the temperature or humidity deviates from the reference state, the correction coefficient is adjusted accordingly to compensate for the environmental impact. In the forgetting factor dynamic adjustment, the value of the forgetting factor is set dynamically according to the grain pile density, and the higher the grain pile density, the larger the value of the forgetting factor.

4. The method of claim 3, wherein, In step (2), the RLS iterative filtering calibrates the covariance matrix based on the initial noise variance of the grain pile, and realizes clutter suppression through the iterative process of error calculation, gain update, and coefficient and covariance update.

5. The method of claim 1, wherein, The specific process of the video image adaptive preprocessing in step (3) is: the multi-scale Retinex enhancement decomposes the image into a light component and a reflection component, the light component is smoothed by using three groups of different Gaussian kernels with sizes of 20-40 pixels, 40-80 pixels and 80-160 pixels respectively to remove background interference, the corresponding reflection components are extracted, and the reflection components are weighted and fused according to the weight proportions of 0.3, 0.5 and 0.2; The dust noise judgment divides the image into a plurality of square regions according to the side length of 16-64 pixels, calculates the gray variance of each region, and presets a gray variance threshold, the gray variance exceeding the threshold is determined as a dust pollution region, and the gray variance not exceeding the threshold is determined as a normal region.

6. The method of claim 5, wherein, The adaptive threshold BM3D filtering in step (3) performs filtering on the dust pollution region, first takes a reference block with a side length of 16-64 pixels, screens similar blocks in a search window with a size of 3-5 times the reference block size, and sorts the similar blocks from large to small to take the first 8-32 similar blocks, then performs 3D-DCT transformation, calculates an adaptive threshold based on the local statistics method of the gray mean and variance in the block, and performs threshold processing on the transformed coefficients; the weighted reconstruction performs weighted summation on the inverse transformed image block according to the similarity weight, and outputs a denoised image.

7. The method of claim 1, wherein, The specific process of the multi-domain feature extraction in step (4) is: the radar signal domain feature extraction adopts a one-dimensional CNN network, the network structure includes an input layer, at least 8 convolution layers, corresponding pooling layers and a full connection layer, and finally outputs 64-256 dimensional radar signal domain features; the radar frequency domain feature extraction adopts one-dimensional FFT to perform sliding window processing on the preprocessed radar signal, the window size is 1s-10s, and the output dimension is 64-256 dimensional radar frequency domain features; the visual image domain feature extraction adopts Yolo series or MobileNet series lightweight convolutional neural network to process the denoised image, and outputs pest contour coordinates, pest type and identification probability as visual image domain features.

8. The method of claim 1, wherein, In step (5): when weighted fusion, the radar confidence is calculated according to SNR, SNR≥20dB takes 1.0, and SNR<0dB takes 0, the radar signal domain, frequency domain and visual domain feature weights are 15%-25%, 15%-25% and 50%-70% respectively, and the fusion sequence is formed by splicing; the multi-head Transformer model is used for classification, after layer normalization, the pest classification probability, count value and average density are output by Softmax.

9. A multi-modal information fusion based intelligent monitoring device for implementing the method of claim 1, characterized in that, It includes a perception module, a front-end processing module and a center processing module, the perception module contains a microwave radar subsystem, a resolution ≥800 million pixel video image subsystem, an environment subsystem containing temperature, humidity, dust concentration, light intensity sensor; the front-end processing module contains a low-power processor with single-precision floating-point operation capability ≥300 GFLOPS, ≥4G memory, ≥256G solid state disk and ≥100Mbps Ethernet communication device, for executing microwave signal and video image adaptive preprocessing; the center processing module contains a high-performance processor with full-core integer operation capability ≥400 GIPS, a GPU processor with single-precision floating-point operation capability ≥40 TFLOPS, ≥64G memory, ≥8T hard disk and ≥1000Mbps network communication device, for executing multi-domain feature extraction and weighted fusion classification output.

Citation Information

Cited By

  • Intelligent pest monitoring system based on microwave matrix and compound sex attractant

    CN122096061A