Agricultural scene adaptive multi-modal feature extraction and fusion method and system

Through multimodal data collection and adaptive feature fusion methods, the problems of low efficiency and misjudgment of multimodal data processing in agricultural production have been solved, efficient and accurate pest and disease prediction has been achieved, and the adaptability of the agricultural monitoring system and precision agriculture decision-making support have been improved.

CN120808085APending Publication Date: 2025-10-17GUANGXI WANJIN NEW ENERGY TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510870781.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In agricultural production, there are difficulties in quickly processing multimodal data and accurately predicting pest and disease targets. Traditional methods are labor-intensive and time-consuming, and the limited spectral coverage of a single sensor and improper collection frequency lead to misjudgments and waste of resources.

Method used

Image sensors, soil sensors, environmental sensors, and acoustic sensors are used for multimodal data collection. ResNet-50, 1D-CNN, and lightweight MobileNetV3 networks are combined to extract features. Adaptive feature fusion is performed through the attention mechanism, and modal weights are dynamically adjusted to generate pest and disease prediction results.

Benefits of technology

It improves the environmental adaptability of multi-source heterogeneous data fusion, reduces the risk of misjudgment, improves the accuracy of pest and disease prediction and the reliability of the system, and reduces manual intervention and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808085A_ABST
    Figure CN120808085A_ABST
Patent Text Reader

Abstract

The invention provides an agricultural scene adaptive multi-modal feature extraction and fusion method and system, and belongs to the technical field of agricultural monitoring, and the method comprises the following steps: obtaining multi-modal data of crops through deploying an image sensor, a soil sensor, an environment sensor and an acoustic sensor in a farmland monitoring area; preprocessing the multi-modal data, wherein preprocessing comprises the steps of performing normalization and noise filtering on the data; visual feature vectors are extracted from the image data through an improved ResNet-50 network, time sequence features are extracted from the environment data through 1D-CNN, and acoustic data are input into a lightweight MobileNetV3 network to extract voiceprint features after being subjected to Mel spectrum conversion; calculating a modal weight based on a feature fusion algorithm of an attention mechanism; and outputting the fusion feature vector, inputting the fusion feature vector into a pest classifier and a growth state regression device, and generating a pest prediction result. Through multi-modal data acquisition, preprocessing, feature extraction and adaptive fusion, the modal weight is dynamically adjusted in combination with an attention mechanism, the multi-source heterogeneous data fusion effectiveness is improved, the disease and pest feature sensitivity is enhanced, and the effect of dynamic weight distribution is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of agricultural monitoring, and in particular to an agricultural scene adaptive multi-modal feature extraction and fusion method and system. BACKGROUND

[0002] In the process of agricultural production, the health of crops and the quality of soil have a decisive influence on the growth and yield of crops. In recent years, with the rapid development of Internet of Things technology, big data analysis and artificial intelligence technology, multi-modal data fusion technology has gradually attracted attention in the field of agriculture.

[0003] Multi-modal data fusion refers to the integration of information from different sensors or data sources to obtain a more complete and accurate description of the agricultural environment. By arranging multiple types of sensors in the agricultural production scene, a large amount of raw data information can be obtained, but if these raw data information is processed manually, it will consume an immeasurable labor and time cost. How to quickly process the data information collected by the sensor and accurately predict the pest target pointed to by the data information is a difficult problem that needs to be solved. Therefore, an agricultural scene adaptive multi-modal feature extraction and fusion method and system are needed. SUMMARY

[0004] The purpose of the present application is to provide an agricultural scene adaptive multi-modal feature extraction and fusion method and system to solve the existing technical problems.

[0005] In order to achieve the above-mentioned purpose, the technical solution adopted by the present application is as follows:

[0006] An agricultural scene adaptive multi-modal feature extraction and fusion method, comprising the following steps:

[0007] Step 1: Multi-modal data acquisition, image sensors, soil sensors, environmental sensors and acoustic sensors are deployed in the farmland monitoring area to obtain image data, environmental data and acoustic data of crops;

[0008] Step 2: Multi-modal data preprocessing, preprocessing the collected data, preprocessing includes normalizing and noise filtering the data;

[0009] Step 3: Multi-modal data feature extraction, image data uses an improved ResNet-50 network to extract visual feature vectors, environmental data uses a 1D-CNN to extract time series features, and acoustic data uses a Mel spectrum conversion to input a lightweight MobileNetV3 network to extract voiceprint features;

[0010] Step 4: Adaptive feature fusion, based on an attention mechanism feature fusion algorithm, calculate the modal weight;

[0011] Step 5: Disease and pest classification decision, output fusion feature vector input disease and pest classifier and growth state regressor, generate disease and pest prediction results.

[0012] Through the above technical solution, the multi-source heterogeneous data fusion problem is effectively solved, and the environmental adaptability of the farmland monitoring system is improved. Through the special feature extraction network, the potential information of each modal data is fully mined, and the dynamic weight adjustment mechanism is combined to ensure the system reliability when part of the sensor data is abnormal. The final fusion feature can accurately reflect the crop health status and provide reliable data support for disease and pest prediction, reducing the risk of misjudgment.

[0013] Further, the image sensor includes a visible light camera and a multispectral camera, and the collection frequency is 1 time per hour.

[0014] Through the above technical solution, the problem of misjudgment of plant health status caused by limited spectral coverage of a single sensor is solved, for example, misjudging leaf curling caused by water shortage as disease infection. At the same time, the waste of storage resources caused by too high collection frequency is avoided, for example, reducing the generation of invalid data during the night period when there is no significant environmental change. The spatio-temporal alignment characteristics of multispectral data and visible light data provide a matching spatio-temporal reference for subsequent feature fusion, ensuring the correspondence of different modal features in the same crop area and collection time point.

[0015] Further, the environmental sensor includes an air temperature and humidity sensor and an illumination sensor, the air temperature and humidity sensor is used to collect the air temperature and humidity in the monitoring area, and the illumination sensor is used to collect the illumination intensity in the monitoring area.

[0016] Through the above technical solution, the core environmental parameters affecting crop growth can be synchronously acquired, and the associated data set of air temperature and humidity and illumination intensity is established. The collaborative collection of multi-dimensional environmental parameters provides a data basis for subsequent analysis of the influence of environmental stress on crop physiological state, for example, identifying transpiration imbalance caused by high temperature and strong light or growth retardation caused by low temperature and weak light, thereby improving the prediction accuracy of the disease and pest warning model.

[0017] Further, the soil sensor is used to collect soil data in the monitoring area, the monitoring depth of the soil sensor is 0-50 cm, and the soil data includes the temperature and humidity, pH value, N / P / K content of the soil.

[0018] Through the above technical solution, full-dimensional data of soil physical state, chemical environment and nutrient supply level can be synchronously acquired, providing multi-scale data support for precise fertilization decision. By matching the monitoring depth with the crop root distribution characteristics, deep soil compaction or salinization risk can be effectively identified. Multi-parameter collaborative analysis can reveal the coupling relationship between soil humidity and nutrient migration, avoiding the problem of excessive irrigation or fertilization caused by misjudgment of a single parameter.

[0019] Further, the acoustic sensor sets a frequency response range of 20Hz-20kHz, and the deployment height is 30±5cm from the ground.

[0020] By the above technical solution, the feature signal missing detection problem caused by the frequency band missing of the acoustic sensor is solved, the sound wave propagation path distortion caused by the installation height deviation is eliminated, and complete and accurate acoustic feature vectors are provided for subsequent multi-modal fusion.

[0021] Further, in step 2, the image data is denoised by bilateral filtering and normalized to [0, 1],

[0022] The calculation formula of image data normalization is:

[0023]

[0024] I norm denotes the normalized image; I denotes the original image data; min(I) denotes the minimum intensity value of the image; max(I) denotes the maximum intensity value of the image;

[0025] The calculation formula of soil / environment data normalization is:

[0026]

[0027] S' denotes the normalized soil feature; S denotes the original soil parameter vector; μ s denotes the mean of the soil parameters; σ s denotes the standard deviation of the soil parameters; E' denotes the normalized environment feature; E denotes the original environment parameter vector; μ e denotes the mean of the environment parameters; σ e denotes the standard deviation of the environment parameters;

[0028] The acoustic data is extracted by extracting the mel frequency cepstral coefficient to generate a 128-dimensional feature vector.

[0029] By the above technical solution, the noise interference caused by the difference between sensors in multi-modal data is effectively eliminated, the numerical dimension of image, soil and environment data is unified, and discriminative acoustic features are extracted, so that the subsequent feature fusion module can accurately calculate the modal weight, thereby improving the processing accuracy of the disease and pest classifier on multi-source heterogeneous data.

[0030] Further, in step 3,

[0031] The calculation formula of image feature extraction is:

[0032] f img= ResNet50(I norm )∈R 2048

[0033] f img denotes image feature vector; ResNet50 denotes deep residual network model; I norm denotes normalized image data; R 2048 denotes 2048-dimensional real vector space;

[0034] The calculation formula of soil / environment feature extraction is:

[0035] f env= 1D-CNN(concat(E')) ∈ R 256

[0036] f env denotes environment fusion feature; 1D-CNN denotes one-dimensional convolutional neural network; concat denotes vector concatenation operation; E' denotes normalized environment feature; R 256 denotes 256-dimensional real number space;

[0037] The calculation formula of acoustic feature extraction is:

[0038] f aud = LSTM(MFCC(A)) ∈ R 512

[0039] f aud denotes acoustic event feature; LSTM denotes long short-term memory network; MFCC denotes mel frequency cepstral coefficient extraction; A denotes original acoustic data.

[0040] Through the above technical solution, efficient feature extraction of multi-modal data in an agricultural scene is realized, and the problem of insufficient feature expression of heterogeneous data is solved. The application of the residual network in the image feature extraction process avoids the degradation of the deep network, the convolution extraction of the soil environment time sequence feature enhances the capture ability of the dynamic correlation between parameters, and the long short-term memory network modeling of the acoustic signal improves the expression ability of the time-frequency feature, thereby providing a high-discrimination fusion feature input for subsequent pest classification.

[0041] Further, in step 4, the calculation formula of the modal weight is:

[0042]

[0043] ω denotes modal weight vector; softmax denotes normalized exponential function; W denotes trainable weight matrix; f img denotes image feature vector; f env denotes environment feature vector; f aud denotes acoustic feature vector; ⊕ denotes feature concatenation operation;

[0044] The calculation formula of the fusion feature is:

[0045] F fusion = ω img f img + ω env f env + ω aud f aud

[0046] F fusion denotes the fusion feature vector; ω img denotes the image feature weight; f img denotes the image feature vector; ω env denotes the environmental feature weight; f env denotes the environmental feature vector; ω aud denotes the acoustic feature weight; f aud denotes the acoustic feature vector.

[0047] Through the above technical solution, dynamic optimization distribution of multi-modal feature weights is realized, effectively solving the poor adaptability problem of fixed fusion strategy in complex agricultural scenes. By suppressing noise modal interference and enhancing the representation ability of key features, the classification accuracy and robustness of the pest prediction model are significantly improved. At the same time, this scheme can automatically adapt to changes in different crop varieties, growth stages and environmental conditions, providing reliable technical support for precision agriculture decision-making.

[0048] Further, in step 5, the calculation formula of pest classification decision is:

[0049] y = argmax (LightGBM (F fusion ))

[0050] y denotes the prediction result output; F fusion denotes the fusion feature vector.

[0051] Through the above technical solution, automatic feature fusion and joint inference of multi-source heterogeneous data are realized, effectively solving the problems of low data processing efficiency and high single sensor false alarm rate in traditional agricultural monitoring. Specifically, the fusion feature vector automatically extracts cross-modal correlation features through end-to-end training, avoiding the computational overhead of manually designed feature combination rules; the parallel structure of the classifier and the regressor can simultaneously output pest type and growth state quantitative indicators, providing double basis for precision agriculture decision-making; the adaptive weight mechanism can adjust the feature fusion strategy according to different crop varieties, for example, increasing the weight of environmental data for greenhouse crops and focusing on soil and acoustic feature analysis for field crops.

[0052] An agricultural scene adaptive multi-modal feature extraction and fusion system, comprising:

[0053] A multi-modal data acquisition module is configured to collect multi-modal data, which includes image data, soil data, environmental data and acoustic data of crops;

[0054] A multi-modal data preprocessing module is configured to normalize and / or filter noise from the image data, soil data, environmental data and acoustic data collected by the multi-modal data acquisition module;

[0055] A multi-modal data feature extraction module is configured to extract features from the data processed by the multi-modal data preprocessing module;

[0056] An adaptive feature fusion module is configured to calculate modal weights based on a feature fusion algorithm of an attention mechanism;

[0057] A disease and pest classification decision module is configured to input the fused feature vector into a disease and pest classifier and a growth state regressor to obtain a disease and pest prediction result.

[0058] The above technical solution solves the problem of low multi-modal data processing efficiency, reduces the need for manual intervention through an automatic feature extraction and fusion process, overcomes the defects of insufficient prediction accuracy of traditional methods, dynamically adjusts the contribution of each mode using an attention mechanism, enables the model to focus on key feature information, reduces data processing costs, and realizes efficient conversion from raw data to decision output through an end-to-end system architecture, thereby providing reliable technical support for precision agriculture.

[0059] The present application has the following advantages due to the use of the above technical solution:

[0060] The present application solves the problems of single data dimension and insufficient feature sensitivity of traditional methods by collecting, preprocessing, extracting features and adaptively fusing multi-modal data, dynamically adjusting modal weights using an attention mechanism, and has the effects of improving the effectiveness of multi-source heterogeneous data fusion, enhancing the sensitivity of disease and pest features, and realizing dynamic weight distribution. BRIEF DESCRIPTION OF DRAWINGS

[0061] Figure 1 is a flowchart of the multi-modal feature extraction and fusion method of the present application;

[0062] Figure 2 is a framework diagram of the multi-modal feature extraction and fusion system of the present application.

[0063] In the drawings, 1 is a multi-modal data acquisition module, 2 is a multi-modal data preprocessing module, 3 is a multi-modal data feature extraction module, 4 is an adaptive feature fusion module, and 5 is a disease and pest classification decision module. DETAILED DESCRIPTION

[0064] For the purposes of the present invention, the technical solutions and advantages thereof are more clearly apparent, further detailed below with reference to the drawings and preferred embodiments. However, it should be noted that many of the details listed in the description are only to enable the reader to have a thorough understanding of one or more aspects of the present invention, and the aspects of the present invention can be implemented even without these specific details.

[0065] As shown in Figure 1 An agricultural scene adaptive multi-modal feature extraction and fusion method, comprising the following steps:

[0066] Step 1: Multi-modal data acquisition, through the deployment of image sensors, soil sensors, environmental sensors and acoustic sensors in the farmland monitoring area, the image data, environmental data and acoustic data of crops are acquired;

[0067] The image sensor includes a visible light camera and a multispectral camera, and the acquisition frequency is 1 time per hour;

[0068] The environmental sensor includes an air temperature and humidity sensor and an illumination sensor, the air temperature and humidity sensor is used to acquire the air temperature and humidity in the monitoring area, and the illumination sensor is used to acquire the illumination intensity in the monitoring area;

[0069] The soil sensor is used to acquire the soil data in the monitoring area, the monitoring depth of the soil sensor is 0-50 cm, and the soil data includes the temperature and humidity, pH value, N / P / K content of the soil;

[0070] The acoustic sensor is set to have a frequency response range of 20 Hz-20 kHz, and the deployment height is 30±5 cm from the ground.

[0071] Step 2: Multi-modal data preprocessing, preprocessing the collected data, the preprocessing includes normalizing and noise filtering the data;

[0072] The image data is denoised by bilateral filtering and normalized to [0, 1],

[0073] The calculation formula of the image data normalization is:

[0074]

[0075] I norm represents the normalized image; I represents the original image data; min(I) represents the minimum intensity value of the image; max(I) represents the maximum intensity value of the image;

[0076] The calculation formula of the soil / environmental data normalization is:

[0077]

[0078] S' denotes normalized soil features; S denotes original soil parameter vector; μ s denotes soil parameter mean; σ s denotes soil parameter standard deviation; E' denotes normalized environment features; E denotes original environment parameter vector; μ e denotes environment parameter mean; σ e denotes environment parameter standard deviation;

[0079] Acoustic data generates 128-dimensional feature vector by extracting Mel-frequency cepstral coefficients.

[0080] Step 3: Multimodal data feature extraction, image data adopts improved ResNet-50 network to extract visual feature vector, environment data extracts time sequence feature through 1D-CNN, acoustic data extracts voiceprint feature after Mel spectrum conversion and inputting into lightweight MobileNetV3 network;

[0081] The calculation formula of image feature extraction is:

[0082] f img= ResNet50(I norm ) ∈ R 2048

[0083] f img denotes image feature vector; ResNet50 denotes deep residual network model; I norm denotes normalized image data; R 2048 denotes 2048-dimensional real vector space;

[0084] The calculation formula of soil / environment feature extraction is:

[0085] f env= 1D-CNN(concat(E')) ∈ R 256

[0086] f env denotes environment fusion feature; 1D-CNN denotes 1D convolutional neural network; concat denotes vector splicing operation; E' denotes normalized environment feature; R 256 denotes 256-dimensional real number space;

[0087] The calculation formula of acoustic feature extraction is:

[0088] f aud= LSTM(MFCC(A)) ∈ R 512

[0089] f aud denotes acoustic event feature; LSTM denotes long short-term memory network; MFCC denotes Mel-frequency cepstral coefficient extraction; A denotes original acoustic data.

[0090] Step 4: Adaptive feature fusion, feature fusion algorithm based on attention mechanism, calculate modal weight;

[0091] The calculation formula of modal weight is:

[0092]

[0093] ω represents the modal weight vector; softmax represents the normalized exponential function; W represents the trainable weight matrix; f img represents the image feature vector; f env represents the environmental feature vector; f aud represents the acoustic feature vector; and represents the feature splicing operation.

[0094] The calculation formula of the fused feature is:

[0095] F fusion = ω img f img + ω env f env + ω aud f aud

[0096] F fusion represents the fused feature vector; ω img represents the image feature weight; f img represents the image feature vector; ω env represents the environmental feature weight; f env represents the environmental feature vector; ω aud represents the acoustic feature weight; f aud represents the acoustic feature vector.

[0097] Step 5: Pest classification decision, output the fused feature vector to the pest classification classifier and the growth state regressor, and generate the pest prediction result;

[0098] The calculation formula of the pest classification decision is:

[0099] y = argmax(LightGBM(F fusion ))

[0100] y represents the prediction result output; F fusion represents the fused feature vector.

[0101] As shown in Figure 2 , an agricultural scene adaptive multi-modal feature extraction and fusion system comprises:

[0102] A multi-modal data acquisition module 1 is configured to collect multi-modal data, including image data, soil data, environmental data, and acoustic data of the crops;

[0103] The image data of the crops is acquired by deploying image sensors in the monitoring area, the image sensors including visible light cameras and multispectral cameras, and the acquisition frequency is 1 time per hour;

[0104] The soil data of the crops is acquired by deploying soil sensors in the monitoring area, the soil sensors being configured to collect soil data in the monitoring area, the soil sensors having a monitoring depth of 0-50 cm, and the soil data including the temperature and humidity, pH value, N / P / K content of the soil;

[0105] The environmental data of the crops is acquired by deploying environmental sensors in the monitoring area, the environmental sensors including air temperature and humidity sensors and light sensors, the air temperature and humidity sensors being configured to collect air temperature and humidity in the monitoring area, and the light sensors being configured to collect light intensity in the monitoring area;

[0106] The acoustic data of the crops is acquired by deploying acoustic sensors in the monitoring area, the acoustic sensors having a frequency response range of 20 Hz-20 kHz and being disposed at a height of 30±5 cm from the ground;

[0107] A multi-modal data preprocessing module 2 is configured to normalize and / or filter noise from the image data, soil data, environmental data, and acoustic data collected by the multi-modal data acquisition module;

[0108] A multi-modal data feature extraction module 3 is configured to extract features from the data processed by the multi-modal data preprocessing module;

[0109] The image data is extracted to a visual feature vector using an improved ResNet-50 network, the soil / environmental data is extracted to a time series feature using a 1D-CNN, and the acoustic data is converted to a Mel spectrum and then input to a lightweight MobileNetV3 network to extract a voiceprint feature;

[0110] An adaptive feature fusion module 4 is configured to calculate modal weights based on a feature fusion algorithm of an attention mechanism;

[0111] A disease and pest classification decision module 5 is configured to input the fused feature vector into a disease and pest classifier and a growth state regressor to obtain a disease and pest prediction result.

[0112] Embodiments

[0113] Taking a wheat planting area (10 mu) of a smart farm in Nanning, Guangxi Zhuang Autonomous Region as an example

[0114] Multi-modal data acquisition

[0115] Sensor deployment: visible light camera (Sony IMX586, resolution 8000x6000) and multispectral camera (Parrot Sequoia+, 5 bands) collect 1 image per hour, for 30 days, resulting in 720 image data.

[0116] Soil sensor (Decagon EC-5) monitors at 0-50 cm depth, with data examples as follows:

[0117] Temperature: 18.5°C (μ=20.3°C, σ=2.1); Humidity: 32.4% (μ=35.1%, σ=4.8) pH: 6.8 (μ=6.5, σ=0.3); NPK content: N=25.6 mg / kg, P=12.3 mg / kg, K=18.7 mg / kg.

[0118] Environmental sensors: air temperature and humidity (SHT31): temperature 25.7°C (μ=26.2°C, σ=1.8), humidity 65.2% (μ=62.4%, σ=5.3); light intensity (BH1750): 1200 Lux (μ=1100, σ=250).

[0119] Acoustic sensor (ReSpeaker Mic Array, sampling rate 44.1 kHz) records insect activity sound patterns (e.g. aphid wingbeat frequency 2-5 kHz).

[0120] Features extracted after preprocessing of multi-modal data are as follows:

[0121] Image features: improved ResNet-50 (remove fully connected layer, add adaptive pooling) outputs 2048-dimensional vector, e.g. f img = [0.45, -0.12,..., 1.23] ∈ R 2048

[0122] Environmental features: 1D-CNN (3 layers of convolution, kernel size=5) processes concatenated soil environment data (S' ⊕ E'), outputs 256-dimensional vector: f env = [0.78, -0.34,..., 0.56] ∈ R 256

[0123] Acoustic features:

[0124] LSTM (hidden layer 128 units) processes MFCC sequence, outputs 512-dimensional vector:

[0125] f aud = [-0.89, 0.23,..., 0.45] ∈ R 512

[0126] Then the attention mechanism (weight matrix W ∈ R3×3072) is used to calculate the weight of each modality:

[0127] ω = softmax([0.32, 0.41, 0.27]) image: 0.32, environment: 0.41, acoustic: 0.27

[0128] After weighted splicing, a 2816-dimensional vector is generated:

[0129] F fusion = 0.32 × fi mg ⊕ 0.41 × f env ⊕ 0.27 × f aud ∈ R 2 815

[0130] Finally, the classification decision and the result

[0131] 1. Disease and pest classifier (3-layer fully connected + Softmax) output:

[0132] Aphid infestation probability: 87.6%

[0133] Rust probability: 12.1%

[0134] Healthy probability: 0.3%

[0135] 2. Growth state regressor (MSE loss) predicts plant height as 82.3 cm (true value 80.5 cm, error 2.2%).

[0136] Effect comparison (30-day data)

[0137] Method Disease and pest accuracy Growth state error Single modality (image) 72.4% 5.8% Traditional multi-modal fusion 83.1% 3.5% The present invention 91.3% 2.2%

[0138] The present application significantly improves the accuracy of disease and pest recognition and reduces the prediction error of growth state by adaptive feature fusion.

[0139] The above is only the preferred embodiment of the present application, it should be noted that for those skilled in the art, without departing from the principles of the present application, can make a number of improvements and refinements, these improvements and refinements should also be considered as the protection scope of the present application.

Claims

1. An adaptive multimodal feature extraction and fusion method for agricultural scenes, characterized by: The following steps are involved: Step 1: Multimodal data acquisition: Deploy image sensors, soil sensors, environmental sensors, and acoustic sensors in the farmland monitoring area to acquire crop image data, environmental data, and acoustic data. Step 2: Multimodal data preprocessing: preprocessing the collected data, including normalization and noise filtering; Step 3: Multimodal data feature extraction: an improved ResNet-50 network is used to extract visual feature vectors from image data, 1D-CNN is used to extract temporal features from soil / environmental data, and acoustic data is converted to Mel spectrum and then input into a lightweight MobileNetV3 network to extract voiceprint features. Step 4: Adaptive feature fusion, a feature fusion algorithm based on the attention mechanism, calculates modal weights; Step 5: Pest and disease classification decision, output fusion feature vector input into the pest and disease classifier and growth status regressor to generate pest and disease prediction results.

2. The method for adaptive multimodal feature extraction and fusion of agricultural scenes according to claim 1, characterized in that: The image sensor includes a visible light camera and a multispectral camera, and the acquisition frequency is once per hour.

3. The method for adaptive multimodal feature extraction and fusion of agricultural scenes according to claim 1, characterized in that: Environmental sensors include air temperature and humidity sensors and light sensors. The air temperature and humidity sensors are used to collect air temperature and humidity in the monitoring area, and the light sensors are used to collect light intensity in the monitoring area.

4. The method for adaptive multimodal feature extraction and fusion of agricultural scenes according to claim 1, characterized in that: The soil sensor is used to collect soil data within the monitoring area. The monitoring depth of the soil sensor is 0-50cm. The soil data includes soil temperature and humidity, pH value, and N / P / K content.

5. The method for adaptive multimodal feature extraction and fusion of agricultural scenes according to claim 1, characterized in that: The acoustic sensor is set to a frequency response range of 20Hz-20kHz and a deployment height of 30±5cm from the ground.

6. The method for adaptive multimodal feature extraction and fusion of agricultural scenes according to claim 1, characterized in that: In step 2, the image data is denoised using bilateral filtering and normalized to [0, 1]. The calculation formula for image data normalization is: I norm Represents the normalized image; I represents the original image data; min(I) represents the minimum intensity value of the image; max(I) represents the maximum intensity value of the image; The calculation formula for normalization of soil / environmental data is: S′ represents the standardized soil characteristics; S represents the original soil parameter vector; μ s represents the mean value of soil parameters; σ s represents the standard deviation of soil parameters; E′ represents the standardized environmental characteristics; E represents the original environmental parameter vector; μ e represents the mean value of environmental parameters; σ e represents the standard deviation of environmental parameters; The acoustic data is extracted by extracting Mel-frequency cepstral coefficients to generate a 128-dimensional feature vector.

7. The method for adaptive multimodal feature extraction and fusion of agricultural scenes according to claim 1, characterized in that: In step 3, The calculation formula for image feature extraction is: f img= ResNet50(I norm )∈R 2048 f img Represents the image feature vector; Res Net50 represents the deep residual network model; I norm Represents the normalized image data; R 2048 Represents a 2048-dimensional real vector space; The calculation formula for soil / environmental feature extraction is: f env= 1D-CNN(concat(E′))∈R 256 f env represents the environmental fusion feature; 1D-CNN represents the dimensional convolutional neural network; concat represents the vector concatenation operation; E′ represents the standardized environmental feature; R 256 Represents 256-dimensional real number space; The calculation formula for acoustic feature extraction is: f aud= LSTM(MFCC(A))∈R 512 f aud Represents acoustic event features; LSTM represents long short-term memory network; MFCC represents Mel-frequency cepstral coefficient extraction; A represents original acoustic data.

8. The method for adaptive multimodal feature extraction and fusion of agricultural scenes according to claim 1, characterized in that: In step 4, the formula for calculating the modal weight is: ω represents the modal weight vector; softmax represents the normalized exponential function; W represents the trainable weight matrix; f img represents the image feature vector; f env represents the environmental feature vector; f aud Represents the acoustic feature vector; ⊕ represents the feature concatenation operation; The calculation formula of fusion features is: F fusion =ω img f img +oh env f env +oh aud f aud F fusion represents the fused feature vector; ω img Represents the image feature weight; f img represents the image feature vector; ω env represents the weight of environmental characteristics; f env represents the environmental feature vector; ω aud represents the acoustic feature weight; f aud represents the acoustic eigenvector.

9. The method for adaptive multimodal feature extraction and fusion of agricultural scenes according to claim 1, characterized in that: In step 5, the calculation formula for pest and disease classification decision is: y=arg max(LightGBM(F fusion )) y represents the prediction result output; F fusion Represents the fused feature vector.

10. An agricultural scene adaptive multimodal feature extraction and fusion system, characterized by: include: A multimodal data acquisition module is used to collect multimodal data, including crop image data, soil data, environmental data, and acoustic data; A multimodal data preprocessing module is used to normalize and / or filter out noise from the image data, soil data, environmental data, and acoustic data collected by the multimodal data acquisition module; A multimodal data feature extraction module is used to extract features from the data processed by the multimodal data preprocessing module; Adaptive feature fusion module, which calculates modal weights based on the feature fusion algorithm of the attention mechanism; The pest and disease classification decision module is used to output the fusion feature vector and input it into the pest and disease classifier and growth status regressor to obtain the pest and disease prediction results.

Citation Information

Cited By

  • Facility agriculture disease and pest prevention and control method and device based on deep learning

    CN121837804A

  • Gold ore component online detection method and system based on machine vision

    CN122193120A

  • A machine vision-based online detection method and system for gold ore composition

    CN122193120B