Method and system for evaluating health status of figs based on multi-modal data fusion

By using a multimodal data fusion method, multimodal feature vectors were constructed and an adaptability evaluation model was used to solve the problems of lagging early identification and low early warning accuracy in fig pest and disease monitoring, thus achieving early identification and accurate classification of fig pests and diseases.

CN121661548BActive Publication Date: 2026-04-21CHENGDU IND VOCATIONAL TECHN COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU IND VOCATIONAL TECHN COLLEGE
Filing Date
2026-02-06
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies for monitoring fig pests and diseases suffer from problems such as delayed early identification, insufficient multimodal data fusion, weak correlation between physiological indicators and pests and diseases, and low early warning accuracy, which cannot meet the needs of 'early detection and early prevention' of major fig pests and diseases.

Method used

A multimodal data fusion method was adopted, which extracts physiological indices, spectroscopic morphological features and leaf texture features by acquiring multispectral images, hyperspectral cubes and visible light images, constructs multimodal feature vectors, and uses an adaptability assessment model to perform deep fusion and feature association, outputting adaptability scores and pest classification.

Benefits of technology

It improves the completeness of feature representation and the accuracy of assessment, enhances the precision and robustness of pest and disease classification, and enables early identification and accurate classification of fig pests and diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661548B_ABST
    Figure CN121661548B_ABST
Patent Text Reader

Abstract

This invention relates to a method and system for assessing the health status of fig trees based on multimodal data fusion. First, a multimodal feature vector is constructed by fusing features from multispectral images, hyperspectral cubes, and visible light images. Then, a multimodal fusion-based adaptability assessment model is proposed, which preprocesses the multimodal feature vectors and performs deep fusion and feature association to obtain fused features. Finally, based on the fused features, an adaptability score and pest / disease classification are output. This significantly enhances the completeness of feature representation, improves assessment accuracy and pest / disease classification precision, and strengthens robustness and anti-interference capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence and its application in agricultural information technology, particularly to a method and system for assessing the health status of figs based on multimodal data fusion. Background Technology

[0002] Currently, fig pest and disease monitoring mainly relies on traditional visual inspection, single-sensor monitoring, machine learning recognition based on visible light images, and conventional chemical detection technologies. Specific technical solutions and their limitations are as follows:

[0003] 1. Traditional visual inspection method: Agricultural technicians conduct regular field inspections to observe whether leaves show visible lesions (such as orange-yellow spore masses for rust, or brown sunken spots for anthracnose) or signs of pest infestation (such as chlorotic spots and spider web residue caused by spider mites). This is combined with experience to determine the type and severity of the pest or disease. This method relies heavily on manual experience, is highly subjective, and can only identify pests or diseases in the mid-to-late stages of development (after lesions have formed or significant damage has been caused). At this point, the physiological damage to the plant is irreversible, and the effectiveness of control measures is greatly reduced.

[0004] 2. Single-sensor monitoring technology: Some studies use ground-based sensors (such as chlorophyll meters and portable spectrometers) or UAV single-modal sensors (such as visible light cameras and single-band multispectral cameras) to collect data. For example, handheld chlorophyll meters are used to measure SPAD values ​​to determine chlorophyll content, but this method can only sample at a single point, with a small coverage area (requiring 50+ sampling points per acre), and is time-consuming and labor-intensive; UAV visible light cameras identify pests and diseases by capturing leaf color changes and surface damage, but visible light images only reflect surface morphological changes and cannot capture internal physiological stresses (such as early chlorophyll decomposition and water loss), with an identification rate of less than 30% for latent pests and diseases (such as the initial stage of spider mite infestation and the early stage of rust inoculation).

[0005] 3. Machine Learning Recognition Methods Based on Visible Light Images: Existing technologies, some studies employ convolutional neural networks (CNNs) to classify pests and diseases in UAV visible light images, such as using ResNet and YOLO models to identify leaf lesions or pest traces. However, these methods are essentially morphological matching of "image-lesion / pest traces," relying on a large number of labeled samples. Early-stage pests and diseases lack obvious morphological features, leading to poor model generalization. For example, for early-stage rust (3-5 days after inoculation) and early-stage spider mite infestation (mite density <5 mites / leaf), the CNN model accuracy is only 52%, and it cannot distinguish between physiological stress and environmental stress (such as yellowing leaves caused by drought or nutrient deficiency).

[0006] 4. Traditional Physiological Indicator Association Methods: A few studies have attempted to associate pests and diseases with single physiological indicators (such as NDVI), but no multi-indicator collaborative assessment mechanism has been established, and major pests are not covered. For example, using only the decrease in NDVI to judge pests and diseases is easily affected by light and soil background (such as NDVI values ​​being generally lower on cloudy days), resulting in a false alarm rate as high as 40%; and it cannot distinguish between different pests and diseases (such as rust, anthracnose, and spider mites, all of which cause a decrease in NDVI, but with different physiological mechanisms), making accurate classification difficult.

[0007] 5. Delayed Response of Early Warning Systems: Most existing early warning systems follow a passive process of "appearance of lesions / insect traces - image recognition - early warning issuance," failing to proactively assess "physiological stress - probability of pests and diseases." For example, in one orchard, the average time from the appearance of lesions / insect traces to the issuance of an early warning is 48 hours. By this time, the pathogen has already spread or the pests have already proliferated, requiring large-scale pesticide application and increasing pesticide usage by more than 30%.

[0008] In summary, existing technologies suffer from problems such as delayed early identification, insufficient multimodal data fusion, weak correlation between physiological indicators and pests and diseases (including pests and diseases), and low early warning accuracy. These technologies cannot meet the needs of "early detection and early prevention" of major pests and diseases of figs, and are technical problems that urgently need to be solved in this field. Summary of the Invention

[0009] To address at least one of the aforementioned technical problems, this invention provides a method for assessing the health status of fig trees based on multimodal data fusion, comprising:

[0010] Collect multispectral, hyperspectral cube, and visible light images of the orchard;

[0011] Based on the multispectral image, physiological index features are extracted to construct a multispectral feature vector; based on the hyperspectral cube, spectroscopic morphological features and dimensionality reduction features are extracted to construct a hyperspectral feature vector; based on the visible light image, leaf texture features are extracted to construct a visible light feature vector; and the multispectral feature vector, hyperspectral feature vector, and visible light feature vector are concatenated to obtain a multimodal feature vector.

[0012] The multimodal feature vectors are input into the trained fitness evaluation model, which consists of an input layer, a feature fusion layer, and an output layer connected in sequence. The input layer is called to preprocess the multimodal feature vectors and use them as model input. The feature fusion layer is called to perform deep fusion and feature association on the model input to obtain fused features. The output layer is called to output the fitness score and pest classification based on the fused features.

[0013] Furthermore, the steps of extracting physiological index features and constructing multispectral feature vectors based on multispectral images include:

[0014] Based on the reflectance of each band in the multispectral image, several physiological index features are calculated to construct a multidimensional multispectral feature vector; the physiological index features include any one or more of the normalized vegetation index, normalized water index, ratio vegetation index, greenness normalized vegetation index, and soil-adjusted vegetation index.

[0015] Furthermore, the steps of extracting spectroscopic morphological features and dimensionality reduction features from the hyperspectral cube to construct a hyperspectral feature vector include:

[0016] Based on the hyperspectral cube, several spectroscopic morphological features are calculated to construct a first hyperspectral feature sub-vector; the spectroscopic morphological features include any one or more of the following: red edge position, absorption valley depth, reflection peak width, flavonoid index, blue edge position, yellow edge position, red edge slope, blue edge slope, reflection peak height, and absorption valley area.

[0017] The hyperspectral cube is reduced in dimensionality by using an autoencoder to learn the low-dimensional representation and outputting multidimensional latent vectors as dimensionality-reduced features to construct the second hyperspectral feature sub-vector.

[0018] The first hyperspectral feature vector and the second hyperspectral feature vector are concatenated to obtain the hyperspectral feature vector.

[0019] Further, the steps of extracting leaf texture features and constructing visible light feature vectors based on visible light images include:

[0020] Based on the visible light image, multidimensional LBP features are extracted to obtain LBP sub-vectors;

[0021] Based on the visible light image, multidimensional GLCM features are extracted to obtain GLCM sub-vectors;

[0022] By concatenating the LBP subvector and the GLCM subvector, the visible light feature vector is obtained.

[0023] Furthermore, the input layer processing steps include:

[0024] The multimodal feature vector is treated as a sequence of a preset length, with each feature dimension being a token in the sequence;

[0025] Different positional codes are constructed for the tokens of multispectral feature vector, hyperspectral feature vector, and visible light feature vector in the multimodal feature vector;

[0026] The mapping relationship between each token and physiological indicators is constructed to obtain the model input.

[0027] Furthermore, the feature fusion layer includes several TransformerEncoder layers;

[0028] Each TransformerEncoder layer includes a multi-head self-attention module and a feedforward network; the multi-head self-attention module includes multiple attention heads and fully connected layers; each attention head takes the model input, calculates the query, key, and value, and obtains a single-head attention output; then the single-head attention outputs are concatenated and passed through a linear layer to obtain the multi-head attention output, in order to learn the association weights between features;

[0029] A feedforward network is used to perform a nonlinear transformation on the multi-head attention output to obtain the transformed features.

[0030] Based on the transformed features, the TransformerEncoder layers are fused to obtain the fused features.

[0031] Furthermore, the output layer, including the dual task header, is as follows:

[0032] The adaptation score regression head is used to output the adaptation score; specifically, it performs global average pooling on the fused features to obtain a global channel average feature vector, which is then mapped to the adaptation score.

[0033] The pest and disease type classification header is used to output pest and disease classifications. Specifically, it outputs the probability distribution of each type based on the global channel average feature vector.

[0034] Model training employs a dual-task joint training approach;

[0035] The total loss function is a weighted sum of the regression loss and the classification loss;

[0036] The regression loss uses mean squared error loss, and the goal is to minimize the difference between the predicted score and the true score.

[0037] The classification loss uses cross-entropy loss, and the goal is to optimize the matching between the probability distribution of each type and the label.

[0038] Furthermore, the steps for constructing the sample set during the training of the suitability assessment model include:

[0039] Based on the pathogenesis of various diseases and pests of figs, a set of physiological indicators reflecting the early characteristics of diseases and pests and the index weights of each physiological indicator were constructed.

[0040] Based on the set of physiological indicators and the preset set of normal thresholds, determine the set of abnormality function that reflects the degree of abnormality of each indicator;

[0041] The anomaly function set and weighting coefficients of each physiological indicator are used to determine the weighted fit score of each physiological stress-disease and pest.

[0042] Based on the threshold range of each physiological stress-disease / pest fit score, disease / pest types are mapped to construct a sample set.

[0043] Furthermore, the classification of diseases and pests includes: healthy plants, rust, anthracnose, and spider mites;

[0044] A set of physiological indicators reflecting rust, anthracnose, and spider mites, including chlorophyll, water, cell structure, and secondary metabolite indicators;

[0045] Physiological indicators reflecting rust are concentrated, with chlorophyll and cell structure indicators having higher weighting coefficients than water and secondary metabolite indicators.

[0046] Physiological indicators reflecting anthrax are concentrated, with water and secondary metabolite indicators having higher weighting coefficients than chlorophyll and cell structure indicators.

[0047] The physiological indicators reflecting spider mites show that the weighting coefficients of chlorophyll, water, and cell structure indicators are higher than those of secondary metabolites.

[0048] On the other hand, the present invention also provides a computer system, including a memory and a processor; the memory stores program code executable by the processor; the program code is used to execute any of the above-described fig health status assessment methods.

[0049] The present invention provides a method and system for assessing the health status of figs based on multimodal data fusion. First, a multimodal feature vector is constructed by fusing features from multispectral images, hyperspectral cubes, and visible light images. Then, a multimodal fusion-based adaptability assessment model is proposed, which preprocesses the multimodal feature vectors and performs deep fusion and feature association to obtain fused features. Finally, based on the fused features, an adaptability score and pest and disease classification are output. This significantly enhances the integrity of feature representation, improves assessment accuracy and pest and disease classification precision, and strengthens robustness and anti-interference ability. Attached Figure Description

[0050] Figure 1 This is a flowchart of an embodiment of the fig health status assessment method based on multimodal data fusion according to the present invention;

[0051] Figure 2 This is a flowchart of another embodiment of the fig health status assessment method based on multimodal data fusion according to the present invention. Detailed Implementation

[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0053] It should be noted that if the embodiments of the present invention involve directional indications, such as up, down, left, right, front, back, etc., these directional indications are only used to explain the relative positional relationships and movement of the components in a specific posture. If the specific posture changes, the directional indications will also change accordingly. Furthermore, if the embodiments of the present invention involve descriptions such as "first," "second," "S1," "S2," "step one," "step two," etc., these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance, or implicitly indicating the number of technical features indicated or the order of method execution. Those skilled in the art will understand that anything that does not violate the inventive concept and is within the scope of the present invention should be included in the protection scope of the present invention.

[0054] like Figures 1 to 2 As shown, this invention provides a method for assessing the health status of figs based on multimodal data fusion:

[0055] S1: Acquire multispectral, hyperspectral cube, and visible light images of the orchard;

[0056] Specifically, multispectral images, hyperspectral cubes, and visible light images can be acquired using drones to form raw multimodal data, represented as: triples. ;in:

[0057] Multispectral images, optional (Spatial resolution) The wavelengths include: 450nm, 550nm, 650nm, 750nm, 850nm, and 1640nm.

[0058] Hyperspectral cube, optional , (Band: 350nm-1000nm, 1nm interval);

[0059] Visible light image (RGB channels), optional .

[0060] S2: Based on the multispectral image, extract physiological index features and construct a multispectral feature vector; based on the hyperspectral cube, extract spectroscopic morphological features and dimensionality reduction features and construct a hyperspectral feature vector; based on the visible light image, extract leaf texture features and construct a visible light feature vector; and concatenate the multispectral feature vector, hyperspectral feature vector and visible light feature vector to obtain a multimodal feature vector.

[0061] Specifically, the optional feature extraction module includes a first feature extraction unit, a second feature extraction unit, a third feature extraction unit, and a splicing unit to implement step S2.

[0062] In a preferred embodiment, step S2, the step of extracting physiological index features based on the multispectral image and constructing a multispectral feature vector, includes:

[0063] S21: Based on the reflectance of each band in the multispectral image, calculate several physiological index features and construct a multidimensional multispectral feature vector; the physiological index features include any one or more of the normalized vegetation index, normalized water index, ratio vegetation index, greenness normalized vegetation index, and soil-adjusted vegetation index.

[0064] Specifically, the first feature extraction unit may optionally include five computational subunits. Based on the reflectance of each band of the multispectral image, it calculates five physiological index features to construct a five-dimensional multispectral feature vector. Specifically, it constructs a five-dimensional multispectral feature vector based on the six bands of the multispectral image. , Optional:

[0065] Normalized Difference Vegetation Index ( wavelength (reflectance at the location)

[0066] Normalized Moisture Index (Calculated directly using the 1640nm band contained in the multispectral data);

[0067] Ratio vegetation index ;

[0068] Greenness Normalized Difference Vegetation Index ;

[0069] Soil-adjusted vegetation index (Optional) (soil background correction coefficient).

[0070] In a preferred embodiment, step S2, the step of extracting spectroscopic morphological features and dimensionality reduction features based on the hyperspectral cube, and constructing a hyperspectral feature vector, includes:

[0071] S22: Based on the hyperspectral cube, calculate several spectroscopic morphological features to construct a first hyperspectral feature sub-vector; the spectroscopic morphological features include: red edge position, absorption valley depth, reflection peak width, flavonoid index, blue edge position, yellow edge position, red edge slope, blue edge slope, reflection peak height, and any one or more of absorption valley area; specifically, 10-dimensional spectroscopic morphological features can be optionally calculated, including:

[0072] Red border position Calculated by interpolation of reflectivity in the 680-750nm band:

[0073] ;

[0074] Absorption Valley Depth The depth of the chlorophyll absorption valley is 670nm. , The reflection peak is on the left side of the absorption valley. This is the minimum value of the absorption valley;

[0075] Reflection peak width The full width at half maximum (FWHM) of the 550nm green peak. ( (The left and right wavelengths corresponding to 50% height of the green peak).

[0076] Flavonoid index (Reflectivity ratio in the 350-500nm band);

[0077] The other 6 dimensions can be selected, including the blue edge position (490-530nm), yellow edge position (560-600nm), red edge slope, blue edge slope, reflection peak height, and absorption valley area, for a total of 10 dimensions. It is worth noting that these 10-dimensional features are only for illustrative purposes, and some of them or other spectroscopic morphological features may be selected.

[0078] S23: Dimensionality reduction of the hyperspectral cube is performed by using an autoencoder to learn a low-dimensional representation, outputting a multi-dimensional latent vector as the dimensionality-reduced feature, and constructing a second hyperspectral feature sub-vector; specifically, this includes an encoder and a decoder:

[0079] The encoder is used to take a 300-dimensional spectral vector as input from the hyperspectral cube, and output a 10-dimensional latent vector as a dimensionality reduction feature through a 3-layer fully connected network (300→128→64→10), which contains key physiological stress information of the hyperspectral data (covering the spectral anomaly features corresponding to three diseases and pests).

[0080] Decoder: 10→64→128→300, used to reconstruct the spectral vector;

[0081] Training objective: Optional: Minimize reconstruction error , It is a hyperspectral cube.

[0082] S24: Concatenate the first and second hyperspectral feature vectors to obtain the hyperspectral feature vector; for example, the hyperspectral feature vector... .

[0083] Preferably, the second feature extraction unit may include 10 computational subunits, an encoder, a decoder, and a first splicing subunit;

[0084] Ten computational sub-units are used to construct the first hyperspectral feature sub-vector; an encoder and a decoder are used to construct the second hyperspectral feature sub-vector; and a first splicing sub-unit is used to splice the first and second hyperspectral feature sub-vectors.

[0085] In a preferred embodiment, step S2, the step of extracting leaf texture features and constructing a visible light feature vector based on the visible light image, includes:

[0086] S25: Extract multi-dimensional LBP features from the visible light image to obtain LBP sub-vectors; specifically, Local Binary Pattern (LBP) can be used. After grayscale conversion of the visible light image, calculate the 8-neighborhood LBP histogram, and take the frequency of the first 5 bins (such as the proportion of uniform mode, speckle density, etc.) to obtain 5-dimensional LBP sub-vectors.

[0087] S26: Based on the visible light image, extract multidimensional GLCM features to obtain GLCM sub-vectors; specifically, calculate the contrast, energy, entropy, and correlation in four directions (0°, 45°, 90°, 135°), take the average, and obtain a 5-dimensional GLCM feature vector.

[0088] S27: Concatenate the LBP subvector and the GLCM subvector to obtain the visible light feature vector; Example: .

[0089] In this embodiment, from Leaf texture features are extracted to reflect changes in surface wrinkles and tiny spots caused by early cell structure damage (rust, anthracnose) and insect infestation (red spider mites). A method combining Local Binary Pattern (LBP) and Gray-Level Co-occurrence Matrix (GLCM) is used to complementarily extract local texture details and global / regional texture statistical features of the image, significantly improving feature representation ability and algorithm robustness.

[0090] Preferably, the third feature extraction unit may include an LBP sub-unit, a GLCM sub-unit, and a second concatenation sub-unit; the LBP sub-unit is used to extract LBP sub-vectors; the GLCM sub-unit is used to extract GLCM sub-vectors; and the second concatenation sub-unit is used to concatenate the LBP sub-vectors and GLCM sub-vectors.

[0091] S28: Concatenate the multispectral feature vector, hyperspectral feature vector, and visible light feature vector to obtain a multimodal feature vector. Specifically, the concatenation unit combines the above 5-dimensional multispectral feature vectors... 20-dimensional hyperspectral eigenvectors 10-dimensional visible light feature vector Concatenate them into a 35-dimensional multimodal feature vector. As input to the model:

[0092] ;

[0093] in, (5-dimensional) (20 dimensions) (10-dimensional) Aligned under the semantics of "physiological stress-adaptability", for example: NDVI corresponding , REP corresponding , FlavR corresponds to .

[0094] S3: Input the multimodal feature vectors into the trained fitness evaluation model. The fitness evaluation model includes an input layer, a feature fusion layer, and an output layer connected in sequence. Call the input layer to preprocess the multimodal feature vectors and use them as model input. Call the feature fusion layer to perform deep fusion and feature association on the model input to obtain fused features. Call the output layer to output the fitness score and pest classification based on the fused features.

[0095] Specifically, based on the processing steps of the adaptability evaluation model, the following explanation will further illustrate the adaptability evaluation model from two aspects: its model structure and training steps. Only preferred embodiments of the model structure and training steps are shown, but they are not limited thereto.

[0096] (1) The processing steps of the input layer include:

[0097] a1: Treat the multimodal feature vector as a sequence of preset length, with each feature dimension being a token in the sequence;

[0098] Specifically: the 35-dimensional multimodal feature vector is considered as the length sequence ,in That is, each feature dimension is a token in the sequence;

[0099] a2: Construct different positional codes for the tokens of the multi-spectral feature vector, hyperspectral feature vector, and visible light feature vector in the multi-modal feature vector;

[0100] Example: Adding modality-specific position encoding To distinguish different modes:

[0101] ;

[0102] in ( Each feature vector (e.g., the first 5 tokens), hyperspectral feature vector (e.g., the middle 20 tokens), and visible light feature vector (e.g., the last 10 tokens) is assigned a different position offset, such as 0.1 / 0.2 / 0.3, to ensure the positional encoding differences between modalities and enhance modal differentiation.

[0103] ;

[0104] a3: Construct the mapping relationship between each token and physiological indicators to obtain the model input; Example: (Multispectral Dimension 1): NDVI (corresponding) ); (Multispectral 2nd Dimension): NDWI (corresponding to) ); (Hyperspectral 2nd dimension): REP (corresponding) ); (Hyperspectral 3rd dimension): FlavR (corresponding to) Through this mapping, the model can explicitly associate input features with the anomaly calculation of physiological indicators. - This will improve the efficiency of utilizing the key characteristics of the three pests and diseases.

[0105] (2) Feature fusion layer, which may include, but is not limited to, several TransformerEncoder layers;

[0106] b1: Each TransformerEncoder layer includes a multi-head self-attention module and a feedforward network; the multi-head self-attention module includes multiple attention heads and fully connected layers; each attention head takes the model input, calculates the query, key, and value, and obtains a single-head attention output; then the single-head attention outputs are concatenated and passed through a linear layer to obtain the multi-head attention output. To learn the association weights between features;

[0107] b2: Feedforward network, used to perform a non-linear transformation on the multi-head attention output to obtain the transformed features. ;

[0108] b3: Based on the transformed features, fuse the TransformerEncoder layers to obtain the fused features. .

[0109] In a preferred embodiment, the feature fusion layer employs a 6-layer TransformerEncoder;

[0110] Each layer contains a multi-head attention module and a feedforward network (FFN); the multi-head attention module can optionally have 8 attention heads to attend to the model input, such as the preprocessed model input. Calculate query ,key ,value The association weights between learned features are used, such as the synergistic effect of NDVI and REP in rust disease, and the synergistic effect of NDWI and FlavR in spider mites, as follows:

[0111] ;

[0112] in, These are the query weight matrix, key weight matrix, and value weight matrix, respectively, and are optional. ;

[0113] The single-head attention output is obtained as follows:

[0114] ;

[0115] The single-head attention outputs are then concatenated and passed through a linear layer to obtain the multi-head attention output. :

[0116] ;

[0117] in, For multi-head attention output, For the attention output of each single head, , , head.

[0118] This multi-head self-attention module can use attention weights Dynamically adjust feature importance, such as NDVI feature in rust samples. REP features weight Increased; NDWI features in spider samples FlavR features weight Increase.

[0119] Feedforward networks (FFNs) are used to perform nonlinear transformations on the multi-head attention output. Each layer of the feedforward network contains residual connections and layer normalization to obtain the transformed features.

[0120] ;

[0121] in, The transformed features, For the feedforward network input, (Hidden layer dimension) ), , For bias.

[0122] By fusing the various TransformerEncoder layers, we obtain the fused feature:

[0123] ;

[0124] ;

[0125] in, Input to the preprocessed model, As an intermediate feature, For fused features; example: after 6 layers of TransformerEncoder, the output fused features. ( ).

[0126] (3) Output layer, which may include a dual-task header (regression + classification), namely: an adaptation score regression header, used to output adaptation scores; and a pest and disease type classification header, used to output pest and disease classifications, including: healthy, rust, anthracnose, and spider mites; Example:

[0127] c1: Adapted score regression head: for fused features Perform global average pooling to obtain the global channel average feature vector. Mapped to adaptation score :

[0128] ;

[0129] ;

[0130] in, , , For the Sigmoid function (compresses the output to [0,1], then ×100 to map to [0,100]).

[0131] c2: Pest and disease type classification head, based on global channel average feature vector. Output the probability distribution of 4 types. (Health, rust, anthrax, spider mites):

[0132] ;

[0133] in , , Indicates that the sample belongs to type The probability of.

[0134] (4) Loss function

[0135] Preferably, model training can employ dual-task joint training:

[0136] d1: The total loss function is a weighted sum of the regression loss and the classification loss.

[0137] ;

[0138] in, For the total loss, To regress the loss, For classifying losses, These are weighting coefficients, which can be selected. This is used to balance the weights of the two tasks.

[0139] d2: Regression loss (fit score prediction):

[0140] The mean squared error (MSE) loss is used, and the objective is to minimize the predicted score. Compared to the actual score Differences:

[0141] ;

[0142] in, For the sample size, For the first The predicted score for each sample. For the first The true score of each sample.

[0143] d3: Classification loss (prediction of pest and disease type):

[0144] Using cross-entropy (CE) loss, the goal is to optimize the matching of four types of probability distributions with labels:

[0145] ;

[0146] in, For the first The unique hot label of each sample ( If the sample belongs to type (otherwise it is 0). (healthy), (Rust) (anthrax), (Red spider) For the first For each sample, the model predicts it belongs to the category. The probability of.

[0147] (5) Sample set construction:

[0148] Based on the aforementioned fit assessment model, this invention can optionally, but not be limited to, transform the early identification of fig pests and diseases into a "multimodal physiological indicator and pest and disease type fit score calculation problem" by using early physiological indicators reflected by common fig pests and diseases, in order to construct a sample set as a benchmark for model training and evaluation.

[0149] Specifically, taking fig rust, anthracnose, and spider mites as examples, a sample set can be constructed based on their early physiological indicators. However, the construction of this sample set is not limited to these; it can also be based on the different responses of other fig diseases and physiological indicators, and sample sets can be constructed based on other physiological indicators. The following is only a preferred embodiment of constructing a sample set based on fig rust, anthracnose, and spider mites, including:

[0150] e1: Based on the pathogenesis of various diseases and pests affecting figs, a set of physiological indicators reflecting the early characteristics of diseases and pests, along with the weights of each indicator, is constructed. Taking the pathogenesis of rust, anthracnose, and spider mites as examples, chlorophyll, water, cell structure, and secondary metabolite indicators are selected to construct a set of physiological indicators reflecting the early characteristics of diseases and pests. ;in:

[0151] Chlorophyll index, which can be represented by the normalized vegetation index (NDVI), reflects changes in chlorophyll content, because rust and spider mites can both lead to chlorophyll decomposition.

[0152] Moisture index, such as the normalized water content index (NDWI), can be used to represent the water content of cells, because anthrax and spider mites can cause water loss.

[0153] Cell structure indicators, such as the hyperspectral red edge position (REP), can be used to characterize the spectral reflectance peak shift caused by cell wall damage, since rust and anthracnose can destroy cell structure.

[0154] Secondary metabolite indicators can be represented by the flavonoid index (FlavR), which reflects the accumulation of flavonoids in the plant's disease and pest resistance response, because rust, anthracnose, and spider mites all induce flavonoid synthesis.

[0155] e2: Determine the set of abnormality functions that reflect the degree of abnormality of each indicator based on the set of physiological indicators and the set of preset normal thresholds;

[0156] Specifically, one can first determine the normal threshold set of various indicators of the canopy of healthy figs through field experiments and laboratory measurements, thus obtaining a preset normal threshold set. :

[0157] ;

[0158] in:

[0159] (NDVI normal threshold): The mean NDVI value was calculated by collecting multispectral data from 1000 healthy fig trees. Standard deviation ,Pick (Ensure that 93% of healthy samples have an NDVI ≥ 0.675);

[0160] (NDWI normal threshold): The moisture content of healthy leaves measured by laboratory drying method, corresponding to the mean NDWI value. Standard deviation ,Pick ;

[0161] (REP normal threshold): Mean value of red edge position in healthy leaves measured by hyperspectral imaging. Standard deviation ,Pick (Red borders shifted to blue indicate cell structural damage);

[0162] (FlavR normal threshold): The flavonoid content of healthy leaves was measured by spectrophotometer, corresponding to the mean FlavR value. Standard deviation ,Pick (Accumulation of flavonoids leads to an increase in FlavR).

[0163] Then, to quantify the degree to which a single indicator deviates from its normal threshold, a set of anomaly function sets reflecting the degree of abnormality of each indicator is determined. , of which Anomalies of each indicator for:

[0164] ;

[0165] Example: To (NDVI, NDWI, REP): When measured values At that time, anomaly degree A positive value indicates a more severe deviation (e.g., NDVI = 0.575). );right (FlavR): When the measured value At that time, anomaly degree When it is positive (e.g., when FlavR=1.5), ); Anomaly 0 indicates normal, and 1 indicates severe abnormality (e.g., when NDVI=0). ).

[0166] e3: Based on the abnormality function set and weight coefficient of each physiological indicator, the weighted matching score of each physiological stress-disease and pest is determined;

[0167] Specifically: To comprehensively assess the anomalies of multiple indicators, a "physiological stress-pest / disease fit score" is defined. Quantify the degree of fit between the samples and rust, anthracnose, and spider mites:

[0168] ;

[0169] in, As the indicator weight, satisfying The indicator weights were determined using principal component analysis (PCA) combined with the physiological mechanisms of the three pests and diseases:

[0170] Rust: Primarily causes chlorophyll decomposition and cell structure damage; indicator weighting ;

[0171] Anthrax: Primarily leads to water loss and accumulation of secondary metabolites; indicator weighting ;

[0172] Red spider mites: primarily cause chlorophyll decomposition, water loss, and flavonoid accumulation; indicator weight. ;

[0173] Healthy samples: All indicators have an anomaly rate of 0. .

[0174] Example: A sample NDVI = 0.575 ( ), REP=704nm ( Other indicators are normal. ), calculated by rust disease weight This indicates low compatibility with rust; if NDVI=0.475 ( ), REP=699nm ( ),but The adaptability increases; if NDWI=0.36 ( ), FlavR=1.6 ( ), calculated by red spider weight Red spiders have low compatibility.

[0175] e4: Map pest and disease types based on the threshold range of each physiological stress-pest and disease adaptation score to construct a sample set.

[0176] Specifically: Define the set of pest and disease types. ( :healthy, Rust disease :anthrax, (Red Spider), based on adaptation score Threshold interval mapping type:

[0177] ;

[0178] The specific threshold range was determined through statistical analysis of experimental field data: 1500 artificially inoculated samples (500 for each pest / disease), and rust samples. Mean 82.32 (standard deviation 8.23), anthrax sample Mean 62.11 (standard deviation 6.32), red spider sample Mean 71.03 (standard deviation 7.65), healthy sample Mean 15.52 (standard deviation 10.45). The "to be observed" interval was further subdivided based on samples from the early stage of rust disease (3-4 days post-inoculation). Mean 45.65 (standard deviation 3.36), early anthrax samples (3-4 days post-inoculation) Mean 35.32 (standard deviation 3.56), early stage of spider mite infestation (1-2 days after inoculation, insect population density <5 mites / leaf) With a mean of 53.45 (standard deviation 4.32), the retesting period for each early type is clearly defined to avoid delays in early warning.

[0179] Specifically, 1500 sample points can be selected, covering different scenarios (healthy orchards, artificially inoculated experimental fields, and naturally diseased orchards). Each sample includes: multimodal data. Ground-based true values: Leaf samples were collected simultaneously, physiological indicators were measured in the laboratory (chlorophyll content was measured using a spectrophotometer, and moisture content was measured using the drying method), and spider mite population density was manually counted. The true score was determined following the steps described above. With real labels (Including four categories: healthy, rust, anthracnose, and red spider mites); A labeled sample set was obtained: .

[0180] More preferably, the sample set construction step also includes:

[0181] e5: Data Preprocessing and Augmentation

[0182] e51: Normalization: Modal features are normalized according to the mean of the training set. and standard deviation standardization: ;

[0183] e52: Data Augmentation

[0184] Hyperspectral: Add Gaussian noise ( ), spectral drift (±5nm wavelength shift), simulating illumination / atmospheric interference;

[0185] Visible light: Random horizontal flip (probability 0.5), rotation (±15°), and addition of tiny speckle noise (simulating early traces of spider mites) to enhance the robustness of texture features;

[0186] Feature perturbation: A small-amplitude random perturbation (±5%) is added to the 35-dimensional feature vector to simulate sensor measurement error.

[0187] (6) Training process

[0188] f1: Optimizer and Learning Rate Scheduling

[0189] Optimizer: AdamW, parameters Weight decay ;

[0190] Learning rate: Linear warm-up cosine decay:

[0191] ;

[0192] in, For the first The learning rate of the step. The current step number. The initial learning rate, To minimize the learning rate, For the number of preheating steps, This represents the total number of steps. Example: Initial learning rate. Minimum learning rate Preheating steps (Learning rate increases linearly from 0 to) Total steps .

[0193] f2: Training Process

[0194] The sample set is divided into three parts: training set 70% (1050 samples), validation set 20% (300 samples), and test set 10% (150 samples).

[0195] Model initialization: Transformer parameters are randomly initialized (following the following rules) ), Hyperspectral autoencoder pre-training (unsupervised, using only hyperspectral data);

[0196] Iterative training: batchsize=32, calculation in each iteration. Backpropagation updates parameters;

[0197] Early stopping strategy: If the validation set If the model does not decrease after 10 consecutive rounds, stop training and save the optimal model (the parameters with the minimum loss on the validation set).

[0198] Model alignment verification: Visualize attention weights to ensure that the model focuses on key metrics (e.g., the attention weight of the NDVI and REP feature pairs in the rust samples is >60%, and the attention weight of the NDWI and FlavR feature pairs in the spider samples is >55%).

[0199] In summary, this invention presents a method for assessing the health status of fig trees based on multimodal data fusion. First, a multimodal feature vector is constructed by fusing features from multispectral images, hyperspectral cubes, and visible light images. Then, a multimodal fusion-based adaptability assessment model is proposed, which preprocesses the multimodal feature vectors and performs deep fusion and feature association to obtain fused features. Finally, based on the fused features, an adaptability score and pest / disease classification are output. The overall solution achieves at least the following technical effects:

[0200] 1. Significantly enhanced feature representation integrity: Feature fusion of multispectral images, hyperspectral cubes and visible light images to construct multimodal feature vectors significantly enhances feature integrity and solves the problem of one-sided features of single modality. For example, visible light alone is easily affected by illumination, and spectral data alone is difficult to locate lesions.

[0201] 2. Improved assessment accuracy and pest classification precision: The adaptability assessment model, through deep fusion and feature association, mines complementary information of different modal features, such as spectral anomaly + texture anomaly = specific disease, avoiding misjudgment based on a single feature, such as leaf discoloration caused by strong light being misjudged as a disease;

[0202] 3. Enhanced robustness and anti-interference ability: The complementarity of multimodal data offsets the noise and environmental interference of single modality, such as uneven lighting, shadows, soil background, and differences in shooting angle; under complex lighting (such as alternating cloudy / sunny days), different growth stages, and different planting environments, it can significantly improve the stability of evaluation results and reduce the misjudgment rate.

[0203] (7) Application process (using a trained model for early warning): Input the UAV multimodal data of the new orchard, and output the early warning results of rust, anthracnose, red spider mites and health status through the following steps:

[0204] g1: Data Acquisition and Preprocessing

[0205] The drone collects orchard data along a preset flight path at a resolution of 1m. 2 / pixel, generating multimodal data Extract 35-dimensional feature vectors Standardization (using the training set) and ).

[0206] g2: Model Inference:

[0207] Will Input a pre-trained Transformer model, output the fit score and the probabilities of the four class types:

[0208] ;

[0209] Type determination and confidence calculation:

[0210] Based on the adaptation score Map the probability distribution of the four types according to the rules. The confidence level is the maximum value of the type probability. .

[0211] g3: Spatial positioning and early warning triggering:

[0212] Orchards are 1m 2 Grid division For each grid :

[0213] like (Warning threshold) )and Trigger warning, output type (Rust / Anthracnose / Red Spider Mite), GPS Coordinates ;

[0214] Heat map of fruit orchard: grid color depth indicates Size (the darker the color, the higher the score), different colors indicate different types of diseases and pests (e.g., red indicates rust, orange indicates anthracnose, purple indicates spider mites).

[0215] In the above preferred embodiment, the present invention provides a fig health status assessment method based on multimodal data fusion, using an end-to-end Transformer model and inputting a 35-dimensional multimodal feature vector. Output adaptation score Probability distribution of pest and disease types (Including healthy plants, rust, anthrax, and spider mites).

[0216] (8) Experimental verification

[0217] h1: Experimental objective

[0218] This study verifies the early warning effect, identification accuracy, environmental generalization ability, and engineering deployment performance of the method of the present invention in the monitoring of three major diseases and pests of fig: rust, anthracnose, and spider mites. It clarifies its core advantages compared with existing technologies and provides data support for the refined management of actual orchards.

[0219] h2: Experimental data

[0220] h21: Sample source: Four experimental areas were selected (A: healthy control orchard, B: artificially inoculated experimental field, C: naturally diseased orchard, D: new orchard across regions), covering the mainstream fig varieties, key growth stages (young fruit stage, fruit expansion stage, maturity stage), and complex environmental conditions (sunlight: sunny, cloudy, partly cloudy; soil type: sandy loam, clay, loam; humidity: 40%-85%), to ensure the diversity and representativeness of the samples.

[0221] h22: Sample size: A total of 1500 sample points were collected, and the sample distribution is as follows:

[0222] 400 healthy samples (all from area A, covering different varieties, growth stages and environmental conditions);

[0223] 700 artificially inoculated samples: 350 rust samples (1-14 days post-inoculation, including incubation period, initial stage, and development stage), and 350 anthrax samples (1-14 days post-inoculation, covering different degrees of infection).

[0224] 300 naturally occurring disease samples (from region C, including 80 rust, 90 anthracnose, and 130 spider mite, simulating real-world disease outbreaks in the field);

[0225] 100 cross-region validation samples (from region D, not used in model training, for generalization testing).

[0226] h3: Data Acquisition

[0227] Drone Platform: The drone is equipped with a multispectral camera (including 6 bands: 450nm, 550nm, 650nm, 750nm, 850nm, and 1640nm), a hyperspectral camera (350nm-1000nm band, 1nm interval), and a visible light camera (20 megapixels, RGB channels).

[0228] Data acquisition parameters: flight altitude 80m, spatial resolution 1m 2 / pixel, flight path overlap rate 80%, GPS coordinates of each sample point are recorded synchronously;

[0229] Ground-based true values ​​were obtained by simultaneously collecting leaf samples from corresponding sample points, measuring core physiological indicators in the laboratory (chlorophyll content was determined using a SPAD-502Plus, moisture content was measured using the drying and weighing method, red edge position was calibrated using a hyperspectral imager, and flavonoid content was measured using a spectrophotometer), manually counting spider mite population density using a magnifying glass counting method, and determining the true score according to the established model. With real labels .

[0230] h4: Experimental Design and Testing Indicators

[0231] h41: Validation of Early Warning Lead Time

[0232] Methods: For artificially inoculated rust and anthracnose samples (350 each), continuous monitoring was conducted until obvious lesions were identified by traditional visual inspection. The time when the first warning of this invention was triggered was recorded. The warning lead time is calculated by comparing the first identification time with the traditional visual inspection method (and the confidence level > 0.85); for 130 naturally occurring red spider mites, the difference between the first warning time of this invention and the time when traces of pests are manually discovered is recorded.

[0233] Test indicator: Early warning lead time (Days), respectively calculate the average advance amount and overall range of the three types of pests and diseases.

[0234] h42: Recognition accuracy verification

[0235] Methods: 1400 samples (including healthy samples, rust samples, anthracnose samples, and spider mite samples) were divided into a training set (980 samples), a validation set (280 samples), and a test set (140 samples) in a 7:2:1 ratio. After training the model according to the training procedure, the performance was evaluated on the test set. The validation metrics were obtained as follows:

[0236] Adaptive score prediction accuracy: root mean square error Mean absolute error ;

[0237] Type classification performance: Overall accuracy Recall rates of various diseases and pests Accuracy ;

[0238] False alarms and false negatives: False alarm rate for healthy samples underreporting rate of pests and diseases .

[0239] h43: Generalization Validation

[0240] Methods: 100 cross-regional samples from region D (not used in model training) were selected. The model performance was tested in three scenarios: no fine-tuning, fine-tuning with 10 samples from region D, and fine-tuning with 20 samples from region D. The changes in accuracy were compared. At the same time, extreme environmental interference (strong light reflection, atmospheric scattering, and soil background mixing) was simulated to test the recognition accuracy of the model under interference conditions.

[0241] Test metric: Accuracy improvement under different fine-tuning scenarios Accuracy retention rate under extreme conditions .

[0242] h44: Spatial Positioning Accuracy Verification

[0243] Method: Twenty pest and disease sample points with known precise locations (GPS coordinate error ≤ 0.5m) were selected in the experimental field. The positioning results were output by the system of this invention, and the positioning deviation was calculated.

[0244] Inspection indicator: Positioning error (m), statistical average positioning error and maximum positioning error.

[0245] h45: Engineering Performance Verification

[0246] Methods: Deploy a compressed model quantized by TensorRT on an edge server (NVIDIA Jetson Xavier NX) and test the time consumption of the entire single-sample inference process; select a 500-acre orchard to simulate the actual operation scenario and count the total time consumption of UAV data collection, data transmission, model inference, and early warning generation; run continuously for 7 days to test system stability.

[0247] Test indicators: single-sample inference delay (total time from data input to early warning output), daily coverage efficiency of 500 mu orchard (total time consumed), and continuous operation stability of the system (percentage of time spent running without failure within 7 days).

[0248] h46: Comparative Experimental Design

[0249] We selected existing mainstream technologies as a control and compared their performance on the same test set (140 samples):

[0250] Comparison Method 1: Traditional manual visual inspection method (joint judgment by 3 senior agricultural technicians);

[0251] Comparison Method 2: YOLOv5 recognition method based on visible light images;

[0252] Comparison Method 3: Single-modal hyperspectral features + SVM classification method;

[0253] Comparison with Method 4: Simple concatenation of multimodal features + CNN method.

[0254] The comparison indicators include early warning lead time, accuracy, false alarm rate, inference delay, and positioning accuracy.

[0255] H5: Experimental Results

[0256] h51: Early Warning Lead Time

[0257] Experimental results show that this invention achieves significant early warning for three major types of pests and diseases, with an overall advance warning time of 7-14 days, meeting the core requirement of "early detection and early prevention" of pests and diseases. The early warning effect for rust is the most effective. (See Table 1 for details.)

[0258] Table 1. Experimental Comparison Results Data

[0259]

[0260] h52: Recognition accuracy

[0261] The evaluation results for the 140 samples in the test set are as follows:

[0262] Adaptive score prediction accuracy: point, The predicted values ​​closely match the actual values; the classification performance is shown in Table 2.

[0263] Table 2: Performance Data by Type Classification

[0264]

[0265] False alarms and false negatives: False alarm rate for healthy samples underreporting rate of pests and diseases All levels are extremely low, effectively avoiding over-application of pesticides and omissions in prevention and control.

[0266] h53: Generalization

[0267] Cross-region generalization: The accuracy was 75% without fine-tuning, and improved to 88% after fine-tuning with 10 samples. After fine-tuning with 20 samples, the accuracy reached 91%. High-precision adaptation can be achieved with only a small number of local samples.

[0268] Adaptability to extreme environments: Under conditions of strong light reflection, atmospheric scattering, and mixed soil background, the accuracy retention rates are 89%, 90%, and 92%, respectively, and the anti-interference ability is significantly better than that of existing technologies.

[0269] h54: Spatial positioning accuracy

[0270] The average positioning error of the 20 test sample points was 0.8m, and the maximum positioning error was 1.2m, which fully meets the 1m requirement.2 The grid-level precision positioning requirement can accurately mark diseased plants or small areas, providing precise guidance for targeted pesticide application.

[0271] h55: Engineering Performance

[0272] Inference efficiency: The average inference latency per sample is 82ms (including data preprocessing, feature extraction, model inference, and result output), which is far lower than the target value of 100ms;

[0273] Coverage efficiency: The total time for data collection, inference, and early warning in a 500-mu orchard is 7.5 hours. A single drone can cover more than 600 mu per day, meeting the daily monitoring needs of large-scale orchards.

[0274] Stability: The system operates without failure for 99.2% of the time, ensuring stable and reliable operation and adaptability to long-term field operations.

[0275] h56: Comparative experimental results are shown in Table 3:

[0276] Table 3: Performance Comparison Data with Existing Technologies

[0277]

[0278] The comparative results show that the present invention surpasses the existing technology in all aspects, including early warning lead time, identification accuracy, false alarm control, inference speed and positioning accuracy, with significant core performance advantages.

[0279] The present invention also provides a computer storage medium storing executable program code; the executable program code is used to execute any of the above-mentioned fig health status assessment methods.

[0280] On the other hand, the present invention also provides a computer system, including a memory and a processor; the memory stores program code executable by the processor; the program code is used to execute any of the above-described fig health status assessment methods.

[0281] For example, the program code can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the program code in the computer system.

[0282] The computer system may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer system may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the computer system may also include input / output devices, network access devices, buses, etc.

[0283] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0284] The memory can be an internal storage unit of a computer system, such as a hard disk or RAM. It can also be an external storage device of the computer system, such as a plug-in hard disk, SmartMediaCard (SMC), Secure Digital (SD) card, or FlashCard. Furthermore, the memory can include both internal and external storage units. The memory is used to store the program code and other programs and data required by the computer system. The memory can also be used to temporarily store data that has been output or will be output.

[0285] The aforementioned computer storage medium and computer system are created based on the aforementioned fig health status assessment method. Their technical functions and beneficial effects will not be elaborated here. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0286] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A method for assessing the health status of fig trees based on multimodal data fusion, characterized in that, include: Collect multispectral, hyperspectral cube, and visible light images of the orchard; Based on multispectral images, physiological index features are extracted to construct multispectral feature vectors; Based on the hyperspectral cube, spectroscopic morphological features and dimensionality reduction features are extracted to construct a hyperspectral feature vector; based on the visible light image, leaf texture features are extracted to construct a visible light feature vector; and the multispectral feature vector, hyperspectral feature vector, and visible light feature vector are concatenated to obtain a multimodal feature vector. The multimodal feature vectors are input into a pre-trained fitness assessment model, which consists of an input layer, a feature fusion layer, and an output layer connected in sequence. The input layer is called to preprocess the multimodal feature vectors before using them as model input. The input layer processing steps include: treating the multimodal feature vectors as a sequence of preset length, with each feature dimension as a token in the sequence; constructing different positional codes for the tokens of multispectral, hyperspectral, and visible light feature vectors; constructing a mapping relationship between each token and physiological indicators to obtain the model input; calling the feature fusion layer to perform deep fusion and feature association on the model input to obtain fused features; and calling the output layer to output the fitness score and pest / disease classification based on the fused features. The output layer includes two task heads: a fitness score regression head for outputting the fitness score and a pest / disease type classification head for outputting the pest / disease classification. The steps for constructing the sample set during the training of the adaptability assessment model include: constructing a set of physiological indicators reflecting the early characteristics of pests and diseases based on the pathogenesis of various fig diseases and pests, and assigning weights to each physiological indicator; determining anomaly function set reflecting the degree of abnormality of each indicator based on the set of physiological indicators and a preset normal threshold set; determining the adaptation score of each physiological stress-pest and disease based on the anomaly function set and weight coefficients of each physiological indicator; and mapping the pest and disease type based on the threshold range of each physiological stress-pest and disease adaptation score to construct the sample set. The set of anomaly functions is represented as follows: ; in, For the first The abnormality of each indicator For the first The measured values ​​of each indicator For the first Normal thresholds for each indicator; The fit score is expressed as: ; in, To adapt the score, The weights are the indicator weights.

2. The method for assessing the health status of figs according to claim 1, characterized in that, The steps for extracting physiological index features and constructing multispectral feature vectors based on multispectral images include: Based on the reflectance of each band in the multispectral image, several physiological index features are calculated to construct a multidimensional multispectral feature vector; the physiological index features include any one or more of the normalized vegetation index, normalized water index, ratio vegetation index, greenness normalized vegetation index, and soil-adjusted vegetation index.

3. The method for assessing the health status of figs according to claim 1, characterized in that, The steps for constructing a hyperspectral feature vector by extracting spectroscopic morphological features and dimensionality reduction features from a hyperspectral cube include: Based on the hyperspectral cube, several spectroscopic morphological features are calculated to construct a first hyperspectral feature sub-vector; the spectroscopic morphological features include: any one or more of the following: red edge position, absorption valley depth, reflection peak width, flavonoid index, blue edge position, yellow edge position, red edge slope, blue edge slope, reflection peak height, and absorption valley area; The hyperspectral cube is reduced in dimensionality by using an autoencoder to learn the low-dimensional representation and outputting multidimensional latent vectors as dimensionality-reduced features to construct the second hyperspectral feature sub-vector. The first hyperspectral feature vector and the second hyperspectral feature vector are concatenated to obtain the hyperspectral feature vector.

4. The method for assessing the health status of figs according to claim 1, characterized in that, The steps for extracting leaf texture features and constructing visible light feature vectors based on visible light images include: Based on the visible light image, multidimensional LBP features are extracted to obtain LBP sub-vectors; Based on the visible light image, multidimensional GLCM features are extracted to obtain GLCM sub-vectors; By concatenating the LBP subvector and the GLCM subvector, the visible light feature vector is obtained.

5. The method for assessing the health status of figs according to claim 1, characterized in that, The feature fusion layer includes several TransformerEncoder layers; Each TransformerEncoder layer includes a multi-head self-attention module and a feedforward network; the multi-head self-attention module includes multiple attention heads and fully connected layers; each attention head takes the model input, computes the query, key, and value, and obtains the single-head attention output; The single-head attention outputs are then concatenated and passed through a linear layer to obtain multi-head attention outputs, in order to learn the correlation weights between features; A feedforward network is used to perform a nonlinear transformation on the multi-head attention output to obtain the transformed features. Based on the transformed features, the TransformerEncoder layers are fused to obtain the fused features.

6. The method for assessing the health status of figs according to claim 1, characterized in that, The adaptation score regression head performs global average pooling on the fused features to obtain a global channel average feature vector, which is then mapped to the adaptation score. The pest and disease type classification head outputs the probability distribution of each type based on the global channel average feature vector.

7. The method for assessing the health status of figs according to claim 1, characterized in that, Model training employs a dual-task joint training approach; The total loss function is a weighted sum of the regression loss and the classification loss; The regression loss uses mean squared error loss, and the goal is to minimize the difference between the predicted score and the true score. The classification loss uses cross-entropy loss, and the goal is to optimize the matching between the probability distribution of each type and the label.

8. The method for assessing the health status of figs according to any one of claims 1 to 7, characterized in that, Pests and diseases are classified as follows: healthy plants, rust, anthracnose, and spider mites; A set of physiological indicators reflecting rust, anthracnose, and spider mites, including chlorophyll, water, cell structure, and secondary metabolite indicators; Physiological indicators reflecting rust are concentrated, with chlorophyll and cell structure indicators having higher weighting coefficients than water and secondary metabolite indicators. Physiological indicators reflecting anthrax are concentrated, with water and secondary metabolite indicators having higher weighting coefficients than chlorophyll and cell structure indicators. The physiological indicators reflecting spider mites are concentrated, with the weighting coefficients of chlorophyll, water, and cell structure indicators being higher than that of secondary metabolite indicators.

9. The method for assessing the health status of figs according to claim 8, characterized in that, Mapping pest and disease types, specifically: ; in, :healthy, Rust disease :anthrax, Red spider.

10. A computer system, characterized in that, It includes a memory and a processor; the memory stores program code that can be executed by the processor; the program code is used to perform the fig health status assessment method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Citrus huanglongbing detection device and detection method based on visual image

    CN119417816A

  • Multi-modal fusion identification method and system for rice leaf type diseases and insect pests

    CN120766129A