Fritillaria variety identification method and device based on hyperspectral image

By collecting and processing multimodal data of Fritillaria samples, constructing a spatiotemporal joint graph structure and applying a dynamic graph convolutional network, we solved the problems of low efficiency in Fritillaria variety identification, large impact of environmental changes and difficulty in anomaly detection, and achieved high-precision and stable classification.

CN120656056APending Publication Date: 2025-09-16HARBIN UNIV OF COMMERCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510707884.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies are inefficient in identifying Fritillaria varieties, easily interfered by human factors, cannot effectively process high-dimensional data, and the classification accuracy decreases when the environment changes. There is a lack of dynamic capture of temporal and spatial information, and anomaly detection is difficult to identify.

Method used

By continuously collecting multimodal data of Fritillaria samples during the flowering period, radiation correction and SIFT algorithm alignment are performed, combined with wavelet packet decomposition and time series weighted fusion, a spatiotemporal joint graph structure is constructed, and dynamic graph convolutional networks are used for variety classification and anomaly detection.

Benefits of technology

It improves the accuracy and stability of Fritillaria variety classification, enhances the adaptability to environmental changes, effectively identifies abnormal samples, and provides a scientific and efficient means of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656056A_ABST
    Figure CN120656056A_ABST
Patent Text Reader

Abstract

The invention provides a fritillary bulb variety identification method and device based on a hyperspectral image, and relates to the technical field of fritillary bulb variety identification. Hyperspectral data, thermal imaging data and RGB images in a fritillary bulb growth period are continuously collected, and the fritillary bulb variety identification method and device based on the hyperspectral image are obtained through radiation correction and SIFT algorithm alignment. Performing frequency domain analysis on the modal data by using wavelet packet decomposition, calculating an energy ratio and generating an adaptive weight coefficient so as to construct a time sequence-space joint graph structure, inputting the joint graph structure of a plurality of varieties into a dynamic graph convolutional network for training, and realizing efficient variety classification by optimizing an adjacent matrix. And finally, based on the output classification result, calculating the Euclidean distance between the to-be-detected sample and the clustering center of the normal sample feature library so as to carry out anomaly detection and accurately identify an abnormal sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Fritillaria variety identification, and in particular to a Fritillaria variety identification method and device based on hyperspectral images. Background Art

[0002] In modern agricultural and plant science research, Fritillaria thunbergii is an important Chinese medicinal material, and the identification and evaluation of its different varieties have important economic and ecological significance. Since the medicinal value of Fritillaria thunbergii is closely related to its variety, accurately identifying different varieties not only helps to improve the quality and safety of medicinal materials, but also provides basic data for scientific research. This identification method has broad application prospects in the fields of medicinal material production, quality control and resource protection. For example, in the market transaction of Chinese medicinal materials, ensuring the true variety of medicinal materials can effectively reduce the circulation of counterfeit and inferior products, thereby protecting the rights and interests of consumers and public health. In addition, with the development of hyperspectral imaging technology, the use of hyperspectral images to extract the spectral characteristics of Fritillaria thunbergii can achieve non-contact identification without destroying the plant samples, providing a new technical approach for plant identification.

[0003] Existing technologies mainly rely on traditional spectral analysis and manual identification methods. These technologies usually measure the spectral reflectance of plants and combine them with machine learning algorithms to classify varieties. Specifically, researchers will obtain the spectral characteristics of plants through data from different spectral bands, and use classification models such as support vector machines and decision trees to distinguish Fritillaria varieties. This method can achieve the purpose of variety identification to a certain extent. However, in practical applications, there are still many limitations. First, traditional spectral analysis often requires manual complex feature engineering, which is inefficient and easily interfered by human factors. Second, existing classification models are often unable to effectively process high-dimensional data, resulting in the loss of feature information, thereby affecting the accuracy of classification.

[0004] Although existing technologies can identify Fritillaria varieties to a certain extent, their shortcomings are particularly obvious. First, when faced with complex environmental changes such as light, temperature and humidity, the classification accuracy of traditional methods often drops significantly and cannot maintain stability; second, existing technologies often rely on fixed feature extraction methods, lack dynamic capture of temporal and spatial information, and cannot fully explore the potential features in Fritillaria samples; in addition, current anomaly detection methods are usually difficult to effectively identify features that are significantly different from normal samples, resulting in some abnormal samples being ignored.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and device for identifying Fritillaria varieties based on hyperspectral images to solve the problems raised in the above background technology.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A method for identifying Fritillaria varieties based on hyperspectral images, comprising the following steps:

[0009] Step 1: Continuously collect multimodal data of the Fritillaria sample to be tested during the flowering period. The multimodal data includes hyperspectral data, thermal imaging data, and RGB images. The multimodal data is preprocessed, including radiometric correction, alignment of RGB images of adjacent time points based on the SIFT algorithm, and simultaneous application of alignment mapping to the hyperspectral and thermal imaging data to eliminate spatial misalignment.

[0010] Step 2: Perform wavelet packet decomposition on the preprocessed multimodal data, calculate the frequency domain energy ratio to generate adaptive weight coefficients, obtain cross-modal features through time-series weighted fusion, and construct a spatiotemporal joint graph structure based on the fused features. Its nodes are pixel-level time-series feature vectors, and its edges are defined by spatial proximity, band correlation, and feature similarity.

[0011] Step 3: Based on the same method, the joint graph structure of multiple Fritillaria samples of known varieties is obtained to construct a training set. The joint graph structure of the Fritillaria samples is used as input, and the varieties are used as labels to train the dynamic graph convolutional network. The network parameters are updated by optimizing the adjacency matrix. The joint graph structure of the Fritillaria samples to be tested is input into the trained dynamic graph convolutional network, and the predicted value of the variety of the Fritillaria samples to be tested is output;

[0012] Step 4: Calculate the Euclidean distance between the predicted value of the variety of the Fritillaria sample to be tested and the cluster center of the same variety in the training set. If the distance exceeds the set threshold, the sample is judged to be abnormal.

[0013] Furthermore, the method of continuously collecting multimodal data of the Fritillaria sample to be tested during the flowering period is as follows: starting to collect data at equal time intervals when the corolla expansion degree of the Fritillaria is greater than or equal to 90%, and stopping collecting data when the petal shedding rate is greater than or equal to 30%, and collecting data a total of k times at a fixed interval of Δt = 5 days;

[0014] The equipment used to collect multimodal data includes a hyperspectral imager with a wavelength range of 400-700nm and a spectral resolution of 2nm or less, a thermal imager with a long-wave infrared wavelength range of 8-14μm and a thermal sensitivity of 30mK or less, and an RGB camera with a spatial resolution of 1024×7.68 million pixels, a resolution of 61MP or more, and macro mode support. A rubidium atomic clock is disciplined using GPS to generate a global trigger pulse, synchronizing the acquisition timing of the three devices.

[0015] The logic for preprocessing hyperspectral data, thermal imaging data, and RGB images is as follows:

[0016] Radiometric correction is performed on the hyperspectral data. A dark reference image D(x, y) is collected in a completely light-shielded environment. Then, a standard whiteboard image W(x, y) with a reflectivity of 99% is collected. The formula for calculating the reflectivity is:

[0017]

[0018] Among them, R raw is the original reflectivity, (x, y) is the pixel coordinate, R cor is the corrected reflectivity;

[0019] For thermal imaging data, the original thermal radiation grayscale value is converted to temperature value according to the following formula:

[0020]

[0021] Among them, G(x,y) is the original grayscale value output by the thermal imager, G offset is the grayscale offset corresponding to the zero temperature of the thermal imager, G gain is the temperature-grayscale linear conversion gain coefficient, T(x,y) is the temperature value;

[0022] According to the temperature data T collected by the temperature and humidity sensor env , Humidity data RH correction plant temperature: T cor (x,y)=T(x,y)-k T ·(T env -25)+0.03·(RH-60%), where k T =0.05 is the compensation coefficient, 25℃ is the calibration reference temperature, T cor is the corrected plant temperature value;

[0023] The SIFT algorithm is used to extract feature points from adjacent temporal RGB images. Reliable matching pairs are screened using a bidirectional nearest neighbor distance ratio of ε=0.6. The RANSAC algorithm is then used to robustly estimate the affine transformation matrix including scaling, rotation, and translation. Finally, bilinear interpolation is used to achieve spatial alignment of images across temporal phases. These images are then synchronously mapped to hyperspectral and thermal imaging data to eliminate spatial misalignment caused by plant growth displacement and perspective offset, ensuring the consistency of multimodal data in both temporal and spatial dimensions.

[0024] Furthermore, the pre-processed multimodal data is decomposed by wavelet packets, and the frequency domain energy ratio is calculated to generate the adaptive weight coefficient. The logic is as follows:

[0025] The DB4 wavelet basis function is used to perform three-layer wavelet packet decomposition on the three modal data of hyperspectral data HSP, thermal imaging data TIR and RGB image, with the decomposition layer number J=3. Each layer of wavelet packet decomposition divides the signal into low-frequency and high-frequency parts. The total number of frequency bands after three-layer decomposition is 2 3 =8:

[0026] Frequency domain energy ratio calculation, the formula for calculating sub-band energy is:

[0027]

[0028] Among them, E WPD (X,s) represents the energy of modal data X in the sth sub-band, X∈{HSP,TIR,RGB} is the modality type identifier, s∈{1,2,…,8} is the sub-band number, is the nth wavelet coefficient of the modal data X in the sth sub-band, N s is the total number of wavelet coefficients in the sth sub-band, n∈{1,2,…,N s} is the wavelet coefficient index;

[0029] Calculate the total energy E total (X) is based on the formula:

[0030]

[0031] Among them, E total (X) represents the overall energy of modal data X;

[0032] The calculation formulas for the weight coefficients α, β and γ are as follows:

[0033]

[0034] Among them, E sum =E total (HSP)+E total (TIR)+E total (RGB);

[0035] The formula for calculating the single-moment fusion feature is:

[0036]

[0037] Among them, F t is the fusion feature vector at the acquisition time t, and the weight coefficient α t ,β t ,γ t is the modal weight at that moment; Normalize(·) means to perform Z-score normalization on the eigenvector; F HSP,tis the feature vector of the hyperspectral data collected at the acquisition time t after wavelet packet decomposition and feature extraction, F TIR,t It is the feature vector of the thermal infrared data collected at the acquisition time t after temperature conversion and feature extraction, F RGB,t It is the feature vector of the RGB image data collected at the acquisition time t after spatial alignment and feature extraction, t is the index of the acquisition time, and t∈[1,k];

[0038] The formula for calculating time series weighted fusion feature data is:

[0039]

[0040] Among them, F fus is the overall temporal weighted feature, w t is the weight at time t, and Where c = 0.7 is the temporal attenuation factor, which is used to enhance the feature contribution of the central time point of the flowering period, and τ = round(k / 2) is the collection time index of the central time point of the flowering period.

[0041] Furthermore, cross-modal features are obtained through time-series weighted fusion. The logic of constructing a spatiotemporal joint graph structure based on the fused features is as follows:

[0042] The node feature is the pixel point of the flower area in the hyperspectral image corresponding to each node, and the node feature is the time series fusion feature vector of the pixel d is the single-moment feature dimension;

[0043] The edge connection rules, based on three types of relationships, namely spatial proximity, band correlation, and feature similarity, jointly characterize the spatiotemporal evolution characteristics of the Fritillaria flowering period:

[0044] For spatial proximity, if the Manhattan distance between two pixels is less than the threshold 5px: ‖(x i ,y i )-(x j ,y j )‖ man =|x i -x j |+|y i -y j |<5px, then an edge is established. This threshold is consistent with the micro texture unit size of the Fritillaria bulb, which is 5×5px;

[0045] For band correlation, if the reflectance correlation coefficient of two pixels in the HSP band of 520-600nm exceeds the threshold: Then the edge is established, and there is significant redundancy when the correlation coefficient between bands is greater than 0.85;

[0046] For feature similarity, the average cosine similarity of the temporal features of nodes i and j is defined as: Then, an edge is established to capture the overall consistency of plant growth during the flowering period;

[0047] The spatiotemporal joint graph structure consists of three parts: a node set, an edge set, and a weight matrix, which is formally defined as: G(t) = {V(t), E(t), A(t)}, where each node corresponds to a pixel in the hyperspectral image, and the node set V(t) = {v x,y (t)|x=1…W,y=1…H}, where W is the image width, H is the height, and the total number of nodes is W×H;

[0048] The construction of the edge set E(t) is achieved through the joint implementation of three types of rules to comprehensively characterize the spatiotemporal evolution characteristics of Fritillaria during the flowering period:

[0049] For the spatial proximity rule, an edge is established if the Manhattan distance between two pixels is less than 5px. This threshold is consistent with the size of the microtexture unit of the Fritillaria bulb and is used to capture the correlation of local growth characteristics. For the band correlation rule, an edge is established if the reflectance correlation coefficient of two pixels in the 520-600nm band is high, reflecting the stability of pigment distribution during the flowering period. For the temporal continuity rule, an edge is established if the interval between two time points is no more than 7 days, covering the adjacent acquisition interval Δt = 5 days and the lag effect of plant growth to ensure the coherence of temporal evolution.

[0050] Weight Matrix Calculated by the above weighted fusion strategy.

[0051] Furthermore, based on the same method, the joint graph structure of multiple Fritillaria samples of known varieties is obtained to construct a training set. The logic of training the dynamic graph convolutional network with the joint graph structure of Fritillaria samples as input and the variety as the label is as follows:

[0052] The input is the spatiotemporal joint graph structure G(t) = {V(t), E(t), A(t)} of the flower area during the flowering period, and the node input feature dimension d in =64, output dimension d out =32, the graph convolution layer of the dynamic graph convolutional network is:

[0053]

[0054] in, is the node feature matrix of the lth layer, is the normalized adjacency matrix, and D is the degree matrix, whose diagonal elements are the node degrees; is the trainable weight matrix of the lth layer, and the adjacency matrix A(l) Decoupling, GELU is a Gaussian error linear unit activation function, and its smooth nonlinear characteristics adapt to the gradual growth pattern of plants during the flowering period;

[0055] During training, the adjacency matrix is ​​adjusted according to the classification loss gradient:

[0056]

[0057] Among them, η = 0.05 is the learning rate, Loss is the cross entropy loss function, the Sigmoid function constrains the adjacent weight range to [0,1], A (l) is the adjacency matrix of the lth layer;

[0058] The node feature transfer across acquisition moments is achieved through gated recurrent units:

[0059] H (t+1) =GRU[H (t) ,Concat(A (t) H (t) W (t) )]

[0060] Among them, H (t) A is the node feature at the acquisition time t, representing the plant state at that time. (t) H (t) W (t) is the spatiotemporal-spectral feature aggregation result, Input GRU update gate, GRU selectively retains historical information through reset gate and update gate to adapt to the feature persistence of different stages of flowering period;

[0061] The output layer generates the variety probability distribution through global mean pooling and fully connected layers:

[0062] P = Softmax <MLP{[MeanPool(H (L) )]}>

[0063] Among them, L is the last layer of the network, H (L) It is the last layer of node feature matrix, which integrates the information of all stages of flowering period; MeanPool aggregates features along the spatial dimension to eliminate the interference of local noise on the classification results.

[0064] Furthermore, the joint graph structure of the Fritillaria sample to be tested is input into the trained dynamic graph convolutional network, and the logic of outputting the Fritillaria variety classification result is as follows:

[0065] The training set includes each sample as a time series graph sequence {G1, G2, ..., G k}, the label is the Fritillaria variety Y = {1, 2, …, C}, C is the total number of categories;

[0066] Input the sample timing diagram sequence G to be tested test , output probability distribution P test :

[0067]

[0068] Among them, P test,C is the output probability distribution of the sample to be tested, and The predicted Fritillaria species category.

[0069] Furthermore, the logic for calculating the Euclidean distance between the predicted value of the variety of the Fritillaria sample to be tested and the cluster center of the same variety in the training set is:

[0070] Extract the second-to-last layer of global features of all known normal samples in the training set:

[0071]

[0072] in, is the node feature matrix of the u-th normal sample in the second-to-last layer, d ′ =64 is the output dimension of the penultimate layer of the dynamic graph convolutional network, v∈{1,2,…,C} is the category label of the known Fritillaria varieties, u=1,2,…,N v is the training sample index of category v, N v is the number of normal samples of category v, and MeanPool(·) is the mean along the spatial dimension W×H;

[0073] Calculate the characteristic mean of normal samples by species:

[0074]

[0075] Among them, ClusterCenter v is the cluster center of category v;

[0076] Input the time-space joint graph structure G of the sample to be tested test Extract global eigenvectors:

[0077]

[0078] in, is the node feature matrix of the sample to be tested in the second to last layer of the model, F test is the global feature vector of the sample to be tested;

[0079] If the classification result is a category Calculate the Euclidean distance between the sample to be tested and the corresponding cluster center:

[0080] S anomaly =‖Ftest -ClusterCenter v ‖2

[0081] Among them, S anomaly is the Euclidean distance between the sample to be tested and the corresponding cluster center;

[0082] The threshold range δ is set based on the distribution of normal sample distances in the training set, and δ=[μ train -3σ train ,μ train +3σ train ], where μ train is the mean distance between all normal samples in the training set and the corresponding cluster center, σ train is the standard deviation of the normal sample distance of the training set;

[0083] If S anomaly >δ, it is determined to be an abnormal sample, and the abnormal label is:

[0084]

[0085] Among them, IsAnomaly is a label indicating whether the sample to be tested is judged to be abnormal. When IsAnomaly=1, that is, S anomlay ∈δ is an abnormal sample that triggers a warning. When IsAnomaly=0, the classification result is obtained.

[0086] The present invention also provides a device for identifying Fritillaria varieties based on hyperspectral images. The system is used to implement the above-mentioned method for identifying Fritillaria varieties based on hyperspectral images, specifically comprising:

[0087] A data acquisition module is used to continuously collect multimodal data of the Fritillaria sample to be tested during the flowering period, the multimodal data including hyperspectral data, thermal imaging data and RGB images, and preprocess the multimodal data. The preprocessing includes radiometric correction, alignment of RGB images of adjacent time points based on the SIFT algorithm, and simultaneous application of alignment mapping to the hyperspectral and thermal imaging data to eliminate spatial misalignment;

[0088] The feature extraction module is used to perform wavelet packet decomposition on the preprocessed multimodal data, calculate the frequency domain energy ratio to generate adaptive weight coefficients, obtain cross-modal features through time-series weighted fusion, and construct a spatiotemporal joint graph structure based on the fused features. Its nodes are pixel-level time-series feature vectors, and its edges are defined by spatial proximity, band correlation, and feature similarity.

[0089] The model construction module is used to obtain the joint graph structure of multiple Fritillaria samples of known varieties based on the same method to construct a training set. The joint graph structure of the Fritillaria samples is used as input, and the varieties are used as labels to train the dynamic graph convolutional network. The network parameters are updated by optimizing the adjacency matrix. The joint graph structure of the Fritillaria samples to be tested is input into the trained dynamic graph convolutional network, and the predicted value of the variety of the Fritillaria samples to be tested is output;

[0090] The classification result module is used to calculate the Euclidean distance between the predicted value of the variety of the Fritillaria sample to be tested and the cluster center of the same variety in the training set. If the distance exceeds the set threshold, the sample is judged to be abnormal.

[0091] Compared with the prior art, the present invention has the following beneficial effects:

[0092] The present invention ensures the accuracy and consistency of the data by continuously collecting hyperspectral data, thermal imaging data and RGB images during the growth cycle, combining radiation correction and SIFT algorithm alignment processing. Such processing reduces the impact of changes in ambient light or other external factors on data quality, ensuring that the acquired modal data can truly reflect the characteristics of the Fritillaria sample; then, wavelet packet decomposition is used to perform frequency domain analysis on the modal data, the energy share is calculated and an adaptive weight coefficient is generated, which can effectively extract important features that affect the classification of Fritillaria. By constructing a time-space joint graph structure, not only the information of different modal data is integrated, but also the model's ability to represent data is enhanced through the construction of nodes of time series vectors and edges based on proximity relationships, band correlation and feature similarity.

[0093] The present invention uses a joint graph structure of multiple varieties of Fritillaria samples as training input to a dynamic graph convolutional network. The model is updated using an optimized adjacency matrix, achieving efficient classification of Fritillaria varieties. This process enables the network to adaptively learn the characteristics of different Fritillaria varieties, improving classification accuracy and the model's generalization ability. Anomaly detection is scored by calculating the Euclidean distance between the sample to be tested and the cluster center of normal samples. This method not only improves the recognition rate of abnormal samples but also effectively distinguishes normal samples from abnormal samples by setting a reasonable threshold.

[0094] The present invention solves the shortcomings of existing technologies in the classification of Fritillaria varieties through multimodal fusion of hyperspectral data, construction of dynamic graph structure and application of dynamic graph convolutional network. This method not only improves the classification accuracy and stability, but also provides a more scientific and efficient technical means for the monitoring and resource protection of Fritillaria varieties. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] Figure 1 Schematic diagram of the overall method flow of the present invention;

[0096] Figure 2 This is a flow chart of the overall system module of the present invention. DETAILED DESCRIPTION

[0097] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to specific embodiments.

[0098] It should be noted that, unless otherwise defined, the technical or scientific terms used in the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.

[0099] Example:

[0100] See also Figure 1 , the present invention provides a technical solution:

[0101] A method for identifying Fritillaria varieties based on hyperspectral images, comprising the following steps:

[0102] Step 1: Continuously collect multimodal data of the Fritillaria sample to be tested during the flowering period. The multimodal data includes hyperspectral data, thermal imaging data, and RGB images. The multimodal data is preprocessed, including radiometric correction, alignment of RGB images of adjacent time points based on the SIFT algorithm, and simultaneous application of alignment mapping to the hyperspectral and thermal imaging data to eliminate spatial misalignment.

[0103] The multimodal data of the Fritillaria sample to be tested were continuously collected during the flowering period as follows: starting from when the corolla expansion degree of the Fritillaria was greater than or equal to 90%, and stopping when the petal shedding rate was greater than or equal to 30%, and collecting k times in total at a fixed interval of Δt = 5 days;

[0104] The equipment used to collect multimodal data includes a hyperspectral imager with a wavelength range of 400-700nm and a spectral resolution of 2nm or less, a thermal imager with a long-wave infrared wavelength range of 8-14μm and a thermal sensitivity of 30mK or less, and an RGB camera with a spatial resolution of 1024×7.68 million pixels, a resolution of 61MP or more, and macro mode support. A rubidium atomic clock is disciplined using GPS to generate a global trigger pulse, synchronizing the acquisition timing of the three devices.

[0105] The logic for preprocessing hyperspectral data, thermal imaging data, and RGB images is as follows:

[0106] Radiometric correction is performed on the hyperspectral data. A dark reference image D(x, y) is collected in a completely light-shielded environment. Then, a standard whiteboard image W(x, y) with a reflectivity of 99% is collected. The formula for calculating the reflectivity is:

[0107]

[0108] Among them, R raw is the original reflectivity, (x, y) is the pixel coordinate, R cor is the corrected reflectivity;

[0109] R cor (x,y) corrected reflectance, which represents the ratio of the light intensity reflected from the plant surface to the incident light intensity at a specific wavelength, R raw (x, y) is the raw reflectance, uncorrected data that is significantly affected by ambient light and sensor noise. D(x, y) is a dark reference image captured in a completely darkened environment, representing the sensor's dark current noise. W(x, y) is a standard whiteboard image with a reflectivity of 99%, which serves as a reference for calibration and represents the ideal reflection intensity.

[0110] By subtracting the dark reference image D(x,y), the inherent noise of the sensor, such as dark current, is removed and multiplied by 99% to match the actual reflectivity of the white board to ensure data comparability; if other parameters are fixed R raw The larger the R cor The smaller the value, the larger the value of D(x,y). cor The smaller the value, the weaker the noise suppression effect; the larger the value of W(x,y), the stronger the R cor The smaller;

[0111] R cor A higher reflectivity indicates a stronger ability of the Fritillaria surface to reflect light in that wavelength band. This may be related to the dense structure of the leaf surface and a high content of pigments, such as chlorophyll or anthocyanins, and indirectly reflects the health status or growth stage. The core of this formula is to eliminate the influence of environmental factors on spectral data, so that subsequent analysis can more realistically reflect the growth status of the plant.

[0112] For thermal imaging data, the original thermal radiation grayscale value is converted to temperature value according to the following formula:

[0113]

[0114] Among them, G(x,y) is the original grayscale value output by the thermal imager, G offset is the grayscale offset corresponding to the zero temperature of the thermal imager, G gain is the temperature-grayscale linear conversion gain coefficient, T(x,y) is the temperature value;

[0115] T(x,y) is the temperature value of the thermal imaging image, reflecting the actual temperature of the Fritillaria plant surface. G(x,y) is the original grayscale value output by the thermal imager, which is proportional to the intensity of the received thermal radiation. offset G is the grayscale offset corresponding to the zero temperature of the thermal imager, which is used to eliminate the baseline deviation of the sensor. gain is the linear conversion gain coefficient between temperature and gray value, which reflects the proportional relationship between gray value and temperature; by subtracting G offset , eliminate the inherent offset of the sensor at zero temperature, such as circuit drift; by dividing by G gain , converting the grayscale value into the actual temperature value; the larger T(x,y) is, the higher the surface temperature of the Fritillaria is, which may be related to active photosynthesis, enhanced transpiration or local diseases, such as fever caused by infection; the larger G(x,y) is, the larger T(x,y) is, which is a positive proportional relationship; G offset The larger the value, the smaller T(x,y); G gain The larger it is, the smaller T(x,y) is;

[0116] The higher the temperature, the more active the plant's physiological activities may be. It may also be related to environmental conditions and health status. This formula can be used to convert grayscale values ​​into actual temperatures, making it easier to analyze the growth status and health level of the plant.

[0117] According to the temperature data T collected by the temperature and humidity sensor env , Humidity data RH correction plant temperature: T cor (x,y)=T(x,y)-k T ·(T env -25)+0.03·(RH-60%), where k T =0.05 is the compensation coefficient, which adjusts the weight of the influence of ambient temperature on plant temperature. 25℃ is the calibration reference temperature, which is usually the standard laboratory ambient temperature. cor The corrected plant temperature value eliminates the interference of ambient temperature and humidity; 60% is the calibration reference humidity, which is usually the appropriate humidity for plant growth;

[0118] When T envWhen the ambient temperature is >25℃, the plant surface temperature may be overestimated due to the increase in ambient temperature. T ·(T env -25) for correction; when RH>60%, high humidity may enhance transpiration, resulting in a decrease in surface temperature, and 0.03·(RH-60%) needs to be added for correction; T cor The larger the value, the higher the actual physiological temperature of the plant, which may reflect active metabolism or local thermal abnormalities, such as diseased areas. When the ambient temperature rises, it needs to be corrected downward, that is, T env The larger the T cor The smaller the RH, the lower the T. cor The bigger;

[0119] Hyperspectral reflectance correction eliminates sensor noise and environmental interference, ensuring that spectral data truly reflects the optical properties of the Fritillaria surface, providing reliable input for subsequent frequency-domain energy analysis. Thermal imaging temperature correction compensates for ambient temperature and humidity to more accurately capture changes in the plant's physiological temperature, avoiding misjudgments such as misinterpreting high ambient temperature as metabolically active plants. All these formulas work together to ensure the consistency of multimodal data (HSP, TIR, and RGB) in the spatiotemporal and physical dimensions, laying the foundation for the construction of spatiotemporal joint graphs and the training of dynamic graph convolutional networks.

[0120] The SIFT algorithm is used to extract feature points from adjacent temporal RGB images. Reliable matching pairs are screened using a bidirectional nearest neighbor distance ratio of ε=0.6. The RANSAC algorithm is then used to robustly estimate the affine transformation matrix including scaling, rotation, and translation. Finally, bilinear interpolation is used to achieve spatial alignment of images across temporal phases. These images are then synchronously mapped to hyperspectral and thermal imaging data to eliminate spatial misalignment caused by plant growth displacement and perspective offset, ensuring the consistency of multimodal data in both temporal and spatial dimensions.

[0121] Step 2: Perform wavelet packet decomposition on the preprocessed multimodal data, calculate the frequency domain energy ratio to generate adaptive weight coefficients, obtain cross-modal features through time-series weighted fusion, and construct a spatiotemporal joint graph structure based on the fused features. Its nodes are pixel-level time-series feature vectors, and its edges are defined by spatial proximity, band correlation, and feature similarity.

[0122] The logic of performing wavelet packet decomposition on the pre-processed multimodal data and calculating the frequency domain energy ratio to generate the adaptive weight coefficient is as follows:

[0123] DB4 wavelet is used to perform three-layer wavelet packet decomposition on the hyperspectral data HSP, thermal imaging data TIR and RGB trimodal data, with the decomposition layer number J = 3, balancing the low-frequency information and high-frequency detail features. Each layer of wavelet packet decomposition divides the signal into low-frequency and high-frequency parts. The total number of frequency bands after three-layer decomposition is 2 3 =8:

[0124] The three-layer decomposition can achieve a balance between retaining the main frequency band information (low frequency) and the detailed features (high frequency). The DB4 wavelet is selected because of its tight support and orthogonality, which is suitable for non-stationary signal analysis.

[0125] Frequency domain energy ratio calculation, the formula for calculating sub-band energy is:

[0126]

[0127] Among them, E WPD (X,s) represents the energy of modal data X in the sth sub-band, X∈{HSP,TIR,RGB} is the modality type identifier, s∈{1,2,…,8} is the sub-band number, is the nth wavelet coefficient of the modal data X in the sth sub-band, N s is the total number of wavelet coefficients in the sth sub-band, n∈{1,2,…,N s} is the wavelet coefficient index;

[0128] E WPD (X,s) is the energy of the modal data X in the sth sub-band, reflecting the strength of the signal in the band, N s is the total number of wavelet coefficients in the sth sub-band, which is determined by the number of wavelet packet decomposition layers. The square sum operation converts the amplitude of the wavelet coefficients into energy, which reflects the intensity of the signal in a specific frequency band. The higher the energy, the more dominant the frequency band is in the signal, which may be related to the specific physiological characteristics of Fritillaria, such as pigment absorption and thermal radiation.

[0129] E WPD The larger the (X,s), the stronger the signal of mode X in frequency band s. For example, high energy in the 520-600 nm band of the hyperspectral spectrum may reflect strong absorption / reflection of anthocyanins or chlorophyll. High energy in a specific frequency band of thermal infrared may indicate abnormal local temperature of the plant. The larger the E WPD The larger (X,s) is;

[0130] Calculate the total energy E total (X) is based on the formula:

[0131]

[0132] Among them, E total (X) represents the overall energy of modal data X, which is the total energy of all 8 sub-bands and represents the overall signal strength of the modality. The higher the total energy, the greater the contribution of the modality, such as HSP, TIR, or RGB, to the multimodal data. It is used to generate adaptive weight coefficients and dynamically adjust the proportion of different modalities in feature fusion.

[0133] Etotal The larger the (X), the stronger the overall signal of the modal data X. For example, the higher the total energy of HSP, the more critical the spectral information is for variety identification; the higher the total energy of TIR, the more sensitive the temperature change is to anomaly detection; the energy of each sub-band E WPD The higher (X,s), the higher the E total (X) is larger;

[0134] The calculation formulas for the weight coefficients α, β and γ are as follows:

[0135]

[0136] Among them, E sum =E total (HSP)+E total (TIR)+E total (RGB);

[0137] α, β and γ are weight coefficients of hyperspectral, thermal infrared and RGB modalities, respectively, ranging from [0,1]. The weight is proportional to the total energy of the modality. The modality with higher energy accounts for a larger proportion in the fusion. The larger α is, the more significant the contribution of hyperspectral data to the fusion feature, which may be due to its rich frequency domain information. total The larger the (HSP), the larger the α; E total The larger (TIR), the larger β; E total The larger (RGB), the larger γ;

[0138] The formula for calculating the single-moment fusion feature is:

[0139]

[0140] Among them, F t is the fusion feature vector at the acquisition time t, and the weight coefficient α t ,β t ,γ t is the modal weight at that moment; Normalize(·) means to perform Z-score normalization on the eigenvector; F HSP,t is the feature vector of the hyperspectral data collected at the acquisition time t after wavelet packet decomposition and feature extraction, F TIR,t It is the feature vector of the thermal infrared data collected at the acquisition time t after temperature conversion and feature extraction, F RGB,t It is the feature vector of the RGB image data collected at the acquisition time t after spatial alignment and feature extraction, t is the index of the acquisition time, and t∈[1,k];

[0141] F t is the fusion feature vector at time t, integrating multimodal information, α t ,β t ,γt is the weight coefficient of each mode at time t; F t The dimension is that the fused feature vector contains multimodal complementary information, for example: HSP feature is the frequency domain energy distribution, TIR feature is the temperature change pattern, and RGB feature is the spatial texture structure; α t The larger the value, the greater the effect of hyperspectral features on F t The stronger the impact;

[0142] The formula for calculating time series weighted fusion feature data is:

[0143]

[0144] Among them, F fus is the overall temporal weighted feature, w t is the weight at time t, and Where c = 0.7 is the temporal attenuation factor, which is used to enhance the feature contribution of the central time point of the flowering period, and τ = round(k / 2) is the collection time index of the central time point of the flowering period;

[0145] F fus Integrate multimodal evolutionary information throughout the flowering period; t is the weight of time t, which obeys a Gaussian decay distribution centered on τ. This formula strengthens the characteristic contribution of the central moment of the flowering period, such as the peak flowering period, through exponential decay, weakens the noise interference at the beginning and end of the period, and captures the key stage changes in the growth process of Fritillaria, such as the continuous characteristic evolution from corolla expansion to withering. t The larger it is, the closer the moment t is to the central moment τ, and the higher its fusion weight is;

[0146] The larger the c is, the more the weight is concentrated on the central moment (faster decay). c = 0.7 balances the near-term and long-term characteristics and adapts to the flowering period of about 30 days of Fritillaria. The smaller |t-τ| is, the more w t The bigger;

[0147] The cross-modal features are obtained through time-series weighted fusion. The logic of constructing the spatiotemporal joint graph structure based on the fused features is as follows:

[0148] The node feature is the pixel point of the flower area in the hyperspectral image corresponding to each node, and the node feature is the time series fusion feature vector of the pixel d is the single-moment feature dimension;

[0149] The node simultaneously carries the fusion information of multiple data (HSP, TIR, RGB) at multiple time points, including the temporal changes of chemical composition (HSP), metabolic activity (TIR), and surface texture (RGB);

[0150] The information carried by node features includes the temporal changes of chemical composition HSP, metabolic activity TIR and surface texture RGB, which can provide multi-dimensional information for subsequent analysis; fus It is the temporal fusion feature vector of each pixel, integrating the multimodal features collected k times during the flowering period. The larger the dimension d of the feature vector, the richer the information carried by each node, and the more comprehensive the description of the plant's growth characteristics. The increase in the number of collections k and the feature dimension d during the flowering period can make the node features more comprehensive, which may improve the performance and analysis capabilities of the model.

[0151] F fus The dynamic temporal integration of hyperspectral, thermal infrared, and RGB imagery was used to characterize the growth pattern of Fritillaria from corolla expansion to withering. A larger k value results in higher temporal resolution and more detailed capture of growth dynamics, but with increased computational complexity. A larger d value results in richer features at each moment, but may introduce redundant information.

[0152] The edge connection rules, based on three types of relationships, namely spatial proximity, band correlation, and feature similarity, jointly characterize the spatiotemporal evolution of the flowering period of Fritillaria:

[0153] For spatial proximity, if the Manhattan distance between two pixels is less than the threshold 5px: ‖(x i ,y i )-(x j ,y j )‖ man =|x i -x j |+|y i -y j |<5px, then an edge is established. This threshold is consistent with the micro texture unit size of the Fritillaria bulb, which is 5×5px;

[0154] 5px reflects the spatial proximity relationship between pixels. That is, if the distance is less than 5 pixels, it is considered adjacent. This proximity can capture the micro-texture changes of plant structures. If the proximity is strong, it may indicate that they are biologically related. The smaller the distance between pixels, the greater the possibility of establishing an edge, indicating that spatially close pixels are more likely to influence each other. The stronger the spatial correlation between pixels, the higher the edge weight may be.

[0155] For band correlation, if the reflectance correlation coefficient of two pixels in the HSP band of 520-600nm exceeds the threshold: Then the edge is established, and there is significant redundancy when the correlation coefficient between bands is greater than 0.85;

[0156] corr(h i ,h j) is the reflectivity sequence of pixels i and j in the 520-600 nm band, reflecting the linear relationship between the reflectivities of the two bands. A high correlation indicates that the two pixels have similar reflectivity characteristics in the corresponding band. A larger correlation coefficient indicates that there is significant redundancy between the two bands, which can indicate the similarity of the two pixels in feature description. The more correlated the reflectivity between the bands, the greater the possibility of establishing an edge, reflecting similarity in feature extraction.

[0157] 520-600nm is the sensitive band for anthocyanins and chlorophyll in Fritillaria thunbergii. The high correlation (>0.85) indicates that the pigment distribution patterns of the two pixels are consistent, reflecting the uniform growth of healthy plants. i ,h j ) is larger, the more consistent the reflectivity change trends of the two pixels are, and the more significant the edge connection is;

[0158] For feature similarity, the average cosine similarity of the temporal features of nodes i and j is defined as: Then, an edge is established to capture the overall consistency of plant growth during the flowering period;

[0159] If the average cosine similarity exceeds 0.7, an edge is established, indicating that the temporal feature evolution patterns of the two pixels are highly similar. It is the temporal fusion feature vector of pixels i and j. The cosine similarity measures the synchronization of feature changes of two pixels throughout the flowering period. A high similarity (>0.7) indicates that their growth rhythms are consistent, such as unfolding or withering at the same time.

[0160] sim(F i ,F j ) reflects the similarity of the temporal features between nodes. The closer it is to 1, the higher the similarity. High similarity may mean that the two nodes are consistent in growth features, which can indicate the growth pattern in the same time period. The higher the similarity of the feature vectors, the greater the probability of establishing an edge, which emphasizes the biological relevance of similar features. i ,F j ) is larger, the more similar the growth patterns of the two pixels are, and the stronger the edge connection is;

[0161] The spatiotemporal joint graph structure consists of three parts: a node set, an edge set, and a weight matrix, which is formally defined as: G(t) = {V(t), E(t), A(t)}, where each node corresponds to a pixel in the hyperspectral image, and the node set V(t) = {v x,y (t)|x=1…W,y=1…H}, where W is the image width, H is the height, and the total number of nodes is W×H;

[0162] The construction of the edge set E(t) is achieved through the joint implementation of three types of rules to comprehensively characterize the spatiotemporal evolution characteristics of Fritillaria during the flowering period:

[0163] For the spatial proximity rule, if the Manhattan distance between two pixels is less than 5px, an edge is established. This threshold is consistent with the size of the microscopic texture unit of the Fritillaria bulb and is used to capture the correlation of local growth characteristics. For the band correlation rule, if the reflectance correlation coefficient of two pixels in the 520-600nm band is greater than the distance between the two pixels, an edge is established, reflecting the stability of the pigment distribution during the flowering period. For the temporal continuity rule, if the interval between two time points is no more than 1 day, an edge is established, covering the adjacent acquisition interval Δt = 5 days and the lag effect of plant growth, ensuring the continuity of temporal evolution; t,t ′ is the collection time index;

[0164] The 7-day threshold covers the adjacent collection interval (5 days) and the lag effect of plant growth (such as developmental delay caused by environmental changes), ensuring the consistency of temporal evolution. For example, the characteristics of the peak flowering period may be affected by the initial flowering period, and information is transmitted through edge connections. The smaller the time interval, the stronger the correlation between time points, and the higher the edge weight may be.

[0165] Weight Matrix The weighted fusion strategy described above is used to calculate the strength of the association between nodes. The weight matrix characterizes the connection strength between nodes and can indicate the relative importance of different nodes in the graph model. The larger the value of A(t), the stronger the degree of association between nodes, which may affect the calculation and results of the fusion features. The distribution of weights can be dynamically adjusted based on the feature correlation, spatial proximity, and temporal similarity of different modalities, reflecting the true biological relationship.

[0166] The larger A(t), the stronger the association between nodes i and j, and the more significant the information transfer in graph convolution; spatial proximity ensures the biological rationality of local structures, such as cell interactions within a bulb unit; band correlation strengthens the stability of spectral features and avoids noise interference; feature similarity captures the global consistency of plant growth and distinguishes the temporal patterns of different varieties; temporal continuity models the dynamic evolution of flowering period and covers growth lag caused by environmental fluctuations; the weight matrix balances the contributions of spatial, spectral and temporal information through weighted fusion, providing interpretable topological relationships for dynamic graph convolution;

[0167] G(t) reflects a dynamic time-space structure that can characterize the relationships and features between pixels and provides a comprehensive graph model for analysis. An increase in the number of nodes V(t) indicates a more complex graph structure, which can capture more growth features. The construction rules of the edge set E(t) ensure that the graph structure can reflect biological correlations. The weight matrix A(t) accurately reflects the mutual influence between nodes by weighting different features, thereby affecting the analysis results of the entire graph structure.

[0168] Step 3: Based on the same method, the joint graph structure of multiple Fritillaria samples of known varieties is obtained to construct a training set. The joint graph structure of the Fritillaria samples is used as input, and the varieties are used as labels to train the dynamic graph convolutional network. The network parameters are updated by optimizing the adjacency matrix. The joint graph structure of the Fritillaria samples to be tested is input into the trained dynamic graph convolutional network, and the predicted value of the variety of the Fritillaria samples to be tested is output;

[0169] Based on the same method, the joint graph structure of multiple known varieties of Fritillaria samples is obtained to construct the training set. The logic of training the dynamic graph convolutional network with the joint graph structure of Fritillaria samples as input and the variety as the label is as follows:

[0170] The input is the spatiotemporal joint graph structure G(t) = {V(t), E(t), A(t)} of the flower area during the flowering period, and the node input feature dimension d in =64, output dimension d out =32, the graph convolution layer of the dynamic graph convolutional network is:

[0171]

[0172] in, is the node feature matrix of the lth layer, which represents the feature representation of each node in the lth layer. is the normalized adjacency matrix, which is used to adjust the structural information of the graph to ensure the balance of information propagation, and D is the degree matrix, whose diagonal elements are the node degrees;

[0173] is the trainable weight matrix of the lth layer, used for feature transformation, and the adjacency matrix A (l) Decoupling, and d in =64,d out =32, GELU is the Gaussian error linear unit activation function, and its smooth nonlinear characteristics adapt to the gradual growth pattern of plants during the flowering period; A (l) Represents the adjacency matrix, optimizes the graph structure, reflects the intrinsic correlation of the data, and optimizes through gradient update, W (l) Only the weight matrix is ​​represented, feature transformation is optimized, high-level semantic information is extracted, and updated through backpropagation;

[0174] H (l+1) It reflects the node features after being processed by the graph convolution layer. It combines the adjacency information and node features and can capture the relationship between nodes. After the convolution operation, the node features can better integrate the information of neighboring nodes, thereby improving the model's ability to classify Fritillaria varieties. The larger the normalized adjacency matrix and feature matrix, the smaller the node feature H. (l+1) The larger the value of , the stronger the weight and features of the node in the graph;

[0175] During training, the adjacency matrix is ​​adjusted according to the classification loss gradient:

[0176]

[0177] Among them, η = 0.01 is the learning rate, which controls the update step size of the adjacency matrix, Loss is the cross entropy loss function, which reflects the gap between the model prediction and the true label, and the Sigmoid function ensures the weight Prevent weight overflow; A (l) is the adjacency matrix of the lth layer, which represents the connection relationship between nodes;

[0178] A (l+1) This reflects the updated adjacency matrix, which enhances the model's learning ability by dynamically adjusting the connection relationships. The updated adjacency matrix can more effectively reflect the similarity and connectivity between nodes, thereby improving the model's performance. The larger the gradient of the loss function, the greater the adjustment of the adjacency matrix, which means that the model needs more learning and optimization in this direction.

[0179] The node feature transfer across acquisition moments is achieved through gated recurrent units:

[0180] H (t+1) =GRU(H (t) ,Concat(A (t) H (t) W (t) ))

[0181] Among them, H (t) A is the node feature at the acquisition time t, representing the plant state at that time. (t) H (t) W (t) It is the spatiotemporal-spectral feature aggregation result, combining node features and adjacency information. Input GRU update gate, GRU selectively retains historical information through reset gate and update gate to adapt to the feature persistence of different stages of flowering period;

[0182] H (t+1) It is the node feature updated by GRU, reflecting the state change between time steps. This feature integrates the information from time t and the dynamic change of adjacency relationship, and can better capture the temporal characteristics of Fritillaria growth. The richer the node features, adjacency matrix and weight, the better the updated feature H. (t+1) It will be more informative and enhance the ability to classify Fritillaria species;

[0183] Back propagation, independent calculation of the weight matrix W (l) and the gradient of the adjacency matrix A (l) , adjust A according to the dynamic update formula (l); Terminate if the validation set loss does not decrease for 10 consecutive rounds;

[0184] The global graph pooling of the output layer is followed by a fully connected layer and Softmax:

[0185] P=Softmax(MLP(MeanPool(H (L) )))

[0186] Where P∈R C is the probability distribution of Fritillaria species, indicating the predicted probability of each category, H (L) It is the feature matrix of the last layer of nodes, which integrates the information of all stages of flowering period. L is the last layer of the network. MeanPool aggregates features along the spatial dimension to eliminate the interference of local noise on the classification results.

[0187] P reflects the model's predicted probability for each Fritillaria species. A larger probability indicates a higher likelihood of that species. The Softmax function converts the output into a probability distribution, ensuring that the sum of the probabilities of all categories is 1, facilitating classification. The higher the quality of the input features, the more accurate the probability distribution of the model output, improving the classification success rate.

[0188] The training set includes each sample as a time series graph sequence {G1, G2, ..., G k}, the label is the Fritillaria variety Y = {1, 2, …, C}, C is the total number of categories;

[0189] Input the sample to be tested timing diagram sequence C test , output probability distribution P test :

[0190]

[0191] Among them, P test,C is the output probability distribution of the sample to be tested and is the predicted Fritillaria species category;

[0192] Is the final output of the classification model, which represents the variety prediction of the sample to be tested. If P test,C The larger the probability, the more confident the model is that the sample belongs to the category, and the more concentrated the probability distribution of the model output is in a certain category. The classification result is based on the probability distribution of the model output. The larger the probability, the stronger the certainty of the classification.

[0193] Step 4: Calculate the Euclidean distance between the predicted value of the species of the Fritillaria sample to be tested and the cluster center of the same species in the training set. If the distance exceeds the set threshold, the sample is judged to be abnormal.

[0194] The logic for calculating the Euclidean distance between the predicted value of the variety of the Fritillaria sample to be tested and the cluster center of the same variety in the training set is:

[0195] Extract the second-to-last layer of global features of all known normal samples in the training set:

[0196]

[0197] in, is the node feature matrix of the u-th normal sample in the second-to-last layer, d ′ =64 is the output dimension of the penultimate layer of the dynamic graph convolutional network, representing the number of features, v∈{1,2,…,C} is the category label of the known Fritillaria varieties, u=1,2,…,N v is the training sample index of category v, N v is the number of normal samples of category v, MeanPool(·) is a global pooling operation, taking the mean along the spatial dimension W×H;

[0198] It reflects the global features of the u-th normal sample and represents the representation of the sample in the feature space. Through the pooling operation, it emphasizes the overall features and ignores the local details, so that the model can better learn the main features of the sample. The richer the content of the node feature matrix, the better the pooling result. It may more accurately reflect the typical characteristics of the sample, thus facilitating subsequent cluster analysis;

[0199] Calculate the characteristic mean of normal samples by species:

[0200]

[0201] Among them, ClusterCenter v is the cluster center of class v, which represents the average value of the sample characteristics of this category;

[0202] ClusterCenter v It reflects the average characteristics of all normal samples in the feature space and can represent the typical characteristics of the category. The calculation of the cluster center enables subsequent anomaly detection to be based on the characteristic distribution of normal samples, more accurately identifying abnormal samples. The more sample feature vectors there are, the more robust the calculation of the cluster center is, and it can better reflect the true characteristics of the category.

[0203] Input the time-space joint graph structure C of the sample to be tested test Extract global eigenvectors:

[0204]

[0205] in, is the node feature matrix of the sample to be tested in the second to last layer of the model, F test is the global feature vector of the sample to be tested;

[0206] F test It reflects the global characteristics of the sample to be tested, which is consistent with the characteristic format of the normal sample. By extracting the characteristics of the sample to be tested, it can be compared with the cluster center of the normal sample to determine whether it is abnormal. The richer the characteristic matrix of the sample to be tested, the better the global characteristics F test It will be more representative and more accurate in judging the status of the sample;

[0207] If the classification result is a category Calculate the Euclidean distance between the sample to be tested and the corresponding cluster center:

[0208] S anomaly =‖F test -ClusterCenter v ‖2

[0209] Among them, S anomaly It is the Euclidean distance between the sample to be tested and the corresponding cluster center, reflecting the degree of deviation between the characteristics of the sample to be tested and the normal sample;

[0210] S anomaly It reflects the distance between the tested sample and the normal sample cluster center in the feature space. The larger the distance, the greater the difference between the tested sample and the normal sample characteristics. The larger the distance value, the more likely it is that the sample has abnormal characteristics and deserves attention. The greater the difference between the tested sample characteristics and the cluster center, the greater the Euclidean distance S. anomaly The larger it is, the more likely the sample is to be considered abnormal;

[0211] The threshold range δ is set based on the distribution of normal sample distances in the training set, and δ=[v train -3σ train ,μ train +3σ train ], where μ train is the mean distance between all normal samples in the training set and the corresponding cluster center, σ train is the standard deviation of the normal sample distance of the training set;

[0212] The threshold δ for anomaly detection is set based on the normal sample distance distribution. It is the critical value used to determine whether a sample is abnormal. The larger the threshold, the stricter the judgment standard. By setting the threshold, normal samples and abnormal samples can be statistically distinguished to ensure the accuracy of detection. The larger the mean and standard deviation of the training set sample distance, the larger the threshold δ may be, which may result in more samples being considered normal.

[0213] If S anomaly >δ, it is determined to be an abnormal sample; the classification label is The exception labels are:

[0214]

[0215] Among them, IsAnomaly is a label indicating whether the sample to be tested is judged to be abnormal. When IsAnomaly=1, that is, S anomaly ∈δ is an abnormal sample that triggers a warning. When IsAnomaly=0, the classification result is obtained;

[0216] IsAnomaly reflects the abnormal state of the sample. A value of 1 indicates abnormality and a value of 0 indicates normality. This judgment enables the system to issue warnings or conduct further analysis based on the degree of deviation of sample characteristics. The larger the Euclidean distance is compared with the threshold, the more likely S anomaly A higher probability of exceeding the threshold leads to an increased likelihood that the sample will be marked as an anomaly.

[0217] See also Figure 2 The present invention also provides a device for identifying Fritillaria varieties based on hyperspectral images. The system is used to implement the above-mentioned method for identifying Fritillaria varieties based on hyperspectral images, specifically comprising:

[0218] A data acquisition module is used to continuously collect multimodal data of the Fritillaria sample to be tested during the flowering period, the multimodal data including hyperspectral data, thermal imaging data and RGB images, and preprocess the multimodal data. The preprocessing includes radiometric correction, alignment of RGB images of adjacent time points based on the SIFT algorithm, and simultaneous application of alignment mapping to the hyperspectral and thermal imaging data to eliminate spatial misalignment;

[0219] The feature extraction module is used to perform wavelet packet decomposition on the preprocessed multimodal data, calculate the frequency domain energy ratio to generate adaptive weight coefficients, obtain cross-modal features through time-series weighted fusion, and construct a spatiotemporal joint graph structure based on the fused features. Its nodes are pixel-level time-series feature vectors, and its edges are defined by spatial proximity, band correlation, and feature similarity.

[0220] The model construction module is used to obtain the joint graph structure of multiple Fritillaria samples of known varieties based on the same method to construct a training set. The joint graph structure of the Fritillaria samples is used as input, and the varieties are used as labels to train the dynamic graph convolutional network. The network parameters are updated by optimizing the adjacency matrix. The joint graph structure of the Fritillaria samples to be tested is input into the trained dynamic graph convolutional network, and the predicted value of the variety of the Fritillaria samples to be tested is output;

[0221] The classification result module is used to calculate the Euclidean distance between the predicted value of the variety of the Fritillaria sample to be tested and the cluster center of the same variety in the training set. If the distance exceeds the set threshold, the sample is judged to be abnormal.

[0222] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0223] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software depends on the specific application and design constraints of the technical solution.

[0224] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment. The above description is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of this application.

Claims

1. A method for identifying Fritillaria varieties based on hyperspectral images, characterized in that: The specific steps include: Step 1: Continuously collect multimodal data of the Fritillaria sample to be tested during the flowering period. The multimodal data includes hyperspectral data, thermal imaging data, and RGB images. The multimodal data is preprocessed, including radiometric correction, alignment of RGB images of adjacent time points based on the SIFT algorithm, and simultaneous application of alignment mapping to the hyperspectral and thermal imaging data to eliminate spatial misalignment. Step 2: Perform wavelet packet decomposition on the preprocessed multimodal data, calculate the frequency domain energy ratio to generate adaptive weight coefficients, obtain cross-modal features through time-series weighted fusion, and construct a spatiotemporal joint graph structure based on the fused features. Its nodes are pixel-level time-series feature vectors, and its edges are defined by spatial proximity, band correlation, and feature similarity. Step 3: Based on the same method, the joint graph structure of multiple Fritillaria samples of known varieties is obtained to construct a training set. The joint graph structure of the Fritillaria samples is used as input, and the varieties are used as labels to train the dynamic graph convolutional network. The network parameters are updated by optimizing the adjacency matrix. The joint graph structure of the Fritillaria samples to be tested is input into the trained dynamic graph convolutional network, and the predicted value of the variety of the Fritillaria samples to be tested is output; Step 4: Calculate the Euclidean distance between the predicted value of the variety of the Fritillaria sample to be tested and the cluster center of the same variety in the training set. If the distance exceeds the set threshold, the sample is judged to be abnormal.

2. The method for identifying Fritillaria varieties based on hyperspectral images according to claim 1, wherein: The method of continuously collecting multimodal data of the Fritillaria sample to be tested during the flowering period is as follows: starting from when the corolla expansion degree of the Fritillaria is greater than or equal to 90%, collecting at equal time intervals until the petal shedding rate is greater than or equal to 30% stops, and collecting k times in total at a fixed interval of Δt = 5 days; The equipment used to collect multimodal data includes a hyperspectral imager with a wavelength range of 400-700nm and a spectral resolution of 2nm or less, a thermal imager with a long-wave infrared wavelength range of 8-14μm and a thermal sensitivity of 30mK or less, and an RGB camera with a spatial resolution of 1024×7.68 million pixels, a resolution of 61MP or more, and macro mode support. A rubidium atomic clock is disciplined using GPS to generate a global trigger pulse, synchronizing the acquisition timing of the three devices. The logic for preprocessing hyperspectral data, thermal imaging data, and RGB images is as follows: Radiometric correction is performed on the hyperspectral data. A dark reference image D(x, y) is collected in a completely light-shielded environment. Then, a standard whiteboard image W(x, y) with a reflectivity of 99% is collected. The formula for calculating the reflectivity is: Among them, R raw is the original reflectivity, (x, y) is the pixel coordinate, R cor is the corrected reflectivity; For thermal imaging data, the original thermal radiation grayscale value is converted to temperature value according to the following formula: Among them, G(x,y) is the original grayscale value output by the thermal imager, G offset is the grayscale offset corresponding to the zero temperature of the thermal imager, G gain is the temperature-grayscale linear conversion gain coefficient, T(x,y) is the temperature value; According to the temperature data T collected by the temperature and humidity sensor env , Humidity data RH correction plant temperature: T cor (x,y)=T(x,y)-k T ·(T env -25)+0.03·(RH-60%), where k T =0.05 is the compensation coefficient, 25℃ is the calibration reference temperature, T cor is the corrected plant temperature value; The SIFT algorithm is used to extract feature points from adjacent temporal RGB images. Reliable matching pairs are screened using a bidirectional nearest neighbor distance ratio of ε=0.

6. The RANSAC algorithm is then used to robustly estimate the affine transformation matrix including scaling, rotation, and translation. Finally, bilinear interpolation is used to achieve spatial alignment of images across temporal phases. These images are then synchronously mapped to hyperspectral and thermal imaging data to eliminate spatial misalignment caused by plant growth displacement and perspective offset, ensuring the consistency of multimodal data in both temporal and spatial dimensions.

3. The method for identifying Fritillaria varieties based on hyperspectral images according to claim 2, characterized in that: The logic of performing wavelet packet decomposition on the pre-processed multimodal data and calculating the frequency domain energy ratio to generate the adaptive weight coefficient is as follows: The DB4 wavelet basis function is used to perform three-layer wavelet packet decomposition on the three modal data of hyperspectral data HSP, thermal imaging data TIR and RGB image, with the decomposition layer number J=3. Each layer of wavelet packet decomposition divides the signal into low-frequency and high-frequency parts. The total number of frequency bands after three-layer decomposition is 2 3 =8: Frequency domain energy ratio calculation, the formula for calculating sub-band energy is: Among them, E WPD (X,s) represents the energy of modal data X in the sth sub-band, X∈{HSR,TIR,RGB} is the modality type identifier, s∈{1,2,…,8} is the sub-band number, is the nth wavelet coefficient of the modal data X in the sth sub-band, N s is the total number of wavelet coefficients in the sth sub-band, n∈{1,2,…,N s } is the wavelet coefficient index; Calculate the total energy E total (X) is based on the formula: Among them, E total (X) represents the overall energy of modal data X; The calculation formulas for the weight coefficients α, β and γ are as follows: Among them, E sum =E total (HSP)+E total (TIR)+E total (RGB); The formula for calculating the single-moment fusion feature is: Among them, F t is the fusion feature vector at the acquisition time t, and the weight coefficient α t ,β t ,γ t is the modal weight at that moment; Normalize(·) means to perform Z-score normalization on the eigenvector; F HSP,t is the feature vector of the hyperspectral data collected at the acquisition time t after wavelet packet decomposition and feature extraction, F TIR,t It is the feature vector of the thermal infrared data collected at the acquisition time t after temperature conversion and feature extraction, F RGB,t It is the feature vector of the RGB image data collected at the acquisition time t after spatial alignment and feature extraction, t is the index of the acquisition time, and t∈[1,k]; The formula for calculating time series weighted fusion feature data is: Among them, F fus is the overall temporal weighted feature, w t is the weight at time t, and Where c = 0.7 is the temporal attenuation factor, which is used to enhance the feature contribution of the central time point of the flowering period, and τ = round(k / 2) is the collection time index of the central time point of the flowering period.

4. The method for identifying Fritillaria varieties based on hyperspectral images according to claim 3, characterized in that: The cross-modal features are obtained through time-series weighted fusion. The logic of constructing the spatiotemporal joint graph structure based on the fused features is as follows: The node feature is the pixel point of the flower area in the hyperspectral image corresponding to each node, and the node feature is the time series fusion feature vector of the pixel d is the single-moment feature dimension; The edge connection rules, based on three types of relationships, namely spatial proximity, band correlation, and feature similarity, jointly characterize the spatiotemporal evolution characteristics of the Fritillaria flowering period: For spatial proximity, if the Manhattan distance between two pixels is less than the threshold 5px: ‖(x i ,y i )-(x j ,y j )‖ man =|x i -x j |+|y i -y j |<5px, then an edge is established. This threshold is consistent with the micro texture unit size of the Fritillaria bulb, which is 5×5px; For band correlation, if the reflectance correlation coefficient of two pixels in the HSP band of 520-600nm exceeds the threshold: Then the edge is established, and there is significant redundancy when the correlation coefficient between bands is greater than 0.85; For feature similarity, the average cosine similarity of the temporal features of nodes i and j is defined as: Then, an edge is established to capture the overall consistency of plant growth during the flowering period; The spatiotemporal joint graph structure consists of three parts: a node set, an edge set, and a weight matrix, which is formally defined as: G(t) = {V(t), E(t), A(t)}, where each node corresponds to a pixel in the hyperspectral image, and the node set V(t) = {v x,y (t)|x=1…W,y=1…H}, where W is the image width, H is the height, and the total number of nodes is W×H; The construction of the edge set E(t) is achieved through the joint implementation of three types of rules to comprehensively characterize the spatiotemporal evolution characteristics of Fritillaria during the flowering period: For the spatial proximity rule, an edge is established if the Manhattan distance between two pixels is less than 5px. This threshold is consistent with the size of the microtexture unit of the Fritillaria bulb and is used to capture the correlation of local growth characteristics. For the band correlation rule, an edge is established if the reflectance correlation coefficient of two pixels in the 520-600nm band is high, reflecting the stability of pigment distribution during the flowering period. For the temporal continuity rule, an edge is established if the interval between two time points is no more than 7 days, covering the adjacent acquisition interval Δt = 5 days and the lag effect of plant growth to ensure the coherence of temporal evolution. Weight Matrix Calculated by the above weighted fusion strategy.

5. The method for identifying Fritillaria varieties based on hyperspectral images according to claim 4, characterized in that: Based on the same method, the joint graph structure of multiple known varieties of Fritillaria samples is obtained to construct the training set. The logic of training the dynamic graph convolutional network with the joint graph structure of Fritillaria samples as input and the variety as the label is as follows: The input is the spatiotemporal joint graph structure G(t) = {V(t), E(t), A(t)} of the flower area during the flowering period, and the node input feature dimension d in =64, output dimension d out =32, the graph convolution layer of the dynamic graph convolutional network is: in, is the node feature matrix of the lth layer, is the normalized adjacency matrix, and D is the degree matrix, whose diagonal elements are the node degrees; is the trainable weight matrix of the lth layer, and the adjacency matrix A (l) Decoupling, GELU is a Gaussian error linear unit activation function, and its smooth nonlinear characteristics adapt to the gradual growth pattern of plants during the flowering period; During training, the adjacency matrix is ​​adjusted according to the classification loss gradient: Among them, η = 0.05 is the learning rate, Loss is the cross entropy loss function, the Sigmoid function constrains the adjacent weight range to [0,1], A (l) is the adjacency matrix of the lth layer; The node feature transfer across acquisition moments is achieved through gated recurrent units: H (t+1) =GRU[H (t) ,Concat(A (t) H (t) W (t) )] Among them, H (t) is the node feature at the acquisition time t, representing the plant state at that time, A (t) H (t) W (t) is the spatiotemporal-spectral feature aggregation result, Input GRU update gate, GRU selectively retains historical information through reset gate and update gate to adapt to the feature persistence of different stages of flowering period; The output layer generates the variety probability distribution through global mean pooling and fully connected layers: P=Softmax<MLP{[MeanPool(H (L) )]}> Among them, L is the last layer of the network, H (L) It is the last layer of node feature matrix, which integrates the information of all stages of flowering period; MeanPool aggregates features along the spatial dimension to eliminate the interference of local noise on the classification results.

6. The method for identifying Fritillaria varieties based on hyperspectral imaging according to claim 5, characterized in that: The joint graph structure of the Fritillaria sample to be tested is input into the trained dynamic graph convolutional network, and the logic for outputting the Fritillaria variety classification result is as follows: The training set includes each sample as a time series graph sequence {G1, G2, ..., G k }, the label is the Fritillaria variety Y = {1, 2, …, C}, C is the total number of categories; Input the sample timing diagram sequence G to be tested test , output probability distribution P test : Among them, P test,C is the output probability distribution of the sample to be tested, and The predicted Fritillaria species category.

7. The method for identifying Fritillaria varieties based on hyperspectral imaging according to claim 6, characterized in that: The logic for calculating the Euclidean distance between the predicted value of the variety of the Fritillaria sample to be tested and the cluster center of the same variety in the training set is: Extract the second-to-last layer of global features of all known normal samples in the training set: in, is the node feature matrix of the u-th normal sample in the second-to-last layer, d ′ =64 is the output dimension of the penultimate layer of the dynamic graph convolutional network, v∈{1,2,…,C} is the category label of the known Fritillaria varieties, u=1,2,…,N v is the training sample index of category v, N v is the number of normal samples of category v, and MeanPool(·) is the mean along the spatial dimension W×H; Calculate the characteristic mean of normal samples by species: Among them, ClusterCenter v is the cluster center of category v; Input the time-space joint graph structure G of the sample to be tested test Extract global eigenvectors: in, is the node feature matrix of the sample to be tested in the second to last layer of the model, F test is the global feature vector of the sample to be tested; If the classification result is a category Calculate the Euclidean distance between the sample to be tested and the corresponding cluster center: S anomaly =‖F test -ClusterCenter v ‖2 Among them, S anomaly is the Euclidean distance between the sample to be tested and the corresponding cluster center; The threshold range δ is set based on the distribution of normal sample distances in the training set, and δ=[μ train -3σ train ,μ train +3σ train ], where μ train is the mean distance between all normal samples in the training set and the corresponding cluster center, σ train is the standard deviation of the normal sample distance of the training set; If S anomaly >δ, it is determined to be an abnormal sample, and the abnormal label is: Among them, IsAnomaly is a label indicating whether the sample to be tested is judged to be abnormal. When IsAnomaly=1, that is, S anomaly ∈δ is an abnormal sample that triggers a warning. When IsAnomaly=0, the classification result is obtained.

8. A device for identifying Fritillaria varieties based on hyperspectral imaging, characterized by: The device is used to execute the Fritillaria variety identification method based on hyperspectral images according to any one of claims 1 to 7, comprising: A data acquisition module is used to continuously collect multimodal data of the Fritillaria sample to be tested during the flowering period, the multimodal data including hyperspectral data, thermal imaging data and RGB images, and preprocess the multimodal data. The preprocessing includes radiometric correction, alignment of RGB images of adjacent time points based on the SIFT algorithm, and simultaneous application of alignment mapping to the hyperspectral and thermal imaging data to eliminate spatial misalignment; The feature extraction module is used to perform wavelet packet decomposition on the preprocessed multimodal data, calculate the frequency domain energy ratio to generate adaptive weight coefficients, obtain cross-modal features through time-series weighted fusion, and construct a spatiotemporal joint graph structure based on the fused features. Its nodes are pixel-level time-series feature vectors, and its edges are defined by spatial proximity, band correlation, and feature similarity. The model construction module is used to obtain the joint graph structure of multiple Fritillaria samples of known varieties based on the same method to construct a training set. The joint graph structure of the Fritillaria samples is used as input, and the varieties are used as labels to train the dynamic graph convolutional network. The network parameters are updated by optimizing the adjacency matrix. The joint graph structure of the Fritillaria samples to be tested is input into the trained dynamic graph convolutional network, and the predicted value of the variety of the Fritillaria samples to be tested is output; The classification result module is used to calculate the Euclidean distance between the predicted value of the variety of the Fritillaria sample to be tested and the cluster center of the same variety in the training set. If the distance exceeds the set threshold, the sample is judged to be abnormal.