A plant class determination method based on multispectral images

CN120599457BActive Publication Date: 2026-09-22INNER MONGOLIA BAZHAO INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510447701.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2026-09-22
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

多光谱图像自身具有高度复杂性,其不同波段间光谱特征存在显著差异

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599457B_ABST
    Figure CN120599457B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a plant category determination method based on a multispectral image, relates to the technical field of image processing, and is convenient for improving plant classification precision.The method comprises the following steps: determining a pixel threshold of each spectral band in a to-be-detected multispectral image; screening the pixels of each spectral band in the to-be-detected multispectral image according to the pixel threshold of each spectral band, to obtain a screened multispectral image; determining local features and global features of the screened multispectral image according to the screened multispectral image; and determining a target plant category included in the to-be-detected multispectral image according to a preset plant classification model, the local features and the global features.The present application is suitable for plant classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method for determining plant categories based on multispectral images. Background Technology

[0002] Multispectral images have significant applications in agriculture, ecological monitoring, and plant classification. However, multispectral images are inherently highly complex, with significant differences in spectral characteristics across different bands. Furthermore, factors such as illumination variations, occlusion, and cluttered backgrounds can significantly impact the accuracy of feature extraction.

[0003] In existing technologies, classification of multispectral images struggles to effectively handle varying lighting conditions, background noise, and differences in image quality. Single-feature extraction methods neglect the overall informational relationships of the plant within the multispectral image, overlooking subtle features that could be crucial for accurate classification. This limitation makes it difficult to meet practical requirements for classification accuracy when faced with complex multispectral image classification tasks. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method for determining plant categories based on multispectral images, which facilitates the improvement of plant classification accuracy.

[0005] In a first aspect, embodiments of the present invention provide a method for determining plant categories based on multispectral images, comprising: determining a pixel threshold for each spectral band in a multispectral image to be detected; filtering pixels in each spectral band of the multispectral image to be detected according to the pixel threshold for each spectral band to obtain a filtered multispectral image; determining local features and global features of the filtered multispectral image according to the filtered multispectral image; and determining the target plant category included in the multispectral image to be detected according to a preset plant classification model, the local features, and the global features.

[0006] According to a specific implementation of an embodiment of this application, determining the pixel threshold of each spectral band in the multispectral image to be detected includes: determining the pixel threshold of the corresponding spectral band based on the mean and standard deviation of each spectral band in the multispectral image to be detected.

[0007] According to a specific implementation of this application, the step of filtering pixels in each spectral band of the multispectral image to be detected based on pixel thresholds of each spectral band to obtain a filtered multispectral image includes: comparing the pixel value of each pixel in each spectral band of the multispectral image to be detected with the pixel threshold of the corresponding spectral band; removing pixels whose pixel values ​​in each spectral band are less than the pixel threshold of the corresponding spectral band; and obtaining the filtered multispectral image based on the remaining pixels in each spectral band excluding the removed pixels.

[0008] According to a specific implementation of an embodiment of this application, determining the local and global features of the selected multispectral image based on the selected multispectral image includes: inputting the selected multispectral image into a dimensionality reduction encoder to obtain a low-dimensional multispectral image; inputting the low-dimensional multispectral image into a reconstruction decoder to obtain a new multispectral image; and determining the local and global features of the new multispectral image based on the new multispectral image.

[0009] According to a specific implementation of an embodiment of this application, determining the local and global features of the new multispectral image based on the new multispectral image includes: extracting the global context features of the new multispectral image based on a preset visual information enhancement module; and extracting the local features of the new multispectral image based on a preset local feature hierarchical network.

[0010] According to a specific implementation of an embodiment of this application, determining the target plant category included in the multispectral image to be detected based on a preset plant classification model, the local features, and the global features includes: fusing the global context features and the local features to obtain enhanced features; and determining the target plant category included in the multispectral image to be detected based on the preset plant classification model and the enhanced features.

[0011] According to a specific implementation of an embodiment of this application, the preset plant classification model is determined according to the following steps: obtaining an unlabeled dataset and a labeled dataset; wherein the dataset includes multiple multispectral images; obtaining an intermediate plant classification model based on the unlabeled dataset and the initial plant classification model; and obtaining a final plant classification model based on the labeled dataset and the intermediate plant classification model.

[0012] According to a specific implementation of an embodiment of this application, the pixel threshold is determined according to the following formula:

[0013] ( ); where μ(λ) is the mean pixel value of band λ; σ(λ) is the standard deviation of pixel values ​​of band λ; ( It is determined by the following formula: ;in, These are weighting coefficients. =1, used to control the contribution of each pixel value in band λ; >0 is the attenuation coefficient, used for adjustment. ( The size of ); n is the number of pixel values ​​in band λ; Let be the pixel value of the i-th pixel in band λ, where i is an integer greater than 0 and less than or equal to n.

[0014] According to a specific implementation of an embodiment of this application, the preset plant classification model includes a fully connected layer; determining the target plant category included in the multispectral image to be detected based on the preset plant classification model, the local features, and the global features includes: obtaining an input feature vector based on the local features and the global features; inputting the input feature vector into the preset plant classification model to obtain multiple probability values ​​corresponding one-to-one with multiple plant categories; determining the target plant category included in the multispectral image to be detected based on the multiple probability values; wherein, the probability value of plant category i is determined according to the following formula: ;in, Let i be the probability value for plant category i; Let be the element in the i-th row and j-th column of the weight matrix of the fully connected layer; C is the total number of rows in the weight matrix of the fully connected layer; D is the total number of columns in the weight matrix of the fully connected layer; i is an integer greater than 0 and less than or equal to C; j is an integer greater than 0 and less than or equal to D; k is an integer greater than 0 and less than or equal to C. It is the j-th element of the input feature vector; is the element in the k-th row and j-th column of the weight matrix of the fully connected layer.

[0015] Secondly, embodiments of the present invention provide a plant category determination device based on multispectral images, comprising: a first determining unit, configured to determine a pixel threshold for each spectral band in a multispectral image to be detected; a filtering unit, configured to filter pixels in each spectral band of the multispectral image to be detected according to the pixel threshold for each spectral band, to obtain a filtered multispectral image; a second determining unit, configured to determine local features and global features of the filtered multispectral image according to the filtered multispectral image; and a third determining unit, configured to determine the target plant category included in the multispectral image to be detected according to a preset plant classification model, the local features, and the global features.

[0016] According to a specific implementation of an embodiment of this application, the first determining unit includes: a first determining module, which determines the pixel threshold of the corresponding spectral band based on the mean and standard deviation of each spectral band in the multispectral image to be detected.

[0017] According to a specific implementation of an embodiment of this application, the filtering unit includes: a comparison module, used to compare the pixel value of each pixel in each spectral band of the multispectral image to be detected with the pixel threshold of the corresponding spectral band; a removal module, used to remove pixels whose pixel value in each pixel of each spectral band is less than the pixel threshold of the corresponding spectral band; and a first obtaining module, used to obtain the filtered multispectral image based on the other pixels in each pixel of each spectral band excluding the removed pixels.

[0018] According to a specific implementation of an embodiment of this application, the second determining unit includes: a first input module, used to input the filtered multispectral image into a dimensionality reduction encoder to obtain a low-dimensional multispectral image; a second input module, used to input the low-dimensional multispectral image into a reconstruction decoder to obtain a new multispectral image; and a second determining module, used to determine the local features and global features of the new multispectral image based on the new multispectral image.

[0019] According to a specific implementation of an embodiment of this application, the second determining module further includes: a first extraction submodule, used to extract global context features of the new multispectral image according to a preset visual information enhancement module; and a second extraction submodule, used to extract local features of the new multispectral image according to a preset local feature hierarchical network.

[0020] According to a specific implementation of an embodiment of this application, the third determining unit includes: a fusion module, used to fuse the global context features and the local features to obtain enhanced features; and a third determining module, used to determine the target plant category included in the multispectral image to be detected based on a preset plant classification model and the enhanced features.

[0021] According to a specific implementation of an embodiment of this application, the third determining unit is specifically used for: acquiring an unlabeled dataset and a labeled dataset; wherein the dataset includes multiple multispectral images; obtaining an intermediate plant classification model based on the unlabeled dataset and an initial plant classification model; and obtaining a final plant classification model based on the labeled dataset and the intermediate plant classification model.

[0022] According to a specific implementation of an embodiment of this application, the third determining unit is further configured to: obtain a first plant classification model based on the labeled dataset and the intermediate plant classification model; determine the loss value of each layer of the network in the first plant classification model based on a preset validation dataset; and adjust the parameters of each layer of the network in the first plant classification model based on the sum of the loss values ​​of each layer of the network to obtain a second plant classification model.

[0023] According to a specific implementation of an embodiment of this application, the third determining unit is further configured to: determine the sparsification loss value of the first plant classification model based on a preset verification dataset; and adjust the parameters of each layer of the first plant classification model according to the sum of the loss functions of each layer of the network to obtain a second plant classification model, comprising: adjusting the parameters of each layer of the first plant classification model according to the sparsification loss value and the sum of the loss functions of each layer of the network to obtain a second plant classification model.

[0024] According to a specific implementation of an embodiment of this application, the pixel threshold is determined according to the following formula:

[0025] ( ); where μ(λ) is the mean pixel value of band λ; σ(λ) is the standard deviation of pixel values ​​of band λ; ( It is determined by the following formula: ;in, These are weighting coefficients. =1, used to control the contribution of each pixel value in band λ; >0 is the attenuation coefficient, used for adjustment. ( The size of ); n is the number of pixel values ​​in band λ; Let be the pixel value of the i-th pixel in band λ, where i is an integer greater than 0 and less than or equal to n.

[0026] According to a specific implementation of an embodiment of this application, the preset plant classification model includes a fully connected layer; the third determining unit includes: a second obtaining module, used to obtain an input feature vector based on the local features and the global features; a third input module, used to input the input feature vector into the preset plant classification model to obtain multiple probability values ​​corresponding one-to-one with multiple plant categories; and a fourth determining module, used to determine the target plant category included in the multispectral image to be detected based on the multiple probability values; wherein, the probability value of plant category i is determined according to the following formula: ;in, Let i be the probability value for plant category i; Let be the element in the i-th row and j-th column of the weight matrix of the fully connected layer; C is the total number of rows in the weight matrix of the fully connected layer; D is the total number of columns in the weight matrix of the fully connected layer; i is an integer greater than 0 and less than or equal to C; j is an integer greater than 0 and less than or equal to D; k is an integer greater than 0 and less than or equal to C. It is the j-th element of the input feature vector; is the element in the k-th row and j-th column of the weight matrix of the fully connected layer.

[0027] Thirdly, embodiments of this application provide an electronic device, including: a housing, a processor, a memory, a circuit board, and a power supply circuit, wherein the circuit board is disposed inside the space enclosed by the housing, and the processor and the memory are disposed on the circuit board; the power supply circuit is used to supply power to various circuits or devices of the above-mentioned electronic device; the memory is used to store executable program code; the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, for executing the plant category determination method based on multispectral images as described in any of the foregoing embodiments.

[0028] Fourthly, embodiments of this application provide a computer-readable storage medium storing one or more computer programs, which, when executed by one or more processors, implement any of the plant category determination methods based on multispectral images described in the foregoing embodiments. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 A flowchart of a plant category determination method provided for embodiments of the present invention;

[0031] Figure 2 A detailed flowchart of a plant category determination method provided for embodiments of the present invention;

[0032] Figure 3 A schematic diagram of a plant category determination device provided in an embodiment of the present invention;

[0033] Figure 4 A schematic diagram of an electronic device provided as an embodiment of the present invention. Detailed Implementation

[0034] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0035] It should be understood that the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0036] In a first aspect, embodiments of the present invention provide a method for determining plant categories based on multispectral images, which facilitates the improvement of plant classification accuracy.

[0037] like Figure 1 As shown, embodiments of the present invention provide a method for determining plant categories based on multispectral images, including:

[0038] S11, determine the pixel threshold for each spectral band in the multispectral image to be detected.

[0039] When determining plant categories based on multispectral images, feature extraction and classification of the multispectral images to be detected are required. First, the features of the original multispectral image need to be preprocessed to provide high-quality input data for subsequent steps. In some embodiments, adaptive thresholding can be used to preprocess the features of the original multispectral image. Adaptive thresholding can dynamically adjust the threshold, intelligently filtering the spectral information of each pixel.

[0040] Multispectral image data contains image information across multiple spectral bands, each including pixel values. Each band exhibits different characteristics under varying lighting conditions, shooting angles, and background environments. The process begins by inputting the raw multispectral image data. During preprocessing, statistical features of each band are calculated to adaptively adjust the pixel values, dynamically determining the pixel threshold for each spectral band in the multispectral image to be detected. This adaptive threshold calculation method effectively addresses issues such as image lighting, occlusion, and noise, ensuring feature consistency in subsequent processing.

[0041] S12, based on the pixel threshold of each spectral band, the pixels of each spectral band in the multispectral image to be detected are filtered to obtain the filtered multispectral image.

[0042] The filtering criteria are determined by comparing the pixel thresholds and pixel values ​​of each spectral band. Pixel values ​​that meet the filtering criteria are retained, while those that do not are set to background values ​​(such as 0 or other preset values). The filtered and retained pixel values ​​for each spectral band are combined to obtain the filtered multispectral image. The multispectral image filtering method based on pixel thresholds can effectively remove irrelevant background noise, retain key feature information, and significantly improve the accuracy and robustness of image processing.

[0043] S13, Based on the filtered multispectral image, determine the local features and global features of the filtered multispectral image.

[0044] Local features refer to the features of specific regions or pixels in an image, typically used to describe detailed information. Global context features refer to features that reflect the overall structure and semantic information of an image. After filtering, the multispectral image can be used to extract global context features through relevant algorithms or models (such as visual information enhancers), and local features can be extracted through relevant algorithms or models (such as local feature hierarchical networks). By extracting both local and global features from the filtered multispectral image, the key information of the image can be comprehensively described, significantly improving the accuracy and robustness of subsequent tasks.

[0045] S14. Based on the preset plant classification model, the local features, and the global features, determine the target plant category included in the multispectral image to be detected.

[0046] Common pre-defined plant classification models include Convolutional Neural Networks (CNNs) and their variants, such as ResNet and VGG. During the training phase, a large amount of multispectral plant image data of different categories needs to be collected to construct a training dataset. The data is preprocessed to ensure consistency and usability, and annotation tools are used to accurately label the plant category of each image.

[0047] The extracted local and global features are input into a pre-defined plant classification model. For some models, further format adjustments or dimensionality transformations of the features may be necessary to adapt to the model's input requirements. Based on the feature patterns and classification rules it has learned, the model processes and predicts the input local and global features to determine the target plant categories included in the multispectral image to be detected.

[0048] The model calculates the probability value for each plant category, reflecting the likelihood that the multispectral image to be detected belongs to each category. Finally, the category with the highest probability value is selected as the final prediction result, determining the plant category included in the multispectral image and outputting the category label. For example, if the model predicts that an image belongs to the "rose" category with a probability of 0.8, which is higher than the probabilities for other plant categories, then the plant in the image is determined to be a rose.

[0049] The plant category determination method based on multispectral images provided by embodiments of the present invention can determine the pixel threshold of each spectral band in a multispectral image to be detected; based on the pixel threshold of each spectral band, the pixels of each spectral band in the multispectral image to be detected are filtered to obtain a filtered multispectral image; based on the filtered multispectral image, the local features and global features of the filtered multispectral image are determined; based on a preset plant classification model, the local features and the global features, the target plant category included in the multispectral image to be detected is determined. Thus, filtering the pixels of each spectral band in the multispectral image to be detected can remove noise, highlight target areas, and improve image quality. Determining the local and global features of the filtered multispectral image can effectively extract global and local features, facilitating improved accuracy in plant category determination.

[0050] In some embodiments, determining the pixel threshold for each spectral band in the multispectral image to be detected includes: determining the pixel threshold for the corresponding spectral band based on the mean and standard deviation of each spectral band in the multispectral image to be detected.

[0051] In the data preprocessing of multispectral images, the statistical characteristics (mean and standard deviation) of each band are calculated. Based on the mean and standard deviation, a pixel threshold is set. In some embodiments, the pixel threshold can be set using a fixed multiplier method, an adaptive method, or a percentage method. This embodiment uses the fixed multiplier method. First, the multispectral image data to be detected is input: let the spectral value of the original image be I(x, y, λ), where x and y are the spatial coordinates of the pixels in the image, and λ is the index of the spectral band, representing different bands in the multispectral image (such as red, green, blue, infrared, etc.). Each band has multiple pixels corresponding to multiple pixel values. The mean and standard deviation are determined based on the multiple pixel values ​​of each band. The mean μ(λ) represents the average value of all pixels in the band, and the standard deviation σ(λ) represents the fluctuation range of pixel values ​​in the band. The pixel threshold for the corresponding spectral band is dynamically determined based on the mean and standard deviation of each spectral band in the multispectral image to be detected. This is an adaptive threshold calculation method. The calculated threshold is used for subsequent pixel filtering. The threshold can be adjusted according to the actual effect to ensure accurate classification results.

[0052] In some embodiments, the step of filtering pixels in each spectral band of the multispectral image to be detected according to the pixel threshold of each spectral band to obtain a filtered multispectral image includes: comparing the pixel value of each pixel in each spectral band of the multispectral image to be detected with the pixel threshold of the corresponding spectral band; removing pixels whose pixel value is less than the pixel threshold of the corresponding spectral band; and obtaining the filtered multispectral image based on the other pixels in each spectral band excluding the removed pixels.

[0053] The pixel threshold for each band is determined using an adaptive threshold calculation method. Then perform pixel filtering.

[0054] For each band λ, iterate through each pixel I(x, y, λ) in the image, and apply the threshold... To determine whether a pixel value meets the filtering criteria, in some embodiments, when determining the pixel value, if pixel I(x, y, λ) ≥ If I(x, y, λ) < Remove pixel I (x, y, λ) from band λ, that is, set this pixel value to a background value (such as 0 or other preset value). Record the filtered pixel value as... (x, y, λ), representing the filtering results for each band. The (x, y, λ) combination yields the filtered multispectral image. This process is called adaptive threshold filtering.

[0055] For example, suppose a multispectral image contains four bands (λ1, λ2, λ3, λ4), and their thresholds are respectively , , and For band λ1, if a pixel value I(x, y, λ1) = 150, and =120, then retain that pixel. For band λ2, if a pixel value I(x, y, λ2) = 80, and If the value is 100, then the pixel value is removed from band λ and set to the background value (e.g., 0). This filtering process effectively removes background noise irrelevant to the target task while retaining key information relevant to classification in the image.

[0056] In some embodiments, determining the local and global features of the filtered multispectral image based on the filtered multispectral image includes: inputting the filtered multispectral image into a dimensionality reduction encoder to obtain a low-dimensional multispectral image; inputting the low-dimensional multispectral image into a reconstruction decoder to obtain a new multispectral image; and determining the local and global features of the new multispectral image based on the new multispectral image.

[0057] Figure 2 A detailed flowchart of the plant category determination method provided for embodiments of the present invention can be found in [reference needed]. Figure 2 After adaptive thresholding, feature extraction and representation are further optimized to obtain a new multispectral image. In some embodiments, an autoencoder can be used to select and enhance features in the multispectral image, retaining the most important regional features and removing background and irrelevant interference. This ensures the quality of the input data for subsequent steps in determining local and global features, thereby further improving the accuracy and robustness of the entire model. Compared to traditional methods, this feature selection and denoising approach provides a better foundation for the efficient operation of other techniques, thus significantly improving classification performance.

[0058] When optimizing the extraction and representation of features from the filtered multispectral images, the filtered multispectral images are input into a dimensionality reduction encoder to obtain low-dimensional multispectral images. In one embodiment, the filtered multispectral images... (x, y, λ) is mapped to a low-dimensional feature space to obtain the compressed feature representation C(x, y, λ), which can be expressed by the formula: C(x, y, λ) = Encoder(I(x, y, λ), WAE), where Encoder is the encoder function, which represents the process of mapping the input data to the low-dimensional feature space; WAE is the weight parameter of the autoencoder network, which is learned through training and represents the mapping relationship in the feature encoding and decoding process.

[0059] The low-dimensional multispectral image is input into the reconstruction decoder to obtain a new multispectral image. For example, the low-dimensional features C(x, y, λ) are mapped back to the original data space to reconstruct the original data. (x, y, λ).

[0060] This can be expressed by a formula: (x, y, λ) = Decoder(C(x, y, λ), WAE), where Decoder is the decoder function, representing the process of mapping low-dimensional features back to the original data space, and WAE are the weight parameters of the autoencoder network, learned through training, representing the mapping relationship between feature encoding and decoding. The goal of an autoencoder is to minimize the input data. (x, y, λ) and reconstructed data The difference between (x, y, λ) is typically represented by the mean squared error (MSE) as the loss function. By optimizing the loss function, the autoencoder can learn a low-dimensional representation of the input data while preserving its key features. By learning this low-dimensional feature representation C(x, y, λ), the autoencoder compresses the input data into a low-dimensional space, retaining key information and removing redundant information. If the input data contains noise, the autoencoder can learn the inherent structure of the data to reconstruct clean output data, thus achieving noise reduction.

[0061] In one specific embodiment, firstly, the weights WAE of the autoencoder are randomly initialized, and the input data... The encoder obtains low-dimensional features C(x, y, λ) from (x, y, λ), and the decoder reconstructs the data from the low-dimensional features. (x, y, λ), calculate the loss function, which is the difference between the input data and the reconstructed data (such as mean square error). The weights WAE can be updated by gradient descent to minimize the loss function. Repeat the above process until the loss function converges to obtain a new multispectral image. Based on the new multispectral image, determine the local and global features of the new multispectral image.

[0062] See Figure 2 In some embodiments, determining the local and global features of the new multispectral image based on the new multispectral image includes: extracting the global context features of the new multispectral image based on a preset visual information enhancement module; and extracting the local features of the new multispectral image based on a preset local feature hierarchical network.

[0063] The pre-defined visual information enhancement module includes a Visual Information Enhancer (VIE), which extracts global contextual features through an Enhanced Information Learning (EIL) mechanism. The EIL mechanism aims to enable the model to learn macroscopic features and structural information of an image over a large scope by introducing global information into the network. In one embodiment, assuming the input image is X, the VIE extracts global features F_VIE through an Enhanced Information Learning module. This process can be represented as: F_VIE = EIL(X, W_VIE), where X represents the original input image data, and W_VIE is the weight parameter in the VIE network, representing the network's ability to adjust for global feature learning. Since local and global features are used in subsequent plant classification, W_VIE can be randomly initialized or pre-trained weights can be used in the initial stage of classification model training. These initial weights provide a starting point for the model to begin learning features specific to the task. When training the model with multispectral plant image data, the loss function is calculated and optimization algorithms, such as SGD and Adam, are used to update the weights. The optimization algorithm will gradually optimize the weights based on the gradient direction and learning rate, making the model's prediction results closer to the actual labels. Techniques such as L1 and L2 regularization may also be used to avoid overfitting and ensure that the model performs well on new and unseen data. At the same time, techniques such as cross-validation are used to adjust and determine the optimal learning rate, batch size, and other hyperparameters, thereby further optimizing W_VIE.

[0064] F_VIE is the global feature extracted after VIE processing. It contains high-level information about the image, such as shape, texture, and overall structure.

[0065] The EIL mechanism enhances the macroscopic features of images through global context learning, helping the model understand the overall structural information of the image. In multispectral images, this global contextual information plays an important role in distinguishing different plant species and growth environments.

[0066] Hierarchical Local Feature Networks (LFNs) focus on extracting local detail features from images, utilizing multiple convolutional operations in Convolutional Neural Networks (CNNs) to extract features from the input image. In this way, LFNs can capture subtle local changes and textures in images, helping models accurately identify features of small objects such as plant leaves and flowers. In one embodiment, given an input image X, the LFN extracts local features F_LFN through multiple convolutional layers. The mathematical expression of this process is: F_LFN = Conv(X, W_LFN), where F_LFN is the local feature extracted through convolution, mainly including the features of details in the image, such as the veins of leaves and the texture of petals; Conv(X, W_LFN) means that the input image X is convolved with the convolution kernel weights W_LFN through convolution operation. W_LFN is the convolution kernel weight in the LFN network, which controls the filter type and size of the convolution operation, similar to W_VIE. W_LFN can also be initialized with initial weights through random initialization or pre-trained model. Since the classification model is a multi-layer convolutional network, in the multi-layer convolutional network, the convolution kernel weights W_LFN of each layer will extract features based on the input feature map it receives. Therefore, the convolution weights of each layer can be adjusted according to the loss function through backpropagation and optimization algorithms to better capture local features. Advanced convolutional neural network techniques, such as residual networks (ResNet) or attention mechanisms, can also be applied to enhance the network's ability to recognize local features.

[0067] LFN extracts high-precision local features from images through stepwise calculations across multiple convolutional layers, ensuring accurate capture of detailed features. This is especially important in plant classification tasks, where details are often crucial for distinguishing different species.

[0068] In some embodiments, determining the target plant category included in the multispectral image to be detected based on a preset plant classification model, the local features, and the global features includes: fusing the global context features and the local features to obtain enhanced features; and determining the target plant category included in the multispectral image to be detected based on the preset plant classification model and the enhanced features.

[0069] like Figure 2 As shown, the fusion of global contextual features and local features is a key step in multispectral image processing. It aims to combine global semantic information and local detail information of an image to improve the expressive power of features and task performance, such as classification, detection, and segmentation. In some embodiments, the fusion method can be feature concatenation, weighted summation, attention mechanisms, and feature pyramid fusion.

[0070] In this embodiment, a weighted summation of global and local features is performed, offering high flexibility and enabling dynamic adjustment of their contributions. Specifically, assuming the global feature extracted by VIE is F_VIE and the local feature extracted by LFN is F_LFN, the fusion process can be expressed as: F_enhanced = α × F_VIE + β × F_LFN, where F_enhanced is the enhanced feature, F_VIE is the global feature containing macroscopic information of the image, and F_LFN is the local feature containing detailed information of the image; α and β are weighting coefficients used to adjust the contribution ratio of global and local features, respectively. In one embodiment, the values ​​of α and β are adjusted according to the needs of the task to ensure a balance between global and local features.

[0071] In some embodiments, α and β can be manually adjusted or dynamically adjusted within the model that fuses features, enabling the model to better integrate global and local features across different tasks. This weighted fusion method ensures complementarity between global and local information, ultimately resulting in an enhanced feature representation F_enhanced, providing more accurate input data for plant classification.

[0072] In some embodiments, the preset plant classification model is determined according to the following steps: obtaining an unlabeled dataset and a labeled dataset; wherein the dataset includes multiple multispectral images; obtaining an intermediate plant classification model based on the unlabeled dataset and the initial plant classification model; and obtaining a final plant classification model based on the labeled dataset and the intermediate plant classification model.

[0073] When preprocessing multiple multispectral images, it is necessary to first obtain unlabeled and labeled datasets. Unlabeled datasets refer to multispectral images without plant species labels. A large number of images can be collected through automated means, such as obtaining a large amount of unlabeled data through multispectral images taken by drones or remote sensing equipment. Labeled data are usually multispectral images with plant species or classification labels already marked. For example, each plant image may have been labeled as "rose", "sunflower", etc.

[0074] Then, an initial plant classification model is constructed, trained using self-supervised learning with unlabeled multispectral images as input. For example... Figure 2 As shown, the model can be trained using self-supervised representation learning (SSL). Contrastive learning can maximize feature similarity. By introducing a contrastive loss function, it ensures that similar image samples are closer in the feature space, while dissimilar samples are further apart, thereby learning effective feature representations.

[0075] For example, assuming the input unlabeled data is X_unlabeled, the weights of the self-supervised representation learning network are W_ssl, and the feature representation is F_ssl, then the feature representation obtained through self-supervised representation learning can be expressed as:

[0076] F_ssl = SSL(X_unlabeled, W_ssl), where: X_unlabeled is the image data in the unlabeled dataset; W_ssl is the weight parameter of the self-supervised representation learning network, which is continuously optimized through training; F_ssl is the preliminary feature representation obtained by the self-supervised representation learning network, which contains the latent features of the image learned from the unlabeled data.

[0077] After obtaining the initial feature representation F_ssl, a contrastive loss function is introduced to calculate the distance between image samples in the feature space, ensuring that similar images are close together and dissimilar images are far apart. For example, suppose there are two images X1 and X2, which can be either positive or negative sample pairs. The contrastive loss function is L_contrastive. The distance between these two images in the feature space is calculated using the following formula:

[0078] L_contrastive = max(0, d(F_ssl(X1), F_ssl(X2)) - margin), where d(F_ssl(X1), F_ssl(X2)) is the distance metric between feature representations X1 and X2, typically using Euclidean distance or cosine similarity. Margin is a hyperparameter used to control the minimum margin between different samples. For positive sample pairs, the loss function minimizes the distance L_contrastive between them, making similar samples close in the feature space.

[0079] For negative sample pairs, the loss function maximizes the distance L_contrastive between them, but does not exceed the margin parameter, so that dissimilar samples are far apart in the feature space.

[0080] After training, an intermediate plant classification model is obtained, which fully utilizes unlabeled data to learn the general features of multispectral images. Then, the intermediate plant classification model is trained using a labeled dataset to enable it to accurately classify plant categories. The trained intermediate plant classification model is loaded, using labeled multispectral images as input. An end-to-end training method is used to train the model, enabling it to learn the classification task based on labeled data, with the goal of minimizing classification error. After training, the final plant classification model is saved.

[0081] Specifically, see Figure 2The process of training an intermediate plant classification model using a labeled dataset is called fine-tuning. Let the input labeled data be X_labeled, and the fine-tuned feature representation be F_finetune. This can be described by the following formula: F_finetune = Fine_tune(X_labeled, W_finetune), where: X_labeled is the labeled data used for fine-tuning; W_finetune is the network weights in the fine-tuning stage, typically optimized based on self-supervised representation learning; F_finetune is the fine-tuned feature representation, containing refined features specific to the task. In one embodiment, fine-tuning updates the weights using a standard loss function (such as cross-entropy loss), which can be expressed as: L_finetune = -Σ[Y × log(P)], where: Y is the true label of the labeled data; P is the predicted probability output by the model; L_finetune is the loss function in the fine-tuning stage, typically cross-entropy loss, used to minimize the difference between the predicted result and the true label. Then, the gradient (i.e., derivative) of the loss function with respect to the model parameters is calculated. Using the chain rule, the gradient of each layer is calculated layer by layer from the output layer forward, and the model weights are updated using gradient descent.

[0082] After self-supervised pre-training and fine-tuning, the resulting feature representation F_finetune can be used for final classification prediction through a fully connected layer. This optimization method, combining self-supervised representation learning (SSL) with end-to-end training, pre-trains on a large amount of unlabeled data through self-supervised learning, enabling the model to autonomously learn effective feature representations from the data. This not only reduces dependence on labeled data but also improves the training efficiency of the model on large-scale datasets. The fine-tuning stage further optimizes the model using a small amount of labeled data, enhancing its classification ability on specific tasks. The synergistic effect of this method is that high-quality feature representations are first obtained through self-supervised representation learning, and then the model is fine-tuned using a small amount of labeled data, improving classification accuracy and the model's generalization ability. In plant classification tasks, especially in multispectral image processing, there is often a lack of labeled data. By combining self-supervised learning with end-to-end training, this method can effectively utilize a large amount of unlabeled data for pre-training, thereby improving the model's accuracy and robustness, especially in real-world application scenarios with significant environmental changes.

[0083] In some embodiments, the pixel threshold is determined according to the following formula: ( ); where μ(λ) is the mean pixel value of band λ; σ(λ) is the standard deviation of pixel values ​​of band λ; ( ) is a constant, determined by the following formula: ;in, These are weighting coefficients. =1, used to control the contribution of each pixel value in band λ; >0 is the attenuation coefficient, used for adjustment. ( The size of ), where n is the number of pixel values ​​in band λ; Let be the pixel value of the i-th pixel in band λ, where i is an integer greater than 0 and less than or equal to n.

[0084] By calculating the mean μ(λ) and standard deviation σ(λ) of each band, the pixel values ​​of the band can be adaptively adjusted, and the threshold Tλ can be dynamically determined by the formula Tλ = μ(λ) + σ(λ) × ( Let μ(λ) be the mean pixel value of band λ, representing the average value of all pixels in that band; and σ(λ) be the standard deviation of pixel values ​​in band λ, representing the range of pixel value fluctuations in that band. ( From the formula Determine, where n is the number of pixel values ​​in band λ; i is the i-th pixel value in band λ; Let be the pixel value of the i-th pixel in band λ. ( The threshold sensitivity can be adjusted, with a value between [1, 2], to control the strictness of subsequent pixel selection; These are weighting coefficients. =1, used to control the contribution of each pixel value in band λ; >0 is the attenuation coefficient, used for adjustment. ( The size of ) and It can be adjusted according to actual needs; These are the pixel values ​​for band λ.

[0085] ( Specifically, when controlling the stringency of subsequent pixel selection, ( The larger the threshold Tλ is, the higher the threshold Tλ will be. This means that only pixels with values ​​significantly greater than the mean can pass the screening and are considered representative features. Therefore, only pixels that are significantly different from the background noise in this band can pass the screening, filtering out more noise and making feature selection more stringent. ( The smaller the threshold Tλ, the lower the threshold will be, and more pixels, such as potential noise and some less obvious features, will be considered valid features and thus pass the screening. Feature screening becomes more lenient, potentially retaining more noise or background interference.

[0086] In some embodiments, obtaining a final plant classification model based on the labeled dataset and the intermediate plant classification model includes: obtaining a first plant classification model based on the labeled dataset and the intermediate plant classification model; determining the loss value of each layer of the network in the first plant classification model based on a preset validation dataset; and adjusting the parameters of each layer of the network in the first plant classification model based on the sum of the loss values ​​of each layer of the network to obtain a second plant classification model.

[0087] When preprocessing multiple multispectral images, in addition to obtaining unlabeled and labeled datasets, a validation dataset is also created. Based on the labeled dataset and the intermediate plant classification model, a first plant classification model is obtained. Since this manually designed model architecture may not be optimal for complex tasks, such as multispectral image classification, and the first plant classification model may perform well on specific datasets but lack generalization ability on other datasets or tasks, further optimization is needed.

[0088] In some embodiments, the optimal network architecture is automatically discovered through an automated search algorithm, thereby obtaining an optimized plant classification model without relying on manual design.

[0089] During model optimization, automated search algorithms automatically discover the optimal network architecture. The core principle is to minimize the overall network loss function by optimizing the structure of each layer. In this process, the loss function of each layer is minimized to obtain the optimal combination of network architectures. See also... Figure 2 In some cases, automated search algorithms can be network architecture search. When performing network architecture search, a separate validation dataset is typically used to evaluate the performance of different network architectures. This data is not used in the training process but is used to validate the effectiveness of each possible network structure in order to select the final model configuration. Specifically, based on a pre-defined validation dataset, the loss values ​​of each layer in the first plant classification model can be determined; based on the sum of the loss values ​​of each layer, the parameters of each layer in the first plant classification model are adjusted to obtain the second plant classification model.

[0090] For example, suppose the model network structure has multiple layers, L1, L2, ..., Ln, each with corresponding convolutional kernel sizes K1, K2, ..., Kn and strides S1, S2, ..., Sn. The overall loss function is defined as: min(L_total) = Σ (L1 + L2 + ... + Ln), where L_total is the overall loss, which is the sum of the loss functions of all layers; Ln is the loss function of the nth layer, typically including classification loss, regularization loss, etc.; and min is the function that minimizes the total loss. Then, an architecture search algorithm, such as reinforcement learning or a genetic algorithm, is used to search for the optimal network structure, and the kernel size K and stride S of each layer are adjusted to obtain the optimal network configuration, resulting in the second plant classification model.

[0091] Since the optimal network architecture may be optimized primarily for feature extraction at a single scale, multi-scale fusion is still needed to further enhance the model's ability to express diverse features, thereby improving the model's performance. In some cases, the optimal network architecture is obtained to extract image features at different scales, and then feature fusion is performed. Feature fusion can be achieved by concatenating the features. The fused features are then used to train the classification model to further optimize the network structure and ensure that the classification model can fully utilize information from different scales of the image.

[0092] For example, the optimized classification model includes convolutional layers and convolutional kernels. Suppose the optimized classification model extracts feature representations of different scales, F_scale1, F_scale2, ..., F_scalen, through convolutional layers of multiple scales, resulting in the final fused feature: F_fused = concat(F_scale1, F_scale2, ..., F_scalen). Here, F_scale represents the features extracted at different scales. For instance, F_scale1 might be obtained from a larger convolutional kernel, capturing broad image information such as overall shape; while F_scale2 might use a smaller convolutional kernel, focusing on finer details such as texture and edges, and so on. These features are extracted from the multispectral data of the same input image, but each layer captures image information with different granularity and perspective. concat is the concatenation function that joins feature vectors from multiple scales together. F_fused is the fused multi-scale feature, containing information from different scales, and can more comprehensively represent the details of the input image. Using F_fused, the optimized classification model is further trained, resulting in the aforementioned pre-defined plant classification model.

[0093] By dynamically adjusting hyperparameters such as network structure, kernel size, and stride, feature learning is performed at multiple scales. Optimization of the network architecture search improves the model's classification ability at different scales. Through joint optimization of multi-scale features, the model can adapt to diverse features in plant images, significantly improving classification performance. Optimization of the multi-scale network architecture ensures the model can handle features at different resolutions and scales, making it more adaptable to multispectral images. During feature fusion, feature information from multiple scales is effectively combined, further improving classification accuracy.

[0094] In some embodiments, the method further includes: determining a sparsification loss value for the first plant classification model based on a preset validation dataset; wherein, adjusting the parameters of each layer of the first plant classification model based on the sum of the loss functions of each layer to obtain a second plant classification model includes: adjusting the parameters of each layer of the first plant classification model based on the sparsification loss value and the sum of the loss functions of each layer to obtain a second plant classification model.

[0095] After determining the optimal network architecture, convolutional layers contain redundant parameters, resulting in high computational cost. To further improve the model's computational efficiency, convolutional layer sparsification techniques can be used to optimize the model. Sparsification adjusts the convolutional layers based on the established network architecture, ensuring that computational cost is reduced without compromising the network's feature extraction capabilities. The main goal is to reduce the number of non-zero weights in the model, thereby reducing model complexity and improving computational efficiency. The specific process is as follows: First, a sparsification objective is set, typically achieved by adding a sparsity penalty term to the loss function. The most common method for choosing a penalty term is L1 regularization, which encourages more weights to become zero by penalizing the absolute value of the weights, thus achieving a sparsity effect. Then, the loss function is integrated, considering the sparsity loss along with other losses such as classification loss in the model's total loss function. This aims to simultaneously optimize the model's predictive performance and sparsity. Next, the weights are updated, with each layer's weight update considering both improving classification accuracy and increasing weight sparsity. This means that weight updates must both reduce prediction error and push weights towards zero. During training, the sparsity strength parameters (such as the L1 regularization coefficient) need to be continuously adjusted to find the optimal balance between performance and sparsity. Simultaneously, the model's performance needs to be monitored to ensure that sparsity does not negatively impact the model's accuracy.

[0096] Specifically, assuming the weights of the nth convolutional kernel are Wn, the sparsification loss function can be expressed as: L_sparse = λ × Σ |Wn|, where L_sparse is the sparsification loss, which achieves sparsity by penalizing the L1 norm of the convolutional kernel weights; λ is a hyperparameter controlling the degree of sparsification; and Σ |Wn| represents the L1 norm of the convolutional kernel weights, indicating the sum of the absolute values ​​of all weights. Weight updates can be performed via gradient descent, and L1 regularization causes some weights to approach 0, thus achieving model sparsity and obtaining the second plant classification model. By introducing convolutional layer sparsification technology, computational efficiency and real-time detection capabilities are further optimized, significantly improving the processing efficiency and real-time performance of the second plant classification model.

[0097] In some examples, the preset plant classification model includes a fully connected layer; determining the target plant category included in the multispectral image to be detected based on the preset plant classification model, the local features, and the global features includes: obtaining an input feature vector based on the local features and the global features; inputting the input feature vector into the preset plant classification model to obtain multiple probability values ​​corresponding one-to-one with multiple plant categories; determining the target plant category included in the multispectral image to be detected based on the multiple probability values; wherein, the probability value of plant category i is determined according to the following formula:

[0098] ;in, Let i be the probability value for plant category i; Let be the element in the i-th row and j-th column of the weight matrix of the fully connected layer; C is the total number of rows in the weight matrix of the fully connected layer; D is the total number of columns in the weight matrix of the fully connected layer; i is an integer greater than 0 and less than or equal to C; j is an integer greater than 0 and less than or equal to D; k is an integer greater than 0 and less than or equal to C. It is the j-th element of the input feature vector; is the element in the k-th row and j-th column of the weight matrix of the fully connected layer.

[0099] like Figure 2 As shown, based on local and global features, a fine-tuned feature vector F_finetune is obtained. This vector is then used for final classification prediction through a fully connected layer of a pre-defined plant classification model. Based on the local and global features, an input feature vector is obtained. This input feature vector is then fed into the pre-defined plant classification model to obtain multiple probability values ​​corresponding to various plant categories. The classification formula is: The pre-defined plant classification model includes fully connected layers. The role of these layers is to map the finely tuned feature representation F_finetune to the category space. The weight matrix of the fully connected layer has a shape of C×D, where C represents the number of plant categories (1 ≤ k ≤ C), and D is the feature dimension, i.e., the length of the input vector F_finetune. The D-dimensional features are mapped to a C-dimensional space (one dimension per category), and the matching degree between the features and the weights of each category is calculated using dot products.

[0100] This refers to the element in the i-th row and j-th column of the weight matrix of the fully connected layer. The j-th element of the feature vector output by the intermediate plant classification model. Let be the probability value of the i-th element, corresponding to the probability value of the i-th category. The total number of plant categories is equal to the total number of rows in the weight matrix. The input feature vector is of the form (A1, A2, ..., AD), containing D elements. The numerator is the exponent of the score of the current category, which is the linear combination of the features and weights. The denominator is a normalization term, representing the sum of the exponents of the scores of all categories, ensuring that the sum of the probabilities of all categories is 1, ultimately yielding the probability value. .

[0101] The input feature vector is fed into a pre-defined plant classification model to obtain multiple probability values ​​corresponding to various plant categories. Based on these probability values, the target plant category included in the multispectral image to be detected is determined. The classification results of the multispectral image classification are category labels and probability distributions. The model outputs a category label for each sample, representing the classification result of the input image. The probability distribution consists of multiple probability values ​​corresponding to various plant categories, representing the confidence level of the classification result. For example, the category label might be "forest," "grassland," or "farmland." The output of the probability distribution might be [0.1, 0.7, 0.2], representing probabilities of belonging to the three categories of 10%, 70%, and 20%, respectively. Based on the results, the image can be determined to belong to the category with the highest probability, i.e., grassland. Further analysis and visualization of the results are possible. During the preprocessing of multiple multispectral images, a test dataset is also created to evaluate the model's performance. Indicators such as accuracy, recall, and F1 score are calculated to verify the model's performance and apply it to real-world scenarios.

[0102] Secondly, embodiments of the present invention provide a plant category determination device based on multispectral images, which facilitates the improvement of plant classification accuracy.

[0103] See Figure 3This invention provides a plant category determination device based on multispectral images, comprising: a first determining unit 31, configured to determine a pixel threshold for each spectral band in a multispectral image to be detected; a filtering unit 32, configured to filter pixels in each spectral band of the multispectral image to be detected according to the pixel threshold for each spectral band, to obtain a filtered multispectral image; a second determining unit 33, configured to determine local features and global features of the filtered multispectral image according to the filtered multispectral image; and a third determining unit 34, configured to determine the target plant category included in the multispectral image to be detected according to a preset plant classification model, the local features, and the global features.

[0104] The plant category determination device based on multispectral images provided in the embodiments of the present invention can determine the pixel threshold of each spectral band in a multispectral image to be detected; filter the pixels of each spectral band in the multispectral image to be detected according to the pixel threshold of each spectral band to obtain a filtered multispectral image; determine the local features and global features of the filtered multispectral image according to the filtered multispectral image; and determine the target plant category included in the multispectral image to be detected according to a preset plant classification model, the local features, and the global features. Thus, filtering the pixels of each spectral band in the multispectral image to be detected can remove noise, highlight the target area, and improve image quality. Determining the local and global features of the filtered multispectral image can effectively extract global and local features, facilitating the improvement of plant category determination accuracy.

[0105] In some examples, the first determining unit includes: a first determining module, which determines the pixel threshold of the corresponding spectral band based on the mean and standard deviation of each spectral band in the multispectral image to be detected.

[0106] In some examples, the filtering unit includes: a comparison module for comparing the pixel value of each pixel in each spectral band of the multispectral image to be detected with the pixel threshold of the corresponding spectral band; a removal module for removing pixels whose pixel value in each pixel of each spectral band is less than the pixel threshold of the corresponding spectral band; and a first obtaining module for obtaining the filtered multispectral image based on the pixels in each pixel of each spectral band excluding the removed pixels.

[0107] In some examples, the second determining unit includes: a first input module for inputting the filtered multispectral image into a dimensionality reduction encoder to obtain a low-dimensional multispectral image; a second input module for inputting the low-dimensional multispectral image into a reconstruction decoder to obtain a new multispectral image; and a second determining module for determining the local features and global features of the new multispectral image based on the new multispectral image.

[0108] In some examples, the second determining module further includes: a first extraction submodule, used to extract global context features of the new multispectral image according to a preset visual information enhancement module; and a second extraction submodule, used to extract local features of the new multispectral image according to a preset local feature hierarchical network.

[0109] In some examples, the third determining unit includes: a fusion module for fusing the global context features and the local features to obtain enhanced features; and a third determining module for determining the target plant category included in the multispectral image to be detected based on a preset plant classification model and the enhanced features.

[0110] In some examples, the third determining unit is specifically used to: acquire an unlabeled dataset and a labeled dataset; wherein the dataset includes multiple multispectral images; obtain an intermediate plant classification model based on the unlabeled dataset and the initial plant classification model; and obtain a final plant classification model based on the labeled dataset and the intermediate plant classification model.

[0111] In some examples, the third determining unit is further configured to: obtain a first plant classification model based on the labeled dataset and the intermediate plant classification model; determine the loss value of each layer of the network in the first plant classification model based on a preset validation dataset; and adjust the parameters of each layer of the network in the first plant classification model based on the sum of the loss values ​​of each layer of the network to obtain a second plant classification model.

[0112] In some examples, the third determining unit is further configured to: determine the sparsification loss value of the first plant classification model based on a preset validation dataset; and adjust the parameters of each layer of the first plant classification model according to the sum of the loss functions of each layer of the network to obtain a second plant classification model, including: adjusting the parameters of each layer of the first plant classification model according to the sparsification loss value and the sum of the loss functions of each layer of the network to obtain a second plant classification model.

[0113] In some examples, the pixel threshold is determined according to the following formula: ( ); where μ(λ) is the mean pixel value of band λ; σ(λ) is the standard deviation of pixel values ​​of band λ; ( It is determined by the following formula: ;in, These are weighting coefficients. =1, used to control the contribution of each pixel value in band λ; >0 is the attenuation coefficient, used for adjustment. ( The size of ); n is the number of pixel values ​​in band λ; Let be the pixel value of the i-th pixel in band λ, where i is an integer greater than 0 and less than or equal to n.

[0114] In some examples, the preset plant classification model includes a fully connected layer; the third determining unit includes: a second obtaining module, used to obtain an input feature vector based on the local features and the global features; a third input module, used to input the input feature vector into the preset plant classification model to obtain multiple probability values ​​corresponding one-to-one with multiple plant categories; and a fourth determining module, used to determine the target plant category included in the multispectral image to be detected based on the multiple probability values; wherein, the probability value of plant category i is determined according to the following formula: ;in, Let i be the probability value for plant category i; Let be the element in the i-th row and j-th column of the weight matrix of the fully connected layer; C is the total number of rows in the weight matrix of the fully connected layer; D is the total number of columns in the weight matrix of the fully connected layer; i is an integer greater than 0 and less than or equal to C; j is an integer greater than 0 and less than or equal to D; k is an integer greater than 0 and less than or equal to C. It is the j-th element of the input feature vector; is the element in the k-th row and j-th column of the weight matrix of the fully connected layer.

[0115] Thirdly, embodiments of this application also provide an electronic device that facilitates improved accuracy in plant classification.

[0116] like Figure 4The electronic device provided in the embodiments of the present invention may include: a housing 51, a processor 52, a memory 53, a circuit board 54, and a power supply circuit 55, wherein the circuit board 54 is disposed inside the space enclosed by the housing 51, and the processor 52 and the memory 53 are disposed on the circuit board 54; the power supply circuit 55 is used to supply power to various circuits or devices of the above-mentioned electronic device; the memory 53 is used to store executable program code; the processor 52 runs a program corresponding to the executable program code by reading the executable program code stored in the memory 53, for executing the process migration method provided in any of the foregoing embodiments.

[0117] For details on the specific execution process of the above steps by the processor 52 and the steps further executed by the processor 52 by running executable program code, please refer to the description of the foregoing embodiments, which will not be repeated here.

[0118] Fourthly, embodiments of this application provide a computer-readable storage medium storing one or more computer programs. When the one or more computer programs are executed by one or more processors, they implement any of the plant category determination methods based on multispectral images described in the foregoing embodiments, thus achieving the corresponding technical effects. This has been described in detail above and will not be repeated here.

[0119] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0120] The various embodiments in this specification are described in a related manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0121] In particular, the device embodiment is basically similar to the method embodiment, so the description is relatively simple. For relevant details, please refer to the description of the method embodiment.

[0122] For ease of description, the above apparatus is described by dividing it into various functional units / modules. Of course, in implementing this invention, the functions of each unit / module can be implemented in one or more software and / or hardware.

[0123] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0124] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for determining plant categories based on multispectral images, characterized in that, This includes determining the pixel threshold for each spectral band in the multispectral image to be detected; Based on the pixel threshold of each spectral band, the pixels of each spectral band in the multispectral image to be detected are filtered to obtain the filtered multispectral image. Based on the filtered multispectral images, determine the local and global features of the filtered multispectral images; Based on the preset plant classification model, the local features, and the global features, the target plant categories included in the multispectral image to be detected are determined; in, The pixel threshold is determined according to the following formula: ( ); where μ(λ) is the mean pixel value of band λ; σ(λ) is the standard deviation of pixel values ​​of band λ; ( It is determined by the following formula: ;in, These are weighting coefficients. =1, used to control the contribution of each pixel value in band λ; >0 is the attenuation coefficient, used for adjustment. ( The size of ); n is the number of pixel values ​​in band λ; Let be the pixel value of the i-th pixel in band λ, where i is an integer greater than 0 and less than or equal to n; The preset plant classification model is determined according to the following steps: obtaining an unlabeled dataset and a labeled dataset; wherein the dataset includes multiple multispectral images; and obtaining an intermediate plant classification model based on the unlabeled dataset and the initial plant classification model. Based on the labeled dataset and the intermediate plant classification model, a first plant classification model is obtained; Based on the preset validation dataset, determine the loss value of each layer of the network in the first plant classification model, and determine the sparsification loss value of the first plant classification model. Based on the sparsification loss value and the sum of the loss functions of each layer of the network, the parameters of each layer of the first plant classification model are adjusted to obtain the second plant classification model.

2. The method according to claim 1, characterized in that, The step of filtering pixels in each spectral band of the multispectral image to be detected based on pixel thresholds for each spectral band to obtain a filtered multispectral image includes: The pixel value of each pixel in each spectral band of the multispectral image to be detected is compared with the pixel threshold of the corresponding spectral band. Remove pixels whose pixel values ​​in each pixel of each spectral band are less than the pixel threshold of the corresponding spectral band; The filtered multispectral image is obtained by considering all pixels in each spectral band except for those that have been removed.

3. The method according to claim 1, characterized in that, The step of determining the local and global features of the selected multispectral images based on the selected multispectral images includes: The filtered multispectral image is input into the dimension reduction encoder to obtain a low-dimensional multispectral image; The low-dimensional multispectral image is input into the reconstruction decoder to obtain a new multispectral image; Based on the new multispectral image, determine the local and global features of the new multispectral image.

4. The method according to claim 3, characterized in that, The step of determining the local and global features of the new multispectral image based on the new multispectral image includes: Based on the preset visual information enhancement module, the global context features of the new multispectral image are extracted; Based on a pre-defined local feature hierarchical network, local features of the new multispectral image are extracted.

5. The method according to claim 1, characterized in that, The preset plant classification model includes a fully connected layer; determining the target plant category in the multispectral image to be detected based on the preset plant classification model, the local features, and the global features includes: Based on the local features and the global features, the input feature vector is obtained; Input the feature vector into the preset plant classification model to obtain multiple probability values ​​that correspond one-to-one with multiple plant categories; Based on multiple probability values, the target plant category included in the multispectral image to be detected is determined; The probability value of plant category h is determined according to the following formula: ; in, This represents the probability value for plant category h. The element in the h-th row and j-th column of the weight matrix of the fully connected layer; C is the total number of rows in the weight matrix of the fully connected layer; D is the total number of columns in the weight matrix of the fully connected layer; h is an integer greater than 0 and less than or equal to C; j is an integer greater than 0 and less than or equal to D; k is an integer greater than 0 and less than or equal to C. It is the j-th element of the input feature vector; is the element in the k-th row and j-th column of the weight matrix of the fully connected layer.

Citation Information

Patent Citations

  • Hyperspectral remote sensing image classification method based on double attention mechanism

    CN113011499A

  • Low-complexity interference identification method based on deep learning

    CN115563485A

  • Hyperspectral image classification method based on spectrum enhanced cyclic consistency Transform

    CN118781384A

  • Mask Transform and contrast learning-based hyperspectral image classification method

    CN119339131A

  • Visual inspection method and system for automobile die surface machining cracks

    CN119693366A