A method and system for identifying imperfect grain kernels

By using the TriMamba three-branch structure of the ConvNeXt framework to extract and fuse features from visible light images and reflectance spectral data of grain grains, the problem of data quality degradation under the influence of external environmental factors is solved, and high-precision and robust grain grain identification is achieved.

CN122115968APending Publication Date: 2026-05-29HENAN UNIVERSITY OF TECHNOLOGY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HENAN UNIVERSITY OF TECHNOLOGY
Filing Date
2026-02-13
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing technologies, external environmental factors can lead to a decline in the quality of grain grain data collection, resulting in low detection stability and accuracy.

Method used

We adopt the TriMamba three-branch structure based on the ConvNeXt framework, combine visible light images and reflectance spectral data, and perform feature extraction and fusion through spectral Mamba branch, spatial Mamba branch and guiding branch. We use the cross-entropy loss function to optimize model training, and achieve deep collaboration and efficient fusion of spectral and spatial features.

Benefits of technology

It improves the accuracy and robustness of grain identification, maintains high-efficiency identification performance under complex backgrounds and varying lighting conditions, and enhances the stability and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115968A_ABST
    Figure CN122115968A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of grain informatization processing, and particularly relates to a kind of imperfect grain kernel identification method and system.Visible light image and reflectance spectrum data of grain kernel are input into identification model;From the features extracted from the visible light image and reflectance spectrum data of grain kernel respectively by the feature extraction module, the input features are obtained;The spectral Mamba branch in the identification model sequentially performs linear layer and convolution processing on the normalized input features, and then inputs the state space model to obtain the output;The spatial Mamba branch performs dimension transposition on the normalized input features, and then sequentially performs linear layer and convolution processing, and then inputs the state space model to obtain the output;The guide branch sequentially performs linear layer processing on the normalized input features, and then inputs the ConvNeXt classifier, and then adjusts the output of the ConvNeXt classifier combined with the SiLU activation function to obtain the output;According to the fusion result of the output of the guide branch, the spectral Mamba branch, the spatial Mamba branch and the input features, the identification result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of grain information processing technology, specifically relating to an imperfect method and system for identifying grain kernels. Background Technology

[0002] With the continuous development of sensing technology, imaging technology, and intelligent vision algorithms, image recognition has demonstrated enormous application potential in the field of grain quality inspection. Especially in the identification of imperfect grain grains, computer vision-based intelligent detection methods are gradually replacing traditional manual screening methods. By acquiring multimodal data of grain grains through hyperspectral and visible light imaging systems, and combining preprocessing techniques such as filtering, denoising, and image enhancement, the clarity and feature discrimination of images can be significantly improved. Based on this, using deep learning models for high-dimensional feature extraction and intelligent recognition of grain images enables automated identification and accurate judgment of imperfect grains such as damaged, insect-eaten, and moldy grains. This method not only effectively improves the efficiency and accuracy of detection but also enhances the robustness of the system under complex lighting conditions and sample differences.

[0003] Compared to traditional grain image acquisition systems, the introduction of intelligent sensing and intelligent detection technologies has enabled a higher degree of integration and automation in the detection process. Such systems typically use acquisition software to uniformly manage and control imaging equipment, achieving efficient acquisition and real-time storage of grain images. Based on the acquired data, the system can further perform tasks such as image processing, feature extraction, and quality assessment, thereby constructing an intelligent detection system that integrates acquisition, analysis, and detection.

[0004] During grain grain image acquisition, factors such as changes in ambient lighting, equipment disturbances, and calibration errors often lead to a decline in data quality, affecting the stability and automation level of the detection system. Traditional quality assessment methods mostly rely on manual experience or semi-automated detection, which suffers from low efficiency, high subjectivity, and insufficient accuracy, making it difficult to meet the needs of large-scale detection.

[0005] Chinese invention patent application CN120563928A discloses a method for detecting abnormal appearance of single wheat grains based on a deep learning model and hyperspectral imaging. The method includes: Step 1, collecting wheat grain samples of different categories; Step 2, acquiring hyperspectral image data of wheat grains using visible-near-infrared and short-wave near-infrared hyperspectral imaging systems, and extracting spectral and image information; Step 3, preprocessing the spectral data and verifying the preprocessing effect; Step 4, screening spectral feature bands, extracting texture and morphological features using a gray-level co-occurrence matrix, and constructing a mid-level data fusion model; Step 5, constructing a deep learning model for spectral feature fusion to achieve high-level fusion of spectral and image features; Step 6, performing pixel-level classification of the hyperspectral images to generate a spatial distribution visualization of wheat grain appearance abnormalities. This invention possesses high precision, non-destructive testing capabilities, and strong generalization ability, and is suitable for wheat quality grading, processing sorting, and warehouse safety management. Summary of the Invention

[0006] The purpose of this invention is to provide an imperfect method and system for identifying grains, which solves the problem in the prior art where the quality of collected data is reduced due to the influence of external environmental factors, resulting in low stability and accuracy of detection.

[0007] To achieve the above objectives, the present invention provides an imperfect method for identifying grain kernels, comprising:

[0008] The visible light image and reflectance spectrum data of the grain are input into the recognition model; the feature extraction module in the recognition model obtains the input features based on the features extracted from the visible light image and reflectance spectrum data of the grain respectively;

[0009] The feature aggregation module in the recognition model includes a spectral Mamba branch, a spatial Mamba branch, and a guiding branch. The spectral Mamba branch processes the normalized input features sequentially through linear layers and convolutions before inputting them into the state space model to obtain the branch's output. The spatial Mamba branch transposes the normalized input features, processes them sequentially through linear layers and convolutions, and then inputs them into the state space model to obtain the branch's output. The guiding branch processes the normalized input features sequentially through linear layers, inputs them into the ConvNeXt classifier, and adjusts the output of the ConvNeXt classifier using the SiLU activation function to obtain the branch's output.

[0010] The recognition result is obtained by fusing the output of the guiding branch, the output of the spectral Mamba branch, the output of the spatial Mamba branch, and the input features; the recognition model is trained using grain sample images.

[0011] Furthermore, methods for extracting features from visible light images of grains include:

[0012] The visible light image of grain grains is downsized using a point-based convolutional layer. The downsized image is then decomposed into different frequency scales using a two-dimensional curvelet transform. The resulting frequency band sub-components are processed by the first convolutional layer and then restored to their original spatial dimensions using a two-dimensional inverse curvelet transform, resulting in a curvelet transform feature map. Features extracted from the visible light image of grain grains are obtained based on the curvelet transform feature map.

[0013] Methods for extracting features from visible light images of grains based on visible light curve transform feature maps include:

[0014] The visible light curve transform feature map is fused with the processing result of the image after the size reduction is processed by the second convolutional layer. Based on the fusion result, the features extracted from the visible light image of the grain are obtained.

[0015] Furthermore, methods for extracting features from the reflectance spectral data of grain grains include:

[0016] The reflectance spectral data of grain grains is subjected to rotational convolution along the spectral dimension for dimensionality reduction. Then, a one-dimensional discrete curvelet transform is used to decompose the dimensionality-reduced spectral features into low-frequency and high-frequency components. Convolution operations are then performed on the low-frequency and high-frequency components separately through channel-dimensional convolution. Finally, a one-dimensional inverse curvelet transform is used to reintegrate the features back into the original spectral dimension, resulting in hyperspectral curvelet transform features. Based on the hyperspectral curvelet transform features, the features extracted from the reflectance spectral data of grain grains are obtained.

[0017] Methods for extracting features from the reflectance spectrum data of grain grains based on hyperspectral curve transform characteristics include:

[0018] The hyperspectral curve transform features and the dimensionality-reduced spectral features are fused together after being processed by the third convolutional layer. Based on the fusion result, the features extracted from the reflectance spectral data of grain grains are obtained.

[0019] Furthermore, the loss value during the training of the recognition model is obtained by fusing the loss corresponding to the real label of the grain sample image and the label predicted by the recognition model, the loss corresponding to the real label of the grain sample image and the label obtained according to the guide branch output, the loss corresponding to the real label of the grain sample image and the label obtained according to the spectral Mamba branch output, and the loss corresponding to the real label of the grain sample image and the label obtained according to the spatial Mamba branch output.

[0020] Furthermore, the losses corresponding to the true labels of grain sample images and the labels predicted by the recognition model, the losses corresponding to the true labels of grain sample images and the labels obtained according to the guide branch output, the losses corresponding to the true labels of grain sample images and the labels obtained according to the spectral Mamba branch output, and the losses corresponding to the true labels of grain sample images and the labels obtained according to the spatial Mamba branch output are all calculated using cross-entropy loss.

[0021] Furthermore, after the recognition model is trained, it is evaluated through the performance evaluation module. If the evaluation result is unsatisfactory, the recognition model is retrained.

[0022] The evaluation methods for the performance evaluation module include:

[0023] The recognition results of the recognition model are evaluated using accuracy, precision, recall, mean precision, and Kappa coefficient as indicators. If the recognition results of the recognition model fail to meet the corresponding indicator thresholds, the evaluation results are deemed unqualified.

[0024] Recall rate is obtained from true positives, true negatives, false positives, and false negatives; precision rate is obtained from true positives and false positives; recall rate is obtained from true positives and false negatives; mean precision is calculated from precision rate and recall rate.

[0025] A true positive is the number of pixels in the sample image where both the label and the recognition result are grain regions; a true negative is the number of pixels in the sample image where both the label and the recognition result are background regions; a false positive is the number of pixels in the sample image where the label is background regions but the recognition result is grain regions; and a false negative is the number of pixels in the sample image where the label is grain regions but the recognition result is background regions.

[0026] Furthermore, the normalization of input features can be achieved by expanding the three-dimensional input features, which include feature map height, feature map width, and channel dimension, into two-dimensional input features, which include channel dimension and batch dimension.

[0027] Furthermore, based on the fusion results of the guiding branch output, spectral Mamba branch output, spatial Mamba branch output, and input features, the recognition results can be obtained in the following ways:

[0028] The sum of the products of the guiding branch output and the spectral Mamba branch output and the spatial Mamba branch output is processed through a linear layer. The processed result is then added to the input features and input into the classifier of the recognition model to obtain the recognition result.

[0029] Furthermore, the visible light image of the grain is obtained by performing preprocessing on the original image of the grain acquired by the visible light camera, including denoising, filtering, enhancement and normalization.

[0030] The reflectance spectral data of grain grains are obtained by scanning the reflectance spectral data cube of grain grain samples line by line in the visible and near-infrared bands. The data is then processed by whiteboard correction to eliminate systematic errors caused by differences in ambient light and instrument response, spectral smoothing algorithm to reduce high-frequency noise, and normalization or standardization of each band.

[0031] The above-described technical solution of the present invention provides a novel method for identifying imperfect grain kernels, the beneficial effects of which include:

[0032] Based on the recognition model obtained by inputting both visible light and hyperspectral data (i.e., reflectance spectral data) of grain grains into the recognition model, a hierarchical interaction and dynamic fusion of spectral features, spatial features, and their joint representations is achieved through a cross-fusion structure of the TriMamba three branches (spectral Mamba branch, spatial Mamba branch, and guiding branch) guided by the ConvNeXt framework. This structure fully leverages the advantages of the Mamba state-space model in long-range dependency modeling, realizing deep collaboration and efficient fusion of spectral and spatial features. As a result, the average recognition accuracy is significantly improved, while maintaining excellent robustness and generalization performance under complex backgrounds and varying illumination conditions.

[0033] This invention also provides an imperfect grain identification system, including a processor storing executable program instructions. These instructions are used to execute the following imperfect grain identification method, specifically including:

[0034] The visible light image and reflectance spectrum data of the grain are input into the recognition model; the feature extraction module in the recognition model obtains the input features based on the features extracted from the visible light image and reflectance spectrum data of the grain respectively;

[0035] The feature aggregation module in the recognition model includes a spectral Mamba branch, a spatial Mamba branch, and a guiding branch. The spectral Mamba branch processes the normalized input features sequentially through linear layers and convolutions before inputting them into the state space model to obtain the branch's output. The spatial Mamba branch transposes the normalized input features, processes them sequentially through linear layers and convolutions, and then inputs them into the state space model to obtain the branch's output. The guiding branch processes the normalized input features sequentially through linear layers, inputs them into the ConvNeXt classifier, and adjusts the output of the ConvNeXt classifier using the SiLU activation function to obtain the branch's output.

[0036] The recognition result is obtained by fusing the output of the guiding branch, the output of the spectral Mamba branch, the output of the spatial Mamba branch, and the input features; the recognition model is trained using grain sample images.

[0037] Furthermore, methods for extracting features from visible light images of grains include:

[0038] The visible light image of grain grains is downsized using a point-based convolutional layer. The downsized image is then decomposed into different frequency scales using a two-dimensional curvelet transform. The resulting frequency band sub-components are processed by the first convolutional layer and then restored to their original spatial dimensions using a two-dimensional inverse curvelet transform, resulting in a curvelet transform feature map. Features extracted from the visible light image of grain grains are obtained based on the curvelet transform feature map.

[0039] Methods for extracting features from visible light images of grains based on visible light curve transform feature maps include:

[0040] The visible light curve transform feature map is fused with the processing result of the image after the size reduction is processed by the second convolutional layer. Based on the fusion result, the features extracted from the visible light image of the grain are obtained.

[0041] Furthermore, methods for extracting features from the reflectance spectral data of grain grains include:

[0042] The reflectance spectral data of grain grains is subjected to rotational convolution along the spectral dimension for dimensionality reduction. Then, a one-dimensional discrete curvelet transform is used to decompose the dimensionality-reduced spectral features into low-frequency and high-frequency components. Convolution operations are then performed on the low-frequency and high-frequency components separately through channel-dimensional convolution. Finally, a one-dimensional inverse curvelet transform is used to reintegrate the features back into the original spectral dimension, resulting in hyperspectral curvelet transform features. Based on the hyperspectral curvelet transform features, the features extracted from the reflectance spectral data of grain grains are obtained.

[0043] Methods for extracting features from the reflectance spectrum data of grain grains based on hyperspectral curve transform characteristics include:

[0044] The hyperspectral curve transform features and the dimensionality-reduced spectral features are fused together after being processed by the third convolutional layer. Based on the fusion result, the features extracted from the reflectance spectral data of grain grains are obtained.

[0045] Furthermore, the loss value during the training of the recognition model is obtained by fusing the loss corresponding to the real label of the grain sample image and the label predicted by the recognition model, the loss corresponding to the real label of the grain sample image and the label obtained according to the guide branch output, the loss corresponding to the real label of the grain sample image and the label obtained according to the spectral Mamba branch output, and the loss corresponding to the real label of the grain sample image and the label obtained according to the spatial Mamba branch output.

[0046] Furthermore, the losses corresponding to the true labels of grain sample images and the labels predicted by the recognition model, the losses corresponding to the true labels of grain sample images and the labels obtained according to the guide branch output, the losses corresponding to the true labels of grain sample images and the labels obtained according to the spectral Mamba branch output, and the losses corresponding to the true labels of grain sample images and the labels obtained according to the spatial Mamba branch output are all calculated using cross-entropy loss.

[0047] Furthermore, after the recognition model is trained, it is evaluated through the performance evaluation module. If the evaluation result is unsatisfactory, the recognition model is retrained.

[0048] The evaluation methods for the performance evaluation module include:

[0049] The recognition results of the recognition model are evaluated using accuracy, precision, recall, mean precision, and Kappa coefficient as indicators. If the recognition results of the recognition model fail to meet the corresponding indicator thresholds, the evaluation results are deemed unqualified.

[0050] Recall rate is obtained from true positives, true negatives, false positives, and false negatives; precision rate is obtained from true positives and false positives; recall rate is obtained from true positives and false negatives; mean precision is calculated from precision rate and recall rate.

[0051] A true positive is the number of pixels in the sample image where both the label and the recognition result are grain regions; a true negative is the number of pixels in the sample image where both the label and the recognition result are background regions; a false positive is the number of pixels in the sample image where the label is background regions but the recognition result is grain regions; and a false negative is the number of pixels in the sample image where the label is grain regions but the recognition result is background regions.

[0052] Furthermore, the normalization of input features can be achieved by expanding the three-dimensional input features, which include feature map height, feature map width, and channel dimension, into two-dimensional input features, which include channel dimension and batch dimension.

[0053] Furthermore, based on the fusion results of the guiding branch output, spectral Mamba branch output, spatial Mamba branch output, and input features, the recognition results can be obtained in the following ways:

[0054] The sum of the products of the guiding branch output and the spectral Mamba branch output and the spatial Mamba branch output is processed through a linear layer. The processed result is then added to the input features and input into the classifier of the recognition model to obtain the recognition result.

[0055] Furthermore, the visible light image of the grain is obtained by performing preprocessing on the original image of the grain acquired by the visible light camera, including denoising, filtering, enhancement and normalization.

[0056] The reflectance spectral data of grain grains are obtained by scanning the reflectance spectral data cube of grain grain samples line by line in the visible and near-infrared bands. The data is then processed by whiteboard correction to eliminate systematic errors caused by differences in ambient light and instrument response, spectral smoothing algorithm to reduce high-frequency noise, and normalization or standardization of each band.

[0057] The technical solution of the imperfect grain identification system described above in this invention can achieve the same beneficial effects as the imperfect grain identification method described above. Attached Figure Description

[0058] Figure 1 This is a schematic diagram of the principle of the identification model of the grain identification method in the imperfect implementation method of the present invention.

[0059] Figure 2 This is a structural example diagram of the feature aggregation module of the identification model in an embodiment of the imperfect grain identification method of the present invention;

[0060] Figure 3 This is an example diagram of the architecture of the system for acquiring the original image of grain and the reflectance spectral data of grain required for the grain identification method in the imperfect implementation of the present invention.

[0061] Figure 4 This is a flowchart of the grain identification method in an imperfect implementation of the present invention.

[0062] Figure 5 This is a schematic diagram illustrating the principle of extracting features from visible light images of grain grains in an imperfect implementation of the grain grain identification method of the present invention.

[0063] Figure 6 This is a schematic diagram illustrating the principle of extracting features from the reflectance spectral data of grain grains in an imperfect implementation of the grain grain identification method of the present invention. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0065] Imperfect implementation method for grain identification

[0066] This embodiment presents a technical solution for identifying imperfect grain kernels. By using multimodal data fusion and multi-scale feature modeling, it achieves high-precision identification of imperfect corn kernels.

[0067] Reference Figure 1 The method specifically includes:

[0068] The visible light image and reflectance spectrum data of the grain are input into the recognition model; the feature extraction module in the recognition model obtains the input features based on the features extracted from the visible light image and reflectance spectrum data of the grain respectively;

[0069] The feature aggregation module in the recognition model includes a spectral Mamba branch, a spatial Mamba branch, and a guiding branch. The spectral Mamba branch processes the normalized input features sequentially through linear layers and 1×1 convolutions before inputting them into the state space model to obtain the branch's output. The spatial Mamba branch transposes the normalized input features, processes them sequentially through linear layers and 1×1 convolutions, and then inputs them into the state space model to obtain the branch's output. The guiding branch processes the normalized input features sequentially through linear layers, inputs them into the ConvNeXt classifier, and adjusts the output of the ConvNeXt classifier using the SiLU activation function to obtain the branch's output.

[0070] The recognition result is obtained by fusing the output of the guiding branch, the output of the spectral Mamba branch, the output of the spatial Mamba branch, and the input features; the recognition model is trained using grain sample images.

[0071] Therefore, based on the recognition model obtained by inputting both visible light data and hyperspectral data (i.e., reflectance spectral data) of grain grains into the recognition model, a hierarchical interaction and dynamic fusion of spectral features, spatial features, and their joint representations is achieved through a cross-fusion structure of the TriMamba three branches (spectral Mamba branch, spatial Mamba branch, and guiding branch) guided by the ConvNeXt framework within the recognition model. This structure fully leverages the advantages of the Mamba state-space model in long-range dependency modeling, realizing deep collaboration and efficient fusion of spectral and spatial features. This results in a significant average improvement in recognition accuracy, while maintaining excellent robustness and generalization performance under complex backgrounds and varying illumination conditions.

[0072] To improve the efficiency of Mamba in long-range dependency modeling, structured preprocessing of features is performed before entering the state-space model. Specifically, the input features are normalized by expanding the three-dimensional input features, which include feature map height W, feature map width H, and channel dimension C, into two-dimensional input features, which include channel dimension C and batch dimension N, to achieve input feature normalization. In one embodiment, the extracted feature blocks are first processed... (i.e., 3D input features), perform a 1D scan operation, and systematically unfold it into a 2D representation. (i.e., two-dimensional input features). This process effectively preserves key information from the spectral and spatial domains while reducing data dimensionality. Compared to directly processing high-dimensional tensors, this transformation compresses three-dimensional features into a unified token sequence, providing a more regular input format for the subsequent Mamba encoder, enabling it to more fully leverage the advantages of state-space models in long-distance dependency modeling and global context capture. Therefore, unlike traditional spectral-spatial models that rely on fixed scanning strategies, the TriMamba structure employs a more flexible representation.

[0073] Specifically, refer to Figure 2 Based on the fusion results of the guiding branch output, spectral Mamba branch output, spatial Mamba branch output, and input features, the recognition results can be obtained in the following ways:

[0074] The sum of the products of the guiding branch output and the outputs of the spectral Mamba branch and spatial Mamba branch, respectively, is processed through a linear layer. The processed result is then added to the input features and input into the classifier of the recognition model to obtain the recognition result. In addition to the spectral and spatial Mamba branches, the nonlinear relationship in the feature representation is enhanced by modulation of the guiding branch using ConvNeXt, allowing it to play a crucial role in the output feature aggregation. The final aggregated features are then used as input to the classifier of the recognition model. This can be expressed by the following formula:

[0075]

[0076] in Indicates linear layer processing; This indicates the output of the Mamba branch of the spectrum; This indicates the output of the Mamba branch in the space; Indicates the output of the guiding branch; This represents the input features.

[0077] In a preferred embodiment, the feature aggregation module in the recognition model mainly consists of TriMamba blocks. The overall structure of the TriMamba block comprises three parallel paths: a spectral Mamba branch, a spatial Mamba branch, and a ConvNeXt guiding branch (i.e., the guiding branch). The input feature map first undergoes normalization preprocessing (i.e., normalization of the input features), and then is fed into the three branches respectively to fully capture spectral, spatial, and global information, and is modeled using a state-space model. Taking the recognition process of corn kernels as an example:

[0078] Reference Figure 2 In the spectral Mamba branch (i.e. Figure 2 The bottommost branch), the normalized input feature sequence First, it passes through a linear layer and a 1×1 convolution (i.e., ... The former is used for feature standardization and recalibration to enhance numerical stability; the latter achieves feature compression and dimensionality transformation through inter-channel information interaction, thereby improving expressive power. This design not only effectively integrates local correlations in the spectral dimension, providing a more compact and discriminative feature representation for selective scanning in the subsequent modeling stage, as shown in the following equation:

[0079]

[0080] Subsequently, the processed feature sequence It is input into a state-space model (SSM) for modeling. In the spectral Mamba branch, the SSM effectively captures the dynamic dependencies between features along the spectral dimension through a parameterized recursive state update mechanism, as shown below (where... (All are parameters of the state-space model)

[0081]

[0082] Reference Figure 2 Spatial Mamba branch (i.e.) Figure 2 The topmost branch (where Permute represents dimensionality transpose) processes similarly to the spectral branch, but with a key difference: before inputting the state space model, the feature sequence undergoes dimensionality transpose (i.e., ...). The data is rearranged from the spectral dimension to the spatial dimension. This allows SSM to establish long-range dependencies between features at different spatial locations, enabling cross-regional contextual modeling. Through this process, the model's ability to perceive global spatial information is significantly enhanced, providing a more expressive spatial representation for subsequent discrimination tasks. The processing of the spatial Mamba branch is specifically represented as follows (where...). (All are parameters of the state-space model)

[0083]

[0084]

[0085] In the pilot branch (i.e.) Figure 2 The middle branch uses the ConvNeXt classifier model and combines it with the SiLU activation function to adjust the output of the state-space model, as shown in the following equation:

[0086]

[0087] Furthermore, to fully leverage the learning potential of each feature extraction branch, a multi-branch loss strategy is designed, with an independent loss function for each branch to optimize the training effect of the recognition model. Specifically, in this embodiment, the loss value during the training of the recognition model is obtained by fusing the loss corresponding to the true label of the grain sample image and the label predicted by the recognition model, the loss corresponding to the true label of the grain sample image and the label output according to the guiding branch, the loss corresponding to the true label of the grain sample image and the label output according to the spectral Mamba branch, and the loss corresponding to the true label of the grain sample image and the label output according to the spatial Mamba branch.

[0088] Throughout the recognition process, the model employs the cross-entropy loss function as the optimization objective. By measuring the difference between the predicted results and the true labels, it effectively guides the network to learn the discriminative features between normal and abnormal grain grain images, thereby improving overall recognition performance. Specifically, the losses corresponding to the true labels of grain grain sample images and the labels predicted by the recognition model, the true labels of grain grain sample images and the labels obtained according to the guiding branch output, the true labels of grain grain sample images and the labels obtained according to the spectral Mamba branch output, and the true labels of grain grain sample images and the labels obtained according to the spatial Mamba branch output, all use cross-entropy loss. In a preferred embodiment, taking the recognition of corn kernels as an example, the loss corresponding to the true labels of grain grain sample images and the labels predicted by the recognition model is expressed as follows:

[0089]

[0090] in, The true label represents the image of a normal corn kernel sample, while This represents the label predicted by the recognition model. When the image contains incomplete corn kernels, it will... If the image shows normal corn kernels, then... .

[0091] The loss corresponding to the spectral feature extraction branch is (Loss corresponding to the true labels of grain sample images and the labels obtained according to the spectral Mamba branch output), the loss corresponding to the spatial image branch is... (The loss corresponds to the true label of the grain sample image and the label obtained according to the spatial Mamba branch output), while the loss corresponding to the guiding branch is... (Loss corresponding to the true labels of grain sample images and the labels obtained according to the output of the guiding branch). Each loss function is specifically designed for the feature dimensions of its respective branch to achieve targeted optimization for different types of features, thereby maximizing the feature representation capability of each branch. The loss of each branch is defined as follows:

[0092]

[0093]

[0094]

[0095]

[0096] in, , and These are the hyperparameters used to balance the loss contributions of each branch. By default, these three hyperparameters are set to 0.001, i.e. This is to ensure that each branch maintains a relatively balanced optimization intensity during joint training.

[0097] In addition, after the recognition model is trained, it is evaluated by the performance evaluation module. If the evaluation result is unsatisfactory, the recognition model is retrained. The evaluation methods of the performance evaluation module include:

[0098] The recognition results of the recognition model are evaluated using accuracy, precision, recall, mean precision, and Kappa coefficient as indicators. If the recognition results of the recognition model fail to meet the corresponding indicator thresholds, the evaluation results are deemed unqualified.

[0099] Recall rate is obtained from true positives, true negatives, false positives, and false negatives; precision rate is obtained from true positives and false positives; recall rate is obtained from true positives and false negatives; mean precision is calculated from precision rate and recall rate.

[0100] A true positive is the number of pixels in the sample image where both the label and the recognition result are grain regions; a true negative is the number of pixels in the sample image where both the label and the recognition result are background regions; a false positive is the number of pixels in the sample image where the label is background regions but the recognition result is grain regions; and a false negative is the number of pixels in the sample image where the label is grain regions but the recognition result is background regions.

[0101] In classification tasks (such as grain identification), accuracy measures the overall prediction correctness of the model across all samples, suitable for balanced sample scenarios; precision focuses on the proportion of samples predicted as positive by the model to be true positives, used to avoid false positives; recall reflects the proportion of true positive samples correctly identified by the model, used to reduce false negatives; mean precision (mAP) is the average of the area under the precision-recall curves for each class, and is a core evaluation metric for multi-class and object detection tasks; the Kappa coefficient eliminates the influence of random guessing, measures the consistency between the predicted results and the true labels, and is more suitable for imbalanced or multi-class tasks. These metrics comprehensively evaluate the performance of the recognition model from different dimensions. Specifically, the evaluation metrics are accuracy (Acc), precision (Pre), recall (Re), mean precision (mAP), and the Kappa coefficient, and the calculation of each metric is as follows:

[0102]

[0103]

[0104]

[0105] Among them, TP (true positive) represents pixels where both the label and the prediction result are in the seed region, TN (true negative) represents pixels where both are in the background region, and FP (false positive) and FN (false negative) correspond to pixels where the label and the prediction result are inconsistent.

[0106]

[0107] Among them, the average accuracy is .

[0108] The Kappa coefficient is used for consistency testing and can also be used to measure recognition accuracy. Kappa is calculated using the following formula:

[0109]

[0110] The Kappa coefficient ranges from [−1, 1]. When k = 1, it indicates perfect agreement; when k = 0, the agreement is comparable to random identification; and when k < 0, the agreement is worse than random identification. For observational consistency (i.e., overall accuracy), it represents the number of correctly identified samples divided by the total number of samples. The expected probability of random consistency is calculated as follows:

[0111]

[0112] Where m represents the total number of categories, and N represents the total number of samples. This represents the number of true samples in class i, while This represents the number of samples predicted from category i. Generally, the higher the Kappa coefficient of the model, the stronger the consistency between its recognition results and the actual situation, and the better its robustness.

[0113] In this embodiment, the visible light image of the grain is obtained by performing preprocessing on the original image of the grain acquired by the visible light camera, which includes denoising, filtering, enhancement and normalization.

[0114] The reflectance spectral data of grain grains are obtained by scanning the reflectance spectral data cube of grain grain samples line by line in the visible and near-infrared bands. The data is then processed by whiteboard correction to eliminate systematic errors caused by differences in ambient light and instrument response, spectral smoothing algorithm to reduce high-frequency noise, and normalization or standardization of each band.

[0115] In a preferred embodiment, taking corn kernels as an example, refer to... Figure 3 The system primarily consists of a hyperspectral camera, a visible light camera, a lens assembly, a fixture, a worktable, a backlight source, a host computer, and corn kernel samples, forming a system for acquiring raw images and reflectance spectral data of grain kernels. The hyperspectral and visible light cameras are fixed to the fixture support, with their lenses pointing vertically downwards towards the worktable. The corn kernel samples are evenly distributed on the worktable to ensure a flat and uniform sample surface during acquisition. The field of view is controlled by adjusting the shifting structure to achieve spatial consistency between the spectral and image acquisition areas. During acquisition, the hyperspectral imaging system first scans the sample line by line in the visible-near-infrared band to obtain a cube of reflectance spectral data. Then, the system switches to the visible light industrial camera to acquire high-resolution RGB images (i.e., raw images of the grain kernels) within the same field of view. By repeatedly changing the sample orientation and quantity, the acquisition process is repeated to form a diverse and highly consistent spectral and image dataset, providing fundamental data support for subsequent feature analysis, model training, and phenotypic recognition.

[0116] The preprocessing of the hyperspectral portion specifically involves: performing whiteboard correction on the original hyperspectral cube data to eliminate systematic errors caused by differences between ambient light and instrument response; secondly, using a spectral smoothing algorithm to reduce high-frequency noise and retain effective spectral information; and finally, normalizing or standardizing each band to ensure that the spectral characteristics of different bands are comparable on the same scale.

[0117] The preprocessing of the visible light image part specifically involves performing noise reduction, filtering, enhancement, and normalization operations on the corn kernel images acquired by the visible light camera in sequence.

[0118] like Figure 4 As shown, features are extracted from the visible light image and reflectance spectrum data of grain grains through corresponding curvelet transforms. The definition of the curvelet function in this embodiment can be found in the prior art paper "Application Research of Curvelet Transform in Airborne Gamma Spectrum Data Processing" published by Li Binghai et al. in the December 2023 issue of World Nuclear Geology Science, Volume 40, No. 4. The curvelet transform is performed on the signal to be processed... With basis functions Perform inner product to achieve sparse representation of the signal. :

[0119]

[0120] in, These correspond to scale, direction, and position, respectively. Two-dimensional curvelet transform decomposes an image into a series of non-overlapping scales and analyzes these scales through local ridge transform.

[0121] To adapt to the needs of digitalization, the discretization of the continuous curvelet transform uses a Cartesian grid as data input and a set of coefficients as output. Assume... The input signal is, where , Let be the spatial coordinates of the input signal, and let the curve transform coefficients be:

[0122]

[0123] In the above formula, are the basis functions of the discrete curvelet transform; j, l, k are the scale parameter, direction parameter, and position parameter, respectively.

[0124] The local window function in the frequency domain is defined as follows:

[0125]

[0126] In the above formula: These are frequency domain parameters; It is a radial function; It is an angle function. Let denote the inner product of a one-dimensional function, and:

[0127]

[0128] Transformation in polar coordinates:

[0129]

[0130] In the above formula: ; These are frequency domain polar coordinates; Given an angular sequence, the discrete curvilinear wavefunction can be defined as:

[0131]

[0132] In the above formula: , For position parameters; , The translation parameter is used. Curveflow transform, as a powerful multi-scale decomposition tool, is suitable for image data with curved edge structures.

[0133] Specifically, methods for extracting features from visible light images of grain grains include:

[0134] The visible light image of grain grains is downsized using a point-based convolutional layer. The downsized image is then decomposed into different frequency scales using a two-dimensional curvelet transform. The resulting frequency band sub-components are processed by the first convolutional layer and then restored to their original spatial dimensions using a two-dimensional inverse curvelet transform, resulting in a curvelet transform feature map. Features extracted from the visible light image of grain grains are obtained based on the curvelet transform feature map.

[0135] Methods for extracting features from visible light images of grains based on visible light curve transform feature maps include:

[0136] The visible light curve transform feature map and the image after the size reduction are processed by the second convolutional layer are fused together. Based on the fusion result, the features extracted from the visible light image of the grain are obtained.

[0137] Reference Figure 5 In a preferred embodiment, the first convolutional layer convolutional kernel Second convolutional layer kernel All images are 3×3 in size, effectively extracting fine local features. Combining curvelet transform theory, for the input image... (Specifically, visible light images) are first processed through point-based convolutional layers. The image size is reduced to generate an optimized feature map. Subsequently, a two-dimensional curvelet transform (i.e., 2D CT) is applied to this reduced feature map via a db2 wavelet transform, achieving multiple decompositions at different frequency scales. This process generates four sub-band components: LL, LH, HL, and HH, representing different frequency bands in the low and high frequencies, respectively. Each frequency band sub-component is then further processed through a first convolutional layer to effectively extract local fine features of the image. After this processing, the feature map is restored to its original spatial dimensions using a two-dimensional inverse wavelet transform, ultimately generating a complete feature map (i.e., a visible light curvelet transform feature map). The calculation formula is as follows:

[0138]

[0139] To enhance model performance, the output of the curvelet decomposition convolution module employs a custom residual connection method. The reduced-size feature map is convolved with another set of convolutional kernels (the second convolutional layer), and then fused with the 2D curvelet transform result (i.e., the visible light curvelet transform feature map) to obtain the final output of the curvelet decomposition convolution module. This refers to features extracted from visible light images of grains. The specific fusion process is as follows:

[0140]

[0141]

[0142] Through the above methods, curvelet transform can capture detailed information in images, effectively reduce computational complexity, and achieve accurate information decomposition in the spatial frequency domain, ultimately improving the efficiency and accuracy of image processing tasks.

[0143] Unlike traditional spatial image processing, hyperspectral images contain a wealth of information along the spectral dimension, with each pixel's response across multiple bands collectively forming its unique spectral profile. To further enhance the feature extraction capability along the spectral dimension in hyperspectral imaging, this implementation introduces spectral curvelet convolution, combining one-dimensional wavelet transform with a convolution kernel to effectively capture spectral features at multiple scales. Specifically, the methods for extracting features from the reflectance spectrum data of grain grains include:

[0144] The reflectance spectral data of grain grains is subjected to rotational convolution along the spectral dimension for dimensionality reduction. Then, a one-dimensional discrete curvelet transform is used to decompose the dimensionality-reduced spectral features into low-frequency and high-frequency components. Convolution operations are then performed on the low-frequency and high-frequency components separately through channel-dimensional convolution. Finally, a one-dimensional inverse curvelet transform is used to reintegrate the features back into the original spectral dimension, resulting in hyperspectral curvelet transform features. Based on the hyperspectral curvelet transform features, the features extracted from the reflectance spectral data of grain grains are obtained.

[0145] Methods for extracting features from the reflectance spectrum data of grain grains based on hyperspectral curve transform characteristics include:

[0146] The hyperspectral curve transform features and the dimensionality-reduced spectral features are fused together after being processed by the third convolutional layer. Based on the fusion result, the features extracted from the reflectance spectral data of grain grains are obtained.

[0147] Reference Figure 6 In a preferred embodiment, for the input image (Specifically, this refers to reflectance spectral data.) A rotational convolution is performed along the spectral dimension to simplify redundant spectral information while preserving key frequency details. The initial number of channels for the rotational convolution is set to 32 to achieve an optimal balance between model complexity and performance. Subsequently, the spectral curvilinear convolution module introduces a one-dimensional discrete curvilinear transform (i.e., 1D CT) along the spectral dimension, decomposing the dimensionality-reduced spectral feature map into low-frequency components X. lf and high-frequency component X hf The low-frequency components characterize the overall trend of the spectral curve, while the high-frequency components capture subtle perturbations and local variations in the spectrum, thus enabling multi-scale modeling of both coarse-grained and fine-grained features. After decomposition, the convolution kernel of the third convolutional layer is used... Perform convolution operations on the low-frequency and high-frequency components separately (i.e. Figure 6 Conv1×1 in Figure 6 (The annotations in the text are abbreviated) to further extract discriminative features at different frequencies. This convolutional kernel can effectively model the correlation between adjacent bands and maintain the expressive power of the model while reducing computational burden and parameter size. In the convolution operation, the image size dimension convolutional kernel is set to 1 and the channel dimension convolutional kernel is set to 3. The core purpose is to efficiently fuse multi-channel features while maintaining spatial resolution, which is suitable for computer vision tasks such as grain recognition that require a balance between accuracy and computational efficiency.

[0148] Finally, the spectral features are reintegrated back into the original spectral dimension through a one-dimensional inverse curvature transform, achieving multi-scale fusion of spectral features. In this way, the spectral curvature convolution module not only highlights global spectral trends but also enhances sensitivity to local details, thereby improving the feature representation capability in hyperspectral imaging tasks. The formula for the hyperspectral curvature transform features is as follows:

[0149]

[0150] Building upon this, the spectral curve convolution module further employs a residual fusion strategy to transform the one-dimensional wavelet feature maps after frequency transformation. Mapping (i.e., hyperspectral curve transform features) with convolution output features Add them together to generate the module's final output. That is, features extracted from the reflectance spectral data of grain grains:

[0151]

[0152]

[0153] Among them, convolution kernel and They are the same, both having a scale of 1x1x3.

[0154] The spectral curvelet convolution module fully leverages the advantages of one-dimensional curvelet transform in spectral frequency decomposition, significantly enhancing the model's sensitivity to complete spectral information. It can capture spectral features more comprehensively while maintaining low computational complexity, thereby improving recognition performance. The frequency domain features extracted using curvelet decomposition provide rich multi-scale contextual information for subsequent Mamba blocks, while the state-space mechanism in the Mamba blocks enhances the model's selective perception of frequency domain features through dynamic parameter adaptation. Adjusting the sensitivity to different frequencies based on input data allows for more accurate extraction of key spectral information.

[0155] Reference Figure 4 Taking the identification of corn kernels as an example, the above identification method process is illustrated below:

[0156] S1: First, connect the hyperspectral imaging equipment and image acquisition sensor to the host computer via connecting cables to ensure effective transmission and processing of spectral and image data. The optical centers of the hyperspectral camera and the visible light camera are perpendicular to the surface of the test platform to ensure accurate imaging of the effective surface of the kernels during shooting. Supplemental lighting is positioned around the camera on the same plane, ensuring the corn kernels are evenly distributed on the test platform, providing clear image contrast with the help of the light sources. With the corn kernels dispersed on the test platform, open the acquisition software on the host computer, calibrate the hyperspectral camera and visible light camera to their optimal shooting states, and then click the acquisition button on the hyperspectral software and the image capture button on the visible light camera to acquire data.

[0157] S2: The hyperspectral data part of the preprocessing module performs whiteboard correction, eliminates systematic errors, filters and denoises, and normalizes the effective data sequentially on the original reflectance spectral data; the visible light image part performs denoising, filtering, enhancement, and normalization operations sequentially on the original corn kernel image; the preprocessed data enters S3 for feature learning;

[0158] S3: Data is input into the spectral curvelet convolution module and the decoupled convolution module to achieve deep fusion of spectral and spatial features. The spectral curvelet convolution combines curvelet decomposition and convolution operations, using multi-scale, multi-directional filters to extract joint spectral and spatial features, enhancing edge and texture representation. The decoupled convolution, through its decomposable convolution structure and Mamba's sequence modeling capabilities, efficiently captures the global dependencies of high-dimensional spectral-spatial sequences, achieving dynamic feature aggregation. Simultaneously, the branching structure design significantly reduces the computational burden in the spatial dimension while preserving key features, thereby improving the overall expressive efficiency and computational performance of the model.

[0159] S4: In the feature aggregation stage, the recognition model dynamically reorganizes and optimizes the spectral and spatial modes through a feature rearrangement mechanism, effectively integrating spectral differences and spatial structure information to improve feature consistency and discriminative power. A TriMamba branch network guided by ConvNeXt is used for cross-modal interactive fusion, achieving efficient representation of the spectral-spatial joint domain. This design not only improves feature extraction and fusion efficiency but also enhances the robustness and accuracy of the model in the corn kernel recognition task.

[0160] S5: Performance evaluation module, which evaluates the quality of the recognition model and provides an effective assessment in terms of objective metrics. The objective evaluation metrics for model parameters are, in order, accuracy (Acc), precision (Pre), recall (Re), mean average precision (mAP), and Kappa coefficient.

[0161] Incomplete implementation of grain grain identification system

[0162] This embodiment provides a technical solution for an imperfect grain identification system, including a processor containing executable program instructions. These instructions are used to execute the following imperfect grain identification method, specifically including:

[0163] The visible light image and reflectance spectrum data of the grain are input into the recognition model; the feature extraction module in the recognition model obtains the input features based on the features extracted from the visible light image and reflectance spectrum data of the grain respectively;

[0164] The feature aggregation module in the recognition model includes a spectral Mamba branch, a spatial Mamba branch, and a guiding branch. The spectral Mamba branch processes the normalized input features sequentially through linear layers and 1×1 convolutions before inputting them into the state space model to obtain the branch's output. The spatial Mamba branch transposes the normalized input features, processes them sequentially through linear layers and 1×1 convolutions, and then inputs them into the state space model to obtain the branch's output. The guiding branch processes the normalized input features sequentially through linear layers, inputs them into the ConvNeXt classifier, and adjusts the output of the ConvNeXt classifier using the SiLU activation function to obtain the branch's output.

[0165] The recognition result is obtained by fusing the output of the guiding branch, the output of the spectral Mamba branch, the output of the spatial Mamba branch, and the input features; the recognition model is trained using grain sample images.

[0166] Therefore, based on the recognition model obtained by inputting both visible light data and hyperspectral data (i.e., reflectance spectral data) of grain grains into the recognition model, a hierarchical interaction and dynamic fusion of spectral features, spatial features, and their joint representations is achieved through a cross-fusion structure of the TriMamba three branches (spectral Mamba branch, spatial Mamba branch, and guiding branch) guided by the ConvNeXt framework within the recognition model. This structure fully leverages the advantages of the Mamba state-space model in long-range dependency modeling, realizing deep collaboration and efficient fusion of spectral and spatial features. This results in a significant average improvement in recognition accuracy, while maintaining excellent robustness and generalization performance under complex backgrounds and varying illumination conditions.

[0167] To improve the efficiency of Mamba in long-range dependency modeling, structured preprocessing of features is performed before entering the state-space model. Specifically, the input features are normalized by expanding the three-dimensional input features, which include feature map height W, feature map width H, and channel dimension C, into two-dimensional input features, which include channel dimension C and batch dimension N, to achieve input feature normalization. In one embodiment, the extracted feature blocks are first processed... (i.e., 3D input features), perform a 1D scan operation, and systematically unfold it into a 2D representation. (i.e., two-dimensional input features). This process effectively preserves key information from the spectral and spatial domains while reducing data dimensionality. Compared to directly processing high-dimensional tensors, this transformation compresses three-dimensional features into a unified token sequence, providing a more regular input format for the subsequent Mamba encoder, enabling it to more fully leverage the advantages of state-space models in long-distance dependency modeling and global context capture. Therefore, unlike traditional spectral-spatial models that rely on fixed scanning strategies, the TriMamba structure employs a more flexible representation.

[0168] Specifically, the recognition results are obtained by fusing the outputs of the guiding branch, the spectral Mamba branch, the spatial Mamba branch, and the input features, including:

[0169] The sum of the products of the guiding branch output and the outputs of the spectral Mamba branch and spatial Mamba branch, respectively, is processed through a linear layer. The processed result is then added to the input features and input into the classifier of the recognition model to obtain the recognition result. In addition to the spectral and spatial Mamba branches, the nonlinear relationship in the feature representation is enhanced by modulation of the guiding branch using ConvNeXt, allowing it to play a crucial role in the output feature aggregation. The final aggregated features are then used as input to the classifier of the recognition model. This can be expressed by the following formula:

[0170]

[0171] in Indicates linear layer processing; This indicates the output of the Mamba branch of the spectrum; This indicates the output of the Mamba branch in the space; Indicates the output of the guiding branch; This represents the input features.

[0172] In a preferred embodiment, the feature aggregation module in the recognition model mainly consists of TriMamba blocks. The overall structure of the TriMamba block comprises three parallel paths: a spectral Mamba branch, a spatial Mamba branch, and a ConvNeXt guiding branch (i.e., the guiding branch). The input feature map first undergoes normalization preprocessing (i.e., normalization of the input features), and then is fed into the three branches respectively to fully capture spectral, spatial, and global information, and is modeled using a state-space model. Taking the recognition process of corn kernels as an example:

[0173] In the spectral Mamba branch, the normalized input feature sequence First, it passes through a linear layer and a 1×1 convolution (i.e., ... The former is used for feature standardization and recalibration to enhance numerical stability; the latter achieves feature compression and dimensionality transformation through inter-channel information interaction, thereby improving expressive power. This design not only effectively integrates local correlations in the spectral dimension, providing a more compact and discriminative feature representation for selective scanning in the subsequent modeling stage, as shown in the following equation:

[0174]

[0175] Subsequently, the processed feature sequence It is input into a state-space model (SSM) for modeling. In the spectral Mamba branch, the SSM effectively captures the dynamic dependencies between features along the spectral dimension through a parameterized recursive state update mechanism, as shown below (where... (All are parameters of the state-space model)

[0176]

[0177] The processing flow of the spatial Mamba branch is similar to that of the spectral branch, but there is a key difference: before inputting the state space model, the feature sequence is first transposed (i.e., ...). The data is rearranged from the spectral dimension to the spatial dimension. This allows SSM to establish long-range dependencies between features at different spatial locations, enabling cross-regional contextual modeling. Through this process, the model's ability to perceive global spatial information is significantly enhanced, providing a more expressive spatial representation for subsequent discrimination tasks. The processing of the spatial Mamba branch is specifically represented as follows (where...). (All are parameters of the state-space model)

[0178]

[0179]

[0180] The pilot branch uses the ConvNeXt classifier model and combines it with the SiLU activation function to adjust the output of the state-space model, as shown in the following equation:

[0181]

[0182] Furthermore, to fully leverage the learning potential of each feature extraction branch, a multi-branch loss strategy is designed, with an independent loss function for each branch to optimize the training effect of the recognition model. Specifically, in this embodiment, the loss value during the training of the recognition model is obtained by fusing the loss corresponding to the true label of the grain sample image and the label predicted by the recognition model, the loss corresponding to the true label of the grain sample image and the label output according to the guiding branch, the loss corresponding to the true label of the grain sample image and the label output according to the spectral Mamba branch, and the loss corresponding to the true label of the grain sample image and the label output according to the spatial Mamba branch.

[0183] Throughout the recognition process, the model employs the cross-entropy loss function as the optimization objective. By measuring the difference between the predicted results and the true labels, it effectively guides the network to learn the discriminative features between normal and abnormal grain grain images, thereby improving overall recognition performance. Specifically, the losses corresponding to the true labels of grain grain sample images and the labels predicted by the recognition model, the true labels of grain grain sample images and the labels obtained according to the guiding branch output, the true labels of grain grain sample images and the labels obtained according to the spectral Mamba branch output, and the true labels of grain grain sample images and the labels obtained according to the spatial Mamba branch output, all use cross-entropy loss. In a preferred embodiment, taking the recognition of corn kernels as an example, the loss corresponding to the true labels of grain grain sample images and the labels predicted by the recognition model is expressed as follows:

[0184]

[0185] in, The true label represents the image of a normal corn kernel sample, while This represents the label predicted by the recognition model. When the image shows abnormal corn kernels, it will be... If the image shows normal corn kernels, then... .

[0186] The loss corresponding to the spectral feature extraction branch is (Loss corresponding to the true labels of grain sample images and the labels obtained according to the spectral Mamba branch output), the loss corresponding to the spatial image branch is... (The loss corresponds to the true label of the grain sample image and the label obtained according to the spatial Mamba branch output), while the loss corresponding to the guiding branch is... (Loss corresponding to the true labels of grain sample images and the labels obtained according to the output of the guiding branch). Each loss function is specifically designed for the feature dimensions of its respective branch to achieve targeted optimization for different types of features, thereby maximizing the feature representation capability of each branch. The loss of each branch is defined as follows:

[0187]

[0188]

[0189]

[0190]

[0191] in, , and These are the hyperparameters used to balance the loss contributions of each branch. By default, these three hyperparameters are set to 0.001, i.e. This is to ensure that each branch maintains a relatively balanced optimization intensity during joint training.

[0192] In addition, after the recognition model is trained, it is evaluated by the performance evaluation module. If the evaluation result is unsatisfactory, the recognition model is retrained. The evaluation methods of the performance evaluation module include:

[0193] The recognition results of the recognition model are evaluated using accuracy, precision, recall, mean precision, and Kappa coefficient as indicators. If the recognition results of the recognition model fail to meet the corresponding indicator thresholds, the evaluation results are deemed unqualified.

[0194] Recall rate is obtained from true positives, true negatives, false positives, and false negatives; precision rate is obtained from true positives and false positives; recall rate is obtained from true positives and false negatives; mean precision is calculated from precision rate and recall rate.

[0195] A true positive is the number of pixels in the sample image where both the label and the recognition result are grain regions; a true negative is the number of pixels in the sample image where both the label and the recognition result are background regions; a false positive is the number of pixels in the sample image where the label is background regions but the recognition result is grain regions; and a false negative is the number of pixels in the sample image where the label is grain regions but the recognition result is background regions.

[0196] In classification tasks (such as grain identification), accuracy measures the overall prediction accuracy of the model across all samples and is suitable for balanced sample scenarios; precision focuses on the proportion of samples predicted as positive by the model to be true positive samples, used to avoid false positives; recall reflects the proportion of true positive samples correctly identified by the model, used to reduce false negatives; mean precision (mAP) is the average of the area under the precision-recall curves for each category and is a core evaluation metric for multi-class and object detection tasks; the Kappa coefficient eliminates the influence of random guessing and measures the consistency between the prediction results and the true labels, making it more suitable for imbalanced or multi-class tasks. These metrics comprehensively evaluate the performance of the recognition model from different dimensions.

[0197] Specifically, the evaluation metrics are accuracy (Acc), precision (Pre), recall (Re), mean precision (mAP), and Kappa coefficient. The calculation of each metric is as follows:

[0198]

[0199]

[0200]

[0201] Among them, TP (true positive) represents pixels where both the label and the prediction result are in the seed region, TN (true negative) represents pixels where both are in the background region, and FP (false positive) and FN (false negative) correspond to pixels where the label and the prediction result are inconsistent.

[0202]

[0203] Among them, the average accuracy is .

[0204] The Kappa coefficient is used for consistency testing and can also be used to measure recognition accuracy. Kappa is calculated using the following formula:

[0205]

[0206] The Kappa coefficient ranges from [−1, 1]. When k = 1, it indicates perfect agreement; when k = 0, the agreement is comparable to random identification; and when k < 0, the agreement is worse than random identification. For observational consistency (i.e., overall accuracy), it represents the number of correctly identified samples divided by the total number of samples. The expected probability of random consistency is calculated as follows:

[0207]

[0208] Where m represents the total number of categories, and N represents the total number of samples. This represents the number of true samples in class i, while This represents the number of samples predicted from category i. Generally, the higher the Kappa coefficient of the model, the stronger the consistency between its recognition results and the actual situation, and the better its robustness.

[0209] In this embodiment, the visible light image of the grain is obtained by performing preprocessing on the original image of the grain acquired by the visible light camera, which includes denoising, filtering, enhancement and normalization.

[0210] The reflectance spectral data of grain grains are obtained by scanning the reflectance spectral data cube of grain grain samples line by line in the visible and near-infrared bands. The data is then processed by whiteboard correction to eliminate systematic errors caused by differences in ambient light and instrument response, spectral smoothing algorithm to reduce high-frequency noise, and normalization or standardization of each band.

[0211] In a preferred embodiment, taking corn kernels as an example, the system mainly consists of a hyperspectral camera, a visible light camera, a lens assembly, a fixture, a worktable, a backlight source, a host computer, and corn kernel samples, forming a system for acquiring raw images and reflectance spectral data of grain kernels. The hyperspectral camera and the visible light camera are fixed on the fixture support, with the lenses pointing vertically downwards towards the worktable. The corn kernel samples are evenly laid out on the worktable to ensure a flat and uniform sample surface during acquisition. The field of view is controlled by adjusting the shifting structure to achieve spatial consistency between the spectral and image acquisition areas. During acquisition, the hyperspectral imaging system first scans the sample line by line in the visible-near-infrared band to obtain a cube of reflectance spectral data; then, the system switches to a visible light industrial camera to acquire high-resolution RGB images (i.e., raw images of the grain kernels) under the same field of view. By repeatedly changing the sample orientation and quantity, the acquisition process is repeated to form a diverse and highly consistent spectral and image dataset, providing basic data support for subsequent feature analysis, model training, and phenotypic recognition.

[0212] The preprocessing of the hyperspectral portion specifically involves: performing whiteboard correction on the original hyperspectral cube data to eliminate systematic errors caused by differences between ambient light and instrument response; secondly, using a spectral smoothing algorithm to reduce high-frequency noise and retain effective spectral information; and finally, normalizing or standardizing each band to ensure that the spectral characteristics of different bands are comparable on the same scale.

[0213] The preprocessing of the visible light image part specifically involves performing noise reduction, filtering, enhancement, and normalization operations on the corn kernel images acquired by the visible light camera in sequence.

[0214] Features are extracted from the visible light image and reflectance spectrum data of grain grains using appropriate curvelet transforms. The definition of the curvelet function in this embodiment can be found in the prior art paper "Application Research of Curvelet Transform in Airborne Gamma Spectrum Data Processing" published by Li Binghai et al. in the December 2023 issue of World Nuclear Geology Science, Volume 40, No. 4. The curvelet transform is applied to the signal to be processed... With basis functions Perform inner product to achieve sparse representation of the signal. :

[0215]

[0216] in, These correspond to scale, direction, and position, respectively. Two-dimensional curvelet transform decomposes an image into a series of non-overlapping scales and analyzes these scales through local ridge transform.

[0217] To adapt to the needs of digitalization, the discretization of the continuous curvelet transform uses a Cartesian grid as data input and a set of coefficients as output. Assume... The input signal is, where , Let be the spatial coordinates of the input signal, and let the curve transform coefficients be:

[0218]

[0219] In the above formula, These are the basis functions of the discrete curvelet transform; their corresponding frequency domain representation is expressed by defining a local window function as follows:

[0220]

[0221] In the above formula: These are frequency domain parameters; It is a radial function; It is an angle function. Let denote the inner product of a one-dimensional function, and:

[0222]

[0223] Transformation in polar coordinates:

[0224]

[0225] In the above formula: ; These are frequency domain polar coordinates; Given an angular sequence, the discrete curvilinear wavefunction can be defined as:

[0226]

[0227] In the above formula: , For position parameters; , The translation parameter is used. Curveflow transform, as a powerful multi-scale decomposition tool, is suitable for image data with curved edge structures.

[0228] Specifically, methods for extracting features from visible light images of grain grains include:

[0229] The visible light image of grain grains is downsized using a point-based convolutional layer. The downsized image is then decomposed into different frequency scales using a two-dimensional curvelet transform. The resulting frequency band sub-components are processed by the first convolutional layer and then restored to their original spatial dimensions using a two-dimensional inverse curvelet transform, resulting in a curvelet transform feature map. Features extracted from the visible light image of grain grains are obtained based on the curvelet transform feature map.

[0230] Methods for extracting features from visible light images of grains based on visible light curve transform feature maps include:

[0231] The visible light curve transform feature map and the image after the size reduction are processed by the second convolutional layer are fused together. Based on the fusion result, the features extracted from the visible light image of the grain are obtained.

[0232] In a preferred embodiment, the first convolutional layer convolutional kernel Second convolutional layer kernel All images are 3×3 in size, effectively extracting fine local features. Combining curvelet transform theory, for the input image... (Specifically, visible light images) are first processed through point-based convolutional layers. The image size is reduced to generate an optimized feature map. Subsequently, a two-dimensional curvelet transform (i.e., 2D CT) is applied to this reduced feature map via a db2 wavelet transform, achieving multiple decompositions at different frequency scales. This process generates four sub-band components: LL, LH, HL, and HH, representing different frequency bands in the low and high frequencies, respectively. Each frequency band sub-component is then further processed through a first convolutional layer to effectively extract local fine features of the image. After this processing, the feature map is restored to its original spatial dimensions using a two-dimensional inverse wavelet transform, ultimately generating a complete feature map (i.e., a visible light curvelet transform feature map). The calculation formula is as follows:

[0233]

[0234] To enhance model performance, the output of the curvelet decomposition convolution module employs a custom residual connection method. The reduced-size feature map is convolved with another set of convolutional kernels (the second convolutional layer), and then fused with the 2D curvelet transform result (i.e., the visible light curvelet transform feature map) to obtain the final output of the curvelet decomposition convolution module. This refers to features extracted from visible light images of grains. The specific fusion process is as follows:

[0235]

[0236]

[0237] Through the above methods, curvelet transform can capture detailed information in images, effectively reduce computational complexity, and achieve accurate information decomposition in the spatial frequency domain, ultimately improving the efficiency and accuracy of image processing tasks.

[0238] Unlike traditional spatial image processing, hyperspectral images contain a wealth of information along the spectral dimension, with each pixel's response across multiple bands collectively forming its unique spectral profile. To further enhance the feature extraction capability along the spectral dimension in hyperspectral imaging, this implementation introduces spectral curvelet convolution, combining one-dimensional wavelet transform with a convolution kernel to effectively capture spectral features at multiple scales. Specifically, the methods for extracting features from the reflectance spectrum data of grain grains include:

[0239] The reflectance spectral data of grain grains is subjected to rotational convolution along the spectral dimension for dimensionality reduction. Then, a one-dimensional discrete curvelet transform is used to decompose the dimensionality-reduced spectral features into low-frequency and high-frequency components. Convolution operations are then performed on the low-frequency and high-frequency components separately through channel-dimensional convolution. Finally, a one-dimensional inverse curvelet transform is used to reintegrate the features back into the original spectral dimension, resulting in hyperspectral curvelet transform features. Based on the hyperspectral curvelet transform features, the features extracted from the reflectance spectral data of grain grains are obtained.

[0240] Methods for extracting features from the reflectance spectrum data of grain grains based on hyperspectral curve transform characteristics include:

[0241] The hyperspectral curve transform features and the dimensionality-reduced spectral features are fused together after being processed by the third convolutional layer. Based on the fusion result, the features extracted from the reflectance spectral data of grain grains are obtained.

[0242] In a preferred embodiment, for the input image (Specifically, this refers to reflectance spectral data.) A rotational convolution is performed along the spectral dimension to simplify redundant spectral information while preserving key frequency details. The initial number of channels for the rotational convolution is set to 32 to achieve an optimal balance between model complexity and performance. Subsequently, the spectral curvilinear convolution module introduces a one-dimensional discrete curvilinear transform (i.e., 1D CT) along the spectral dimension, decomposing the dimensionality-reduced spectral feature map into low-frequency components X. lf and high-frequency component X hf The low-frequency components characterize the overall trend of the spectral curve, while the high-frequency components capture subtle perturbations and local variations in the spectrum, thus enabling multi-scale modeling of both coarse-grained and fine-grained features. After decomposition, the convolution kernel of the third convolutional layer is used... Convolution operations are performed separately on low-frequency and high-frequency components to further extract discriminative features at different frequencies. This convolution kernel can effectively model the correlation between adjacent bands while maintaining the model's expressive power, reducing computational burden and parameter size. The convolution operation is set to 1 kernel for the image size dimension and 3 kernels for the channel dimension. The core objective is to efficiently fuse multi-channel features while maintaining spatial resolution, making it suitable for computer vision tasks such as grain recognition that require a balance between accuracy and computational efficiency.

[0243] Finally, the spectral features are reintegrated back into the original spectral dimension through a one-dimensional inverse curvature transform, achieving multi-scale fusion of spectral features. In this way, the spectral curvature convolution module not only highlights global spectral trends but also enhances sensitivity to local details, thereby improving the feature representation capability in hyperspectral imaging tasks. The formula for the hyperspectral curvature transform features is as follows:

[0244]

[0245] Building upon this, the spectral curve convolution module further employs a residual fusion strategy to transform the one-dimensional wavelet feature maps after frequency transformation. Mapping (i.e., hyperspectral curve transform features) with convolution output features Add them together to generate the module's final output. That is, features extracted from the reflectance spectral data of grain grains:

[0246]

[0247]

[0248] Among them, convolution kernel and They are the same, both having a scale of 1x1x3.

[0249] The spectral curvelet convolution module fully leverages the advantages of one-dimensional curvelet transform in spectral frequency decomposition, significantly enhancing the model's sensitivity to complete spectral information. It can capture spectral features more comprehensively while maintaining low computational complexity, thereby improving recognition performance. The frequency domain features extracted using curvelet decomposition provide rich multi-scale contextual information for subsequent Mamba blocks, while the state-space mechanism in the Mamba blocks enhances the model's selective perception of frequency domain features through dynamic parameter adaptation. Adjusting the sensitivity to different frequencies based on input data allows for more accurate extraction of key spectral information.

[0250] Taking the identification of corn kernels as an example, the above identification method process is illustrated as follows:

[0251] S1: First, connect the hyperspectral imaging equipment and image acquisition sensor to the host computer via connecting cables to ensure effective transmission and processing of spectral and image data. The optical centers of the hyperspectral camera and the visible light camera are perpendicular to the surface of the test platform to ensure accurate imaging of the effective surface of the kernels during shooting. Supplemental lighting is positioned around the camera on the same plane, ensuring the corn kernels are evenly distributed on the test platform, providing clear image contrast with the help of the light sources. With the corn kernels dispersed on the test platform, open the acquisition software on the host computer, calibrate the hyperspectral camera and visible light camera to their optimal shooting states, and then click the acquisition button on the hyperspectral software and the image capture button on the visible light camera to acquire data.

[0252] S2: The hyperspectral data part of the preprocessing module performs whiteboard correction, eliminates systematic errors, filters and denoises, and normalizes the effective data sequentially on the original reflectance spectral data; the visible light image part performs denoising, filtering, enhancement, and normalization operations sequentially on the original corn kernel image; the preprocessed data enters S3 for feature learning;

[0253] S3: Data is input into the spectral curvelet convolution module and the decoupled convolution module to achieve deep fusion of spectral and spatial features. The spectral curvelet convolution combines curvelet decomposition and convolution operations, using multi-scale, multi-directional filters to extract joint spectral and spatial features, enhancing edge and texture representation. The decoupled convolution, through its decomposable convolution structure and Mamba's sequence modeling capabilities, efficiently captures the global dependencies of high-dimensional spectral-spatial sequences, achieving dynamic feature aggregation. Simultaneously, the branching structure design significantly reduces the computational burden in the spatial dimension while preserving key features, thereby improving the overall expressive efficiency and computational performance of the model.

[0254] S4: In the feature aggregation stage, the recognition model dynamically reorganizes and optimizes the spectral and spatial modes through a feature rearrangement mechanism, effectively integrating spectral differences and spatial structure information to improve feature consistency and discriminative power. A TriMamba branch network guided by ConvNeXt is used for cross-modal interactive fusion, achieving efficient representation of the spectral-spatial joint domain. This design not only improves feature extraction and fusion efficiency but also enhances the robustness and accuracy of the model in the corn kernel recognition task.

[0255] S5: Performance evaluation module, which evaluates the quality of the recognition model and provides an effective assessment in terms of objective metrics. The objective evaluation metrics for model parameters are, in order, accuracy (Acc), precision (Pre), recall (Re), mean average precision (mAP), and Kappa coefficient.

[0256] It should be understood that the above-described specific embodiments of the present invention are merely illustrative or explanatory of the principles of the present invention, and do not constitute a limitation thereof.

Claims

1. An imperfect method for identifying grain kernels, characterized in that, include: The visible light image and reflectance spectrum data of the grain are input into the recognition model; the feature extraction module in the recognition model obtains the input features based on the features extracted from the visible light image and reflectance spectrum data of the grain respectively; The feature aggregation module in the recognition model includes a spectral Mamba branch, a spatial Mamba branch, and a guiding branch. The spectral Mamba branch processes the normalized input features sequentially through linear layers and convolutions before inputting them into the state space model to obtain the branch's output. The spatial Mamba branch transposes the normalized input features, processes them sequentially through linear layers and convolutions, and then inputs them into the state space model to obtain the branch's output. The guiding branch processes the normalized input features sequentially through linear layers, inputs them into the ConvNeXt classifier, and adjusts the output of the ConvNeXt classifier using the SiLU activation function to obtain the branch's output. The recognition result is obtained by fusing the output of the guiding branch, the output of the spectral Mamba branch, the output of the spatial Mamba branch, and the input features; the recognition model is trained using grain sample images.

2. The imperfect grain identification method according to claim 1, characterized in that, Methods for extracting features from visible light images of grains include: The visible light image of grain grains is downsized using a point-based convolutional layer. The downsized image is then decomposed into different frequency scales using a two-dimensional curvelet transform. The resulting frequency band sub-components are processed by the first convolutional layer and then restored to their original spatial dimensions using a two-dimensional inverse curvelet transform, resulting in a curvelet transform feature map. Features extracted from the visible light image of grain grains are obtained based on the curvelet transform feature map. Methods for extracting features from visible light images of grains based on visible light curve transform feature maps include: The visible light curve transform feature map is fused with the processing result of the image after the size reduction is processed by the second convolutional layer. Based on the fusion result, the features extracted from the visible light image of the grain are obtained.

3. The imperfect grain identification method according to claim 1, characterized in that, Methods for extracting features from the reflectance spectral data of grain kernels include: The reflectance spectral data of grain grains is subjected to rotational convolution along the spectral dimension for dimensionality reduction. Then, a one-dimensional discrete curvelet transform is used to decompose the dimensionality-reduced spectral features into low-frequency and high-frequency components. Convolution operations are then performed on the low-frequency and high-frequency components separately through channel-dimensional convolution. Finally, a one-dimensional inverse curvelet transform is used to reintegrate the features back into the original spectral dimension, resulting in hyperspectral curvelet transform features. Based on the hyperspectral curvelet transform features, the features extracted from the reflectance spectral data of grain grains are obtained. Methods for extracting features from the reflectance spectrum data of grain grains based on hyperspectral curve transform characteristics include: The hyperspectral curve transform features and the dimensionality-reduced spectral features are fused together after being processed by the third convolutional layer. Based on the fusion result, the features extracted from the reflectance spectral data of grain grains are obtained.

4. The imperfect grain identification method according to any one of claims 1-3, characterized in that, The loss value during the training of the recognition model is obtained by fusing the loss corresponding to the real label of the grain sample image and the label predicted by the recognition model, the loss corresponding to the real label of the grain sample image and the label obtained according to the guide branch output, the loss corresponding to the real label of the grain sample image and the label obtained according to the spectral Mamba branch output, and the loss corresponding to the real label of the grain sample image and the label obtained according to the spatial Mamba branch output.

5. The imperfect grain identification method according to claim 4, characterized in that, The losses corresponding to the true labels of grain sample images and the labels predicted by the recognition model, the losses corresponding to the true labels of grain sample images and the labels obtained according to the guide branch output, the losses corresponding to the true labels of grain sample images and the labels obtained according to the spectral Mamba branch output, and the losses corresponding to the true labels of grain sample images and the labels obtained according to the spatial Mamba branch output are all calculated using cross-entropy loss.

6. The imperfect grain identification method according to any one of claims 1-3, characterized in that, After the recognition model is trained, it is evaluated through the performance evaluation module. If the evaluation result is unsatisfactory, the recognition model is retrained. The evaluation methods for the performance evaluation module include: The recognition results of the recognition model are evaluated using accuracy, precision, recall, mean precision, and Kappa coefficient as indicators. If the recognition results of the recognition model fail to meet the corresponding indicator thresholds, the evaluation results are deemed unqualified. Recall rate is obtained from true positives, true negatives, false positives, and false negatives; precision rate is obtained from true positives and false positives; recall rate is obtained from true positives and false negatives; mean precision is calculated from precision rate and recall rate. A true positive is the number of pixels in the sample image where both the label and the recognition result are grain regions; a true negative is the number of pixels in the sample image where both the label and the recognition result are background regions; a false positive is the number of pixels in the sample image where the label is background regions but the recognition result is grain regions; and a false negative is the number of pixels in the sample image where the label is grain regions but the recognition result is background regions.

7. The imperfect grain identification method according to any one of claims 1-3, characterized in that, One way to normalize input features is to expand the three-dimensional input features, which include feature map height, feature map width, and channel dimension, into two-dimensional input features that include channel dimension and batch dimension, so as to achieve input feature normalization.

8. The imperfect grain identification method according to any one of claims 1-3, characterized in that, Based on the fusion results of the guiding branch output, spectral Mamba branch output, spatial Mamba branch output, and input features, the recognition results can be obtained in the following ways: The sum of the products of the guiding branch output and the spectral Mamba branch output and the spatial Mamba branch output is processed through a linear layer. The processed result is then added to the input features and input into the classifier of the recognition model to obtain the recognition result.

9. The imperfect grain identification method according to any one of claims 1-3, characterized in that, The visible light image of grain is obtained by performing preprocessing on the original grain image acquired by the visible light camera, including denoising, filtering, enhancement and normalization. The reflectance spectral data of grain grains are obtained by scanning the reflectance spectral data cube of grain grain samples line by line in the visible and near-infrared bands. The data is then processed by whiteboard correction to eliminate systematic errors caused by differences in ambient light and instrument response, spectral smoothing algorithm to reduce high-frequency noise, and normalization or standardization of each band.

10. An imperfect grain identification system, comprising a processor, wherein the processor stores executable program instructions, characterized in that, The executable program instructions are executed to implement the imperfect grain identification method according to any one of claims 1-9.