A Low-Cost Prediction Method for Amino Acids in Lotus Seeds Based on Multimodal Feature Importance Analysis and Key Feature Screening

CN122575551APending Publication Date: 2026-08-14JIANGSU UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

当前莲子氨基酸定量检测主流方法为高效液相色谱法(HPLC)、液相色谱-串联质谱法(LC-MS/MS),虽定量精度高,但存在样品前处理繁琐、检测周期长、试剂与仪器成本高、检测过程破坏样本、对操作人员专业要求高等缺陷,无法满足莲子产业端批量样本的快速、低成本、无损检测需求

Benefits of technology

[0023]1.本发明通过多模态数据融合技术,整合了NIRS光谱的化学组成信息与GLCM纹理的物理特征信息,弥补了单一模态数据信息维度不足的缺陷,双模态模型预测性能显著优于单一模态,三模态模型进一步丰富信息维度,结合贝叶斯优化,实现了莲子复杂基质下氨基酸含量的高精度定量预测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575551A_ABST
    Figure CN122575551A_ABST
Patent Text Reader

Abstract

This invention discloses a low-cost prediction method for lotus seed amino acids based on multimodal feature importance analysis and key feature screening. The method acquires NIRS, hyperspectral images, and amino acid data from lotus seed samples of different origins, and preprocesses them to obtain a feature dataset. Multiple sets of three-modal full-feature fusion prediction models are constructed. After training and testing using the feature dataset, Bayesian hyperparameter optimization is performed, and the optimal hyperparameters are selected to determine the optimal three-modal baseline model. Global feature importance analysis and non-physicochemical key feature screening are performed on this baseline model to determine the NIRS+GLCM key feature subset, which is used to train, validate, and verify the performance of the constructed low-cost prediction model. Then, the core amino acid content of lotus seed samples from unknown origins is predicted. This method is adaptable to lotus seeds from multiple origins, possessing both high accuracy and strong generalization, providing a new technical path for the non-destructive and rapid detection of functional components in food-medicine homologous agricultural products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of non-destructive testing of agricultural product quality, spectral analysis and machine learning technology, and specifically relates to a low-cost prediction method for amino acids in lotus seeds based on multimodal feature importance analysis and key feature screening. Background Technology

[0002] Lotus seeds are a traditional Chinese agricultural product that is both food and medicine. Amino acids, as their core nutritional and functional active ingredients, are the core indicators for lotus seed quality grading and nutritional value evaluation. Currently, the mainstream methods for quantitative detection of amino acids in lotus seeds are high-performance liquid chromatography (HPLC) and liquid chromatography-tandem mass spectrometry (LC-MS / MS). Although these methods offer high quantitative accuracy, they suffer from drawbacks such as cumbersome sample pretreatment, long detection cycles, high reagent and instrument costs, sample destruction during the detection process, and high professional requirements for operators. These methods cannot meet the needs of the lotus seed industry for rapid, low-cost, and non-destructive testing of large batches of samples.

[0003] Near-infrared spectroscopy (NIRS) can rapidly acquire chemical composition information of samples, while gray-level co-occurrence matrix (GLCM) texture analysis can characterize the physical structure of samples. Both have the advantages of being non-destructive and fast. However, existing lotus seed quality detection models based on a single modality suffer from problems such as limited information dimensions and insufficient feature complementarity, making it difficult to adapt to the high-precision quantitative prediction of trace amino acids in the complex matrix of lotus seeds. Existing multimodal fusion models mostly focus on classification tasks, with few studies on quantitative prediction of lotus seed amino acids. They generally rely on amino acid physicochemical detection data as model input, making it impossible to achieve low-cost prediction without physicochemical detection. At the same time, they lack global analysis of the contribution of multimodal features and accurate screening of redundant features, which limits the model's generalization ability and practicality. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention proposes a low-cost prediction method for lotus seed amino acids based on multimodal feature importance analysis and key feature screening. By integrating multimodal data, Bayesian hyperparameter optimization, global feature importance analysis, and screening of non-physicochemical key features, the method completely eliminates the model's input dependence on amino acid physicochemical detection data. High-precision, low-cost, and non-destructive rapid prediction of lotus seed amino acids can be achieved using only NIRS and hyperspectral image data.

[0005] The present invention achieves the above-mentioned technical objectives through the following technical means.

[0006] A low-cost method for predicting amino acids in lotus seeds based on multimodal feature importance analysis and key feature screening:

[0007] Near-infrared spectral data and hyperspectral image data of lotus seed samples from different origins were collected, and amino acid metabolite data of the lotus seed samples were measured; gray-level co-occurrence matrix texture feature data were obtained from the hyperspectral image data; and core amino acids were screened from the amino acid metabolite data.

[0008] Preprocessing of near-infrared spectral data, gray-level co-occurrence matrix texture feature data, and amino acid metabolite data yields the corresponding feature datasets.

[0009] Multiple sets of trimodal full-feature fusion prediction models are built. After training and testing using the feature dataset, Bayesian hyperparameter optimization is performed to select the optimal hyperparameters of the model and thus determine the optimal trimodal benchmark model.

[0010] Global feature importance analysis and non-physicochemical key feature screening are performed on the optimal trimodal benchmark model. The selected spectral key features and texture key features constitute the NIRS+GLCM key feature subset.

[0011] A low-cost prediction model that does not require amino acid physicochemical data input is constructed. The same algorithm and optimal hyperparameter combination as the optimal three-modal benchmark model are used. Based on the NIRS+GLCM key feature subset, the model training, validation and performance verification are completed.

[0012] For new lotus seed samples from different origins, near-infrared spectral data and hyperspectral image data were collected. After preprocessing, the key feature subset of NIRS+GLCM was determined and input into a validated low-cost prediction model to output the core amino acid content.

[0013] Furthermore, the amino acid metabolite data of the lotus seed sample to be tested were determined by liquid chromatography-tandem mass spectrometry, and the core amino acids included L-carnosine, γ-aminobutyric acid, and phosphate ethanolamine.

[0014] Furthermore, the near-infrared spectral data is preprocessed, specifically by using one of three preprocessing methods: smoothing the first derivative, multivariate scattering correction, or standard normal variable transformation.

[0015] Furthermore, the amino acid metabolite data were preprocessed, specifically by performing missing value imputation, ... The process involves transformation, probability quotient normalization, Z-score standardization, robust principal component analysis for quality control, and then feature correlation analysis.

[0016] Furthermore, the missing value filling uses the larger of 1 / 2 and 1% of the minimum non-missing value of each amino acid metabolite column as the filling value.

[0017] Furthermore, the multi-group trimodal full-feature fusion prediction model is constructed based on five algorithms: XGBoost, Bagged Trees, SVR, PLSR, and CNN.

[0018] Furthermore, the Bayesian hyperparameter optimization includes the number of decision trees, maximum tree depth, learning rate, and minimum number of leaf node samples.

[0019] Furthermore, the Bagged Trees trimodal full feature fusion model, which has the highest prediction determination coefficient and relative analysis error and the lowest prediction root mean square error, is selected as the optimal trimodal benchmark model.

[0020] Furthermore, the global feature importance analysis of the optimal three-modal benchmark model is specifically performed by grouping and calculating the total contribution of three modalities: near-infrared spectroscopy, gray-level co-occurrence matrix texture features, and amino acid metabolites.

[0021] Furthermore, the non-physicochemical key feature screening rule is as follows: taking the cross-validation prediction performance deviation ratio (RPD) for amino acid prediction modeling as the core evaluation index, and determining the candidate feature space to ensure the predictive ability of the model through variable projection importance analysis based on the RPD values ​​of different modal features; then, within the candidate feature space, calculating the Pearson correlation coefficient between each modal feature and the target variable one by one, retaining only the features that are positively correlated with the target variable, and eliminating redundant features that are negatively correlated or have no significant correlation.

[0022] The beneficial effects of this invention are:

[0023] 1. This invention integrates the chemical composition information of NIRS spectra and the physical feature information of GLCM texture through multimodal data fusion technology, which makes up for the lack of information dimension of single-modal data. The prediction performance of the dual-modal model is significantly better than that of the single-modal model, and the trimodal model further enriches the information dimension. Combined with Bayesian optimization, it realizes high-precision quantitative prediction of amino acid content under the complex matrix of lotus seeds.

[0024] 2. This invention presents a low-cost prediction method for lotus seed amino acids based on multimodal feature importance analysis and key feature screening. It innovatively conducts trimodal global feature importance analysis, completely eliminating the model's input dependence on amino acid physicochemical detection data, and screening out a subset of key spectral and texture features to construct a lossless, low-cost prediction model. This reduces the cost of lotus seed amino acid prediction and shortens the prediction cycle. Attached Figure Description

[0025] Figure 1 This is a graph showing the analysis of abnormal samples in the amino acid data of this invention;

[0026] Figure 2 This is a GLCM data extraction diagram from the present invention;

[0027] Figure 3 This is a screening chart of the top 10 most relevant target amino acids in this invention;

[0028] Figure 4 This invention presents a heatmap of Pearson correlation coefficients and a metabolic association network diagram of amino acids.

[0029] Figure 5 This is a screening diagram of the correlation intensity and positive correlation features of near-infrared spectroscopy, amino acids, and gray-level co-occurrence matrix features in this invention;

[0030] Figure 6 This is a flowchart illustrating the low-cost prediction process for lotus seed amino acids based on multimodal feature importance analysis and key feature screening, as described in this invention. Detailed Implementation

[0031] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but the scope of protection of the present invention is not limited thereto.

[0032] This embodiment provides a low-cost prediction method for lotus seed amino acids based on multimodal feature importance analysis and key feature screening, such as... Figure 6 As shown, the specific implementation process is as follows:

[0033] I. Experimental Samples and Multi-Source Data Acquisition, and Truth Value Calibration

[0034] 1. Experimental Samples: Pure lotus seed samples were selected from the three major lotus seed producing areas in my country (Xiangtan, Honghu and Jianning), with 24 samples from each area, for a total of 72 samples. After collection, impurities and moldy particles were removed from the samples, and they were stored at 4℃ for later use, covering the quality differences in different producing areas and ensuring the generalization ability of the model.

[0035] 2. NIRS Data Acquisition: Spectral data were acquired using a Shanghai Fuxiang Optics NIR1700 near-infrared spectrometer. The infrared spectral wavelength range was 900~1700nm, the spectral resolution was 7.8nm, the number of scans was 64, the integration time was 22ms, and the acquisition temperature was 24±2℃. Before acquisition, blank background calibration was performed, and then the lotus seed samples were placed in the sample cell. The diffuse reflectance spectral data of the samples were acquired. Each sample was acquired three times consecutively and the average value was taken to reduce the influence of spatial heterogeneity. Finally, NIRS spectral data of 72 samples were obtained.

[0036] 3. Hyperspectral Image Data Acquisition and GLCM Texture Feature Extraction: Visible / near-infrared hyperspectral imaging equipment from Taiwan Wuling Optics was used to acquire hyperspectral data in the 400–1000 nm wavelength range in reflectance mode. The ambient temperature during acquisition was 20–25℃, and the relative humidity was 30%–40%. After acquisition, spectral and image correction was performed using a standard PTFE black and white plate. The correction formula is as follows: In the formula: This is the corrected relative reflection information; This refers to the original sample information; The background information is completely dark, and its reflectivity is approximately 0%. For standard black and white board reflectance information, the corrected hyperspectral data is converted into grayscale images and standardized in size. The gray-level co-occurrence matrix (GLCM) is calculated based on 8 levels of grayscale quantification in four directions (0°, 45°, 90°, and 135°). Six texture features—mean, variance, contrast, entropy, correlation, and energy—are extracted to construct a GLCM texture feature dataset of 72 samples. Figure 2 As shown.

[0037] 4. Amino acid metabolite data determination and target amino acid screening: SCIEX QTRAP was used. The amino acid metabolite content of the sample was determined using a 6500+ liquid chromatography-tandem mass spectrometry (LC-MS / MS) instrument. The target amino acids were then screened using the following process: 50 mg ± 2.5 mg of lotus seed sample was accurately weighed into a centrifuge tube, and 500 μL of 70% methanol-water pre-cooled to -20°C and 400 μL of chloroform were added. The mixture was vortexed for 3 min. The mixture was then centrifuged at 12000 r / min for 10 min at 4°C. 250 μL of the supernatant was collected and allowed to stand at -20°C for 30 min. The mixture was then centrifuged again under the same conditions (4°C, 12000 r / min) for 10 min. 200 μL of the supernatant was processed using a protein precipitation plate before being analyzed. A total of 52 amino acid metabolites were detected. The Top 10 amino acids significantly correlated with the entire NIRS spectrum were screened, ultimately identifying L-carnosine-Acid, γ-aminobutyric acid (GABA), and phosphate ethanolamine as the three core target amino acids.

[0038] II. Multi-source data adaptation preprocessing

[0039] Targeted preprocessing was performed on the three types of collected data to construct a three-modal feature dataset:

[0040] 1. Amino acid metabolite data preprocessing: Perform missing value imputation, ... Logarithmic transformation, PQN (probability quotient) normalization, and Z-score standardization were performed. Missing values ​​were imputed using the larger of half the minimum non-missing value in the amino acid metabolite column and the 1% quantile. In this example, the original data had a missing value ratio of 0.00%. Robust principal component analysis was used for quality control, with a Hotelling T² threshold of 69.39. Four potential outliers were identified, and no severely outliers exceeded the T² threshold by more than 1.5 times. The amino acid metabolite dataset met the modeling requirements. Figure 1 As shown. Figure 3 , 4As shown, this invention uses Pearson correlation analysis to screen out the 10 amino acids and related metabolites with the strongest correlation to near-infrared spectral responses, and analyzes the variation of their correlation coefficients in the wavelength range of 900~1700nm. Among them, γ-aminobutyric acid (γ-Aminobutyric-Acid), phosphate ethanolamine (Phosphorylethanolamine), and L-carnosine (L-Carnosine) are the top three core metabolites with the highest correlation. These three not only show extremely strong correlation responses with near-infrared spectroscopy, but also have the highest inter-group Pearson correlation among the 10 amino acids, and can be used as core markers for linking spectral information and sample metabolic characteristics.

[0041] 2. GLCM texture feature preprocessing: Perform Z-score standardization on all texture features to eliminate the dimensional differences between different features and obtain a standardized texture feature dataset.

[0042] 3. NIRS data preprocessing: One of three preprocessing methods, namely smoothing the first derivative (SG-1st), multivariate scattering correction (MSC), or standard normal variable transformation (SNV), is used to suppress noise and enhance features in the original spectral data to obtain the preprocessed spectral dataset.

[0043] III. Construction of the Three-Modal Fusion Model and Bayesian Hyperparameter Optimization

[0044] 1. Dataset partitioning: The 72 preprocessed sample datasets (i.e., the amino acid metabolite dataset, texture feature dataset, and spectral dataset fused together after preprocessing) were divided into training and testing sets in a 7:3 ratio. Four-fold cross-validation was used to train the model and conduct initial screening of its generalization ability.

[0045] 2. Model Building: Based on five algorithms—XGBoost, Bagged Trees, SVR, PLSR, and CNN—three-modal full-feature fusion prediction models were built, resulting in a total of five candidate models. The basic configurations for each algorithm are as follows:

[0046] XGBoost: It uses LSBoost as the ensemble method, decision trees as weak learners, with an initial number of 50 trees, a maximum tree depth of 6, a learning rate of 0.1, and a minimum leaf node size of 5.

[0047] SVR: Uses a Gaussian kernel function, automatically optimizes the initial kernel scale, sets the box constraint to 10, performs a log1p transformation on the target value, and removes outliers based on the IQR rule;

[0048] Bagged Trees: A parallel bagged ensemble regression strategy is adopted, with an initial number of 100 trees, a minimum leaf node size of 22, a maximum number of splits of 20, and out-of-bag prediction enabled;

[0049] PLSR: It uses the one standard deviation rule to select the optimal number of principal components, and the pre-data diagnosis module alleviates the collinearity problem;

[0050] CNN: A lightweight dual-convolutional block regression network is used, with the following structure: input layer → convolutional layer 1 → batch normalization layer → ReLU activation function → max pooling layer → convolutional layer 2 → batch normalization layer → ReLU activation function → global average pooling layer → dropout layer → fully connected layer → regression output layer. The pooling kernel size is 4×1, and the padding method is same.

[0051] Five candidate models were trained on the training set and then tested on the test set.

[0052] 3. Bayesian Hyperparameter Optimization: A Bayesian optimization algorithm is introduced to automatically optimize the hyperparameters (including the number of decision trees (num Trees), maximum tree depth (max Depth), learning rate (learn Rate), and minimum leaf size) of five candidate models. The algorithm iterates 50 times, with the optimization objectives being to maximize the Rp² (prediction determination coefficient) and RPD (relative analysis error) on the test set and minimize the RMSEP (root mean square error of prediction). A dedicated optimization parameter space is set for different models. The mapping relationship between hyperparameters and model performance is fitted using a Gaussian process surrogate model, and the optimal hyperparameter combination is iteratively selected. In this embodiment, the optimal hyperparameter combination is: num Trees = 50, max Depth = 5, learn Rate = 0.10166, and min Leaf Size = 6.

[0053] 4. Optimal Benchmark Model Selection: Using Rp², RPD, and RMSEP as the core evaluation indicators, the model with the highest Rp² and RPD, and the lowest RMSEP, was selected as the optimal trimodal benchmark model. After optimization, the Bagged Trees trimodal fusion model showed the best overall performance, with Rp² values ​​of 0.96055, 0.98647, and 0.98067 for L-carnosine, γ-aminobutyric acid, and phosphoethanolamine, respectively; RPD values ​​of 4.7901, 8.2666, and 6.7476, respectively; and RMSEP values ​​of 90.680, 4.2159 × 10⁻⁶, and 4.2159 × 10⁻⁶, respectively. 1.3026× .

[0054] IV. Analysis of the Importance of Modal Features and Screening of Non-physicochemical Key Features

[0055] Global feature importance weight analysis was performed on the optimal Bagged Trees three-modal benchmark model. The total contribution was calculated by grouping the models into three modalities: NIRS, GLCM, and amino acid metabolites. Figure 5As shown, the results indicate that the combined contribution of NIRS and GLCM features to the prediction of target amino acids exceeds 90%, with NIRS features having a higher contribution rate and amino acid physicochemical features having a very low contribution rate. Therefore, NIRS is identified as the core contributing factor and GLCM as the complementary contributing factor.

[0056] Based on the contribution analysis results, the physicochemical characteristics of amino acid modalities were completely eliminated, and the following screening rules were applied:

[0057] Using the cross-validation prediction performance bias ratio (RPD) for amino acid prediction modeling as the core evaluation index, and based on the RPD values ​​of different features (including three types of features: near-infrared spectroscopy (NIRS), amino acid composition (AA), and gray-level co-occurrence matrix (GLCM), variable projection importance (VIP) analysis is used to determine the candidate feature space that can guarantee the model's predictive ability. Subsequently, within this candidate feature space, the Pearson correlation coefficient between each modal feature and the target variable (i.e., the predicted amino acid) is calculated one by one. Only features that are positively correlated with the target variable (spectral key features and texture key features) are retained, while negatively correlated and redundant features with no significant correlation are eliminated, thus completing the variable screening. The screened spectral key features and texture key features together constitute the NIRS+GLCM key feature subset.

[0058] V. Construction and Validation of a Low-Cost Prediction Model Without Amino Acid Input

[0059] A low-cost prediction model that does not require amino acid physicochemical data input is constructed. The Bagged Trees algorithm and the optimal hyperparameter combination, consistent with the optimal three-modal benchmark model, are used. Based on the selected NIRS+GLCM key feature subset, the model training, 4-fold cross-validation and generalization performance verification are completed.

[0060] The low-cost prediction model was tested and found to have the following predictive performance for the three core target amino acids: L-carnosine: Rp²=0.9221, RPD=3.5705, RMSEP=121.65; γ-aminobutyric acid: Rp²=0.9486, RPD=4.3327, RMSEP=804380.6; ethanolamine phosphate: Rp²=0.9421, RPD=4.1207, RMSEP=213307.3. All target amino acids met the accuracy requirements of Rp²≥0.92 and RPD>3.5, which is in line with the quantitative detection standards for agricultural product components.

[0061] VI. Low-cost, non-destructive prediction of amino acids in test samples

[0062] For new lotus seed samples from different production areas (Xiangtan, Honghu, Jianning), only NIRS data and hyperspectral image data need to be collected. After preprocessing (the data processing method in Part 2), the key feature subset of NIRS+GLCM is extracted (using the screening principles in Part 4, based only on the RPD values ​​of two types of features: near-infrared spectral NIRS and gray-level co-occurrence matrix GLCM). This subset is then input into the trained low-cost prediction model, which can directly output the predicted content of three core amino acids—L-carnosine, γ-aminobutyric acid, and phosphoethanolamine—in the sample without any amino acid physicochemical testing. The single-sample detection cycle is shortened by more than 95% compared to the traditional LC-MS / MS method, and the detection cost is reduced by more than 90%, truly achieving low-cost, non-destructive, and rapid detection of lotus seed amino acids.

[0063] The method in this embodiment can be directly adapted to the amino acid detection needs of lotus seeds from different producing areas such as Xiangtan, Honghu, and Jianning. It can also be extended to non-destructive quantitative prediction of amino acids in other food and medicinal agricultural products such as fox nuts and lilies.

[0064] The embodiments described above are preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Any obvious improvements, substitutions or modifications that can be made by those skilled in the art without departing from the essence of the present invention shall fall within the protection scope of the present invention.

Claims

1. A low-cost prediction method for lotus seed amino acids based on multimodal feature importance analysis and key feature screening, characterized in that: Near-infrared spectral data and hyperspectral image data of lotus seed samples from different origins were collected, and amino acid metabolite data of the lotus seed samples were measured; gray-level co-occurrence matrix texture feature data were obtained from the hyperspectral image data; and core amino acids were screened from the amino acid metabolite data. Preprocessing of near-infrared spectral data, gray-level co-occurrence matrix texture feature data, and amino acid metabolite data yields the corresponding feature datasets. Multiple sets of trimodal full-feature fusion prediction models are built. After training and testing using the feature dataset, Bayesian hyperparameter optimization is performed to select the optimal hyperparameters of the model and thus determine the optimal trimodal benchmark model. Global feature importance analysis and non-physicochemical key feature screening are performed on the optimal trimodal benchmark model. The selected spectral key features and texture key features constitute the NIRS+GLCM key feature subset. A low-cost prediction model that does not require amino acid physicochemical data input is constructed. The same algorithm and optimal hyperparameter combination as the optimal three-modal benchmark model are used. Based on the NIRS+GLCM key feature subset, the model training, validation and performance verification are completed. For new lotus seed samples from different origins, near-infrared spectral data and hyperspectral image data were collected. After preprocessing, the key feature subset of NIRS+GLCM was determined and input into a validated low-cost prediction model to output the core amino acid content.

2. The low-cost prediction method for lotus seed amino acids according to claim 1, characterized in that, The amino acid metabolite data of the lotus seed samples were determined by liquid chromatography-tandem mass spectrometry. The core amino acids included L-carnosine, γ-aminobutyric acid, and ethanolamine phosphate.

3. The low-cost prediction method for lotus seed amino acids according to claim 1, characterized in that, The near-infrared spectral data are preprocessed using one of three methods: smoothing the first derivative, multivariate scattering correction, or standard normal variable transformation.

4. The low-cost prediction method for lotus seed amino acids according to claim 1, characterized in that, The amino acid metabolite data were preprocessed, specifically by performing missing value imputation, ... The process involves transformation, probability quotient normalization, Z-score standardization, robust principal component analysis for quality control, and then feature correlation analysis.

5. The low-cost prediction method for lotus seed amino acids according to claim 4, characterized in that, The missing value filling uses the larger of 1 / 2 and 1% of the minimum non-missing value in each amino acid metabolite column as the filling value.

6. The low-cost prediction method for lotus seed amino acids according to claim 1, characterized in that, The multi-group trimodal full-feature fusion prediction model is constructed based on five algorithms: XGBoost, Bagged Trees, SVR, PLSR, and CNN.

7. The low-cost prediction method for lotus seed amino acids according to claim 6, characterized in that, The Bayesian hyperparameter optimization includes the number of decision trees, maximum tree depth, learning rate, and minimum number of leaf node samples.

8. The low-cost prediction method for lotus seed amino acids according to claim 7, characterized in that, The Bagged Trees trimodal full feature fusion model, which has the highest prediction determination coefficient and relative analysis error and the lowest prediction root mean square error, is selected as the optimal trimodal benchmark model.

9. The low-cost prediction method for lotus seed amino acids according to claim 1, characterized in that, The global feature importance analysis of the optimal three-modal benchmark model is specifically performed by grouping and calculating the total contribution of three modalities: near-infrared spectroscopy, gray-level co-occurrence matrix texture features, and amino acid metabolites.

10. The low-cost prediction method for lotus seed amino acids according to claim 9, characterized in that, The non-physicochemical key feature screening rule is as follows: taking the cross-validation prediction performance deviation ratio (RPD) for amino acid prediction modeling as the core evaluation index, and determining the candidate feature space to ensure the predictive ability of the model through variable projection importance analysis based on the RPD values ​​of different modal features; then, within the candidate feature space, calculating the Pearson correlation coefficient between each modal feature and the target variable one by one, retaining only the features that are positively correlated with the target variable, and eliminating redundant features that are negatively correlated or have no significant correlation.