A method and system for rapid nondestructive detection of lipid content and deterioration degree of red pine seeds based on hyperspectral imaging and deep learning

By combining hyperspectral imaging with deep learning, an improved one-dimensional dilated convolutional network with a dynamic weight allocation module was designed to construct a deep learning model. This solved the problem of rapid and non-destructive detection of lipid content and oxidative deterioration in red pine nuts, achieving efficient, non-destructive, and rapid detection results.

CN121207914BActive Publication Date: 2026-05-05NORTHEAST AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHEAST AGRICULTURAL UNIVERSITY
Filing Date
2025-10-29
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing methods for detecting lipid content and oxidative deterioration in red pine nuts suffer from problems such as long detection cycles, cumbersome operation, difficulty in large-scale detection, high dependence on organic solvents, and destructive testing of the sample, making it difficult to meet the needs for rapid, green, high-coverage, and non-destructive testing.

Method used

By employing a method based on hyperspectral imaging and deep learning, near-infrared spectral data of red pine nuts are collected. An improved one-dimensional dilated convolutional network with a dynamic weight allocation module is designed to construct a deep learning model, enabling rapid and non-destructive detection of lipid content and the degree of oxidative deterioration.

Benefits of technology

This method enables efficient, non-destructive, and rapid detection of lipid content and oxidative deterioration in red pine nuts, improving detection efficiency and intelligence, and solving problems such as strong nonlinearity, weak signal, and complex background in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121207914B_ABST
    Figure CN121207914B_ABST
Patent Text Reader

Abstract

This invention relates to a rapid, non-destructive testing method and system for lipid content and deterioration degree in red pine nuts based on hyperspectral imaging and deep learning. The invention belongs to the field of non-destructive testing technology for agricultural and forestry product quality and safety. It addresses the problems of existing methods for detecting lipid content and oxidation degree in red pine nut kernels, which generally suffer from complex operation procedures, high costs for large-scale testing, long processing times, and difficulty in achieving full-process testing during storage and transportation. Key technical points: The original near-infrared spectral data of pine nut samples are obtained through hyperspectral acquisition; the true values ​​of lipid reference content and reference degree of oxidation deterioration of pine nut samples at different sampling times are determined, and a database is established based on these values. The original near-infrared spectral data of red pine nuts are preprocessed, and the data from different batches of red pine nut samples are divided into training and validation sets. An improved one-dimensional dilated convolutional network based on a dynamic weight allocation module is designed to construct a deep learning model suitable for the spectral feature analysis of red pine nut kernels. The constructed model is trained using a backpropagation algorithm, and the network weight parameters are updated in reverse by minimizing the loss function. When the constructed model meets the requirements for rapid detection, it can be used for efficient and non-destructive simultaneous detection of lipid content and oxidative rancidity in red pine nuts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of non-destructive testing technology for the quality and safety of agricultural and forestry products, specifically relating to a rapid and non-destructive testing method and system for the synergistic detection of lipid content and degree of oxidative deterioration in the kernel of Korean pine seeds based on hyperspectral imaging and deep learning. Background Technology

[0002] With increasing public emphasis on food quality and safety, precise control of nutritional components, and sustainable development, the lipid content and acidification level of red pine nuts, as a high-oil, high-nutrient economic agricultural and forestry product, have become key indicators requiring close monitoring in the industry chain. Furthermore, lipids, as an important substrate for oxidative metabolism, exhibit an inverse relationship with oxidative metabolic (acidic) products. Red pine nuts are rich in unsaturated fatty acids, vitamin E, and other functional lipids and trace elements, possessing high nutritional value and market potential, and are widely used in food processing and vegetable oil extraction. However, during storage, transportation, and processing, they are prone to oxidative decomposition, producing peroxides and potent carcinogens, seriously threatening food safety and consumer health. Therefore, establishing an efficient and accurate method for detecting the lipid content and oxidative deterioration level of red pine nuts is of significant practical importance and application value for improving product quality control, promoting the deep processing of red pine nuts, and ensuring the sustainable development of the industry. How to achieve simultaneous quality detection of lipids and oxidation levels in red pine nuts without destructive testing by integrating spectral analysis and deep learning intelligent modeling technology, meeting the urgent needs of modern agricultural product and food safety supervision and high-throughput quality evaluation, is a pressing requirement.

[0003] Existing methods for detecting lipid content and oxidative rancidity levels in red pine nuts mainly include Soxhlet extraction, acid hydrolysis + Soxhlet extraction, acid value method, peroxide value method (POV), TBARS method, carbonyl value determination, colorimetry, and gas chromatography-mass spectrometry (GC-MS). Although these methods have high accuracy, they generally suffer from problems such as long detection cycles, cumbersome operation, difficulty in large-scale and high-frequency detection, high dependence on organic solvents, and destructive testing of the tested samples. These limitations restrict their application in rapid, batch detection for quality grading and screening of pine nut raw materials during storage and transportation, and make it difficult to meet the practical needs of efficient, non-destructive, high-coverage, and green detection.

[0004] Near-infrared spectroscopy (NIRS), a non-destructive optical analysis method based on the molecular-level vibrational characteristics of the sample, boasts advantages such as fast detection speed, simple operation, no need for chemical pretreatment, and suitability for online real-time monitoring. It has been widely used for the quantitative and qualitative analysis of lipids, proteins, moisture, toxins, and other material indicators in agricultural products and food. However, the spectral response sites of lipid content and oxidation markers in pine nuts are complex and somewhat similar. Overlapping spectral bands between samples and background interference significantly affect detection accuracy. In other words, existing detection technologies for detecting lipid content and oxidation marker levels in pine nuts have not effectively solved problems such as strong nonlinear relationships, weak signals, and complex backgrounds. Furthermore, the limited receptive field and insufficient multi-scale spectral feature extraction capabilities of one-dimensional convolutional network NIRS models, coupled with the difficulty in modeling the mapping between high-dimensional spatial structures and material components, make NIRS spectral modeling a technical bottleneck.

[0005] Therefore, it is imperative to find a way to construct a neural network model with a multi-layered structure through algorithms to automatically extract and effectively express nonlinear features and high-order information in hyperspectral data, and to solve the problems of strong nonlinear relationships, weak signals, and complex backgrounds in the detection of lipid content and oxidation marker levels in red pine nuts as soon as possible. Summary of the Invention

[0006] To address the problems of existing methods for detecting lipid content and oxidation level in red pine nuts, such as complex operation procedures, high cost for large-scale testing, long time consumption, difficulty in detecting throughout the storage and transportation process, and the need to use organic solvents or chemical reagents, which cannot meet the requirements of rapid, green, high-coverage, and non-destructive testing, this invention provides a method and system for simultaneous and rapid detection of lipid content and oxidation level in red pine nuts based on near-infrared spectroscopy and deep learning models.

[0007] To address the aforementioned problems, this invention provides a rapid and non-destructive method for detecting lipid content and deterioration degree in pine nuts based on hyperspectral imaging and deep learning. In short, it is a rapid method for detecting lipid content and oxidative rancidity in the edible parts of pine nuts. This invention acquires raw near-infrared spectral data of pine nut samples through hyperspectral acquisition; determines the true lipid reference content and reference degree of oxidative deterioration of pine nut samples at different sampling times, and establishes a database based on these. The acquired raw near-infrared spectral data of pine nuts are preprocessed, and pine nut sample data collected in different batches are divided into training and validation sets. An improved one-dimensional dilated convolutional network based on a dynamic weight allocation module is designed to construct a deep learning model suitable for the spectral feature analysis of pine nut kernels. The constructed model is trained using a backpropagation algorithm, and the network weight parameters are updated in reverse by minimizing the loss function. When the performance of the constructed model meets the requirements for rapid detection, it is used for efficient and non-destructive simultaneous detection of lipid content and oxidative rancidity in pine nut kernels.

[0008] This invention acquires the original near-infrared spectral data of pine nut samples through hyperspectral collection, and determines the true value of lipid reference content and the reference degree of oxidative deterioration of pine nut samples at different sampling times according to national standard methods or commonly used laboratory methods (Soxhlet extraction to measure lipid content, and colorimetric detection of oxidative rancidity using acid value / peroxide value test paper). A database is then established based on this data to ensure correspondence with the original spectral data measured in S2. The degree of deterioration refers to the degree of oxidative rancidity.

[0009] This invention also provides a rapid and non-destructive detection system for lipid content and deterioration degree of pine nuts based on hyperspectral imaging and deep learning. Specifically, it includes: a raw data acquisition unit for pine nut samples, used to acquire hyperspectral data of pine nut samples at different times and extract raw near-infrared spectral data, oxidation degree, and lipid content of the pine nut samples, and establish a database based on the aforementioned accurate data; a data preprocessing unit, used to preprocess and extract features from the acquired raw near-infrared spectral data of the pine nut samples, and divide it into training and validation sets; and a deep learning model building unit, which constructs a multi-task neural network model through a convolutional neural network, incorporating a dilated convolution mechanism and a dynamic weight allocation module, to fit the near-infrared spectral data and lipid content data of the training set pine nut samples, and establish a database. A multi-task learning model I for lipid content in pine nuts; a multi-task learning model II for detecting the degree of oxidative rancidity in pine nuts by classifying the near-infrared spectral data and oxidative rancidity data of the training set samples; a model validation and optimization unit, which minimizes the loss function through backpropagation algorithm, updates the weight parameters of model I and model II in reverse, and uses the validation set to validate and optimize the hyperparameters of the model to improve the prediction accuracy and performance of the model. After multiple validations and optimizations, the final detection model is obtained; a rapid detection unit for pine nut samples, which obtains the original spectral data of the sample to be tested, inputs the original spectral data into the multi-task learning models I and II for lipid content and oxidative rancidity in pine nuts, and outputs the predicted values ​​of lipid content and oxidative rancidity in pine nuts.

[0010] This invention offers the following beneficial technical effects: It establishes a rapid detection method for lipid content and oxidative rancidity in pine nuts based on near-infrared spectroscopy and deep learning. First, it collects raw near-infrared spectral data of pine nut samples at different times, including the degree of oxidative rancidity and lipid content under corresponding conditions. After preprocessing, the samples are randomly divided into training and validation sets. The training set spectral features are automatically extracted using a convolutional neural network, and a dilated convolutional neural network is introduced to expand the model's receptive field and enhance its multi-scale spectral feature extraction capability. The model also incorporates a dynamic weight allocation module to perform feedback adjustment, enhancing the trade-off between dynamic selection and fine-tuning. Finally, a weighted feature representation is output, reflecting the relative importance of different spectral channels, which is particularly helpful in identifying key features of lipid content and oxidative rancidity in near-infrared spectral data. This weighted feature representation is then subjected to nonlinear transformation and feature abstraction through a feedforward neural network layer, outputting predicted values. The final detection model is obtained by minimizing the loss function through backpropagation and iterating and optimizing continuously. This model can be used for the rapid detection of lipid content and oxidation degree in pine nuts.

[0011] The present invention designs an improved one-dimensional dilated convolutional network based on a dynamic weight allocation module for near-infrared spectral modeling. It effectively solves the problems of limited receptive field and insufficient multi-scale spectral feature extraction capability in near-infrared spectral models based on one-dimensional convolutional networks. At the same time, near-infrared spectroscopy has the problem of difficulty in modeling the mapping between high-dimensional spatial structure and material composition.

[0012] Furthermore, this invention simultaneously improves the balance between dynamic selection and fine-tuning of effective bands. Specifically, it proposes a module based on a gated unit coupled with a dynamic weight allocation mechanism, called the Dynamic Information Regulation Unit (Dynamic Information Regulation Unit). This unit mainly includes the following four stages: First, a feature encoding mechanism is designed to extract spatial structure features in the near-infrared spectrum, improving the model's perception of spectral spatial distribution; second, a nonlinear mapping relationship between spatial structure features and the internal composition of pine nuts is established using a fully connected network; third, a generative network module is introduced to predict discriminative sampling points in the high-dimensional spectral space based on spatial structure features, achieving an effective correspondence between the input space and the spectral feature space; finally, a spectral resampling mechanism is designed to extract effective band information from the spectrum at the predicted sampling points. Simultaneously, this method introduces a dilated convolution structure to extract spectral sequence features at different receptive fields and scales with different dilation rates. Adaptive selection of spectral feature bands is achieved based on the dynamic weight mechanism.

[0013] This invention introduces a deep learning (DL) algorithm to construct a multi-layered neural network model, enabling the automatic extraction and effective representation of nonlinear features and higher-order information in hyperspectral data. The method is specifically designed for high-throughput near-infrared spectral data of red pine kernels, which are characterized by high dimensionality, high noise, and multivariate coupling. It possesses excellent feature recognition and parameter analysis capabilities, accurately identifying key spectral bands related to lipid content and oxidative degradation. This allows for non-destructive, rapid, and high-precision simultaneous prediction of these indicators, significantly improving detection efficiency and intelligence.

[0014] This invention designs a network structure based on the collected raw near-infrared spectral data of red pine nuts. In other words, this invention fully considers the unique characteristics of the near-infrared spectrum of red pine nuts and designs a targeted network structure to achieve the specific purpose of this invention. It effectively solves the problems of strong nonlinear relationship, weak signal and complex background in the detection of lipid content and oxidation marker level of red pine nut kernels in the prior art.

[0015] In summary, the present invention, by employing the above-mentioned technical means, effectively solves the problems of existing methods for detecting lipid content and oxidation degree in red pine nuts, which generally involve complex operation procedures, high cost for large-scale testing, long time consumption, and difficulty in achieving full-process testing during storage and transportation. Attached Figure Description

[0016] The present invention will be further described below with reference to the accompanying drawings and specific embodiments:

[0017] Figure 1 This is a flowchart of the method for detecting the lipid content and oxidative rancidity of pine nuts according to the present invention (the steps on the right side of the flowchart are further refinements of the corresponding steps on the left side).

[0018] Figure 2 These are the original spectral curves of 500 pine nut samples from Example 1;

[0019] Figure 3 These are the spectral curves of 500 pretreated pine nut samples from Example 1;

[0020] Figure 4 This is a diagram of the dilated convolutional multitasking network structure based on the dynamic weight allocation module of the present invention.

[0021] Figure 5 This is a training scatter plot of the pine nut lipid content determined by the national standard method and the predicted values ​​by the spectral model in Example 2;

[0022] Figure 6 This is a confusion matrix of the classification results of the acid content range of pine nuts in Example 1;

[0023] Figure 7 Performance comparison of three model structures (ordinary convolution, dilated convolution, and dilated convolution combined with dynamic weight allocation module) in ablation experiments. Detailed Implementation

[0024] To more clearly and comprehensively illustrate the technical solution of the present invention, the following will describe it in conjunction with specific embodiments and examples, as well as the appendix. Figure 1-7 This invention will be described in detail below. It should be understood that the embodiments described are merely illustrative of the technical content of this invention and do not constitute a limitation on the scope of protection of this invention. Any equivalent improvements, modifications, or substitutions made by those skilled in the art based on the concept of this invention without inventive effort should be considered to fall within the scope of protection of this invention.

[0025] The technical solution adopted by this invention to solve the above-mentioned detection problem is: a rapid detection method for lipid content and oxidative rancidity in pine nut products, comprising:

[0026] S1 places pine nut samples from the same harvest batch in the sample bin for proper storage to ensure that the oxidation level of pine nut samples at the same sampling time point can be kept in a uniform state.

[0027] S2. Pine nut samples were taken out every 2 days and placed on the experimental table, ensuring that there was no stacking or obstruction between the pine nut samples, and then hyperspectral scanning was performed.

[0028] S3 acquires the original near-infrared spectral data of pine nut samples through hyperspectral acquisition; the true value of lipid reference content and the reference degree of oxidative deterioration of pine nut samples at different sampling times are determined according to national standard methods or commonly used laboratory methods (Soxhlet extraction to measure lipid content; acid value / peroxide value test paper for colorimetric detection of rancidity), and a database is established based on this to ensure that it corresponds with the original spectral data measured in S2.

[0029] S4 preprocesses the raw near-infrared spectral data of red pine nuts collected, and divides the red pine nut sample data collected in different batches into training set and validation set. The division process conforms to the random sampling principle and the general dataset setting principle of machine learning modeling research.

[0030] Based on a one-dimensional convolutional neural network, this invention designs an improved one-dimensional dilated convolutional network with a dynamic weight allocation module to construct a deep learning model suitable for analyzing the spectral features of red pine nut kernels. The model enhances its ability to model long-range dependencies in spectral sequences by introducing a dilated convolution mechanism and achieves adaptive adjustment of feature channels by combining a dynamic weight mechanism. Using this model, near-infrared spectral data of training samples are fitted and regressed with corresponding lipid contents. Simultaneously, the degree of oxidative rancidity of the samples is modeled hierarchically, thereby constructing a multi-task learning framework capable of simultaneously predicting lipid content and identifying oxidation levels, thus achieving efficient and collaborative detection of quality indicators for red pine nut kernels.

[0031] S6 uses the backpropagation algorithm to train the constructed Model I and Model II, and updates the network weight parameters in reverse by minimizing the loss function; the model performance is evaluated using the validation set partitioned in step S4, and the key hyperparameters of the model are tuned based on the evaluation results to improve the model's prediction accuracy and robustness; through multiple rounds of cross-validation and parameter optimization, a convergent and stable final detection model is obtained; when the model's coefficient of determination R on the validation set... 2 When the root mean square error (RMSE) is greater than or equal to 0.75 and less than or equal to 0.2, the constructed model is considered to meet the requirements for rapid detection and can be used for efficient and non-destructive simultaneous detection of lipid content and oxidative rancidity in red pine nuts.

[0032] S7 prepares the pine nut samples to be tested, places them in the sample preparation chamber, takes out the samples at regular intervals to scan the hyperspectral images and obtain the raw spectral data, and then performs Soxhlet extraction experiments and acid value / peroxide value tests; inputs the raw spectral data into the final detection model in S6, and the output results simultaneously obtain the predicted results of pine nut lipid content and oxidative rancidity.

[0033] This invention aims to train a model and perform non-destructive testing on the lipid content and oxidative degradation level of the edible portion (peeled and kernel-less) of red pine nuts. The technical solution adopted includes the following steps:

[0034] S1 Lighting Treatment and Environmental Control

[0035] Shells were removed from red pine nuts harvested from the same batch, and the pine kernel samples were placed in an environmentally controlled sample chamber and allowed to stand under sunlight. Throughout the experiment, the temperature was controlled at 25±2℃ and the humidity at 45±5%, and environmental parameter data were collected every 15 minutes using a HOBO U23-001 temperature and humidity recorder to ensure that the sample preparation was carried out in a suitable environment.

[0036] S2 Sample Acquisition Array Construction and Hyperspectral Acquisition

[0037] The selected samples were evenly arranged on a perforated black acrylic plate to avoid sample shaking caused by platform movement during the acquisition process. They were arranged in a 5×5 array (plate size 15cm×15cm) to optimize the regularity of the hyperspectral scanning acquisition. Hyperspectral images of the pine nut samples were scanned, exported, and saved. The samples were then placed in a sealed bag and stored properly in the dark.

[0038] S3 spectral data acquisition; physicochemical testing and label value construction

[0039] Regions of interest (ROIs) were selected from the exported hyperspectral images of pine nut samples, and near-infrared spectral data were acquired for these regions. Immediately after hyperspectral acquisition in S2, physicochemical analysis was performed on these samples to construct training labels. Random, sufficient samples were taken from each acrylic plate and divided into two portions. The first portion was crushed and hexane was added to extract lipids. After standing for 10 minutes, colorimetric detection was performed using MLbio acid value / peroxide value test paper. The reaction time of the acid value test paper was controlled within 5 minutes to form a mapping dataset between the spectrum and lipid oxidative rancidity indicators. The second portion underwent Soxhlet extraction according to national standard methods to form a mapping dataset between the spectrum and lipid content indicators. The acidification degree and lipid content of the pine nut samples at different oxidation times were measured to establish a database. To ensure a one-to-one correspondence with the original spectral data measured in S2, the lipid content and oxidation level data obtained from physicochemical testing were labeled and matched with the corresponding hyperspectral image data to construct a complete training dataset.

[0040] S4 Data Preparation and Modeling

[0041] The raw near-infrared spectral data of the collected pine nut samples were preprocessed and randomly divided into training and validation sets to provide basic data support for subsequent deep learning modeling and research on the lipid content and oxidative deterioration degree of pine nut kernels.

[0042] Establishment of S5 Neural Network Model

[0043] A multi-task network was used to construct a deep learning model by introducing a module architecture based on dynamic weight allocation and a dilated convolution mechanism. Using near-infrared spectral data and lipid content data of pine nut samples in the training set, a multi-task detection model I for pine nut lipid content was established. Using near-infrared spectral data and oxidation degree level data of pine nut samples in the training set, a multi-task detection model II for oxidative rancidity in pine nuts was established.

[0044] S6 Loss Function

[0045] By minimizing the loss function using the backpropagation algorithm and updating the weight parameters in reverse, two models with different weight parameters were obtained after multiple iterations and optimizations. The prediction accuracy and performance of the two detection models were then validated using near-infrared spectral data, lipid content data, and oxidation level data from the validation set samples. The validated models, if R... 2 With a mean square error (RMSE) greater than or equal to 0.75 and a root mean square error (RMSE) less than or equal to 0.2, this model can be used for the rapid detection of lipid content and oxidative rancidity in pine nuts.

[0046] As an improvement to the above technical solution, in step S1, a sufficient amount of pine nut kernel samples are placed in an artificially controlled environment chamber with supplemental lighting function, and periodic sunlight and UV-C light treatment is carried out: when natural light cannot meet the requirements, fluorescent lamps are used for supplemental lighting, and UV-C irradiation is carried out once every 2 hours, followed by a 1-hour pause, and 4 consecutive cycles of irradiation are completed to form an intermittent irradiation scheme to avoid quality changes caused by heat accumulation, so that the samples can reach a regular gradient rancidity level in a short time, thereby shortening the oxidative deterioration experimental cycle and accelerating the overall research process.

[0047] As an improvement to the above technical solution, in step S2, the designed shape of the holes in the black background acrylic plate ensures the stability of the pine nut sample, and the side with the larger area is used as the acquisition surface. After spectral scanning, the region of interest (ROI) is selected, and the average spectral value in the ROI is used as the original spectral data of this pine nut sample. To ensure the accuracy and robustness of the model, at least 500 sets of spectral data are collected. The experiment uses the HyperSpec® VNIR-A hyperspectral imaging system from Headwall Technologies, equipped with a CCD imaging lens, an adjustable halogen light source, a lifting platform, and a stepper motor-controlled conveyor platform. The acquired spectral band range is 400–1000 nm, containing 203 bands, with a spectral resolution of 2–3 nm, a sampling interval of 3 nm, and a spatial resolution of 0.15 nm. The system is controlled by HyperSpec III software to acquire sample images and spectral information.

[0048] As an improvement to the above technical solution, in step S3, starting from the initial moment of the experiment (Day 0), sample spectra are collected every 48 hours, and Soxhlet extraction experiments and acid value / peroxide value colorimetric tests are performed (the lipid content in the sample is determined according to the extraction method in GB5009.5—2010). The sampling experiment continues until the end of Day 30. The 30-day time series data collection simulates the long storage and transportation cycle of red pine nuts, which is closer to the actual transportation situation in forest areas. At the same time, it makes the sample preparation gradient and the physicochemical indicators more differentiated.

[0049] As an improvement to the above technical solution, in step S4, the acquired raw near-infrared spectral data is preprocessed to eliminate background interference, improve the signal-to-noise ratio, and enhance the model's ability to extract key features. The preprocessing methods include, but are not limited to, one or more combinations of first-derivative processing, second-derivative processing, wavelet transform, Z-score normalization, Savitzky-Golay (SG) smoothing, multivariate scattering correction (MSC), standard normal variable transformation (SNV), baseline drift correction, trend term removal, moving average filtering, median filtering, and minimum-maximum normalization. The selected preprocessing algorithm can be flexibly adjusted according to spectral characteristics and sample type to effectively extract key information and suppress redundant information in the spectral data, thereby improving the accuracy and stability of subsequent modeling.

[0050] S4: Data Preprocessing

[0051] S4.1: Data Collection and Cleaning. Before undertaking any machine learning task, data collection and cleaning are essential. For near-infrared spectral data, spectral information is typically obtained from different time points or under different processing conditions. Since spectral data often contains noise, missing values, and outliers, data cleaning is a necessary step. Common methods for handling missing values ​​include: Linear interpolation: For the location of a missing value, the linear relationship between the missing value and the sample mean is used to estimate the missing value. Imputation: When the proportion of missing data is small and interpolation is not feasible, the mean of the corresponding feature can be used for imputation.

[0052] S4.2: Perform data normalization. To improve the stability and convergence speed of model training, the input data is usually standardized. Standardization methods include Z-Score standardization.

[0053]

[0054] in, It is the raw data. It is the mean of the data. This is the standard deviation of the data. Standardized data has zero mean and unit variance, which helps prevent training instability caused by excessively large numerical ranges for certain features.

[0055] S4.3: For processing one-dimensional sequence data, especially in near-infrared spectroscopy analysis, this implementation uses a sliding window technique to segment the time series data. Assume the input data is... The sliding window operation then divides the data into multiple overlapping subsequences:

[0056]

[0057] in, This refers to the window size. Each window's data serves as input to the model.

[0058] As an improvement to the above technical solution, S5 includes the following specific steps: S5.1 Using the SPXY algorithm to randomly divide the samples into training and validation sets; S5.2 Establishing a deep learning-based regression model; S5.3 Automatically extracting the spectral features of the preprocessed training set through a dilated convolutional neural network to expand the model's receptive field and improve its multi-scale spectral feature extraction capability; S5.4 Introducing a dynamic weight allocation module to perform feedback adjustment, enhancing the model's trade-off between dynamic selection and fine-tuning, and finally outputting predicted values ​​for pine nut lipid content and oxidation degree; S5.5 Further providing a multi-task learning-based pine nut component prediction method, by introducing a dynamic weight mechanism and an absorption peak position mapping mechanism, to achieve accurate modeling and collaborative prediction of multiple band features in the input spectrum to multiple output tasks.

[0059] As an improvement to the above technical solution, step S5.3 expands the receptive field of the convolution kernel by introducing a dilation rate, effectively capturing a wider range of dependencies without increasing computational cost. One-dimensional dilated convolution is a special type of convolution operation. This method performs particularly well when processing near-infrared spectral signals. The following is the overall workflow based on one-dimensional dilated convolution.

[0060] S5.3.1: Design Process of Dilated Convolutional Layers and Construction of One-Dimensional Dilated Convolutional Network Model: The core advantage of dilated convolution lies in its ability to utilize the void ratio... This expands the receptive field of the convolution kernel, thereby capturing dependencies over longer periods or over a wider range.

[0061] S5.3.2: For one-dimensional dilated convolution, assume the input signal is... The kernel size is The output calculation formula for dilated convolution is:

[0062]

[0063] in, The void ratio, The weights of the convolution kernel, For bias terms, It is the output index, and This represents the spacing between elements of the convolution kernel. By choosing an appropriate dilation rate *r*, the receptive field can be effectively increased without increasing the size of the convolution kernel, thus avoiding a surge in computational complexity. Dilated convolution has the ability to capture a wider range of feature dependencies without increasing computational cost. This makes it particularly suitable for processing the near-infrared spectra of pine nuts with long-term dependencies.

[0064] S5.3.3: Stacked dilated convolutional layers are used to construct a deep feature extraction network in order to further expand the receptive field. Each convolutional layer extracts feature information at different scales based on its dilation rate. For example, the first layer has a dilation rate of 1, the second layer has a dilation rate of 2, and the third layer has a dilation rate of 4. As the dilation rate increases, the receptive field expands exponentially, thereby capturing temporal dependencies over longer distances.

[0065] S5.3.4: Activation Function Selection. A non-linear activation function is typically applied after each convolutional layer. The most commonly used is ReLU (Rectified Linear Unit), which is defined as:

[0066]

[0067] The ReLU function introduces non-linearity into the network, enabling the model to learn more complex patterns. It also effectively alleviates the vanishing gradient problem and accelerates the training process of deep networks.

[0068] As an improvement to the above technical solution, the goal of the S5.4 dynamic weight allocation module is to automatically identify key absorption peaks in the input spectrum and establish a mapping relationship with the corresponding band response positions of the output task. Ultimately, it establishes a correspondence between the absorption peak and wavelength position in the pine nut output characteristic band and the absorption peak and wavelength position in the input characteristic band. This allows the network to automatically focus on spectral positions related to functional group characteristic vibrations (chemical bond vibration characteristics) or molecular absorption peaks during training (ensuring the network automatically focuses on spectral positions with clear physical meaning during training). Functional group characteristic vibrations (chemical bond vibration characteristics) can be obtained from the original near-infrared spectral data of pine nut samples acquired through hyperspectral analysis; a specific region of the near-infrared spectrum is called a spectral band. Functional group characteristic vibrations refer to the specific vibrational frequencies of different functional groups in the infrared spectrum.

[0069] S5.4.1: As an improvement to the above technical solution, S5.4.1 designs a characteristic spectral localization network structure.

[0070] S5.4.1.1: Define the output characteristic spectrum as , which represents the wavelength position after feature encoding.

[0071] The feature spectrum is mapped to the input feature using a multilayer perceptron for feature encoding, and the corresponding input feature spectrum is calculated as follows. :

[0072]

[0073] S5.4.1.2: Based on Design of a Localization Network for Multichannel Feature Encoding:

[0074] This module extracts from the input spectrum. Extract characteristic band information related to pine nut mold and regress it. Multilayer perceptrons encode spectral information features in the form of a function:

[0075]

[0076] in, It refers to a learnable function (such as a CNN with fully connected layers), which, in the context of this patent, means the use of a multilayer perceptron. Feature encoding of spectral bands:

[0077] S5.4.1.3: (Process A, when the lengths of different channels in the MLP network are consistent): Based on the spliced ​​MLP:

[0078]

[0079] GAP is an abbreviation for Global Average Pooling. It uses a global average pooling layer to filter characteristic bands from the output bands and uses a (Batch Normalization) normalization layer to merge the pooled characteristic bands.

[0080] S5.4.1.3: (Process B, when the lengths of different channels in the MLP network are inconsistent): Extract features independently for each branch:

[0081] ,

[0082] S5.4.2: As an improvement to the above technical solution, S5.4.2 designs a characteristic spectrum fusion structure.

[0083] S5.4.2.1: Determine if the dimensions of the characteristic bands are consistent. If the characteristic dimensions are inconsistent, use... Convolution for dimension alignment:

[0084]

[0085] S5.4.2.2: Using gated weighted fusion of spectral bands:

[0086]

[0087] S5.4.3: As an improvement to the above technical solution, S5.4.3 designs a generator structure.

[0088] This is used to calculate the sampling position in the input spectrum for each near-infrared spectral output. That is, to generate a sampling spectral network:

[0089]

[0090] in This represents the set of characteristic spectra of the target output.

[0091] S5.4.4: As an improvement to the above technical solution, S5.4.4 designs a multi-channel characteristic spectrum sampler structure:

[0092] For input characteristic spectrum In position Sampling is performed on the sample, typically using bilinear interpolation.

[0093] For the The output is for one near-infrared spectral band:

[0094]

[0095] Where H is the maximum number of bands, W is the maximum number of channels, n is the current number of bands, and m is the current number of channels, which can also be written as:

[0096]

[0097] Kernel function It is a bilinear interpolation kernel.

[0098] The specific implementation process of step S5.5 is as follows:

[0099] S5.5.1: Task Decoupling Modeling

[0100] Two specific subnetworks are set up for lipid content and oxidation level. Each task-specific subnetwork contains an independent parameter structure to learn the feature representation required for that task. The output features of the dynamic weight allocation module are fed into multiple task-specific subnetworks, enabling parallel modeling of multiple tasks.

[0101] S5.5.2: Shared Feature Extraction

[0102] The raw spectral data of pine nuts is input into a shared feature extraction subnetwork. A unified latent feature representation is extracted through several convolutional layers, fully connected layers, or transform coding structures to capture the global spectral information and the distribution of the main absorption peaks in the input band.

[0103] S5.5.3: Modeling the position of absorption peaks

[0104] Based on task decoupling, an absorption peak correspondence mechanism is further introduced. By calculating the correlation between the absorption peak positions of the input band and the significant response bands of each output task, the mapping between the band and the characteristic vibrations of functional groups (chemical bond vibration characteristics) is realized, thereby improving the band interpretability of the model (realizing the band mapping in a physical sense).

[0105] S5.5.4: Joint Loss Optimization

[0106] The predicted results output by each task sub-network are compared with the ground truth labels to calculate the loss, and a joint loss function is constructed by combining dynamically assigned task weights. Through end-to-end joint training, the shared layer, task layer, and dynamic weight module are simultaneously optimized to achieve collaborative learning and feature fusion across multiple tasks.

[0107] As a further optimization of the above technical solution, in step S6, the backpropagation algorithm is used to iteratively train the model. The error calculation is based on the mean squared error (MSE) loss function, and the model weight parameter update uses the adaptive moment estimation (Adam) optimizer to improve convergence efficiency and stability. In terms of model performance evaluation, the coefficient of determination (R²) is introduced to characterize the overall goodness of fit of the model, and the root mean square error (RMSE) is introduced as a measure of the model's prediction accuracy. Through multiple rounds of performance verification on the validation set, a final stable and reliable detection model is obtained. When the model satisfies that R² is greater than or equal to 0.75 and RMSE is less than or equal to 0.2, it indicates that the model has high prediction accuracy and robustness, and can be directly applied to the rapid non-destructive detection of lipid content and oxidative deterioration degree of pine nut kernels.

[0108] S6: The specific process of model training and optimization is as follows:

[0109] S6.1: Loss function selection. During training, a suitable loss function is selected to optimize the model parameters.

[0110] For regression problems, the mean squared error (MSE) is used, defined as:

[0111]

[0112] For classification problems, the cross-entropy loss function is used:

[0113]

[0114] in, For real labels, The class probabilities predicted by the model. This represents the number of categories.

[0115] S6.2: During the model training phase, to achieve efficient optimization of network parameters, the Adam optimizer is selected to update the model parameters. The Adam optimization algorithm combines the ideas of momentum and adaptive learning rate, dynamically adjusting the parameter update step size by simultaneously estimating the first moment (mean) and the second moment (variance). Its parameter update rules are as follows:

[0116]

[0117] The above formula is used for calculation. The first moment of the gradient of the model at time step [time]. The objective loss function. This represents the gradient of the loss function with respect to the model parameters. The first moment of the gradient is obtained by calculating the exponentially weighted moving average of the gradient.

[0118]

[0119] The above formula is used for calculation. The second moment of the gradient of the model at time step. Wherein, and These are the exponential decay coefficients of the first and second moments, respectively; The second moment of the gradient is obtained by calculating the exponentially weighted moving average of the squared gradient.

[0120]

[0121] The above formula is based on The first and second moments of the gradient of the model at each time step are used for update calculations. The parameters of the time-matter model. Among them, For the first The model parameter vector at the next iteration; These are the model parameters from the previous iteration. The learning rate controls the step size for each parameter update. To prevent small constants with denominators of zero, a typical value is taken as... .

[0122] Through the above optimization process, the Adam optimizer can dynamically adjust the learning rate of each parameter while maintaining gradient smoothness, thereby improving the convergence stability and training efficiency of the model in complex non-convex loss spaces.

[0123] This part is strongly related to neural networks and is an optimization algorithm for neural networks. It is separated from the near-infrared spectral data by one layer. This invention designs a network for spectral data and then uses a general algorithm to optimize the neural network, ultimately achieving the purpose of this invention.

[0124] Example

[0125] like Figure 1 As shown, this invention proposes a rapid detection method for the lipid content and oxidative rancidity level of pine nuts, comprising:

[0126] S1 places pine nut samples from the same harvest batch in the sample bin for proper storage to ensure that the oxidation level of pine nut samples at the same sampling time point can be kept in a uniform state.

[0127] S2 samples pine nuts were taken out every 2 days and placed on the experimental table, ensuring that there was no stacking or obstruction between the pine nut samples, and then hyperspectral scanning was performed.

[0128] S3 acquires the original near-infrared spectral data of pine nut samples through hyperspectral acquisition; determines the reference level of oxidative deterioration and the true value of lipid reference content of pine nut samples at different sampling times according to national standard methods or commonly used laboratory methods (colorimetric detection of rancidity using acid value / peroxide value test paper; lipid content measured by Soxhlet extraction), and establishes a database based on this to ensure correspondence with the original spectral data determined in S2.

[0129] S4 preprocesses the acquired raw near-infrared spectral data and randomly divides it into training and validation sets.

[0130] Based on a one-dimensional convolutional neural network, S5 designs an improved one-dimensional dilated convolutional network with a dynamic weight allocation module to construct a deep learning model. This model fits the near-infrared spectral data and lipid content data of the training set samples. It also classifies the near-infrared spectral data and oxidation degree of the training set samples to establish a multi-task learning model for detecting the lipid content and oxidation degree of pine nuts.

[0131] S6 minimizes the loss function using the backpropagation algorithm, updates the weight parameters of Model I and Model II in reverse, and uses the validation set from S4 to validate and optimize the model's hyperparameters to improve the model's prediction accuracy and robustness. After multiple validations and optimizations, the final detection model is obtained; if the coefficient of determination R... 2 A root mean square error (RMSE) greater than or equal to 0.75 and less than or equal to 0.2 indicates that the model meets the requirements and can be used for rapid detection of lipid content and oxidative deterioration in pine nuts.

[0132] S7 prepares the pine nut samples to be tested, places them in the sample chamber, and removes the samples at regular intervals to collect raw spectral data. Soxhlet extraction and acid value / peroxide value tests are then performed. The raw spectral data is input into the final detection model in S6, and the output results are the predicted results of the lipid content and oxidative rancidity of the pine nut products. This model can be used for the rapid detection of pine nut lipid reference content and oxidative rancidity reference level.

[0133] Prepare pine nut samples to be tested, place them in the sample chamber, and take out the samples at regular intervals to collect raw spectral data. Then, perform Soxhlet extraction experiments and acid value / peroxide value tests. Input the raw spectral data into the final detection model, and the output results are the predicted results of lipid content and oxidative rancidity of pine nut products.

[0134] Example 1

[0135] This embodiment proposes a method for preparing graded oxidative rancid pine nut samples, and for determining their lipid content, acidification level, and corresponding near-infrared spectral data, specifically including:

[0136] In this experiment, pine nut samples were placed in an artificially controlled environment chamber equipped with supplemental lighting, and periodic sunlight and UV-C light treatment was implemented: four cycles were completed daily with 2 hours of UV-C irradiation followed by a 1-hour interval; fluorescent lighting was used to supplement insufficient natural light. The samples were placed on a specially made black acrylic plate with a porous structure to ensure sample stability and optimize the collection surface. A hyperspectral imaging system (HyperSpec® VNIR-A) from Headwall Corporation, USA, was used to perform spectral scanning in the range of 400–1000 nm, extracting regions of interest (ROIs) and calculating the average value as the raw spectral data. No fewer than 500 sets of data were collected to enhance the robustness of the model. Starting from Day 0, samples were collected every 48 hours for lipid content determination (based on the extraction method of GB5009.5—2010 standard) and acid value / peroxide value colorimetric testing, continuing until Day 30. A time-series dataset covering the entire period, including hyperspectral and physicochemical indicators, was constructed to ensure accurate sample information and standardized numbering. The lipid content and acidification level labels of different batches of pine nut samples corresponded one-to-one with their near-infrared spectral data, thus establishing a standard pine nut sample database. The original spectral curves of the samples are shown below. Figure 2 As shown.

[0137] The raw spectral data in the sample database is preprocessed. Based on the final prediction result parameters, the most suitable spectral data preprocessing scheme is selected. The pine nut spectral preprocessing result after processing by SNV and SG convolution smoothing algorithms is shown below. Figure 3 As shown.

[0138] Example 2

[0139] This embodiment is used to extract features and establish a deep learning model from the near-infrared spectral data of pine nut samples collected and preprocessed in Example 1. The specific process includes: dividing the database samples into a training set and a validation set, wherein there are 500 training set samples and 100 validation set samples.

[0140] First, the pine nut spectra are batch-wise input into the input layer of a one-dimensional neural network according to a tensor format. The shape of the spectral tensor is then modified to create a data structure with a three-dimensional shape (batch size, features, channels). The proposed one-dimensional convolutional neural network is then used to automatically extract spectral features from the training set, such as... Figure 4 As shown, each layer consists of three convolutional layers working together, sliding the convolutional kernel along the feature dimension and performing dot product operations to mine the intrinsic features of the spectral data. A portion of these features are input to a shared layer, and the other portion is input to a gated function unit, enabling multi-task learning. The model training curve is shown below. Figure 5As shown. In addition, the ReLU activation function is used to enhance the non-linear expressive power of the model, enabling the network to capture more complex data patterns. Maxpooling is used for downsampling to reduce computational complexity and capture key feature information. Dropout layer, as a regularization technique, prevents overfitting and enhances the generalization ability of the model by randomly dropping some network connections.

[0141] The scatter plot results of the regression model are as follows Figure 6 As shown, the coefficient of determination R0 for the predicted and detected lipid content values ​​of the training set samples is... 2 The root mean square error (RMSE) was 0.12, and the coefficient of determination (R²) for the predicted and detected lipid content values ​​in the validation set was 0.83. 2 The error is 0.79, and the root mean square error (RMSE) is 0.13; the confusion matrix results for the classification model are as follows. Figure 7 As shown, the overall classification accuracy exceeds 80%.

[0142] The deep learning model constructed in this embodiment innovatively combines the feature processing capabilities of a one-dimensional dilated convolutional neural network with the efficient learning capabilities of a gated, multi-task learning mechanism and a feedforward neural network. This model first expands the one-dimensional near-infrared spectral data into a three-channel near-infrared spectral tensor, fully considering the complex relationships between spectral wavelengths and the high-dimensionality of the data. Through the combination of CNN and a multi-task learning mechanism, the model can selectively extract features from the spectral data according to different tasks, thereby establishing a more accurate multi-task prediction model. During training, the backpropagation algorithm is used to continuously optimize the weight parameters of the multi-task model for pine nut lipid content and oxidation degree, significantly improving prediction accuracy.

[0143] Validation results show that the model performs exceptionally well on the validation set, demonstrating extremely high reliability. Compared to traditional models, this model not only automates the training process, eliminating the need for complex feature extraction, but also achieves a significant improvement in prediction accuracy, far surpassing existing traditional algorithms.

[0144] Based on the method proposed in this invention, a rapid non-destructive testing system for the lipid content and deterioration degree of pine nuts based on hyperspectral imaging and deep learning has been developed using a programming language. This system has program modules corresponding to the steps of the aforementioned technical solution, and executes the steps in the aforementioned rapid non-destructive testing method for the lipid content and deterioration degree of pine nuts based on hyperspectral imaging and deep learning. The developed system (software) computer program is stored on a computer-readable storage medium, and the computer program is configured to implement the steps of the method when called by a processor. In other words, this invention is materialized on a carrier, becoming a computer program product.

[0145] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, application-specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0146] The computational programs (also referred to as programs, software, software applications, or code) of this invention include machine instructions of a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device PLD) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0147] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.

Claims

1. A rapid and non-destructive method for detecting lipid content and deterioration degree in red pine nuts based on hyperspectral imaging and deep learning, characterized in that, The method includes: S1. Place pine nut samples from the same harvest batch in a sample storage bin to ensure that the degree of deterioration of pine nut samples at the same sampling time point can be kept in a uniform state; the pine nut samples are red pine nut kernel samples. S2. Pine nut samples are taken out every 1-3 days and placed on the experimental table, ensuring that there is no stacking or obstruction between the pine nut samples, and then hyperspectral scanning is performed. S3. Obtain the original near-infrared spectral data of pine nut samples through the collected hyperspectral data; determine the true value of lipid reference content and the reference degree of oxidative deterioration of pine nut samples at different sampling times, and establish a database based on this to ensure that it corresponds with the original spectral data determined in S2. S4. Preprocess the raw near-infrared spectral data of red pine nuts collected, and divide the red pine nut sample data collected in different batches into training set and validation set. The division process conforms to the principles of random sampling and machine learning modeling dataset setting. S5. Based on a one-dimensional convolutional neural network, an improved one-dimensional dilated convolutional network with a dynamic weight allocation module is designed to construct a deep learning model suitable for the spectral feature analysis of red pine nut kernels. The deep learning model enhances the ability to model long-distance dependencies in spectral sequences by introducing a dilated convolution mechanism and achieves adaptive adjustment of feature channels by combining a dynamic weight mechanism. The model is used to fit and regress the near-infrared spectral data of the training set samples with the corresponding lipid content, and at the same time, it performs hierarchical modeling of the degree of deterioration of the samples, thereby constructing a multi-task learning framework that can simultaneously complete lipid content prediction and deterioration degree identification. S6. The deep learning model is trained using the backpropagation algorithm, and the network weight parameters are updated in reverse by minimizing the loss function; the model's performance is evaluated using the validation set partitioned in step S4, and the key hyperparameters of the model are tuned based on the evaluation results to improve the model's prediction accuracy and robustness; through multiple rounds of cross-validation and parameter optimization, a convergent and stable final detection model is obtained; when the model's coefficient of determination R on the validation set... 2 If the value is greater than or equal to 0.75 and the root mean square error (RMSE) is less than or equal to 0.2, the constructed model is considered to meet the requirements for rapid detection. S7. Place the pine nut sample to be tested in the sample preparation chamber, take out the sample at regular intervals to scan the hyperspectral image and obtain the raw spectral data, and then perform Soxhlet extraction experiment and acid value / peroxide value test; input the raw spectral data into the final detection model in S6, and the output result is the prediction result of pine nut lipid content and deterioration degree at the same time. The goal of the dynamic weight allocation module is to automatically identify key absorption peaks in the input spectrum and establish a mapping relationship with the corresponding band response positions of the output task. Ultimately, it establishes a correspondence between the absorption peak and wavelength position in the pine nut output characteristic band and the absorption peak and wavelength position in the input characteristic band. This allows the model to automatically focus on spectral positions related to functional group characteristic vibrations or molecular absorption peaks during training. This process includes the following four stages: First, a feature encoding mechanism is designed to extract spatial structure features in the near-infrared spectrum, improving the model's perception of spectral spatial distribution. Second, a nonlinear mapping relationship is established between spatial structure features and the internal material composition of the pine nut, relying on a fully connected network. Then, a generative network module is introduced to predict discriminative sampling points in the high-dimensional spectral space, using spatial structure features as a condition, achieving an effective correspondence between the input space and the spectral feature space. Finally, a spectral resampling mechanism is designed to extract effective band information from the spectrum at the predicted sampling points.

2. The rapid and non-destructive detection method for lipid content and deterioration degree of red pine nuts based on hyperspectral imaging and deep learning according to claim 1, characterized in that, In step S1, light treatment and environmental control are carried out, specifically: the shells of red pine nuts from the same harvest batch are removed, the pine nut samples are placed in an environmentally controllable sample chamber, and static treatment is carried out under sunlight. The temperature is controlled at 25±2℃ and the humidity at 45±5% throughout the experiment, and environmental parameter data are collected every 15 to 30 minutes using a HOBO U23-001 temperature and humidity recorder. In step S2, the sample acquisition array construction and hyperspectral acquisition are specifically as follows: the selected samples are evenly arranged on a black plate with holes to avoid sample shaking caused by platform movement during the acquisition process. The samples are arranged in an array according to the plate size to optimize the regularity of hyperspectral scanning acquisition. The hyperspectral images of the pine nut samples are scanned and acquired and exported and saved. The samples are placed in a sealed bag and properly stored in the dark.

3. A rapid and non-destructive method for detecting lipid content and deterioration degree of red pine nuts based on hyperspectral imaging and deep learning, as described in claim 1 or 2, characterized in that... In step S3, The region of interest (ROI) was selected from the exported hyperspectral image of the pine nut sample, and the near-infrared spectral data of the region was obtained. Immediately after hyperspectral acquisition, physicochemical analysis experiments were performed on this portion of the samples to construct training labels. Random and sufficient samples were taken from each plate and divided into two portions. The first portion was crushed and hexane was added to extract lipids. After standing for 10-15 minutes, colorimetric detection was performed using MLbio acid value / peroxide value test paper. The reaction time of the acid value test paper was controlled within 5 minutes to form a mapping dataset between the spectrum and lipid oxidative rancidity indicators. The second portion was subjected to Soxhlet extraction experiments according to national standard methods to form a mapping dataset between the spectrum and lipid content indicators. The degree of deterioration and lipid content of pine nut samples at different oxidation times were measured to establish a database. To ensure a one-to-one correspondence with the original spectral data measured in S2, the lipid content and oxidation level data obtained from physicochemical tests were labeled and matched with the corresponding hyperspectral image data to construct a complete training dataset.

4. The rapid and non-destructive detection method for lipid content and deterioration degree of red pine nuts based on hyperspectral imaging and deep learning according to claim 1, characterized in that, Step S4 is used to complete data cleaning and modeling preparation, specifically as follows: The acquired raw near-infrared spectral data is preprocessed to eliminate background interference, improve the signal-to-noise ratio, and enhance the model's ability to extract key features. The selected preprocessing algorithm is flexibly adjusted according to spectral characteristics and sample type to effectively extract key information from the spectral data and suppress redundant information, thereby improving the accuracy and stability of subsequent modeling. The data preprocessing process is as follows: S4.1: Perform data collection and cleaning; S4.2: Perform data normalization to improve the stability and convergence speed of model training. The normalization method is Z-Score: in, It is the raw data. It is the mean of the data. It is the standard deviation of the data. Standardized data has zero mean and unit variance, which is used to prevent training instability caused by the excessively large range of values ​​for certain features. S4.3: In near-infrared spectroscopy analysis, the sliding window technique is used to segment time series data. Assume the input data is... The sliding window operation then divides the data into multiple overlapping subsequences: And so on, in, It is the window size, and the data in each window serves as the input to the model; The raw near-infrared spectral data of the preprocessed pine nut samples were randomly divided into training and validation sets to provide basic data for subsequent deep learning modeling and identification of lipid content and oxidative deterioration in pine nut kernels.

5. A rapid and non-destructive method for detecting lipid content and deterioration degree of red pine nuts based on hyperspectral imaging and deep learning, as described in claim 1 or 4, characterized in that... Step S5 specifically includes: S5.1 uses the SPXY algorithm to randomly divide the samples into training and validation sets; S5.2 Establish a deep learning-based regression model; S5.3 automatically extracts spectral features from the preprocessed training set through a dilated convolutional neural network, thereby expanding the model's receptive field and improving its ability to extract multi-scale spectral features. S5.4 introduces a dynamic weight allocation module to perform feedback adjustment, enhancing the model's balance between dynamic selection and fine-tuning, and ultimately outputting predicted values ​​for pine nut lipid content and deterioration degree; S5.5 proposes a multi-task learning-based method for predicting pine nut composition. By introducing a dynamic weighting mechanism and an absorption peak position mapping mechanism, it achieves accurate modeling and collaborative prediction of multiple band features in the input spectrum to multiple output tasks. The constructed deep learning model uses a multi-task network. Using near-infrared spectral data and lipid content data of pine nut samples in the training set, a multi-task detection model I for pine nut lipid content is established; using near-infrared spectral data and deterioration level data of pine nut samples in the training set, a multi-task detection model II for the degree of deterioration in pine nuts is established.

6. A rapid and non-destructive method for detecting lipid content and deterioration degree of red pine nuts based on hyperspectral imaging and deep learning, as described in claim 5, is characterized in that... S5.3 expands the receptive field of the convolution kernel by introducing a dilation rate, effectively capturing a wider range of dependencies without increasing computational cost. One-dimensional dilated convolution is used to process near-infrared spectral signals. The specific process of step S5.3 is as follows: S5.3.1: Design process of dilated convolutional layers and construction of one-dimensional dilated convolutional network model: Dilated convolution uses the dilation rate... This expands the receptive field of the convolution kernel, capturing dependencies over longer periods or over a wider range. S5.3.2: For one-dimensional dilated convolution, assume the input signal is... The kernel size is The output calculation formula for dilated convolution is: in, The void ratio, The weights of the convolution kernel, For bias terms, It is the output index, and The spacing between convolution kernel elements is represented by the dilation rate r. By selecting an appropriate dilation rate r, the receptive field can be effectively increased without increasing the size of the convolution kernel, thus avoiding a surge in computational complexity. Dilated convolution has the ability to capture a wider range of feature dependencies without increasing computational cost, and can be used to process the near-infrared spectrum of pine nuts with long-term dependencies. S5.3.3: Stacked dilated convolutional layers are constructed by stacking multiple dilated convolutional layers to form a deep feature extraction network to expand the receptive field. Each convolutional layer extracts feature information at different scales according to its dilation rate. As the dilation rate increases, the receptive field expands exponentially, thereby capturing temporal dependencies at greater distances. S5.3.4: Activation function selection. A non-linear activation function, ReLU, is applied after each convolutional layer. Its definition is: The ReLU function helps introduce non-linearity into the network, enabling the model to learn more complex patterns, which can alleviate gradient vanishing and accelerate the training process of deep networks. The specific process of step S5.4 is as follows: S5.4.1: Design of the characteristic spectral localization network structure: S5.4.1.1: Define the output characteristic spectrum as , representing the wavelength position after feature encoding; The feature spectrum is mapped to the input feature using a multilayer perceptron for feature encoding, and the corresponding input feature spectrum is calculated as follows. : S5.4.1.2: Based on Design of a Localization Network for Multichannel Feature Encoding: This module takes the input spectrum Extract characteristic band information related to pine nut mold and regress it. Multilayer perceptrons encode spectral information features in the form of a function: in, It is a learnable function, representing the use of a multilayer perceptron. Feature encoding of spectral bands; S5.4.1.3: When the lengths of different channels in the MLP network are the same, the result is obtained based on the spliced ​​MLP. : GAP is an abbreviation for Global Average Pooling. It uses a global average pooling layer to filter feature bands from the output bands and uses a batch normalization layer to merge the pooled feature bands. S5.4.1.3: When the lengths of different channels in an MLP network are inconsistent, each branch extracts features independently. , S5.4.2: Design of Characteristic Spectral Fusion Structure S5.4.2.1: Determine if the dimensions of the characteristic bands are consistent. If the characteristic dimensions are inconsistent, use... Convolution for dimension alignment: S5.4.2.2: Using gated weighted fusion of spectral bands: S5.4.3: Designing the Generator Structure For each near-infrared spectral output, calculate its sampling position within the input spectrum. That is, to generate a sampling spectral network: in This represents the set of characteristic spectra of the target output; S5.4.4: Design a multi-channel characteristic spectral sampler structure: For input characteristic spectrum In position Sampling is performed on the sample using bilinear interpolation; For the The output is for one near-infrared spectral band: Where H is the maximum number of bands, W is the maximum number of channels, n is the current number of bands, and m is the current number of channels, denoted as: Kernel function It is a bilinear interpolation kernel; The specific process of step 5.5 is as follows: S5.5.1: Task Decoupling Modeling Two specific subnetworks are set up for lipid content and deterioration degree. Each task subnetwork contains an independent parameter structure to learn the feature representation required for the task. The output features of the dynamic weight allocation module are sent to multiple task-specific subnetworks to achieve parallel modeling of multiple tasks. S5.5.2: Shared Feature Extraction The raw spectral data of pine nuts is input into a shared feature extraction subnetwork. A unified latent feature representation is extracted through several convolutional layers, fully connected layers or transform coding structures to capture the global spectral information and the distribution of the main absorption peaks in the input band. S5.5.3: Modeling the position of absorption peaks Based on task decoupling, an absorption peak correspondence mechanism is introduced. By calculating the correlation of absorption peak positions between the input band and the significant response bands of each output task, the mapping between the band and the characteristic vibration of functional groups is realized, and the band interpretability of the model is improved. S5.5.4: Joint Loss Optimization The predicted results output by each task sub-network are compared with the real labels to calculate the loss. Combined with dynamically allocated task weights, an overall joint loss function is constructed. Through end-to-end joint training, the shared layer, task layer and dynamic weight module are optimized simultaneously to achieve collaborative learning and feature fusion of multiple tasks.

7. A rapid and non-destructive method for detecting lipid content and deterioration degree of red pine nuts based on hyperspectral imaging and deep learning, as described in claim 6, is characterized in that... Step S6 uses the backpropagation algorithm to iteratively train the model, where error calculation is based on the mean squared error loss function, and the model weight parameters are updated using an adaptive moment estimation optimizer to improve convergence efficiency and stability; the specific process of model training and optimization is as follows: S6.1: Select a loss function during training to optimize model parameters; For regression problems, the mean squared error (MSE) is used, defined as: For classification problems, the cross-entropy loss function is used: in, For real labels, The class probabilities predicted by the model. Number of categories; S6.2: During the model training phase, the Adam optimizer is used to update the model parameters to achieve efficient optimization of the network parameters. The Adam optimization algorithm combines the momentum method and the adaptive learning rate, and dynamically adjusts the parameter update step size by simultaneously estimating the first and second moments. Its parameter update rules are as follows: The above formula is used for calculation. The first moment of the gradient of the model at time step is given by, where, The objective loss function; This represents the gradient of the loss function with respect to the model parameters. The first moment of the gradient is obtained by calculating the exponentially weighted moving average of the gradient. The above formula is used for calculation. The second moment of the gradient of the model at time step; where, and These are the exponential decay coefficients of the first and second moments, respectively; The second moment of the gradient is obtained by calculating the exponentially weighted moving average of the squared gradient. The above formula is based on The first and second moments of the gradient of the model at each time step are used for update calculations. The parameters of the time-matter model, where, For the first The model parameter vector at the next iteration; These are the model parameters from the previous iteration. The learning rate controls the step size for each parameter update; To prevent small constants with denominators of zero, the value is taken as... ; Through the above optimization process, the Adam optimizer can dynamically adjust the learning rate of each parameter while maintaining gradient smoothness, thereby improving the convergence stability and training efficiency of the model in complex non-convex loss spaces.

8. A rapid and non-destructive detection system for lipid content and deterioration degree of red pine nuts based on hyperspectral imaging and deep learning, characterized in that, The system is implemented based on the detection method according to any one of claims 1-7, comprising: The raw data acquisition unit for pine nut samples is used to acquire hyperspectral data of pine nut samples at different times and extract raw near-infrared spectral data, the degree of deterioration of pine nut samples and lipid content, and to establish a database based on this data. The data preprocessing unit is used to preprocess and extract features from the raw near-infrared spectral data of the acquired pine nut samples, and divide them into training and validation sets. The deep learning model building unit uses a convolutional neural network to construct a multi-task neural network model by introducing a dilated convolution mechanism and a dynamic weight allocation module. It fits the near-infrared spectral data and lipid content data of pine nut samples in the training set to establish a multi-task learning model for lipid content in pine nuts; and classifies the near-infrared spectral data and deterioration degree data of the training set samples to establish a multi-task learning model for detecting the deterioration degree in pine nuts. The model validation and optimization unit is used to minimize the loss function through the backpropagation algorithm, back-update the weight parameters of the multi-task learning model for lipid content in pine nuts and the multi-task learning model for detecting the degree of deterioration in pine nuts, and use the validation set to validate and optimize the hyperparameters of the model to improve the prediction accuracy and performance of the model. After multiple validations and optimizations, the final detection model is obtained. The rapid detection unit for pine nut samples acquires the original spectral data of the sample to be tested, inputs the original spectral data into the multi-task learning model for lipid content and deterioration degree in pine nuts, and outputs the predicted values ​​of lipid content and deterioration degree in pine nuts. The goal of the dynamic weight allocation module is to automatically identify key absorption peaks in the input spectrum and establish a mapping relationship with the corresponding band response positions of the output task. Ultimately, it establishes a correspondence between the absorption peak and wavelength position in the pine nut output characteristic band and the absorption peak and wavelength position in the input characteristic band. This allows the model to automatically focus on spectral positions related to functional group characteristic vibrations or molecular absorption peaks during training. This process includes the following four stages: First, a feature encoding mechanism is designed to extract spatial structure features in the near-infrared spectrum, improving the model's perception of spectral spatial distribution. Second, a nonlinear mapping relationship is established between spatial structure features and the internal material composition of the pine nut, relying on a fully connected network. Then, a generative network module is introduced to predict discriminative sampling points in the high-dimensional spectral space, using spatial structure features as a condition, achieving an effective correspondence between the input space and the spectral feature space. Finally, a spectral resampling mechanism is designed to extract effective band information from the spectrum at the predicted sampling points.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program configured to, when invoked by a processor, execute the corresponding steps in the rapid non-destructive detection method for lipid content and deterioration degree of red pine nuts based on hyperspectral imaging and deep learning as described in any one of claims 1-7.

10. A rapid, non-destructive testing device for the lipid content and deterioration degree of red pine nuts, characterized in that: The detection device includes at least one processor, a memory communicatively connected to the at least one processor, and a device for acquiring raw near-infrared spectral data of pine nut samples using hyperspectral imaging; wherein, the memory stores instructions executable by the at least one processor, enabling the at least one processor to perform the corresponding steps in the rapid non-destructive detection method for lipid content and deterioration degree of red pine nuts based on hyperspectral imaging and deep learning as described in any one of claims 1-7, thereby realizing the detection of lipid content and deterioration degree of red pine nut kernels.

Citation Information

Patent Citations

  • Multi-task learning model construction method based on attention mechanism and deformable convolution

    CN113554156A

  • Hyperspectral image super-resolution reconstruction method based on multiple attention mechanisms

    CN117114985A