Terahertz spectrum image intelligent detection method for agricultural product producing area identification, electronic equipment and storage medium

By adopting the intelligent detection method of terahertz spectral images, through dynamic weight filtering, dimensionality reduction processing and local feature extraction and image-level majority voting method, local feature extraction and image-level majority voting mechanism, combined with dynamic weight filtering, dimensionality reduction processing, local feature extraction and lightweight model deployment, the problems of environmental interference, imaging distortion, large data volume and low processing efficiency in terahertz detection are solved, and high-precision agricultural product origin identification is achieved.

CN120635591APending Publication Date: 2025-09-12TERAHERTZ TECH APPL (GUANGDONG) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510981086.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing terahertz detection technology faces problems such as severe environmental interference, imaging distortion of curved objects, large data volume, low processing efficiency, and lack of a unified image processing and classification framework, making it difficult to meet the needs of rapid non-destructive testing for identifying the origin of agricultural products.

Method used

The terahertz spectral image intelligent detection method is adopted, and high-precision origin identification is achieved through the dynamic weight filtering algorithm, local feature extraction and image-level majority voting method, local feature extraction and image-level majority voting method, local feature extraction and image-level majority voting method, local feature extraction and image-level majority voting mechanism, combined with dynamic weight filtering, dimensionality reduction processing, local feature extraction and lightweight model deployment.

Benefits of technology

It improves the accuracy and robustness of agricultural product origin identification, realizes non-contact and non-destructive testing, is suitable for fragile or precious samples, has high-precision origin identification capabilities, strong anti-interference ability, supports edge deployment, and is suitable for nuts and grain agricultural products. The system has a high degree of system integration and meets the needs of rapid on-site testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635591A_ABST
    Figure CN120635591A_ABST
Patent Text Reader

Abstract

The invention provides a terahertz spectrum image intelligent detection method for agricultural product producing area identification, electronic equipment and a storage medium, and the method comprises the steps: collecting a to-be-detected agricultural product image, carrying out the spectrum data extraction, inputting the spectrum data into a classification model, and carrying out the detection, and finally outputting a prediction result of the producing area of the to-be-detected agricultural product. The method further comprises the step of training the classification model, and the training step comprises the following steps: S1, carrying out image acquisition and preprocessing on an image to obtain an absorption spectrum image; s2, calculating a weighted mean value by using a dynamic weight filtering algorithm, and optimizing a frequency band so as to obtain a dimensionality-reduced image; and S3, carrying out local feature extraction on the dimensionality-reduced image and constructing a data set, then training a classification model by using the data set, and finally carrying out edge deployment on the trained model to run in an embedded platform. Therefore, non-contact and non-destructive monitoring of agricultural products is realized, and high-precision production place identification is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of terahertz nondestructive testing and food quality control, and more specifically to a terahertz spectral image intelligent detection method, electronic equipment, and storage medium for identifying the origin of agricultural products. Background Art

[0002] As consumers' demand for food traceability grows, traditional chemical testing and DNA fingerprinting methods, due to their high cost, complex operation, and high destructiveness, are unable to meet the needs of rapid, non-destructive testing. In recent years, terahertz imaging has shown potential in food testing due to its advantages such as non-contact, strong penetration, and ability to obtain the frequency domain characteristics of materials. However, terahertz testing currently faces the following technical bottlenecks:

[0003] 1. Severe environmental interference: Water vapor in the air produces a strong absorption peak in the 0.16-0.73 THz frequency band, causing the sample characteristic peak to shift, affecting the recognition accuracy;

[0004] 2. Distortion of curved surface imaging: Nut samples have curved surfaces, and the reflected signal is out of focus, resulting in blurred textures.

[0005] 3. Large data volume and low processing efficiency: The three-dimensional spectral cube has a large amount of computation and is difficult to meet the real-time deployment requirements of embedded devices;

[0006] 4. Lack of a unified image processing and classification framework: Existing research mostly focuses on a single frequency band or global features, and fails to fully utilize local structural information.

[0007] In addition, although some improvement methods have been proposed in existing literature (such as the dynamic filtering algorithm in CN33894339A and the Raman coupling module in CN113418891A), there is no complete solution that combines terahertz image preprocessing, local feature extraction, lightweight model deployment and origin classification. Summary of the Invention

[0008] The present invention aims to overcome at least one defect (shortcoming) of the above-mentioned prior art and provide a terahertz spectral image intelligent detection method, electronic device and storage medium for identifying the origin of agricultural products, which are used to solve the problems in the prior art such as severe environmental interference, distortion of surface object imaging, large data volume, low processing efficiency, and lack of a unified image processing and classification framework.

[0009] The technical solution adopted by the present invention is a terahertz spectral image intelligent detection method for identifying the origin of agricultural products. The method includes: collecting images of the agricultural products to be tested and extracting spectral data, then inputting the spectral data into a classification model for detection, and finally outputting a prediction result of the origin of the agricultural products to be tested;

[0010] Preferably, the method further comprises training the classification model, and the training step comprises:

[0011] S1: perform image acquisition and image preprocessing to obtain an absorption spectrum image;

[0012] S2: Use the dynamic weight filtering algorithm to calculate the weighted mean and optimize the frequency band to obtain the image after dimensionality reduction;

[0013] S3: Extract local features from the reduced-dimensional image and construct a dataset. The dataset is then used to train the classification model. Finally, the trained model is deployed to the edge on the embedded platform for operation.

[0014] The terahertz spectral image intelligent detection method for agricultural product origin identification proposed in the present invention improves the clarity and accuracy of the image through a dynamic weight filtering algorithm and dimensionality reduction processing, and effectively reduces the influence of environmental interference and noise. After obtaining the absorption spectrum image in the image preprocessing stage, local feature extraction can accurately capture the microscopic differences related to the origin, thereby improving the recognition accuracy of the classification model. The trained classification model is deployed at the edge on an embedded platform to achieve efficient, low-latency real-time detection, meeting the demand for fast and convenient detection in agricultural production. This method not only improves the accuracy and robustness of agricultural product origin identification, but also has strong adaptability and flexibility. It can be widely used in intelligent detection in other fields and has important practical application value.

[0015] Preferably, the step S1 includes:

[0016] S11: Use terahertz time-domain spectrometer to collect the time domain signal of the sample;

[0017] S12: Obtain the frequency domain response through Fourier transform and calculate the absorption coefficient;

[0018] S13: Construct a three-dimensional spectral cube based on the absorption coefficient, where each layer corresponds to an absorption spectrum image of a frequency point.

[0019] This step uses a terahertz time-domain spectrometer to collect the sample's time-domain signal, converts it into a frequency-domain response through Fourier transform, calculates the absorption coefficient, and further constructs a three-dimensional spectral cube, with each layer corresponding to an absorption spectrum image at a different frequency point. This process can eliminate the influence of water vapor in the air on the terahertz absorption coefficient, effectively improve the imaging clarity of curved objects, accurately capture the absorption characteristics of the sample at each frequency, present spectral information in a three-dimensional form, and provide richer analytical data, thereby improving the accuracy and robustness of detection. In this way, a more refined data foundation can be provided for the subsequent identification of the origin of agricultural products, improving the accuracy of intelligent detection and multi-dimensional analysis capabilities.

[0020] Preferably, the step S2 includes:

[0021] S21: For each pixel x i Select its area point and get {x i-3 ,...,x i+3};

[0022] S22: Then calculate the weighted mean based on the domain points. The calculation formula is:

[0023]

[0024] Among them, ω j is the weight, which is obtained by dynamic adjustment according to the ambient humidity and water vapor concentration.

[0025] This step effectively reduces the impact of noise on the image and enhances the smoothness and continuity of the image by selecting the domain points for each pixel and calculating the weighted mean based on these domain points. By dynamically adjusting the weights to account for changes in ambient humidity and water vapor concentration, image processing under different environmental conditions can be more accurate and robust. Dynamic adjustment of weights enables the algorithm to adapt to different environmental interferences, improving the stability and accuracy of detection. This method can maintain high image quality under different sampling environments, enhance the accuracy of subsequent image analysis and classification, and ensure the reliability of detection results.

[0026] Preferably, the step S3 includes:

[0027] S31: Binarize the image after dimensionality reduction and extract the foreground area;

[0028] S32: Divide the fixed size into non-overlapping rectangular windows;

[0029] S33: extracting the average spectrum curve within each rectangular window as a local feature;

[0030] S34: Construct a training set and use several classifiers for modeling;

[0031] S35: Introduce several evaluation indicators to evaluate the model and verify the model performance.

[0032] By binarizing the image after dimensionality reduction, the foreground area can be effectively extracted, thereby removing background interference and focusing on important information in the image. By dividing the image into rectangular windows of fixed size and ensuring that they do not overlap, the image can be divided into multiple small areas, which facilitates the extraction of local features. The average spectrum curve extracted in each rectangular window is used as a local feature, which helps to capture the spectral information of different areas in the image and enhances the ability to analyze details. Next, constructing a training set and using multiple classifiers for modeling can ensure the adaptability and robustness of the model in different scenarios. Introducing several evaluation indicators to evaluate the model can quantify the performance of the model and verify its accuracy and stability. Overall, this step optimizes the image processing and classification modeling process through multiple means, improves the accuracy, reliability and generalization ability of the model, and provides strong support for subsequent applications.

[0033] Preferably, the step S31 includes:

[0034] S311: Use dimensionality reduction techniques to compress data dimensions;

[0035] S312: performing binary segmentation on the dimensionality-reduced image by using a global threshold method or a local adaptive threshold method;

[0036] S313: Remove noise and fill holes through morphological operations;

[0037] S314: Use connected domain analysis or contour detection to extract the foreground area, and screen valid targets based on features such as area and distance.

[0038] By adopting dimensionality reduction technology to compress data dimensions, the complexity of the data is reduced, the efficiency of subsequent processing is improved, and the computational burden is effectively reduced. By performing binarization segmentation on the dimensionality-reduced image through the global threshold method or the local adaptive threshold method, the image can be quickly and accurately segmented into foreground and background, facilitating the extraction of key features. In the binarized image, morphological operations are used to eliminate noise and fill holes, further improving the image quality and making the foreground area more coherent and clear. By extracting the foreground area through connected domain analysis or contour detection, the target object can be accurately identified, and the target can be screened by features such as area and distance, thereby eliminating irrelevant parts and ensuring that the extracted target is effective and representative. Overall, this step effectively improves the accuracy of image segmentation, optimizes the extraction process of the foreground area, and enhances the reliability and accuracy of subsequent analysis through the combination of multiple image processing technologies.

[0039] Preferably, the step S33 includes:

[0040] S331: Performing window processing on the original signal or image, dividing the data into multiple local areas using a sliding window;

[0041] S332: Calculate the average spectrum curve of the data in each window;

[0042] S333: Combine the multi-scale window strategy to capture the local features of different frequency ranges, and finally concatenate the average spectrum curves of each window to form a global feature vector.

[0043] By performing window processing on the original signal or image and using a sliding window to divide the data into multiple local areas, the image can be effectively processed locally to avoid information loss caused by over-smoothing. By calculating the average spectral curve of the data in each window, the local features within the window can be extracted, which helps to capture local changes and patterns in the image or signal. Combined with a multi-scale window strategy, by using windows of different sizes, local features in different frequency ranges can be captured simultaneously, thereby comprehensively analyzing the detailed information of the image and improving the diversity and adaptability of the features. Finally, the average spectral curves of each window are connected in series to form a global feature vector, so that the model can consider local and global feature information at the same time, improving the comprehensive understanding and recognition ability of the image or signal. Overall, this step enhances the accuracy of feature extraction and improves the expressiveness and robustness of the model through the combination of window processing and multi-scale strategy. It is particularly suitable for processing image data with complex local features.

[0044] Preferably, the step S34 includes:

[0045] S341: Preprocess the training set, including feature normalization and class imbalance processing;

[0046] S342: training random forest and SVM classifiers;

[0047] S343: Use network search or cross-validation to tune each classifier;

[0048] S344: Fusion of the advantages of each classifier through ensemble learning, and use of validation set to evaluate model performance.

[0049] The performance of the classification model was significantly improved through a combination of data preprocessing, model training, tuning, and ensemble learning. First, the training data was optimized through feature standardization and class imbalance processing, avoiding bias caused by varying feature scales and uneven class distribution. Next, random forest and SVM classifiers were trained, leveraging their respective strengths to extract complex patterns in the data. Classifiers were tuned through network searches or cross-validation to further optimize hyperparameters and improve model accuracy. Finally, ensemble learning was used to combine the strengths of multiple classifiers, enhancing the stability and robustness of the model. Model performance was evaluated using a validation set to ensure its generalization ability on unknown data. Overall, this process effectively improved the accuracy and reliability of the classification results.

[0050] Preferably, in step S35, the plurality of evaluation indicators include two evaluation indicators: single spectrum curve recognition accuracy and image-level majority voting result recognition accuracy.

[0051] By introducing two evaluation metrics, single spectral curve recognition accuracy and image-level majority voting result recognition accuracy, the performance of the model can be effectively evaluated from different perspectives. The single spectral curve recognition accuracy can measure the model's recognition accuracy for a single spectral curve, ensuring that the model can accurately extract and classify local features in the signal or image, and reflecting the model's recognition ability at the detail level. The image-level majority voting result recognition accuracy takes into account the voting results of multiple classifiers and measures the model's recognition accuracy at the overall image level. In particular, in the context of ensemble learning, the synergy of multiple classifiers improves the reliability and stability of the final recognition results. The combination of these two evaluation metrics allows for a comprehensive evaluation of the model's performance, taking into account both the accuracy of local features and the robustness and consistency of the overall classification, thereby optimizing the model's overall performance.

[0052] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the terahertz spectral image intelligent detection method for agricultural product origin identification as described above.

[0053] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the terahertz spectral image intelligent detection method for identifying the origin of agricultural products as described above.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] 1. Non-contact, non-destructive testing: suitable for fragile or precious samples;

[0056] 2. High-precision origin identification: Using local feature extraction and image-level majority voting, OA-Voted reaches 98.7%;

[0057] 3. Strong anti-interference ability: Dynamic weight filtering effectively suppresses water vapor interference, and the characteristic peak offset error is <0.01THz;

[0058] 4. Support edge deployment: The model's lightweight design is compatible with embedded devices, meeting the needs of rapid on-site detection;

[0059] 5. Wide range of applications: suitable for various nuts and grain agricultural products;

[0060] 6. High system integration: It integrates hardware probes, data processing, and classification output, making it easy to implement in engineering. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 This is a flow chart of the model training steps provided by the present invention.

[0062] Figure 2 This is the terahertz absorption coefficient spectrum of pistachios provided by the present invention.

[0063] Figure 3 The present invention provides a terahertz absorption coefficient spectrum image layer of almonds and pistachios at different frequency points.

[0064] Figure 4 This is a pseudo-color image of the terahertz absorption coefficient spectrum image layer of almonds and pistachios in the range of 0.11 to 1.10 THz provided by the present invention.

[0065] Figure 5 This is the average spectrum curve of the terahertz absorption coefficient spectrum image of almonds and pistachios provided by the present invention.

[0066] Figure 6 The first two principal component graphs of the average spectrum curve data of almonds and pistachios provided by the present invention.

[0067] Figure 7 This is a statistical chart of the number of pistachio batches produced by the present invention.

[0068] Figure 8 The present invention provides a confusion matrix for identifying the origin of pistachios using the RUSBoost algorithm.

[0069] Figure 9 This is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0070] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting the present invention. To better illustrate the following embodiments, some components in the accompanying drawings may be omitted, enlarged, or reduced in size, and do not represent actual product dimensions. Those skilled in the art will appreciate that some well-known structures and their descriptions may be omitted from the accompanying drawings.

[0071] Example 1

[0072] This embodiment provides a terahertz spectral image intelligent detection method for identifying the origin of agricultural products. The method includes: collecting images of the agricultural products to be detected and extracting spectral data; then inputting the spectral data into a classification model for detection; and finally outputting a prediction result of the origin of the agricultural products to be detected;

[0073] Specifically, in this embodiment, relevant hardware is set to extract and process the spectral data of the image of the agricultural product to be tested, and a surface adaptive probe is set to detect the surface of the agricultural product to be tested. By integrating a six-axis robotic arm and a structured light positioning module, it can be ensured that the probe is always perpendicular to the surface of the sample to be tested; secondly, a dual-mode detector is also set, which is formed by parallel connection of a strange electronic device (0.1~3THz) and a UTC-PD optoelectronic module (response time <1ps), which can realize wide-band high-resolution detection; in addition, a pressure feedback module is also set, and a miniature pressure sensor (range 0~10N) is integrated at the end to ensure that the distance between the probe and the sample is constant (2mm±0.1mm).

[0074] like Figure 1-8 As shown, the method further includes training the classification model, and the training step includes:

[0075] Step S1: Capture an image and pre-process the image to obtain an absorption spectrum image. In this application, a pistachio origin identification experiment is used as an example: 65 pistachio samples from Iran, the United States, and other places are collected;

[0076] Further preferably, the step S1 includes:

[0077] Step S11: using a terahertz time-domain spectrometer to collect the time domain signal of the sample in the frequency range of 0.01 to 2 THz;

[0078] Step S12: Obtain frequency domain response through Fourier transform and calculate absorption coefficient;

[0079] Step S13: construct a three-dimensional spectral cube (length×width×depth) according to the absorption coefficient, with each layer corresponding to an absorption spectrum image of a frequency point.

[0080] By using a terahertz time-domain spectrometer to collect the sample's time-domain signal, converting it into a frequency-domain response through Fourier transform, calculating the absorption coefficient, and further constructing a three-dimensional spectral cube, each layer corresponds to an absorption spectrum image at a different frequency point. This process can eliminate the influence of water vapor in the air on the terahertz absorption coefficient, effectively improving the imaging clarity of curved objects, accurately capturing the sample's absorption characteristics at various frequencies, presenting spectral information in a three-dimensional form, and providing richer analytical data, thereby improving the accuracy and robustness of detection. In this way, a more refined data foundation can be provided for subsequent identification of the origin of agricultural products, improving the accuracy of intelligent detection and multi-dimensional analysis capabilities.

[0081] Step S2: using a dynamic weight filtering algorithm to calculate the weighted mean and optimize the frequency band to obtain a dimensionality-reduced image, thereby eliminating water vapor interference;

[0082] Preferably, the step S2 includes:

[0083] S21: For each pixel x i Select its area point and get {x i-3 ,...,x i+3};

[0084] S22: Then calculate the weighted mean based on the domain points. The calculation formula is:

[0085]

[0086] Among them, ω j is the weight, which is obtained by dynamic adjustment according to the ambient humidity and water vapor concentration, and can effectively suppress the interference peak offset.

[0087] This step effectively reduces the impact of noise on the image and enhances the smoothness and continuity of the image by selecting the domain points for each pixel and calculating the weighted mean based on these domain points. By dynamically adjusting the weights to account for changes in ambient humidity and water vapor concentration, image processing under different environmental conditions can be more accurate and robust. The dynamic adjustment of weights enables the algorithm to adapt to different environmental interferences, thereby eliminating water vapor interference and improving the stability and accuracy of detection. This method can maintain high image quality under different sampling environments, improve the accuracy of subsequent image analysis and classification, and ensure the reliability of detection results.

[0088] In addition, the optimization of the frequency band in step S2 further includes:

[0089] Principal component analysis (PCA) is used to screen out a frequency band with high discrimination, for example, in this embodiment, the frequency band screened is 0.16 to 0.73 THz;

[0090] Then, within this range, the calculated absorption coefficient of agricultural products is >0.25cm -1 , with good classification ability.

[0091] Step S3: Extract local features of the reduced-dimensional image and construct a dataset, then use the dataset to train the classification model. Finally, deploy the trained model to the edge of the embedded platform for operation.

[0092] Preferably, the step S3 includes:

[0093] Step S31: binarizing the image after dimensionality reduction to extract the foreground area;

[0094] Specifically, step S31 includes:

[0095] S311: Use dimensionality reduction techniques such as principal component analysis (PCA) to compress data dimensions;

[0096] S312: Binarize and segment the reduced-dimensional image using a global thresholding method (such as the Otsu algorithm) or a local adaptive thresholding method. The Otsu algorithm automatically determines the optimal threshold by maximizing the inter-class variance. For complex scenes, a deep learning-based semantic segmentation network can be used to divide the foreground, background, and suspected areas using dual thresholding.

[0097] S313: Remove noise and fill holes through morphological operations (such as opening and closing operations);

[0098] S314: Use connected domain analysis or contour detection to extract foreground areas, and screen valid targets based on features such as area and distance. Furthermore, for low-quality images, a threshold array system combined with stochastic resonance processing can be used to enhance noise immunity.

[0099] By adopting dimensionality reduction technology to compress data dimensions, the complexity of the data is reduced, the efficiency of subsequent processing is improved, and the computational burden is effectively reduced. By performing binarization segmentation on the dimensionality-reduced image through the global threshold method or the local adaptive threshold method, the image can be quickly and accurately segmented into foreground and background, facilitating the extraction of key features. In the binarized image, morphological operations are used to eliminate noise and fill holes, further improving the image quality and making the foreground area more coherent and clear. By extracting the foreground area through connected domain analysis or contour detection, the target object can be accurately identified, and the target can be screened by features such as area and distance, thereby eliminating irrelevant parts and ensuring that the extracted target is effective and representative. Overall, this step effectively improves the accuracy of image segmentation, optimizes the extraction process of the foreground area, and enhances the reliability and accuracy of subsequent analysis through the combination of multiple image processing technologies.

[0100] Step S32: Divide the fixed-size windows into non-overlapping rectangular windows;

[0101] In this embodiment, the fixed size is set to 17×17 pixels, and those skilled in the art can select specific size parameters according to actual needs.

[0102] Step S33: extract the average spectrum curve in each rectangular window as a local feature, and obtain a total of 419 samples;

[0103] Specifically, step S33 includes:

[0104] S331: Performing window processing on the original signal or image, using a sliding window (such as a Hamming window of fixed length) to divide the data into multiple local areas;

[0105] S332: Calculate the average spectrum curve of the data in each window, usually by converting the time domain signal into the frequency domain through Fast Fourier Transform (FFT), and then taking the average spectrum amplitude of all data points in the window;

[0106] S333: To improve robustness, a multi-scale window strategy (such as superposition of sliding windows of different lengths) can be combined to capture local features in different frequency ranges. Finally, the average spectrum curves of each window are connected in series to form a global feature vector.

[0107] By performing window processing on the original signal or image and using a sliding window to divide the data into multiple local areas, the image can be effectively processed locally to avoid information loss caused by over-smoothing. By calculating the average spectrum curve of the data in each window, the local features within the window can be extracted, which helps to capture local changes and patterns in the image or signal. Combined with a multi-scale window strategy, by using windows of different sizes, local features in different frequency ranges can be captured simultaneously, thereby comprehensively analyzing the detailed information of the image and improving the diversity and adaptability of the features. Finally, the average spectrum curves of each window are connected in series to form a global feature vector, so that the model can consider local and global feature information at the same time, improving the comprehensive understanding and recognition ability of the image or signal. Overall, this step enhances the accuracy of feature extraction and improves the expressiveness and robustness of the model through the combination of window processing and multi-scale strategies. It suppresses noise through local smoothing while retaining the peak and distribution characteristics of the spectrum. It is particularly suitable for processing image data with complex local features and for tasks such as signal classification or pattern recognition.

[0108] Step S34: construct a training set and use several classifiers for modeling;

[0109] Further preferably, the step S34 includes:

[0110] S341: Preprocess the training set, including feature standardization and class imbalance processing. For example, the RUSBoost algorithm solves the sample imbalance problem by integrating random undersampling with AdaBoost.

[0111] S342: Training classifiers such as random forest (by building multiple decision trees for parallel voting) and SVM (using kernel functions to find the optimal classification hyperplane);

[0112] S343: Use network search or cross-validation to tune the key parameters of each classifier, such as the number of trees in random forest and the penalty coefficient of SVM;

[0113] S344: Use ensemble learning (such as bagging or stacking) to combine the strengths of individual classifiers. Use a validation set to evaluate model performance and select the model with the best metrics, such as F1 score or AUC, for prediction. For small sample data, transfer learning can be used to improve generalization capabilities.

[0114] The performance of the classification model was significantly improved through a combination of data preprocessing, model training, tuning, and ensemble learning. First, the training data was optimized through feature standardization and class imbalance processing, avoiding bias caused by varying feature scales and uneven class distribution. Next, random forest and SVM classifiers were trained, leveraging their respective strengths to extract complex patterns in the data. Classifiers were tuned through network searches or cross-validation to further optimize hyperparameters and improve model accuracy. Finally, ensemble learning was used to combine the strengths of multiple classifiers, enhancing the stability and robustness of the model. Model performance was evaluated using a validation set to ensure its generalization ability on unknown data. Overall, this process effectively improved the accuracy and reliability of the classification results.

[0115] Step S35: Introduce several evaluation indicators to evaluate the model and verify the model performance.

[0116] Specifically, in step S35, the plurality of evaluation indicators include two evaluation indicators: single spectrum curve recognition accuracy OA-Patch and image-level majority voting result recognition accuracy OA-Voted.

[0117] The purpose of introducing two evaluation metrics, OA-Patch and OA-Voted, is to assess model performance at different granularities: OA-Patch measures the classification accuracy of a single spectral curve, reflecting the model's ability to recognize local features; OA-Voted calculates image-level classification accuracy by counting the majority votes of all local prediction results within the image, assessing the model's global consistency. This dual-metric evaluation system can both analyze the model's sensitivity to fine-grained features (such as anomaly detection) and verify its reliability in overall decision-making, making it particularly suitable for tasks that require simultaneous attention to local anomalies and global discrimination.

[0118] Further preferably, the method further includes: performing feature compression and edge deployment;

[0119] By deploying the feature compression model to the Cortex-M7 microcontroller and using lightweight convolution kernels (3×3 depthwise separable convolution) to compress spectral data into 32-dimensional feature vectors, the model parameter size is reduced to less than 500KB. It can run on embedded platforms such as ARM Cortex-M7, achieving a processing speed of 30fps and power consumption of less than 0.5W.

[0120] Specifically, in this embodiment, according to the above method, a pistachio origin identification experimental process and confusion matrix generation steps are also provided:

[0121] First, data preparation and preprocessing:

[0122] Data source: Pistachio samples were collected from Iran, Australia, and the United States, and spectral data (such as near-infrared spectra or hyperspectral images) were extracted.

[0123] Feature extraction: Preprocess each spectral data (denoising, smoothing, normalization), and extract features (such as the average spectral curve, features after principal component analysis dimensionality reduction).

[0124] Label assignment: Each sample is labeled with its true origin (Iran, Australia, United States).

[0125] Training set and test set division: Randomly divide the data set in a ratio (such as 7:3 or 8:2) to ensure a balanced distribution of samples in each category.

[0126] Perform model training (RUSBoost algorithm):

[0127] RUSBoost principle: Combining random undersampling (RUS) and AdaBoost ensemble learning to solve the problem of category imbalance.

[0128] Random Undersampling (RUS): In each iteration, the majority class samples are randomly downsampled to make the number of samples of each class close to balance.

[0129] AdaBoost enhancement: Build an integrated model based on decision trees (weak classifiers) and adjust sample weights to make the model pay more attention to difficult-to-classify samples.

[0130] Hyperparameter tuning: Parameters such as decision tree depth, learning rate, and number of iterations were optimized through cross-validation (e.g., 5-fold CV). As shown in Table 1, the experimental results of 5-fold cross-validation for pistachio origin identification can be obtained.

[0131]

[0132] Table 1

[0133] Model evaluation and confusion matrix generation:

[0134] Test set prediction: Use the trained RUSBoost model to predict the test set samples and output the predicted label of each sample (Iran, Australia, United States).

[0135] Confusion matrix calculation:

[0136] like Figure 8 As shown, the matching of true labels (rows) and predicted labels (columns) is counted to form a 3×3 matrix:

[0137] Line: Real origin (Iran, United States).

[0138] Column: Origin predicted by the model.

[0139] The diagonal elements represent the number of correctly classified samples, and the off-diagonal elements represent the number of misclassified samples.

[0140] Indicator calculation:

[0141] OA-Patch (single sample accuracy): the proportion of correctly classified samples in all test samples.

[0142] OA-Voted (image-level voting accuracy): If a batch (image) contains multiple spectra, majority voting is used to determine the predicted label for the batch, and then the accuracy is calculated.

[0143] Analysis of the Identification Situation between the United States and Iran

[0144] Extract from the confusion matrix:

[0145] The recognition rate of the origin of the United States is: the number of samples correctly classified as the United States / the total number of samples that are actually the United States.

[0146] Identification rate of Iranian origin: number of samples correctly classified as Iranian / total number of samples that are actually Iranian.

[0147] Confusion situations: For example, the proportion of American samples mistakenly identified as Iranian, or the proportion of Iranian samples mistakenly identified as American.

[0148] Finally, according to the results, we can get Figure 8 The resulting visualization is shown, where:

[0149] Figure 7 (Origin Batch Statistics): The bar chart shows the distribution of sample batches from Iran and the United States in the training or test set.

[0150] Figure 8 (Confusion Matrix): This matrix is ​​presented as a heatmap, showing the classification accuracy and confusion level for each category, with a focus on the recognition performance of the United States and Iran.

[0151] The confusion matrix provides a visual representation of the model's ability to discriminate between different origins, with a particular focus on the recognition rate and confusion between the United States and Iran. Dual metrics (OA-Patch and OA-Voted) assess the model's classification reliability at the single-sample and batch levels, respectively.

[0152] Example 2

[0153] According to the terahertz spectral image intelligent detection method for agricultural product origin identification described in Example 1, this embodiment also provides an electronic device, such as Figure 9 As shown, Figure 9Schematic diagram of the structure of the electronic device provided by this solution, which may include: a processor (processor) 910, a communication interface (Communications Interface) 920, a memory (memory) 930 and a communication bus 940, wherein the processor 910, the communication interface 920, and the memory 930 communicate with each other through the communication bus 940. The processor 910 can call the logic instructions in the memory 930 to execute the terahertz spectral image intelligent detection method for agricultural product origin identification. The method includes: collecting images of agricultural products to be detected and extracting spectral data, then inputting the spectral data into a classification model for detection, and finally outputting a prediction result of the origin of the agricultural products to be detected; and executing training of the classification model. The training steps include: S1: collecting images and preprocessing the images to obtain absorption spectrum images; S2: using a dynamic weight filtering algorithm to calculate the weighted mean and optimize the frequency band to obtain a reduced-dimensional image; S3: extracting local features of the reduced-dimensional image and constructing a data set, and then using the data set to train the classification model, and finally deploying the trained model to the edge of the embedded platform for operation.

[0154] In addition, the logic instructions in the above-mentioned memory 930 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of this solution, or the part that contributes to the existing technology, or the part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this solution. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0155] On the other hand, this embodiment also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the terahertz spectral image intelligent detection method for agricultural product origin identification provided by the above methods. The method includes: collecting images of agricultural products to be detected and extracting spectral data, then inputting the spectral data into a classification model for detection, and finally outputting a prediction result of the origin of the agricultural products to be detected; and executing training of the classification model, the training steps including: S1: performing image acquisition and preprocessing the image to obtain an absorption spectrum image; S2: calculating the weighted mean using a dynamic weight filtering algorithm, and optimizing the frequency band to obtain a reduced-dimensional image; S3: extracting local features of the reduced-dimensional image and constructing a data set, and then using the data set to train the classification model, and finally deploying the trained model to the edge of an embedded platform for operation.

[0156] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0157] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0158] Obviously, the above embodiments of this solution are merely examples for the purpose of clarifying this solution and are not intended to limit the implementation of this solution. Those skilled in the art will be able to make other variations or modifications based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this solution shall be included within the scope of protection of the claims of this solution.

Claims

1. A terahertz spectral image intelligent detection method for identifying the origin of agricultural products, the method comprising: Collect images of the agricultural products to be tested and extract spectral data, then input the spectral data into a classification model for testing, and finally output a prediction result of the origin of the agricultural products to be tested; Characterized in that the method further comprises training the classification model, and the training step comprises: S1: perform image acquisition and image preprocessing to obtain an absorption spectrum image; S2: Use the dynamic weight filtering algorithm to calculate the weighted mean and optimize the frequency band to obtain the image after dimensionality reduction; S3: Extract local features from the reduced-dimensional image and construct a dataset. The dataset is then used to train the classification model. Finally, the trained model is deployed to the edge on the embedded platform for operation.

2. The terahertz spectral image intelligent detection method for agricultural product origin identification according to claim 1 is characterized in that: The step S1 includes: S11: Use terahertz time-domain spectrometer to collect the time domain signal of the sample; S12: Obtain the frequency domain response through Fourier transform and calculate the absorption coefficient; S13: Construct a three-dimensional spectral cube based on the absorption coefficient, where each layer corresponds to an absorption spectrum image of a frequency point.

3. The terahertz spectral image intelligent detection method for agricultural product origin identification according to claim 2 is characterized in that: The step S2 includes: S21: For each pixel x i Select its area point and get {x i-3 ,...,x i+3 }; S22: Then calculate the weighted mean based on the domain points. The calculation formula is: Among them, ω j is the weight, which is obtained by dynamic adjustment according to the ambient humidity and water vapor concentration.

4. The terahertz spectral image intelligent detection method for agricultural product origin identification according to claim 3 is characterized in that: The step S3 includes: S31: Binarize the image after dimensionality reduction and extract the foreground area; S32: Divide the fixed size into non-overlapping rectangular windows; S33: extracting the average spectrum curve within each rectangular window as a local feature; S34: Construct a training set and use several classifiers for modeling; S35: Introduce several evaluation indicators to evaluate the model and verify the model performance.

5. The terahertz spectral image intelligent detection method for agricultural product origin identification according to claim 4 is characterized in that: The step S31 includes: S311: Use dimensionality reduction techniques to compress data dimensions; S312: performing binary segmentation on the dimensionality-reduced image by using a global threshold method or a local adaptive threshold method; S313: Remove noise and fill holes through morphological operations; S314: Use connected domain analysis or contour detection to extract the foreground area, and screen valid targets based on features such as area and distance.

6. The terahertz spectral image intelligent detection method for agricultural product origin identification according to claim 5 is characterized in that: The step S33 includes: S331: Performing window processing on the original signal or image, dividing the data into multiple local areas using a sliding window; S332: Calculate the average spectrum curve of the data in each window; S333: Combine the multi-scale window strategy to capture the local features of different frequency ranges, and finally concatenate the average spectrum curves of each window to form a global feature vector.

7. The terahertz spectral image intelligent detection method for agricultural product origin identification according to claim 6, characterized in that: The step S34 includes: S341: Preprocess the training set, including feature normalization and class imbalance processing; S342: training random forest and SVM classifiers; S343: Use network search or cross-validation to tune each classifier; S344: Fusion of the advantages of each classifier through ensemble learning, and use of validation set to evaluate model performance.

8. The terahertz spectral image intelligent detection method for agricultural product origin identification according to claim 7, characterized in that: In step S35, the plurality of evaluation indicators include two evaluation indicators: single spectrum curve recognition accuracy and image-level majority voting result recognition accuracy.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the terahertz spectral image intelligent detection method for agricultural product origin identification according to any one of claims 1 to 8 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the terahertz spectral image intelligent detection method for agricultural product origin identification according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Terahertz ground detection system for detecting safety of bottom of vehicle and detection method thereof

    CN113418891A