A non-destructive testing method and system for formaldehyde residue in aquatic products

By combining Raman spectroscopy with the InceptionTime model of deep learning, the problems of slow detection speed and low accuracy of formaldehyde in aquatic products have been solved, achieving rapid, accurate, and non-destructive testing, which is suitable for food safety supervision and supply chain management.

CN119804410BActive Publication Date: 2026-01-30FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411904897.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2026-01-30
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing technologies for formaldehyde detection in aquatic products are slow and inaccurate, making it difficult to achieve rapid and accurate formaldehyde residue detection in complex environments.

Method used

By combining Raman spectroscopy with deep learning, an InceptionTime model was constructed. By classifying a dataset of formaldehyde residues in aquatic products, optimizing the model structure, and performing spectral preprocessing, the accuracy and speed of detection were improved.

Benefits of technology

It enables rapid, accurate, and non-destructive detection of formaldehyde residues in aquatic products, improves detection efficiency, and is suitable for food safety supervision and supply chain management, with broad application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119804410B_ABST
    Figure CN119804410B_ABST
Patent Text Reader

Abstract

This invention discloses a non-destructive testing method and system for formaldehyde residue in aquatic products, belonging to the field of formaldehyde detection technology. The method includes the following steps: Raman spectroscopy is performed on aquatic product samples soaked in formaldehyde solutions of different concentrations to determine the formaldehyde residue and construct a formaldehyde residue dataset; different models are used to classify the Raman spectral formaldehyde residue data, and the model performance is evaluated by classification accuracy, with the InceptionTime model selected as the classification model; Raman spectroscopy is performed on aquatic product samples soaked in formaldehyde-like substances, and a selective dataset is constructed using mixed data; the selective dataset is used to verify the selectivity of the classification model for formaldehyde; the optimal spectral preprocessing method is selected to optimize the model structure; and the classification results of the classification model are verified and analyzed. This invention uses deep learning to provide an accurate, simple, and economical method for detecting formaldehyde in aquatic products, with broader application prospects in food safety.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of formaldehyde detection, and particularly relates to a nondestructive detection method and system for formaldehyde residues in aquatic products. BACKGROUND

[0002] Formaldehyde (FA) is a colorless, flammable, and highly reactive gas that readily polymerizes under normal conditions. It has a pungent and distinct odor, and high concentration exposure can cause a burning sensation in the eyes, nasal cavity, and lungs. Long-term excessive exposure to formaldehyde can irritate the mucous membranes and cause inflammation of the liver and kidneys. Formaldehyde can also cause damage to the gastrointestinal tract, kidneys, liver, and lungs, and poses a risk of carcinogenesis. Formaldehyde has been listed in the directory of non-edible substances and food additives prohibited from use in food, and is prohibited from use in food production and operation. However, in reality, there are still cases of illegal use of formaldehyde in the breeding and sales of aquatic products, mainly for the purpose of achieving the effects of preservation and sterilization. Eating food containing formaldehyde can cause a variety of health problems, and the internationally recognized safety standard is 5mg / kg. However, researchers have found that the content of formaldehyde in seafood illegally added with formaldehyde as a preservative generally exceeds 300mg / kg, and in extreme cases, it can even reach 4250mg / kg.

[0003] Compared with other formaldehyde detection methods, surface-enhanced Raman spectroscopy can detect formaldehyde at the nmol level, and is simple to operate and fast in detection. In addition, a variety of small portable spectrometers have been developed, which are a potential excellent on-site rapid analysis technology in the detection of formaldehyde in aquatic products in the field. However, the current detection and analysis method still faces the challenge of time efficiency, especially in the reaction stage of formaldehyde and the reactant, which usually takes at least 20min for the subsequent detection step. In view of this, developing a new detection strategy to shorten the analysis time is the key to improving the detection efficiency.

[0004] In recent years, machine learning algorithms have been widely used in the field of spectral signal analysis. Machine learning is a branch of artificial intelligence, which corresponds to a system that can acquire knowledge by extracting features from raw data, and then use this newly acquired knowledge to solve real-world problems by making decisions. When dealing with complex background interference, although portable Raman spectrometers are more in line with actual detection scenarios, their sensitivity is relatively low, making it difficult to judge the value of the information visually, highlight the characteristics of the target detection material, and conduct quantitative analysis. Therefore, it is an urgent need to use artificial intelligence technology to deeply mine the deep information of Raman spectrum and significantly improve the accuracy of the judgment of the detected material. Combining portable Raman spectroscopy technology with artificial intelligence can meet the needs of specific application scenarios and realize the possibility of real-time monitoring. In addition, by integrating or connecting professional spectral analysis software, not only can the user's use difficulty be reduced, but also complex problems such as mixture identification and similar substance analysis can be effectively solved. SUMMARY

[0005] The purpose of the present application is to provide a non-destructive detection method and system for formaldehyde residues in aquatic products, to solve the problems of slow speed and low detection accuracy in the formaldehyde detection process of existing technologies in the background art.

[0006] To achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows:

[0007] The first aspect of the present application proposes a non-destructive detection method for formaldehyde residues in aquatic products, comprising the following steps:

[0008] S1, establishing a formaldehyde residue data set; collecting Raman spectra of aquatic product samples soaked in formaldehyde solutions of different concentrations, determining the formaldehyde residue in the aquatic product samples, and dividing the collected Raman spectrum data into different groups according to the formaldehyde residue determination results, and constructing a Raman spectrum formaldehyde residue data set;

[0009] S2, constructing a classification model; using random forest, support vector machine, XGBoost, convolutional neural network for comparison and InceptionTime model to classify the Raman spectrum formaldehyde residue data, evaluating the model performance through classification accuracy, and selecting InceptionTime model as the classification model;

[0010] S3, establishing a selective data set; collecting Raman spectra of aquatic product samples soaked in formaldehyde similar substances, using the same label for the data of aquatic products soaked in formaldehyde similar substances, using the same label for the data of aquatic products soaked in water and formaldehyde, and constructing a selective data set using the mixed data;

[0011] S4, verifying the selectivity of the classification model to formaldehyde; using the InceptionTime model to select the mixed data to determine the accuracy of formaldehyde identification;

[0012] S5, optimizing the structure of the classification model; using different spectral pretreatment methods to process the Raman spectrum formaldehyde residual data set, selecting the optimal spectral pretreatment method; based on the pretreated Raman spectrum formaldehyde residual data set, optimizing the structure of the InceptionTime model;

[0013] S6, verifying and analyzing the classification results of the classification model.

[0014] Preferably, before data collection, the aquatic products are soaked in different concentrations of formaldehyde, and the formaldehyde preservation effect is evaluated by observing the appearance change and the content change of volatile nitrogen.

[0015] Preferably, the formaldehyde residual amount in the aquatic product sample in S1 is determined by the spectrophotometry method specified in the SC-T 3025-2006 standard. In order to verify the accuracy of the spectrophotometry method, part of the determination results are compared and verified with the determination results of HPLC. The formaldehyde residual amount of each aquatic product determined by the water vapor distillation method combined with the spectrophotometry method is used as the formaldehyde residual amount label of the corresponding spectrum, and the data is divided into six groups according to the formaldehyde residual amount determination results: ≥1000mg / kg, 500-1000mg / kg, 100-500mg / kg, 5-100mg / kg, >100mg / kg and >5mg / kg.

[0016] Preferably, the InceptionTime model in S2 is as follows:

[0017] The InceptionTime model is composed of residual blocks, and each residual block contains three Inception sub-modules; when processing an M-dimensional multivariate time series, the data is first passed through the bottleneck layer of the Inception sub-module, which uses m filters with a length of 1 and a step of 1 to convert the M-dimensional signal to m-dimensional;

[0018] After the residual block, the InceptionTime model uses a global average pooling layer to average process the multivariate time series in the time dimension; then a standard fully connected Softmax layer is used for classification, the number of neurons of which matches the number of categories in the data set, so as to output the classification result.

[0019] Further, 20% of all Raman spectra data of aquatic product individuals were randomly selected from the aquatic product data set to form a test set. The remaining aquatic product spectrum data was used as a training set. In order to improve the reliability of the results, the present application adopts three repeated tests, and the average results of the three tests are taken as the final evaluation index. Two deep learning models, InceptionTime and Cmp_CNN, and traditional machine learning models RF, XGBoost and SVM are used to classify the data set, and the performance is compared according to the classification ACC.

[0020] The model performance is evaluated by the classification accuracy (ACC), and the applicability and efficiency of machine learning in the classification of non-destructive Raman spectrum data set are analyzed by combining the confusion matrix results.

[0021] Preferably, the S3 is specifically as follows:

[0022] Formaldehyde-like substances are selected as methanol, ethanol and acetaldehyde, and Raman spectrum collection is performed on aquatic product samples soaked in methanol, ethanol and acetaldehyde. The aquatic product data soaked in methanol, ethanol and acetaldehyde are marked with the same label.

[0023] Further, the Raman spectrum of the surface of the aquatic product soaked in methanol (20%), ethanol (20%) and acetaldehyde (7%) for 1 hour is marked as "1", and the Raman spectrum of the surface of the aquatic product soaked in water and formaldehyde for 1 hour is marked as "0". In order to keep the balance of the data, the "0" data is divided into layers according to the aquatic product individuals, and these data are used for classification test.

[0024] Preferably, the S4 is specifically as follows:

[0025] According to the detection limit category, a corresponding classifier is designed in the InceptionTime model; four substances including formaldehyde are distinguished by the InceptionTime model to judge the selectivity of the model for formaldehyde.

[0026] Preferably, the different spectrum pretreatment methods in S5 are specifically as follows:

[0027] Different spectrum pretreatment methods include multivariate scatter correction, standard normal variable transformation, wavelet transformation, SG and baseline correction; the Raman spectrum formaldehyde residue data set is pretreated by multivariate scatter correction, standard normal variable transformation, wavelet transformation, SG and baseline correction respectively, and then model verification is performed to evaluate the influence of different spectrum pretreatment methods on the classification results of the classification model.

[0028] Preferably, the structure improvement and optimization of the InceptionTime model in S5 are specifically as follows:

[0029] The number of residual blocks is adjusted, the number of residual blocks of the InceptionTime model is set to 1, 2 and 3 respectively, and other model parameters remain unchanged, the classification results of the pretreated Raman spectrum formaldehyde residual data set are determined, and the best residual block configuration is determined; the InceptionTime model structure is optimized and is composed of two residual blocks.

[0030] Further, when determining the best residual block configuration, the Attention, LSTM, GRU and Transformer modules are added to compare the model performance, and the TimesNet model is compared with the InceptionTime model.

[0031] Preferably, the S6 is specifically as follows: the model classification result is analyzed in combination with the Shap model classification weight visualization result.

[0032] Further, the model classification result is analyzed in combination with the Shap model classification weight visualization result, and the specific process is as follows:

[0033] A Shap value is assigned to each feature of each prediction sample by using the Shap model, which represents the contribution degree of the feature to the prediction, and the weights of different features are projected on the spectrum for visualization; the classification basis of the model is explained and verified by using the Shapley additive explanation method, so that the transparency and credibility of the model decision process are enhanced.

[0034] The second aspect of the present application provides a nondestructive detection system for formaldehyde residues in aquatic products, which comprises the InceptionTime model after the InceptionTime model structure is optimized by the nondestructive detection method for formaldehyde residues in aquatic products.

[0035] Compared with the prior art, the present application has the following advantages:

[0036] (1) The method successfully realizes the nondestructive detection of formaldehyde residues in aquatic products by combining Raman spectrum with deep learning, which proves that the method is fast, accurate, nondestructive and noninvasive. The method has a broad application prospect in the field of food safety. In the future, this technology can be extended to the rapid detection of other foods and related products, such as identifying harmful substances (such as heavy metals and pesticide residues) or detecting food adulteration and deterioration. In terms of market supervision, portable devices combining Raman spectrum with deep learning algorithms can realize real-time detection, improve supervision efficiency, reduce costs, and provide fast feedback at various stages of the supply chain. With further optimization of detection range and sensitivity, combined with big data and Internet of Things technology, the method is expected to support the intelligent development of food safety monitoring systems.

[0037] (2), In the application, a deep neural network model InceptionTime suitable for water product surface Raman spectrum data set is designed and constructed, and corresponding classifiers are designed according to different formaldehyde residual amount groups and detection limit categories, and two types of data are trained and tested respectively. The results of the application prove that the model has excellent detection performance in different formaldehyde residual concentration groups. When the detection limit is 100 mg / kg and 5 mg / kg, the accuracy is 85.17% and 84.40% respectively. Weight visualization shows that the model pays special attention to peaks related to amino acids, lipids and nucleotides. The research of the application shows that deep learning provides an accurate, simple and economical method for detecting formaldehyde in aquatic products, and has a wider application prospect in food safety.

[0038] (3), In the application, InceptionTime model and other models are used to classify and analyze Raman spectrum data sets processed by different pretreatment methods. On the Raman spectrum data set processed by different pretreatment methods, the classification performance of InceptionTime model is the best. Compared with other pretreatment methods, SG pretreatment method can further improve the classification performance. When the residual block number is configured as 2, the accuracy of InceptionTime model on the test data set is the highest, and this optimal performance may be attributed to the balance achieved under this specific setting. The comprehensive comparison of model performance shows that neither the improved InceptionTime model nor the alternative architecture (such as TimesNet) can match the high accuracy obtained by the InceptionTime model with two residual blocks. This finding clearly indicates that in the current experimental environment, the InceptionTime model with 2 residual blocks represents the best choice. The model structure is optimized to improve the classification accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 The flowchart of the nondestructive detection method of formaldehyde residue in aquatic products in the application;

[0040] Figure 2 The schematic diagram of the data set construction process in the application;

[0041] Figure 3 The analysis diagram of the formaldehyde preservative effect in the application in terms of appearance change and volatile nitrogen content change (Fig. a is the appearance change diagram, and Fig. b is the volatile nitrogen content change diagram);

[0042] Figure 4 The architecture diagram of the InceptionTime model in the application;

[0043] Figure 5The schematic diagram for classification performance comparison of different pretreatment methods on formaldehyde residual amount in the application (Fig. a-Fig. c are respectively the accuracy rate change curve, loss change curve and test set classification accuracy rate result of the model in the training process of 500-1000 mg / kg data set after treatment by different pretreatment methods);

[0044] Figure 6 The spectral analysis in the application Figure 1 (Fig. a is a distribution diagram of negative and positive aquatic products containing formaldehyde residues in non-destructive Raman spectrum data, and Fig. b is a difference diagram of negative data and positive data);

[0045] Figure 7 The spectral analysis in the application Figure 2 (Fig. c is a spectrum diagram of negative data and different groups of positive data, and Fig. d is a difference diagram of the mean value of negative data and the mean value of positive data);

[0046] Figure 8 The spectral analysis in the application Figure 3 (Fig. e is a diagram in which the SHAP interpreter focuses on the first 20 characteristics of the InceptionTime model, and Fig. f is a diagram for visualizing all feature weights on the frequency spectrum. The color depth represents the weight of each feature: the darker the color, the higher the feature weight, indicating greater importance to the model prediction result). DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.

[0048] Embodiment 1

[0049] In this embodiment, the aquatic product is prawn, and the formaldehyde residue detection of prawn is specifically performed. To solve the problems in the prior art, such as Figure 1 As shown in the drawings, the application provides a non-destructive detection method for formaldehyde residue, which specifically comprises:

[0050] Step 1, soak prawns in different concentrations of formaldehyde, and evaluate the formaldehyde preservative effect by observing the appearance change and the content change of volatile basic nitrogen.

[0051] Shrimp were purchased from a supermarket, placed on crushed ice and stunned, washed with water and air-dried. The shrimp were soaked in different concentrations of formaldehyde solution and stored at 4°C for 7 days. The freshness indicator of shrimp, total volatile basic nitrogen, was determined by microdiffusion method. Specifically, the head, shell and gut of shrimp were removed from the shrimp, and the shrimp meat was chopped and ground into a uniform paste with a pre-cooled mortar. 2 grams of shrimp meat was mixed with 10 mL of 10% trichloroacetic acid solution. Stir for 1 min, stand for 30 min, centrifuge at 8000 rpm for 5 min, and separate the supernatant. Add 1 mL of boric acid solution and 1 drop of mixed indicator to the central chamber of the diffusion disc. In the outer chamber, place 1 mL of supernatant, cover with ground glass, and mix quickly with 1 mL of saturated potassium carbonate solution. Seal the culture dish tightly, rotate gently, and incubate at 37°C for 2 h. After cooling to room temperature, remove the cover and titrate the inner circle solution with 0.01 mol / L hydrochloric acid. The amount of hydrochloric acid consumed is used to quantify the volatile basic nitrogen. Reagent blank is included, and each treatment is analyzed in triplicate.

[0052] It was found that the control group showed obvious corruption on the fourth day, with blackening of the head and joints. In contrast, the shrimp soaked in formaldehyde solution with a concentration greater than 1 mol / L showed a significantly reduced rate of decay and no obvious signs of corruption. The volatile basic nitrogen content in these samples remained below the threshold value of 20 mg / 100 g.

[0053] Step two, collect Raman spectra of shrimp soaked in different concentrations of formaldehyde solution, collect data from multiple batches of shrimp samples, and construct a Raman spectrum dataset for non-destructive detection of shrimp, as shown in Figure 2-3 .

[0054] Shrimp were soaked in different solutions (0, 0.1, 0.25, 0.5, 1, 2 mol / L) for 1 h, wiped clean, and the Raman spectrum of the shrimp surface was collected directly. The Raman spectrum of the sample was obtained using a handheld Raman spectrometer (RMS1000, Oceanhood) equipped with a 785 nm laser, with a laser power of 300 mW and an integration time of 1 s. For each shrimp sample, at least 30 points were collected from the surface of the shrimp, covering a range of 300-3500 cm -1 Figure 2 ) using the "findpeaks" function in Python. Abnormal values were excluded based on standards such as peak position, peak intensity, and full width at half maximum. A total of 482 shrimp surface Raman spectra were collected, resulting in a dataset of 19504 Raman spectra.

[0055] ​The actual residual amount of formaldehyde in shrimp samples was determined by spectrophotometry specified in the SC-T 3025-2006 standard. Five grams of crushed shrimp meat was mixed with 20 mL of distilled water, stirred and then allowed to stand for 30 min. Then 10 mL of phosphoric acid solution was added, and distilled to 200 mL of distillate, including a blank control. A certain volume (1-10 mL) of distillate was diluted with water to a final volume of 10 mL, and then 1 mL of acetylacetone solution was added. Then the mixture was heated in a boiling water bath for 10 min, and cooled to room temperature. The absorbance was measured at 413 nm using a microplate reader (SYNERGY-H1, BioTek, USA). According to the different residual amounts of formaldehyde, the positive data set was subdivided into 5-100 mg / kg, 100-500 mg / kg, 500-1000 mg / kg and ≥1000 mg / kg groups, as shown in Table 1. Before model training, each spectral data point was normalized to improve training speed.

[0056] Table 1 shows the formaldehyde residual amount data set of shrimp

[0057]

[0058] Step three, selection, construction and performance comparison of classification model.

[0059] Random forest (RF), support vector machine (SVM), XGBoost, compare convolutional neural network (Cmp_CNN) and InceptionTime model were used to classify shrimp samples with different formaldehyde concentration levels. The model performance was evaluated by accuracy (ACC), and the applicability and efficiency of machine learning in the classification of non-destructive Raman spectral data set were analyzed combined with the confusion matrix results. At the same time, the classification performance was further improved through spectral data preprocessing and model structure optimization.

[0060] In this embodiment, InceptionTime was used to analyze the Raman spectral data.

[0061] The model structure of InceptionTime is composed of two different residual blocks (Residual block), each of which contains three Inception sub-modules, replacing the traditional fully connected layer, such as Figure 4The Inception module aims to process time series data in parallel through multiple filters, thus being able to extract different features. This design allows the input to be directly passed to the input of the next residual block through a shortcut connection, effectively alleviating the problem of gradient disappearance while maintaining the prediction accuracy and improving the scalability of the model. Multiple filters are applied to extract features. When processing an M-dimensional multivariate time series, the data is first passed through the bottleneck layer of the Inception module, which uses m filters of length 1 and stride 1 to convert the M-dimensional signal to m dimensions. After the residual block, the network uses a global average pooling layer to average the multivariate time series in the time dimension. Finally, a standard fully connected Softmax layer is used for classification, with the number of neurons matching the number of classes in the dataset, thus outputting the final classification result.

[0062] In the implementation of the present application, four convolution kernels are used in each Inception module, with sizes of 1, 5, 11 and 23. 20% of the Raman spectra from the shrimp samples were randomly selected as the test set, and the remaining spectra were used for model training. To ensure the reliability of the results, the process was repeated for more than 10 iterations, and the average results were used as the final performance indicators. During the training process, the training set was further divided into two subsets, 80% for model training and 20% as a validation set to evaluate the training process and the reliability of the results.

[0063] All models were subjected to 5-fold cross-validation, and the best-performing weights in training were used for the final testing phase. To accurately classify the Raman spectrum data, the performance of InceptionTime was compared with other models, including Cmp_CNN, XGBoost, RF and SVM, and the results are shown in Table 2.

[0064] Table 2 Classification accuracy of different models

[0065] Group RF SVM XGBoost Cmp_CNN InceptionTime ≥1000mg / kg 0.7688 0.8101 0.8170 0.8457 0.8911 500-1000mg / kg 0.7815 0.8096 0.8216 0.8380 0.8702 100-500mg / kg 0.7209 0.7505 0.7556 0.7592 0.8032 5-100mg / kg 0.7206 0.7124 0.7625 0.7714 0.7888 >100mg / kg 0.7845 0.8171 0.8055 0.8174 0.8517 >5mg / kg 0.7871 0.8017 0.8026 0.8206 0.8440

[0066] Table 2 shows that in all concentration groups, the InceptionTime model is always superior to the other four models (RF, SVM, XGBoost and Cmp_CNN), achieving the highest classification accuracy, with a classification accuracy of 0.8911 in the ≥1000 mg / kg group and 0.8702 in the 500-1000 mg / kg group. Even in the lower concentration groups of 100-500 mg / kg and 5-100 mg / kg, InceptionTime still maintains a high accuracy (0.8032 and 0.7888, respectively).

[0067] Step four, classification of single concentration groups and different detection limits for the dataset.

[0068] The Raman spectra of the surface of shrimps soaked in methanol (20%), ethanol (20%), and acetaldehyde (7%) for 1 h were collected by the present application and labeled as "1", while the Raman spectra of the surface of shrimps soaked in water and formaldehyde for 1 h were labeled as "0". The detailed information of the data set is shown in Table 1. In order to keep the balance of the data, the "0" data was stratified according to the individual shrimps. Then these data were used for classification test.

[0069] Step five, the selectivity of the model to formaldehyde was verified.

[0070] In this embodiment, the InceptionTime model was used to distinguish four substances including formaldehyde, so as to judge the selectivity of the model to formaldehyde.

[0071] Table 3 Identification results of formaldehyde in similar substances by different models

[0072] Similar substance RF SVM XGBoost Cmp_CNN InceptionTime Acetaldehyde 0.9442 0.9540 0.9587 0.9710 0.9734 Ethanol 0.8772 0.9104 0.8970 0.9241 0.9372 Methanol 0.8688 0.9057 0.9086 0.9139 0.9170

[0073] As shown in Table 3, the InceptionTime model has a high identification accuracy for three substances, and the identification classification performance for acetaldehyde is 0.9734. This indicates that the InceptionTime model can effectively distinguish the Raman spectra of shrimps soaked in formaldehyde from shrimps soaked in other similar substances.

[0074] Step six, the model structure is optimized.

[0075] In this embodiment, a variety of pretreatment methods were used, including multiplicative scatter correction (MSC), standard normal variate transform (SNV), Hilbert transform (HT), Savitzky-Golay (SG), and baseline correction. After the data were processed by MSC, SNV, HT, SG, and baseline correction, the model was verified to evaluate the influence of pretreatment on the model results.

[0076] As one of the commonly used pretreatment methods, MSC aims to solve the problem of spectral data scattering caused by individual differences of samples. The spectral changes caused by scattering are often greater than the changes of the sample composition itself, and MSC can effectively correct these changes to obtain more true spectral information. This pretreatment algorithm focuses on processing errors caused by scattering. It assumes that the spectral intensity distribution is approximately normally distributed, and adjusts the spectral data through transformation to reduce the scattering effect.

[0077] The process of MSC conversion can be represented as:

[0078]

[0079] where X ij is the element in the original data matrix X, X msc is the element in the MSC transformed data matrix, b i and k i are the intercept and slope obtained by linear regression of the ith row in the original data matrix X and the average spectrum of the whole data matrix.

[0080] SNV is a preprocessing algorithm that focuses on handling errors caused by scattering. It assumes that the spectral intensity distribution is approximately normally distributed and adjusts the spectral data by transformation to reduce scattering effects.

[0081] The process of SNV transformation can be represented as:

[0082]

[0083] where X snv is the element in the SNV transformed data matrix, X i is the average of the ith row in the original data matrix X. The denominator part is the standard deviation of the ith row in the original data matrix X.

[0084] For wavelet transform, given that the effective signal in spectral information often appears in the form of low frequency, while noise usually appears as high frequency components, the present application uses HT to process the original spectral data. Through this method, the spectrum is decomposed into components of different frequencies. Subsequently, the present application uses thresholding algorithm, especially the hilbert function in the signal module of the scipy library, to effectively filter out the decomposed high-frequency noise, thereby achieving the effect of noise reduction.

[0085] SG smoothing uses the savgol_filter function in the SciPy library to apply the Savitzky-Golay filter for smoothing and denoising. The length of the filter window is set to 21, considering the data of 21 points (including the point itself) before and after the point for smoothing processing; the order of the polynomial used to fit each window subset is 5, using a 5th order polynomial to fit the data within the window. This technique can reduce errors caused by random and mechanical noise, making the wavelength information smoother, which helps to eliminate the impact of noise on spectral data.

[0086] The polyfit function of NumPy is used to perform polynomial fitting on the data for baseline correction, setting the fitting order to 5. The process of polynomial fitting to remove baseline can be represented as:

[0087] X corrected = X ij - P i (j)

[0088] where X is a spectral data matrix, each row i (i = 1, 2, …, m) represents a sample, and each column j (j = 1, 2, …, n) represents a spectral variable. corrected is the element in the baseline-corrected data matrix, P i (j) is the baseline obtained by polynomial fitting of the i-th row in the original data matrix X.

[0089] The purpose of baseline correction is to eliminate the baseline shift phenomenon in the spectral signal caused by environmental factors, instrument response non-uniformity, or sample scattering, etc. Baseline shift will cause distortion of the shape of the spectral curve and shift of the spectral peak position, thereby affecting the accuracy of qualitative and quantitative analysis of spectral data. Through baseline correction, the baseline part in the spectral data will be adjusted to approach the zero baseline, reducing or eliminating the interference of the baseline on the true spectral peak, and ensuring that the recovered spectral signal is closer to the true absorption characteristics of the sample.

[0090] As shown in Figure 5 , each group of models began to approximate convergence after 40 steps of training. After reaching stability, the accuracy of the original model, the standard normal variable transformation (SNV) model, the Hilbert transform (HT) model, and the Savitzky-Golay (SG) model were comparable. However, the Multiplicative scatter correction (MSC) and baseline correction methods showed relatively poor performance. This trend is consistent with the change of loss function. Different models were used to classify and analyze the Raman spectral data sets processed by different preprocessing methods. From Figure 5 , it can be seen that the classification performance of the InceptionTime model is optimal on the Raman spectral data sets processed by different preprocessing methods. In addition, compared with other methods, the SG preprocessing method can further improve the classification performance. Therefore, the subsequent classification work of the present application is carried out on the Raman spectral data set preprocessed by the SG method.

[0091] Table 4 Verification results of structure optimization of InceptionTime model

[0092]

[0093] The data analysis of Table 4 shows that the InceptionTime model has the highest accuracy of 0.8797 on the test dataset when the number of residual blocks is configured to 2. This optimal performance can be attributed to the balance achieved in this particular setting. With only 1 residual block, the model lacks sufficient depth to capture a sufficient amount of information. Conversely, increasing the number of residual blocks to 3 can make the model too deep, leading to possible learning bias and consequent decrease in accuracy. Therefore, configuring the number of residual blocks to 2 can effectively balance the risk of feature extraction and overfitting, while also saving computational power. The comprehensive comparison of model performance shows that neither the improved InceptionTime model nor the alternative architectures such as TimesNet can match the high precision achieved by the InceptionTime model with two residual blocks. This finding explicitly indicates that the InceptionTime model with 2 residual blocks represents the best choice in the current experimental environment.

[0094] Step seven, analyze the model classification results by combining the Shap model classification weight visualization results.

[0095] The present application uses a SHapley Additive explanations (Shap) interpreter to visualize the classification weights of the InceptionTime model, as shown in Figure 6-8 The Shap analysis explicitly emphasizes the most emphasized features of the model. It determines that 1003 cm -1 , 1492 cm -1 , and 1150 cm -1 are the most critical. In addition, the focus of the model is mainly concentrated on the shoulders of these three mountain peaks. In particular, the wave number close to 1003 corresponds to 7 of the top 20 features. In addition, the fluorescence band in the range of 300-400 cm -1 is also highlighted.

[0096] A Shap value is assigned to each feature of each predicted sample by the Shap model, indicating the degree of contribution of the feature to the prediction, and the weights of different features are projected on the spectrum for visualization. The SHapley Additive explanations (SHAP) method is used to explain and verify the classification basis of the model, thereby enhancing the transparency and credibility of the model decision-making process.

[0097] In this study, the present invention demonstrated the feasibility of using Raman spectroscopy for non-destructive detection of formaldehyde residues in shrimp. By collecting Raman data from shrimp samples and utilizing spectrophotometric measurements of formaldehyde residues as classification labels, the present invention evaluated the classification performance of models across different concentration groups. InceptionTime model consistently outperformed other models, including Cmp_CNN, RF, XGBoost, and SVM, and maintained the highest accuracy across various detection limits in all concentration groups.

[0098] Further analysis revealed that the InceptionTime model exhibited more stable performance across different test sets compared to other models, highlighting its robustness in handling different data conditions. Additionally, the present invention's investigation of spectral preprocessing methods showed that SG preprocessing improved the model's performance. This finding underscores the benefit of specific spectral preprocessing methods for classification tasks. During the optimization of the InceptionTime model structure, the present invention observed that the model's performance peaked in two residual blocks, indicating that this configuration was optimal for the present invention's dataset. The addition of other modules, such as Attention, LSTM, GRU, and Transformer, did not significantly improve the model's performance, suggesting that the added complexity of these modules might not be necessary or beneficial for this particular application. Furthermore, the present invention found that the TimesNet model consistently underperformed compared to InceptionTime, highlighting the effectiveness of the InceptionTime architecture in capturing relevant features of the data. These findings suggest that for the classification of shrimp Raman spectra, simpler structures and fewer residual blocks can provide sufficient accuracy without the need for more complex modules or alternative models. Future work can explore fine-tuning existing structures or combining them with lightweight enhancements to further optimize performance without sacrificing computational efficiency.

[0099] The above merely explains the method of the present invention and its core idea, but the protection scope of the present invention is not limited thereto. Any equivalent replacement or change of the technical solutions and the inventive concept of the present invention within the technical scope disclosed by the present invention, according to the technical solutions and the inventive concept of the present invention, should be covered within the protection scope of the present invention. In summary, the content of the present specification should not be understood as a limitation of the present invention.

Claims

1. A non-destructive method for detecting formaldehyde residues in aquatic products, characterized in that, Comprising the following steps: S1, establishing a formaldehyde residue dataset; collecting Raman spectra of aquatic product samples soaked in formaldehyde solutions of different concentrations, determining the formaldehyde residue in the aquatic product samples, dividing the collected Raman spectrum data into different groups according to the formaldehyde residue determination results, and constructing a Raman spectrum formaldehyde residue dataset; S2, constructing a classification model; using random forest, support vector machine, XGBoost, a convolutional neural network for comparison and InceptionTime model to classify the Raman spectrum formaldehyde residue data, evaluating the model performance through classification accuracy, and selecting InceptionTime model as the classification model; The InceptionTime model is composed of residual blocks, and each residual block contains three Inception sub-modules; when processing an M-dimensional multivariate time series, the data is first passed through the bottleneck layer of the Inception sub-module, which uses m filters with a length of 1 and a step of 1 to convert the M-dimensional signal to m dimensions; After the residual block, the InceptionTime model uses a global average pooling layer to average the multivariate time series in the time dimension; then a standard fully connected Softmax layer is used for classification, the number of neurons in the layer matches the number of categories in the dataset, and the classification result is output; S3, establishing a selective data set; Collecting Raman spectra of aquatic product samples soaked in formaldehyde analogs, using the same label for aquatic product data soaked in formaldehyde analogs, and using the same label for aquatic product data soaked in water and formaldehyde, and constructing a selective data set using mixed data; S4, verifying the selectivity of the classification model to formaldehyde; using the InceptionTime model to select mixed data to judge the accuracy of formaldehyde identification; S5, optimizing the structure of the classification model; using different spectral pretreatment methods to process the Raman spectrum formaldehyde residue dataset, and selecting the optimal spectral pretreatment method; Based on the pretreated Raman spectrum formaldehyde residue dataset, the InceptionTime model structure is optimized; S6, verifying and analyzing the classification results of the classification model; The classification results of the model are analyzed in combination with the Shap model classification weight visualization results; A Shap value is assigned to each feature of each prediction sample using the Shap model, indicating the contribution of the feature to the prediction, and the weights of different features are projected on the spectrum for visualization; the classification basis of the model is explained and verified using the Shapley additive explanation method. The S3 is specifically as follows:

2. The non-destructive method for detecting the residual formaldehyde in aquatic products according to claim 1, characterized in that, Formaldehyde analogs are selected from methanol, ethanol and acetaldehyde, and Raman spectra of aquatic product samples soaked in methanol, ethanol and acetaldehyde are collected, and the aquatic product data soaked in methanol, ethanol and acetaldehyde are labeled with the same label. The S4 is specifically as follows: according to the detection limit category, a corresponding classifier is designed in the InceptionTime model.

3. The non-destructive method for detecting the residual formaldehyde in aquatic products according to claim 2, characterized in that, The different spectral pretreatment methods in S5 are specifically as follows:

4. The non-destructive method for detecting the residual formaldehyde in aquatic products according to claim 1, characterized in that, ​ Different spectral preprocessing methods include multiple scattering correction, standard normal variate transformation, wavelet transformation, SG and baseline correction; the Raman spectrum formaldehyde residual data set is preprocessed by multiple scattering correction, standard normal variate transformation, wavelet transformation, SG and baseline correction respectively, and model verification is carried out, so as to evaluate the influence of different spectral preprocessing methods on the classification results of the classification model.

5. The non-destructive method for detecting the residual formaldehyde in aquatic products according to claim 4, characterized in that, The InceptionTime model structure is improved and optimized in S5, and the specific process is as follows: The number of residual blocks is adjusted, the number of residual blocks of the InceptionTime model is set to 1, 2 and 3 respectively, and other model parameters remain unchanged, the classification results of the preprocessed Raman spectrum formaldehyde residual data set are determined to determine the best residual block configuration; after the InceptionTime model structure is optimized, it is composed of two residual blocks.

6. A non-destructive testing system for detecting formaldehyde residues in aquatic products, characterized in that, The InceptionTime model after the InceptionTime model structure is optimized in the nondestructive detection method of the water product formaldehyde residue according to any one of claims 1-5.

Citation Information

Patent Citations

  • Device and method for detecting formaldehyde

    CN102749318A

  • Correction method in Raman spectroscopy quantitative detection under temperature fluctuation condition

    CN103674927A