On-line monitoring method for shewanella spirosea biofilm based on hyperspectral technology
By constructing a qualitative classification model in a pure culture system using hyperspectral technology, the problems of non-destructiveness, dynamism, and visualization in existing biofilm monitoring technologies have been solved. This enables accurate classification and spatial distribution monitoring of biofilm growth stages, and is suitable for online monitoring in food processing and storage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIMEI UNIV
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies cannot achieve non-destructive, dynamic, and visual online monitoring of biofilms, especially in complex food matrices where it is difficult to distinguish different growth stages and reflect microstructures, thus failing to meet the needs for rapid and automated monitoring in food processing and storage.
Hyperspectral technology was used to collect microscopic data of biofilms in a pure culture system. Through feature screening and qualitative classification model construction, a visual distribution map was generated to achieve accurate classification and spatial distribution monitoring of biofilm growth stages.
It enables automated and precise classification of biofilm growth stages, provides rich monitoring information, meets the online and real-time monitoring needs of food processing and storage, and is suitable for complex food matrices.
Smart Images

Figure CN121994787A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of food testing technology, and in particular relates to an online monitoring method for Baltic Shewanella biofilm based on hyperspectral technology. Background Technology
[0002] In the field of food testing technology, particularly for monitoring pathogenic bacterial contamination in perishable foods such as aquatic products, biofilm monitoring is a crucial step in ensuring food safety and extending shelf life. Current technologies typically employ traditional microbiological and biochemical methods for the detection and quantification of biofilms, with crystal violet staining being the most commonly used. This method, through staining, elution, and photometric measurement of the biofilm, can reliably quantify the total biomass of the biofilm and is currently one of the standard methods for assessing the ability and dynamic changes of bacterial biofilm formation in both laboratory and industrial settings. Due to its standardized operation, relatively low cost, and compatibility with traditional microbial detection methods such as culture and counting, this method has certain applicability in scientific research and routine testing, providing fundamental data support for understanding the static formation patterns of biofilms.
[0003] However, the aforementioned existing technologies have significant shortcomings. First, chemical detection methods such as crystal violet assays are in vitro and destructive, requiring the termination of sample culture and sample consumption. This makes continuous, non-invasive online monitoring of the same object difficult, hindering the capture of the dynamic process and spatial heterogeneity of biofilm evolution on food surfaces over time. Second, these methods only provide a single indicator of the total amount of biofilm, failing to distinguish different growth stages (such as adhesion, maturation, and diffusion) or reflect the microstructure and distribution characteristics on the food matrix surface. This results in limited monitoring information, failing to meet the needs for precise intervention and process control of biofilms. Furthermore, existing methods are susceptible to interference in complex food matrices (such as fish and meat surfaces), potentially reducing detection sensitivity and specificity. The procedures are also cumbersome and time-consuming, making them unsuitable for the rapid and automated online monitoring requirements of food processing and storage. Therefore, there is an urgent need to develop an online monitoring technology that can achieve non-destructive, dynamic, and visual monitoring, accurately identifying the growth stages of biofilms. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention proposes an online monitoring method for Baltic Shewanella biofilm based on hyperspectral technology, thereby resolving the issues present in the prior art.
[0005] To achieve the above objectives, in a first aspect, the present invention provides an online monitoring method for Baltic Shewanella biofilm based on hyperspectral technology, comprising:
[0006] S1. Under a pure culture system, collect microscopic hyperspectral data of Baltic Shewanella biofilm at different growth stages and perform preprocessing.
[0007] S2. Perform feature screening on the preprocessed microscopic hyperspectral data. The feature screening includes preliminary screening using analysis of variance, and secondary screening using at least one of the following: minimum absolute shrinkage and selection operator, continuous projection algorithm and variable projection importance analysis method, to obtain key feature wavelengths.
[0008] S3. Based on the selected key feature wavelengths, construct a dataset and train at least one qualitative classification model to classify different growth stages of biofilms in a pure culture system.
[0009] S4. The qualitative classification model trained and validated in step S3 is transferred and applied to the biofilm hyperspectral data collected in the food system to classify the growth stages of the biofilm on the food surface.
[0010] S5. The classification results are fused with the spatial coordinate information of the hyperspectral image to generate a visualized distribution map reflecting the spatial distribution of the biofilm growth stages.
[0011] Preferably, in step S1, the preprocessing includes Savitzky-Golay smoothing filtering and first-order difference processing.
[0012] Preferably, in step S2, the minimum absolute shrinkage and selection operator employs L1 regularization, and its loss function is:
[0013]
[0014] in, For sample labels, For the first The first sample One characteristic, For the first The coefficients of each characteristic, For regularization parameters, For the sample size, It is the characteristic number.
[0015] Preferably, in step S2, the continuous projection algorithm iteratively selects variables by calculating the orthogonal projection residuals of the remaining variables on the selected feature subspace. The formula for calculating the orthogonal projection residuals is:
[0016]
[0017] in, For the selected eigenvectors, form an orthogonal basis spanning the space. The modulus represents the variable. The amount of information not covered by the selected features.
[0018] Preferably, in step S2, the formula for calculating the variable importance index VIP value of the variable projection importance analysis method is as follows:
[0019]
[0020] in, For the total number of variables, The cumulative explanatory power of all potential components for the dependent variable. For the number of potential components, For the first The explanatory power of each latent component for the dependent variable For the first The variable in the first... The weights of each potential component, and the VIP value, represent the overall contribution of the variable to the model's explanatory power.
[0021] Preferably, in step S3, the qualitative classification model includes support vector machine, random forest, nearest neighbor model, and linear discriminant analysis.
[0022] Preferably, in step S4, the food system is the surface of aquatic products, specifically the surface of large yellow croaker pieces.
[0023] Preferably, after acquiring the hyperspectral images in steps S1 and S4, black and white correction processing is performed.
[0024] Preferably, in step S5, the visualized distribution map uses different colors to represent different growth stages and content levels of the biofilm.
[0025] In a second aspect, the present invention also discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.
[0026] Compared with the prior art, the present invention has the following advantages and technical effects:
[0027] This invention first constructs and optimizes a qualitative classification model in a controlled pure culture system, and then directly applies the validated model to real food systems. This complete technical approach ensures that the established monitoring method is not limited to the laboratory environment, but can be adapted to and reliably applied to complex and ever-changing real food production and storage scenarios, solving the problems of poor applicability and unstable results of existing technical models in real matrices.
[0028] This invention employs a composite feature wavelength screening strategy that includes preliminary screening and secondary fine screening, and trains a qualitative classification model based on the screened key features. This technical solution can effectively extract discriminative information strongly correlated with the growth stage of biofilm from hyperspectral data, thereby achieving automated and accurate classification of multiple growth stages such as "adhesion" and "maturation," overcoming the shortcomings of existing technologies that can only provide the total amount of biofilm but cannot distinguish its dynamic development stage.
[0029] This invention fuses the prediction results of a classification model for each pixel with the spatial coordinate information of a hyperspectral image to generate a visual distribution map that intuitively reflects the spatial distribution of biofilm growth stages. This technical solution elevates monitoring results from single numerical values or overall judgments to spatial images that reveal the specific location, distribution range, and developmental heterogeneity of biofilms on the sample surface, providing richer and more intuitive monitoring information than traditional methods.
[0030] This invention is based on hyperspectral image acquisition technology. The entire monitoring process does not require damage or consumption of samples, allowing for continuous and repeated measurements of the same monitoring object. This technical solution enables dynamic, in-situ tracking of the entire process of biofilm formation and development, meeting the urgent need for online, real-time monitoring of key control points in food processing and distribution. Attached Figure Description
[0031] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0032] Figure 1 This is a flowchart of the online monitoring method for Baltic Shewanella biofilm based on hyperspectral technology according to an embodiment of the present invention;
[0033] Figure 2 The diagram shows the microscopic hyperspectral instrument (A) and the hyperspectral imaging instrument (B) according to an embodiment of the present invention.
[0034] Figure 3 This is a schematic diagram of the crystal violet quantitative experiment for the pure system diagram (C) and the food system diagram (D) of an embodiment of the present invention;
[0035] Figure 4 A visualization of the feature variable selection after filtering variables in this embodiment of the invention;
[0036] Figure 5 This is a confusion matrix diagram of the qualitative model in an embodiment of the present invention;
[0037] Figure 6 This is a visualization of the four stages of biofilm based on VNIR spectroscopy in an embodiment of the present invention. Detailed Implementation
[0038] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0039] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0040] Example 1
[0041] like Figure 1 As shown, this embodiment provides an online monitoring method for Baltic Shewanella biofilm based on hyperspectral technology, including:
[0042] S1. Under a pure culture system, collect microscopic hyperspectral data of Baltic Shewanella biofilm at different growth stages and perform preprocessing.
[0043] Furthermore, in step S1, the preprocessing includes Savitzky-Golay smoothing filtering and first-order difference processing.
[0044] Specifically, step S1 includes:
[0045] S101: Activated Baltic Shewanella strain, cultured at 28℃ and 200 rpm / min until the optical density (OD) is around 0.45, at which point the bacterial count is approximately 7 LogCFU / mL;
[0046] S102: Place a sterile round glass slide in a 24-well plate, then add 1 mL of fresh 2216E medium, followed by 10 μL of activated bacterial solution, mix thoroughly, and seal in a 28°C incubator.
[0047] S103: After the specified time, aspirate the supernatant from the well plate, then wash it three times with sterile PBS solution and dry it.
[0048] S104: Add 1 mL of 0.1% crystal violet to each well, let stand at room temperature for 15 min, then rinse three times with sterile PBS to remove any remaining crystal violet, and dry.
[0049] S105: Add 1 mL of 95% ethanol to each well to elute the crystal violet on the biofilm, and let stand for 15 minutes.
[0050] S106: OD595 was measured using an ELISA reader, with 95% ethanol as a blank control to remove background interference. The absorbance of each bacterium was equal to the average of its three replicates minus the average of the three blank controls (containing no bacteria). Schematic diagrams of the crystal violet quantitative experiment are shown in the pure system diagram (C) and the food system diagram (D), as follows. Figure 3 As shown.
[0051] The experiment of detecting biofilms using microscopic hyperspectral imaging in a pure system includes the following steps:
[0052] S107: Follow the steps in S101, S102, and S103, then prepare the round-bottomed glass slide and perform black and white correction before acquiring the image using microscopic hyperspectral imaging;
[0053] In this embodiment, four growth stages of Baltic Shewanella biofilm samples were used: 55 samples in the first stage, 62 samples in the second stage, 32 samples in the third stage, and 28 samples in the fourth stage, for a total of 177 samples. Each category was assigned a different identification number.
[0054] Among them, Baltic Shewanella (MCCC 1A12412) is preserved by the Marine Microbial Culture Collection Center.
[0055] When using microscopic hyperspectral acquisition operations, such as Figure 2 As shown in (A), the microscopic hyperspectral imaging system used in this embodiment consists of a hyperspectral camera (imec Snapscan VNIR), a color camera, a computer, and an inverted fluorescence microscope (BM-38X). The hyperspectral imaging part relies on the hyperspectral camera, combined with a spectroscope and a color camera. With the help of the lens tube and objective lens of the inverted fluorescence microscope, and combined with the transmitted light illumination method, hyperspectral images of the sample can be acquired. The computer is used for system control and data processing. The parameters of the spectral camera are adjusted to the optimal value, with an integration time of 8 ms and a white plate reference reflectance of 40%. The imaging light path of the acquired biofilm passes through three parts: the slide, the sample, and the coverslip. Refraction and reflection of the light path will occur at the slide and coverslip, which will cause the sensitive hyperspectral imaging system to produce brightness imbalance in the acquired hyperspectral pathological images. Therefore, a blank sample needs to be collected around the tissue sample when acquiring hyperspectral images. The formulas (1)-(3) used for black and white correction processing are:
[0056] (1)
[0057] (2)
[0058] (3)
[0059] in, For wavelengths in specific bands during hyperspectral image acquisition, The intensity of the incident light. The transmitted light intensity is when the sample is present. The transmitted light intensity when it is a blank sample. , , The transmittance values are for the glass slide, coverslip, and sample, respectively.
[0060] S108: The extent of the region of interest (ROI) is determined based on the adhesion characteristics of the biofilm on the circular glass slide. The ROI of the corrected hyperspectral image is extracted using ENVI 5.3, and the average spectral reflectance of all pixels within the ROI is taken as the average spectrum of the sample.
[0061] S109: Perform data preprocessing on the raw spectral data (RAW) obtained in S108, and then further process it to divide it into training set and prediction set;
[0062] This implementation employs Savitzky-Golay smoothing filtering (SG) and first-difference preprocessing on the average spectrum to obtain hyperspectral data with smoothed spectral information after noise removal. ANOVA-based preliminary dimensionality reduction is then performed: by calculating the variance difference (p-value) of each band among different class samples, and using a significance level α=0.05 as a threshold, bands that significantly contribute to class differentiation are selected. The changes in the number of variables after different spectral preprocessing are shown in Table 1.
[0063] To ensure the reliability and generalization ability of the model training, the preprocessed data was split, balanced, and standardized. The training and test sets were divided in an 8:2 ratio to avoid data leakage. Because the number of samples in the four classes is not completely consistent, to prevent the model from biasing towards the majority class due to differences in class sample size, classes with fewer samples than the maximum number of classes in the training set were randomly and repeatedly sampled until the number of samples in all classes was equal to the maximum number of classes. The data was standardized based on the mean and standard deviation of the training set to eliminate the differences in units across different bands.
[0064] Table 1
[0065]
[0066] S2. Feature screening is performed on the preprocessed microscopic hyperspectral data. This feature screening includes initial screening using analysis of variance, and secondary screening using at least one of the following: minimum absolute shrinkage and selection operator, continuous projection algorithm, and variable projection importance analysis, to obtain key feature wavelengths. A visualization of the selected feature variables is shown below. Figure 4As shown.
[0067] Specifically, ANOVA was used for initial feature screening, followed by secondary feature screening using three methods: Least Absolute Shrinkage and Selection Operator (LASSO), Continuous Projection Algorithm (SPA), and Variable Projection Importance Analysis (PLS-VIP), resulting in four feature subsets.
[0068] ANOVA, by comparing the mean differences of samples from different categories on the same feature, filters variables that significantly contribute to class distinction, helping us understand whether variables between different groups have a significant impact on the outcome variable. In this study, the ANOVA significance level α=0.05 was used as the initial screening result for bands; one-way ANOVA was used to calculate the results for the spectral data after first differencing. Finally, by comparing the p-values of all bands, redundant variables with insignificant differences between groups were eliminated, laying the foundation for subsequent secondary screening.
[0069] LASSO applies a penalty constraint to the model coefficients through L1 regularization, compressing the coefficients of redundant variables to 0, thereby achieving feature selection and dimensionality reduction. In this study, the key parameters of LASSO are set as follows: 10-fold cross-validation is used to determine the optimal regularization parameter, and the parameter corresponding to 1 standard error (1SE) is selected as the final solution, balancing generalization ability while ensuring model simplicity; among them, the loss function calculation formula (4) for L1 regularization is as follows:
[0070] (4)
[0071] in, For sample labels, For the first The first sample One characteristic, For the first The coefficients of each characteristic, This is a regularization parameter (controlling the intensity of the penalty). For the sample size, The number of features is denoted by . This formula achieves feature selection by introducing a term that shrinks the coefficients of unimportant features to 0.
[0072] SPA (Special Feature Projection) selects the feature subset with the lowest redundancy and richest information by calculating the projection vectors between feature variables. In this study, the maximum number of candidate features for SPA is set to 15. The cumulative contribution is calculated using the variable projection importance (VIP) index, and the variable with the largest projection vector is iteratively selected until the preset maximum number of features is reached or the projection gain stabilizes. SPA assumes the selected feature vectors are... , ,… The orthogonal projection residuals of the remaining variables on this subspace are:
[0073] (5)
[0074] in, For the selected eigenvectors, form an orthogonal basis spanning the space. The modulus represents the variable. The information content not covered by the selected features is considered more significant when the modulus is larger, indicating a more substantial independent information contribution from that variable, and thus it is preferentially included in the feature subset. This process effectively eliminates redundant variables by maximizing the information complementarity between features.
[0075] PLS-VIP calculates the overall contribution of a variable to the model's explanatory power (VIP value) to select features that significantly influence the dependent variable (class label), taking into account the correlation between the variable and both the independent and dependent variables. In this study, PLS-VIP determines the optimal number of latent components using 5-fold cross-validation. Using the features selected by ANOVA as input, it calculates the VIP value for each band, retaining bands with VIP ≥ 1 as the final feature subset, achieving secondary dimensionality reduction of the features. Its core principle is that in the PLS model, the VIP value of each variable integrates the variable's loading on all latent components and the explanatory power of the latent components on the dependent variable, as shown in the following formula:
[0076] (6)
[0077] in, For the total number of variables, The cumulative explanatory power of all potential components for the dependent variable. For the number of potential components, For the first The explanatory power of each latent component for the dependent variable For the first The variable in the first... The weights of each potential component. The larger the VIP value, the more important the variable is to the model.
[0078] S3. Based on the selected key feature wavelengths, construct a dataset and train at least one qualitative classification model to classify different growth stages of biofilms in a pure culture system.
[0079] Further, in step S3, the qualitative classification model includes support vector machine, random forest, nearest neighbor model, and linear discriminant analysis. Its qualitative model confusion matrix is shown below. Figure 5 As shown.
[0080] Specifically, each feature subset was input into four qualitative classification models: Support Vector Machine (SVM), Random Forest (RF), Nearest Neighbor (KNN), and Linear Discriminant Analysis (LDA). By comparing the models' accuracy, precision, recall, and F1 score, the combination scheme that simplifies the model structure, eliminates redundant variables, and has the best robustness and generalization ability was selected. The model performance comparison results are shown in Table 2.
[0081] SVM classifies samples by finding the optimal hyperplane, and for nonlinear data, it can map to a high-dimensional space using a kernel function. In this study, the parameters of the SVM are determined through Bayesian optimization: with the accuracy of 5-fold cross-validation as the optimization objective, the parameter combination that optimizes the performance on the validation set is finally selected. Classification performance is evaluated using metrics such as accuracy and F1 score. By using a kernel function, finding the optimal classification decision boundary becomes:
[0082] (7)
[0083] (8)
[0084] Where C is a regularization parameter, used to control the trade-off between the classification margin and the penalty for misclassification.
[0085] Randomization (RF) is an ensemble learning model based on multiple decision trees. It constructs multiple trees through sampling and outputs the final classification result using a voting mechanism. This effectively reduces the risk of overfitting and has strong adaptability to high-dimensional data. Assume there is a training set D, from which a subset D is generated by sampling with replacement. b Used for training each tree; when splitting at each node, a subset of size is randomly selected from the feature set, and the node with the best feature is chosen for splitting. Ultimately, each decision tree... The output is the prediction result; the final result of the random forest is the prediction result of all decision trees. The decision trees were generated through voting. In this study, the number of decision trees in the RF algorithm was set to 200, the number of features randomly selected when splitting nodes was the square root of the total number of features, and the minimum number of leaf nodes was set to 5. The generalization ability was finally verified using classification metrics on the test set.
[0086] (9)
[0087] in, It is the first The predicted results for each tree. This represents the total number of decision trees. For classification problems, the output is the final prediction based on the majority class selected through a voting process.
[0088] KNN is an instance-based lazy learning algorithm that calculates the distance between a sample to be classified and the k nearest samples in the training set, and uses the majority class as the prediction result. It is simple in principle and does not require pre-training of the model. The similarity of samples is measured using formula (10). For each sample to be classified, the k nearest training samples are selected, and its class is determined by voting. The value of k determines the smoothness of the model.
[0089] (10)
[0090] in It is the characteristic number.
[0091] The basic idea of LDA is to reduce the dimensionality of data by maximizing the ratio of the scatter matrix between classes to the scatter matrix within classes, making it suitable for multi-class linearly separable problems. In this study, the LDA projection dimension is set to 3, combined with PCA preprocessing (to reduce the collinearity of the original features), using the selected feature subset as input, and finally outputting the classification results of the test set and calculating the evaluation index. Among them: the scatter matrix within classes... The scatter matrix reflects the degree of dispersion of data within the same category; the scatter matrix between categories. This reflects the degree of dispersion of data across different categories. The core principle is to find a projection matrix such that the projected data satisfies... .
[0092] (11)
[0093] (12)
[0094] (13)
[0095] in The mean of class c, This is the overall average.
[0096] Table 2
[0097]
[0098] S4. The qualitative classification model trained and validated in step S3 is transferred and applied to the biofilm hyperspectral data collected in the food system to classify the growth stages of the biofilm on the food surface.
[0099] Furthermore, in step S4, the food system is the surface of aquatic products, specifically the surface of large yellow croaker fish pieces.
[0100] Specifically, step S4 includes:
[0101] S401: Clean the surface of high-quality chilled yellow croaker, take the muscle from the back of the fish, about 5g per piece, treat it with 75% alcohol and ultraviolet light in a clean bench, and after sterilization, place each piece of fish meat in a sealed box and store it at 4℃ for later use.
[0102] S402: Follow the steps in S101, then dilute with sterile saline to a concentration of 10. 3 CFU / mL, inoculate each fish fillet sample with 1mL of bacterial solution, shake well in a sealed box, cover the surface of the fish meat with a round glass slide, and incubate at 4℃; then remove the glass slide after the specified time and place it in a 24-well plate, and complete the steps S103, S104, S105, and S106.
[0103] S403: Perform inoculation according to steps S401 and S402, collect hyperspectral images of the biofilm on the surface of large yellow croaker at different time periods, and perform black and white correction.
[0104] In this embodiment, four growth stages of Baltic Shewanella biofilm samples were used: 42 samples in the first stage, 29 samples in the second stage, 32 samples in the third stage, and 28 samples in the fourth stage, for a total of 131 samples. Each category was assigned a different identification number.
[0105] When performing hyperspectral image acquisition operations, such as Figure 2 As shown in (B) of this embodiment, the hyperspectral imaging system used includes a hyperspectral imager, a platform control system, and a computer. The SWIR-HIS (900-1700 nm) system consists of an imaging spectrometer (N17E, Specim, Finland) with a spectral resolution of 4 nm, a camera with a resolution of 640×512, a camera lens (640 mini Raptor, Northern Ireland), two tungsten lamps (LS-150, Wuling Optics), a mobile platform driven by a stepper motor (HSIM-800, Wuling Optics), and a computer with software. Hyperspectral acquisition software.
[0106] The entire acquisition process was conducted in a dark box to prevent ambient light from affecting the acquired hyperspectral images. The parameters before acquiring the hyperspectral images were: exposure time 43ms, platform movement time 4.5cm / s, and the angle between the two 150W tungsten lamps and the platform was 50 degrees.
[0107] In a preferred embodiment, the black-and-white correction formula (14) used in the black-and-white correction processing is:
[0108] (14)
[0109] in, The image shows the corrected reflectance hyperspectral image, expressed in relative reflectance (%). Represents the original hyperspectral image; For dark images (0% reflectance). White reference image (100% reflectivity).
[0110] S404: The region of interest (ROI) is determined based on the surface of the fish piece. ENVI 5.3 is used to extract the ROI from the corrected hyperspectral image, and the average spectral reflectance of all pixels within the ROI is taken as the average spectrum of the sample.
[0111] S405: Follow steps S109, S2, and S3.
[0112] S5. The classification results are fused with the spatial coordinate information of the hyperspectral image to generate a visualized distribution that reflects the spatial distribution of the growth stages of the biofilm.
[0113] Furthermore, in step S5, the visualized distribution map uses different colors to represent different growth stages and content levels of the biofilm.
[0114] Specifically, based on the visualization of the four stages of biofilm based on VNIR spectroscopy, the spectral values corresponding to pixels in the hyperspectral image of the prepared large yellow croaker fish fillet sample are extracted. The selected spectral values are substituted into the optimal model to predict the biofilm information at each pixel in the hyperspectral image of the fish fillet. Finally, through information fusion using hyperspectral imaging technology, the distribution map of the biofilm on the surface of the fish fillet is reconstructed on a plane based on the pixel coordinate information and its corresponding biofilm stage. Figure 6 As shown in the distribution map, the colors range from blue to red. A redder color indicates a higher biofilm content, while a bluer color indicates a lower biofilm content.
[0115] Example 2
[0116] This embodiment also discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in Embodiment 1.
[0117] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for online monitoring of Baltic Shewanella biofilm based on hyperspectral technology, characterized in that, Includes the following steps: S1. Under a pure culture system, collect microscopic hyperspectral data of Baltic Shewanella biofilm at different growth stages and perform preprocessing. S2. Perform feature screening on the preprocessed microscopic hyperspectral data. The feature screening includes preliminary screening using analysis of variance, and secondary screening using at least one of the following: minimum absolute shrinkage and selection operator, continuous projection algorithm and variable projection importance analysis method, to obtain key feature wavelengths. S3. Based on the selected key feature wavelengths, construct a dataset and train at least one qualitative classification model to classify different growth stages of biofilms in a pure culture system. S4. The qualitative classification model trained and validated in step S3 is transferred and applied to the biofilm hyperspectral data collected in the food system to classify the growth stages of the biofilm on the food surface. S5. The classification results are fused with the spatial coordinate information of the hyperspectral image to generate a visualized distribution map reflecting the spatial distribution of the biofilm growth stages.
2. The method according to claim 1, characterized in that, In step S1, the preprocessing includes Savitzky-Golay smoothing filtering and first-order difference processing.
3. The method according to claim 1, characterized in that, In step S2, the minimum absolute shrinkage and selection operator employs L1 regularization, and its loss function is: in, For sample labels, For the first The first sample One characteristic, For the first The coefficients of each characteristic, For regularization parameters, For the sample size, It is the characteristic number.
4. The method according to claim 1, characterized in that, In step S2, the continuous projection algorithm iteratively selects variables by calculating the orthogonal projection residuals of the remaining variables on the selected feature subspace. The formula for calculating the orthogonal projection residuals is as follows: in, For the selected eigenvectors, form an orthogonal basis spanning the space. The modulus represents the variable. The amount of information not covered by the selected features.
5. The method according to claim 1, characterized in that, In step S2, the formula for calculating the variable importance index VIP value of the variable projection importance analysis method is as follows: in, For the total number of variables, The cumulative explanatory power of all potential components for the dependent variable. For the number of potential components, For the first The explanatory power of each latent component for the dependent variable For the first The variable in the first... The weights of each potential component, and the VIP value, represent the overall contribution of the variable to the model's explanatory power.
6. The method according to claim 1, characterized in that, In step S3, the qualitative classification model includes support vector machine, random forest, nearest neighbor model, and linear discriminant analysis.
7. The method according to claim 1, characterized in that, In step S4, the food system refers to the surface of aquatic products, specifically the surface of large yellow croaker pieces.
8. The method according to claim 1, characterized in that, After acquiring hyperspectral images in steps S1 and S4, black and white correction processing is performed.
9. The method according to claim 1, characterized in that, In step S5, the visualized distribution map uses different colors to represent different growth stages and content levels of the biofilm.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1-9.