Monitoring system, method, corresponding device, storage medium, and product
By employing a feature-level spectral reconstruction module and unsupervised continuous learning, the high cost and insufficient accuracy of existing technologies for monitoring substances of interest in food have been addressed, enabling low-cost, high-efficiency histamine monitoring with environmental adaptability.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE HONG KONG UNIV OF SCI & TECH
- Filing Date
- 2025-11-06
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies are costly and inaccurate when monitoring the content of substances of interest in food, especially histamine in sashimi, which is difficult to detect automatically, passively, and cost-effectively. Furthermore, existing spectral reconstruction methods cannot effectively filter out redundant and irrelevant information, resulting in reconstruction errors and inaccurate predictions.
By employing a feature-level spectral reconstruction module, combined with a spectral base model and a convolutional neural network, and through the feature-level spectral reconstruction module and unsupervised continuous learning, the ability to extract spectral features is enhanced, enabling precise monitoring of the content of substances of interest in food ingredients.
It improves the correlation between spectral reconstruction data and substances of interest, reduces hardware costs, enables efficient and accurate monitoring of substances such as histamine, and has robustness to different environmental changes.
Smart Images

Figure CN2025133128_15052026_PF_FP_ABST
Abstract
Description
Systems, methods, and related equipment, storage media, and products for monitoring.
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 717,285, filed November 7, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of food monitoring, and more specifically, to systems and methods for monitoring the content of substances of interest in food, as well as corresponding computer equipment, computer storage media and computer products. Background Technology
[0004] Over time, ingredients change, and indicators of these changes (such as freshness) can be indicated by the content of specific substances within the ingredients. Therefore, the changes in these indicators over time can be assessed by monitoring the content of specific substances in the ingredients. However, existing technologies have some drawbacks in this regard, such as high monitoring costs or insufficient accuracy in assessment.
[0005] Taking sashimi (raw fish slices) as an example, histamine is the most critical indicator of freshness. If the sample is not handled or stored properly (usually requiring low temperature), its histamine content can rapidly increase to toxic levels within a few hours (see references
[0010] and
[0014] ). However, in environments such as sushi restaurants or fresh food stores, the temperature of the display case is always fluctuating, and the freshness and safety of sashimi samples cannot be guaranteed (see references
[0032] ,
[0035] , and
[0048] ). Therefore, due to the regulatory requirements for histamine in response to health risks and the perceived profit needs of businesses, it is necessary to adopt a passive, automatic, and cost-effective histamine monitoring method in display cases.
[0006] Existing histamine detection methods are typically costly and complex, posing a challenge to small sushi shops and fresh food stores. Most operations require expensive equipment, such as hyperspectral imaging (HSI) cameras (see references [6] and
[0013] ), or time-consuming laboratory experiments, such as high-performance liquid chromatography (see references
[0020] and
[0042] ), to detect histamine. Low-cost solutions, such as enzyme-linked immunosorbent assay kits (see references
[0022] and
[0034] ), are also labor-intensive and disruptive, failing to enable automated and passive monitoring. Recent work (see references
[0019] and
[0037] ) has attempted to achieve cost-effective automated monitoring of daily food using spectral reconstruction (SR) technology, making it possible for low-cost spectroscopic cameras to achieve sensing performance similar to laboratory-grade HSI cameras.
[0007] However, since the histamine level in sashimi samples is quite low, for example, the Codex Alimentarius Commission's safety standard is 20 mg / 100g (see reference [8]), its spectral characteristics are easily masked by other components such as proteins, which poses a challenge to accurately predicting histamine levels using coarse-grained or highly redundant spectral data. Unfortunately, previous solutions (see references
[0019] ,
[0037] ) tend to reconstruct complete hyperspectral data without filtering out highly redundant or irrelevant information, and even inevitably introduce reconstruction errors in the reconstruction stage, thus making histamine prediction impossible. Summary of the Invention
[0008] The purpose of this application is to provide a solution for monitoring the content of substances of interest in food ingredients, so as to solve or alleviate at least some of the problems of the prior art.
[0009] In a first aspect, this application provides a system for monitoring the content of a substance of interest in food, wherein the content of the substance of interest in the food indicates an indicator of the change of the food over time. The system includes: a feature-level spectral reconstruction module configured to: reconstruct a spectral feature representation related to the content of the substance of interest in the food based on multispectral image data of the food, wherein the feature-level spectral reconstruction module includes a spectral base model, and the encoder of the spectral base model is trained to enhance the encoder's ability to extract spectral features related to the substance of interest; and an estimation module configured to: estimate the content of the substance of interest in the food based on the reconstructed spectral feature representation.
[0010] In a second aspect, this application provides a method for monitoring the content of a substance of interest in food, wherein the content of the substance of interest in the food indicates an indicator of the change of the food over time. The method includes: acquiring multispectral image data of the food; reconstructing a spectral feature representation related to the content of the substance of interest in the food using a feature-level spectral reconstruction module based on the acquired multispectral image data, wherein the feature-level spectral reconstruction module includes a spectral base model, and the encoder of the spectral base model is trained to enhance the encoder's ability to extract spectral features related to the substance of interest; and estimating the content of the substance of interest in the food based on the reconstructed spectral feature representation.
[0011] In a third aspect, this application provides a computer device including a memory and a processor, wherein the memory stores computer instructions that, when executed by the processor, cause the method described in the second aspect to be performed.
[0012] In a fourth aspect, this application provides a non-transitory storage medium having computer instructions stored thereon, which, when executed by a processor, cause the method described in the second aspect to be performed.
[0013] In a fifth aspect, this application provides a computer program product including computer instructions that, when executed by a processor, cause the method described in the second aspect to be performed.
[0014] The scheme of this application can enhance the ability to extract spectral features associated with the monitored substance of interest, thereby enhancing the correlation between the reconstructed data and the substance of interest. This will improve the effectiveness in predicting the content of the substance of interest and realize related benefits such as reducing the requirements for hardware such as cameras, thereby reducing costs. Attached Figure Description
[0015] Non-limiting and non-exhaustive embodiments of this application are described by way of example with reference to the following figures, wherein:
[0016] Figure 1 shows a schematic diagram of an example working scenario to which the solution of this application can be applied;
[0017] Figure 2 shows a schematic diagram of the structure of the FreshSpec proposed in this application;
[0018] Figure 3 is a schematic diagram illustrating the extraction of the region of interest (ROI);
[0019] Figure 4A illustrates the correlation matrix of each band in the original HSI image, Figure 4B illustrates the correlation matrix of each band in the corresponding reconstructed HSI image from the state-of-the-art (SOTA) reconstruction model, and Figure 4C illustrates the correlation matrix of each band in the corresponding reconstructed features from FreshSpec.
[0020] Figure 5 schematically illustrates the architecture of the feature-level spectral reconstruction algorithm according to this application;
[0021] Figure 6 schematically illustrates the training scheme of the feature-level spectral reconstruction algorithm according to this application;
[0022] Figure 7 schematically illustrates a convolutional neural network (CNN) model for histamine regression according to this application;
[0023] Figure 8 schematically illustrates the generalization gap in histamine prediction between seen and unseen samples;
[0024] Figure 9 schematically illustrates the monotonic accumulation characteristics of histamine and the corresponding unlabeled MSI image;
[0025] Figures 10A and 10B schematically illustrate the actual working scenario and prototype of FreshSpec;
[0026] Figure 11 schematically illustrates the sashimi sample and histamine ground truth collection;
[0027] Figures 12A and 12B schematically illustrate the histamine prediction results of FreshSpec and the baseline (SOTA), respectively.
[0028] Figure 13 schematically illustrates the comparison of root mean square error (RMSE) between FreshSpec and the baseline;
[0029] Figures 14A and 14B schematically illustrate the T-SNE embedding distribution of FreshSpec and the T-SNE embedding distribution of the baseline (SOTA), respectively.
[0030] Figure 15 schematically illustrates FreshSpec's generalization performance across a variety of ambient temperatures;
[0031] Figure 16 schematically illustrates FreshSpec's generalization performance for a variety of ambient lighting conditions;
[0032] Figure 17 schematically illustrates the generalization performance of FreshSpec for various sample locations;
[0033] Figure 18 schematically illustrates FreshSpec's generalization performance for various sample-device distances;
[0034] Figure 19 schematically illustrates FreshSpec's generalization performance for various sashimi sample sizes;
[0035] Figure 20 schematically illustrates FreshSpec's generalization performance for various sashimi sample thicknesses. Detailed Implementation
[0036] To make the above and other features and advantages of this application clearer, the application is further described below in conjunction with the accompanying drawings. The drawings form part of this application and, together with the embodiments of this application, serve to illustrate the application. For clarity and simplicity, detailed descriptions of the known functions and structures of the devices, apparatuses, and / or devices described herein will be omitted where they might obscure the subject matter of this application. It should be understood that the specific embodiments given herein are for the purpose of explanation to those skilled in the art and are exemplary only, not restrictive.
[0037] The features described herein may be embodied in different forms and should not be construed as being limited to the embodiments described herein. Rather, the embodiments described herein are provided merely to illustrate some of the many possible ways of implementing the methods, apparatus, and / or systems described herein, which will become apparent upon understanding the disclosure of this application.
[0038] As used herein, the term “and / or” includes any one of the associated listed items and any combination of any two or more of the associated listed items.
[0039] Although terms such as “first,” “second,” and “third” may be used herein to describe various components, parts, sections, or elements, these components, parts, sections, or elements are not limited by these terms. Rather, these terms are used only to distinguish one component, part, section, or element from another. Therefore, without departing from the teachings of this application, a first component, part, section, or element referred to herein may also be referred to as a second component, part, section, or element.
[0040] The terminology used herein is for describing various embodiments only and is not intended to limit the scope of this disclosure. Unless the context clearly indicates otherwise, "a," "an," and "the" are intended to also include plural forms. The terms "comprising," "including," and "having" specify the presence of the stated features, operations, components, elements, and / or combinations thereof, but do not exclude the presence or addition of one or more other features, operations, components, elements, and / or combinations thereof.
[0041] In the following description, numerous specific details are set forth to provide a thorough understanding of this application. However, it will be apparent to those skilled in the art that these specific details are not required to practice this application. In other instances, well-known steps or operations have not been described in detail to avoid obscuring this application.
[0042] To overcome the shortcomings of existing technologies, the inventors have focused on filtering out redundant and irrelevant information during the SR process and enhancing the ability to extract spectral features related to the monitored substance of interest, thereby increasing the correlation between the reconstructed data and the substance of interest and improving the effectiveness of SR in predicting the content of the substance of interest.
[0043] When the monitored substance of interest is histamine, given the weak spectral characteristics of histamine, fine-grained feature extraction capabilities are needed to significantly enhance the correlation between reconstructed data and histamine values, thereby improving the effectiveness of SR in histamine prediction. However, traditional feature selection methods such as principal component analysis (see reference
[0040] ) cannot adequately address this issue. Inspired by the success of foundational models in various fields (see references [9],
[0017] ,
[0018] , and
[0033] )—which demonstrate their ability to extract deep features—the inventors discovered that a spectral foundation model (SFM) can be introduced into the SR process to obtain more informative features.
[0044] However, several challenges remain in turning this idea into a practical and feasible one:
[0045] Challenge 1: Extracting information-rich histamine-related spectral features. Although previous SFMs (see references [5],
[0018] , and
[0027] ) have achieved excellent performance in remote sensing applications, their training data are remote sensing images, such as trees and fields. The spectral features of remote sensing images are significantly different from those of sashimi and are unrelated to histamine. When the encoder of the SFM is directly introduced into spectral reconstruction, it is difficult to extract information-rich histamine-related spectral features, resulting in poor reconstruction performance and unsatisfactory histamine prediction results.
[0046] Challenge 2: Generalization gap for new samples. Due to differences in bacterial colonies and enzymes within individual sashimi samples, there can be significant differences in spectral characteristics between different sashimi samples. Baseline regression models struggle to obtain accurate predictions for new samples collected over time.
[0047] In response to this, this application proposes a solution.
[0048] Specifically, this application provides a system for monitoring the content of a substance of interest (SIO) in food ingredients, wherein the content of SIO in the food ingredients indicates an indicator of the change of the food ingredients over time. The system includes a feature-wise spectral reconstruction (FSR) module and an estimation module. The FSR module, also referred to as an FSR model, is configured to: reconstruct a spectral feature representation related to the content of SIO in the food ingredients based on multispectral image data of the food ingredients. The FSR module includes a spectral base model, the encoder of which is trained to enhance the encoder's ability to extract spectral features related to SIO. The estimation module is configured to: estimate the content of SIO in the food ingredients based on the reconstructed spectral feature representation.
[0049] In one embodiment, the system provided in this application further includes a multispectral imaging (MSI) camera. In this case, the multispectral image data can be acquired through the MSI camera. For example, the MSI camera can capture raw multispectral images of the monitored food ingredients, and then the captured raw multispectral images can be preprocessed, including, but not limited to, ROI extraction, cropping, and resizing, to obtain the multispectral image data.
[0050] In one embodiment, the encoder of the spectral base model is trained by iteratively performing a first training process using a first training set, the first training set including multiple hyperspectral images of multiple food samples of the food ingredient, the multiple food samples having different contents of the substance of interest. The first training process includes: acquiring a set of hyperspectral images selected from the multiple hyperspectral images, the set of hyperspectral images including a target sample, a negative sample corresponding to the target sample, and a positive sample corresponding to the target sample, wherein the content difference of the substance of interest between the positive sample and the target sample is less than a first content threshold, the content difference of the substance of interest between the negative sample and the target sample is greater than a second content threshold, and the second content threshold is greater than the first content threshold; and training the encoder of the spectral base model using the set of hyperspectral images to reduce the feature distance between the target sample and the positive sample and increase the feature distance between the target sample and the negative sample.
[0051] In one embodiment, training the encoder of the spectral base model using the set of hyperspectral images to reduce the feature distance between the target sample and the positive sample and increase the feature distance between the target sample and the negative sample includes: adjusting the encoder of the spectral base model using the set of hyperspectral images to minimize the following loss.
[0052] Where I represents the target sample, I - I represents the negative sample corresponding to the target sample. + Let Θ(I) represent the positive sample corresponding to the target sample, and let Θ(I) represent the spectral features of the target sample extracted by the encoder. - ) represents the spectral features of the negative sample corresponding to the target sample extracted by the encoder, Θ(I + ) represents the spectral features of the positive sample corresponding to the target sample extracted by the encoder, ξ(Θ(I + ), Θ(I)) means Θ(I + The mean square distance between ξ(Θ(I) and Θ(I) - ), Θ(I)) means Θ(I - The mean square distance between ) and Θ(I).
[0053] In one embodiment, the FSR module further includes a spectral reconstruction model whose output is connected to the input of the spectral base model. The spectral reconstruction model is trained by iteratively performing a second training process using a second training set, which includes multiple pairs of training images. Each pair of training images includes a hyperspectral image and a multispectral image of the same food sample corresponding to the same time point. The second training process includes: inputting the hyperspectral image from a pair of training images into the trained encoder and obtaining a first spectral feature representation at the output of the trained encoder; inputting the multispectral image from the pair of training images into the spectral reconstruction model and obtaining a second spectral feature representation at the output of the trained encoder; and adjusting the spectral reconstruction model to minimize the difference between the first spectral feature representation and the second spectral feature representation.
[0054] In one embodiment, adjusting the spectral reconstruction model to minimize the difference between the first spectral feature representation and the second spectral feature representation includes: adjusting the spectral reconstruction model such that the following loss Minimize:
[0055] Among them, I hsi and I msi Let Θ(I) represent the hyperspectral image and multispectral image in the pair of training images, respectively. hsi ) represents the first spectral feature representation, Ψ(I msi ) represents the second spectral feature.
[0056] In one embodiment, each pair of training images includes a hyperspectral image and a multispectral image that are pixel-aligned and are obtained by preprocessing an initial hyperspectral image and an initial multispectral image of the same food sample captured at the same time point.
[0057] In one embodiment, the preprocessing includes: determining a first image region representing a corresponding food sample in the original hyperspectral image, and a second image region representing a corresponding food sample in the original multispectral image; determining a first bounding box and a first minimum bounding rectangle of the first image region, and a second bounding box and a second minimum bounding rectangle of the second image region; aligning the first image region and the second image region by overlapping the first bounding box and the first minimum bounding rectangle of the first image region with the second bounding box and the second minimum bounding rectangle of the second image region, respectively; identifying a maximum inscribed rectangle for the aligned first image region and the second image region, and selecting corresponding portions of the first image region and the second image region within the maximum inscribed rectangle as pixel-aligned hyperspectral and multispectral images, respectively.
[0058] The spectral base model can be selected from existing spectral base models. In one embodiment, the spectral base model is SpectralGPT, and the second training process is performed with the Transformer of SpectralGPT, including the trained encoder, frozen.
[0059] The spectral reconstruction model can be selected from existing spectral reconstruction models. In one embodiment, the spectral reconstruction model is MST++.
[0060] In one embodiment, the estimation module includes a regression model based on a convolutional neural network, which is configured to receive the reconstructed spectral feature representation as input and output the content of the substance of interest in the food ingredient.
[0061] In one embodiment, the content of the substance of interest in the food increases monotonically over time. In this case, the regression model based on the convolutional neural network can be obtained as follows:
[0062] An initial regression model is obtained, and the initial regression model is trained using a third training set to obtain a trained initial regression model as a base regression model. The third training set includes multiple labeled multispectral images of multiple food samples of the food.
[0063] Unsupervised continuous learning is performed on the base regression model. This unsupervised continuous learning iteratively performs an unsupervised adaptation process, which includes: acquiring multiple multispectral images of the currently monitored food ingredient captured at multiple consecutive time points; generating a regression model prediction value for each of the multiple multispectral images using the current version of the base regression model, wherein for each of the multiple multispectral images, its spectral feature representation is reconstructed by the FSR module based on the multispectral image, and its regression model prediction value represents the content of the substance of interest in the food ingredient at the time point when the multispectral image was captured, predicted by the current version of the base regression model based on the reconstructed spectral feature representation; and updating the current version of the base regression model based on the monotonic cumulative characteristic of the content of the substance of interest in the food ingredient increasing monotonically over time, utilizing the magnitude relationship between the regression model prediction values of the multiple multispectral images captured at multiple consecutive time points. Thus, the updated base regression model obtained after performing unsupervised continuous learning can serve as the regression model based on the convolutional neural network.
[0064] In one embodiment, based on the monotonic cumulative characteristic of the content of the substance of interest in the food ingredient increasing monotonically over time, the current version of the base regression model is updated using the magnitude relationship between the regression model prediction values of the multiple multispectral images captured at multiple consecutive time points, including: when at a later time point t... i+1 The second multispectral image captured The regression model predicts a value no greater than the value at the preceding time point t, which is adjacent to and before the following time point. i The first multispectral image captured When calculating the predicted values of the regression model, adjust the current version of the base regression model to achieve the following error difference. Minimize:
[0065] Where i is 0 or a positive integer, This represents the reconstructed spectral feature representation obtained by the FSR module based on the first multispectral image. This represents the reconstructed spectral feature representation obtained by the FSR module based on the second multispectral image. This represents the regression model prediction value of the first multispectral image. This represents the regression model prediction value of the second multispectral image.
[0066] Advantageously, the reconstructed spectral feature representation provided to the regression model by the FSR module is a 3D feature representation, rather than a 2D feature representation, to better ensure the integrity of the spatial-spectral information. For example, this can be achieved by reconstructing the spectral feature representation from the non-last layer output of the encoder of the FSR module into the regression model.
[0067] It is understood that obtaining the convolutional neural network-based regression model can include two stages: an initial training stage, in which an initially trained basic regression model is obtained from the initial regression model; and an unsupervised continuous learning stage, which is performed on the initially trained basic regression model. The initial regression model can be selected from existing CNN regression models. In the initial training stage, a supervised learning method is used to train the selected initial regression model. The training data used is labeled multispectral image data (from known samples in the database, the training set); that is, each multispectral image sample used for training is labeled, indicating the content of its corresponding substance of interest. Subsequently, in the unsupervised continuous learning stage, the initially trained basic regression model undergoes unsupervised adaptation (unknown samples, the test set) with each time stamp (as new multispectral image data is acquired); that is, the basic regression model is updated at each time stamp (time point).
[0068] Optionally, obtaining the convolutional neural network-based regression model may further include: after performing the unsupervised adaptation process once or multiple times, retraining the current version of the base regression model using the third training set. During retraining, a supervised learning method is employed again, utilizing the training data from the initial training phase to retrain the current version of the base regression model. This helps avoid or at least mitigate the catastrophic forgetting that may result from unsupervised continuous learning.
[0069] The multispectral images in the third training set can be multiple multispectral images of multiple known food samples, such as multispectral images from the second training set mentioned above, but this application is not limited to this.
[0070] This application also provides a method for monitoring the content of substances of interest in food using the above-described system.
[0071] Furthermore, this application also provides a method for monitoring the content of a substance of interest (SIO) in food ingredients, wherein the content of SIO in the food ingredients indicates an indicator of the change of the food ingredients over time. The method includes: acquiring multispectral image data of the food ingredients; reconstructing a spectral feature representation related to the content of the SIO in the food ingredients using an FSR module based on the acquired multispectral image data, wherein the FSR module includes a spectral base model, and the encoder of the spectral base model is trained to enhance the encoder's ability to extract spectral features related to the SIO; and estimating the content of the SIO in the food ingredients based on the reconstructed spectral feature representation.
[0072] In one embodiment, estimating the content of the substance of interest in the food ingredient based on the reconstructed spectral feature representation includes: inputting the reconstructed spectral feature representation into a regression model based on a convolutional neural network, so that the regression model outputs the content of the substance of interest in the food ingredient.
[0073] The various aspects of the system description above relating to this application, including the training of the encoder of the spectral base model, the spectral reconstruction model included in the FSR module and its training, and the regression model based on the convolutional neural network and its acquisition, training and updating, are all applicable to the method proposed in this application, and will not be repeated here.
[0074] The ingredients targeted by the scheme in this application can be various possible fresh ingredients, including, but not limited to, meat such as fish, and the estimated content of the substance of interest can indicate the freshness of the ingredients.
[0075] The scheme described in this application is particularly advantageous for histamine monitoring.
[0076] The following section provides a more detailed description of the concept, background, and various aspects of this application in conjunction with histamine monitoring.
[0077] 1. Overview
[0078] To address the aforementioned challenges associated with histamine monitoring, the inventors propose a passive, low-cost approach for accurate histamine monitoring, along with an example system embodying this approach (hereinafter referred to as "FreshSpec"). As the first passive, low-cost system for accurate histamine monitoring, it requires no human intervention and utilizes only a commercially available multispectral imaging (MSI) camera deployed in the field (as shown in Figure 1). Specifically, FreshSpec achieves this through two novel designs. First, instead of reconstructing the complete HSI containing irrelevant and redundant data, FreshSpec utilizes a novel feature-level SR framework that focuses solely on key, histamine-related spectral features, thereby significantly improving the quality of the reconstructed data used for subsequent histamine regression. To achieve this, the inventors introduce an SFM encoder into the reconstruction process and train it using a contrastive learning scheme to enhance its ability to extract histamine-related features from sashimi samples. Second, noting the monotonically accumulating nature of histamine over time, the inventors designed an unsupervised model improvement scheme based on this characteristic. This scheme ensures that the model adapts to new samples using only a small amount of unlabeled data, thus minimizing the generalization gap. Specifically, FreshSpec utilizes unsupervised continuous learning, constraining pseudo-label relationships to enable the model to continuously improve itself for new sashimi samples during actual deployment.
[0079] The inventors prototyped FreshSpec using a commercially available MSI device costing less than $100 and evaluated its histamine regression performance on 240 sashimi samples covering various sashimi types (including salmon, tuna, and sea bream) and different histamine levels ranging from 0 to 80 mg / 100g. The results showed that FreshSpec achieved a mean coefficient of determination (R-squared, R²) of 0.9319 and an RMSE of 3.101 mg / 100g, comparable to laboratory-grade HSI and significantly superior to the baseline protocol (RMSE reduced by 2.744 mg / 100g, R² improved by 0.1631). Furthermore, FreshSpec demonstrated robustness to various environmental variations, including shooting location, altitude, lighting, sashimi size, and thickness.
[0080] In summary, the main contributions of this application specifically to histamine monitoring include:
[0081] - Introduced FreshSpec, the first MSI-based passive, low-cost system designed to accurately monitor the freshness levels (i.e., histamine) of sashimi in display cases.
[0082] - A novel feature-level spectral reconstruction scheme is proposed, which effectively reconstructs histamine-related spectral features by utilizing a finely tuned spectral baseline model. Simultaneously, an unsupervised model improvement scheme is proposed to adapt the regression model to unseen sashimi samples appearing over time.
[0083] - The datasets include a reconstructed dataset containing 712 pairs of MSI-HSI images and a regression dataset containing 1,440 sets of MSI-histamine data, which are publicly available in reference
[0011] .
[0084] 2. Background
[0085] First, some general background is introduced. Histamine and its detection methods via hyperspectral imaging are then presented. The challenges of cost-effective spectroscopic solutions for histamine detection are then discussed, and the underlying spectroscopic model used in this application is introduced.
[0086] Histamine Management: Histamine is a toxic metabolite produced during spoilage and fermentation caused by certain bacteria. Since harmful levels of histamine do not affect the taste or appearance of food, control measures must be implemented throughout the food chain. The most effective methods for controlling histamine production are time and temperature management, such as refrigeration and freezing. Without proper temperature control, histamine levels can rise rapidly. For example, in fish stored at 20°C or higher, toxic levels of histamine can form within 2 to 3 hours (see reference
[0010] ). In settings such as sushi restaurants or fresh food stores where maintaining refrigeration or freezing conditions is difficult, the freshness and safety of sashimi samples cannot be guaranteed.
[0087] Histamine monitoring via HSI: HSI technology is the best option for providing passive automated histamine monitors, requiring no sample preparation or chemical manipulation. Theoretically, HSI can identify -CH bonds generated during the decarboxylation reaction in histamine production (see references
[0030] and
[0031] ). However, existing HSI-based histamine monitoring products (see reference
[0031] ) rely on complex and extremely expensive hyperspectral cameras (costing, for example, over $10,000), hindering their widespread adoption by ordinary users at the end of the food chain.
[0088] The Challenges of Economical HSI Solutions: Spectral reconstruction is a promising method for addressing high costs, demonstrating good performance in routine food monitoring tasks such as organic fruit classification (see reference
[0038] ) or macronutrient estimation (see reference
[0019] ). However, the high redundancy and unrelated information in the reconstructed HSI, along with reconstruction errors, negatively impact the model's prediction accuracy, rendering previous solutions ineffective in histamine scenarios. Specifically:
[0089] (1) Irrelevant data: Certain spectral bands unrelated to chemical composition can introduce noise, leading to inaccurate predictions and potential overfitting. In other words, the model captures noise rather than data patterns.
[0090] (2) Redundant data: Highly correlated bands can produce multicollinearity and mask the correlated signals, making it complicated to evaluate the independent contributions in the regression algorithm.
[0091] (3) Reconstruction error: Errors can appear randomly in all bands of the output data. Given that the spectral characteristics of trace substances are weak, these errors may obscure valuable spectral information.
[0092] Therefore, this application assumes that reconstructing the complete spectrum is counterproductive for subsequent applications, as reconstruction errors will distract the model from valuable information.
[0093] Spectral foundational models are renowned for their ability to effectively capture complex patterns and representations across various domains (see references [9],
[0017] ,
[0021] , and
[0033] ). In the field of hyperspectral imaging, these models are used to extract meaningful knowledge representations from complex spatial-spectral mixed data to address the challenges of remote sensing (see references [5],
[0018] , and
[0027] ). Notably, SpectralGPT (see reference
[0018] ) has achieved superior performance on multiple remote sensing tasks thanks to its innovative 3D generative pre-trained Transformer architecture. This architecture was trained on over one million spectral images and has over 600 million parameters. SpectralGPT can effectively extract information from spatial-spectral coupled tokens while mitigating hyperspectral redundancy through a 3D masking strategy. Therefore, utilizing the feature extraction capabilities of SpectralGPT to assist spectral reconstruction algorithms is promising. However, given that its training samples come from the field of remote sensing, such as trees and fields, and their spectral characteristics and spatial distribution are completely different from those of sashimi samples, SpectralGPT cannot be directly used for histamine-related spectral feature extraction.
[0094] 3 System Design
[0095] This section details the design of FreshSpec. Figure 2 illustrates the structure of FreshSpec, which comprises three main parts: data preprocessing, spectral reconstruction (SR), and unsupervised continuous histamine regression. We will first explain how to process and pair MSI and HSI data (in Section 3.1). Then, we will describe the design principles and training scheme of the proposed application-related feature-wise spectral reconstruction (AFSR) (in Section 3.2). Next, we will introduce the histamine regression model and the histamine-specific unsupervised continuous learning scheme (HUCL) for adaptive regression model learning (in Section 3.3).
[0096] 3.1 Data Preprocessing and Region of Interest (ROI) Extraction
[0097] Training a spectral reconstruction model requires a large amount of paired MSI and HSI data. Paired data means that data from two devices need to be pixel-aligned, but due to the diversity of hardware parameters and shooting settings, physical acquisition is impractical. Previous work has used downsampling to generate large-scale data, but this has been shown to cause a severe performance degradation on real test datasets (see reference
[0038] ). To address this issue, the inventors propose a sashimi sample preprocessing workflow for extracting sashimi sample regions while aligning MSI images with corresponding HSI images.
[0098] As shown in Figure 3, the process begins with ROI extraction. After black and white image calibration, the Segment Anything model (SAM, see reference
[0021] ) is used to accurately identify and separate ROIs. This process simultaneously generates a binary mask to highlight the selected regions. Next, outlier filtering is performed to refine the masked regions by removing potentially noisy regions, ensuring that only the sashimi region is retained. The subsequent stage is rotation and alignment, which involves calculating the minimum bounding rectangle and adjusting the orientation of the two images to 0 degrees. That is, the bounding boxes of the contours and the minimum bounding matrix completely overlap. Therefore, even if the two images are taken at different angles, they can be aligned to the same angle. After this, cropping and resizing are performed, i.e., the maximum inscribed rectangle is identified to crop the image and extract the relevant ROI from the original image. Then, to avoid shadows that may occur during shooting, 90% of the cropped inner region is taken as the final extraction region. Finally, the size of the cropped MSI and HSI regions is adjusted to meet the size requirements of subsequent analysis. Based on the above process, it can be ensured that the processed MSI and HSI images can be well matched, regardless of the shooting position, angle or resolution of the two cameras.
[0099] In Figure 3, the yellow rectangle represents the smallest bounding rectangle of the sashimi area after using SAM, the blue rectangle represents the rotated area, the green rectangle represents the found inscribed rectangle, and the red rectangle inside it represents the finally extracted area.
[0100] 3.2 Characteristic-level spectral reconstruction (FSR)
[0101] The redundancy of hyperspectral images has been validated in previous work (see references
[0025] and
[0040] ). From a spectral perspective, redundancy refers both to the high correlation between different wavelength bands and the inclusion of bands irrelevant to histamine detection. The correlation matrices in Figures 4A-4C show high correlations in some adjacent wavelengths, which may be useless for subsequent histamine detection. However, this highly correlated information is preserved in the spectrum reconstructed using the SOTA spectral reconstruction (SR) model. Furthermore, since the histamine concentration in sashimi samples is quite low (i.e., on the order of mg / 100g), the spectral absorption characteristics of histamine are easily masked by the absorption bands of other substances (such as proteins). Moreover, reconstruction errors are consistently distributed across various frequency bands of the reconstructed data. This consistent distribution of errors across non-critical bands complicates the extraction of valuable information from the reconstructed dataset.
[0102] The inventors attempted to address the aforementioned problems by proposing FSR. As shown in Figure 5, unlike previous spectral reconstruction algorithms, the FSR model proposed in this application aims to reconstruct the most informative spectral features, rather than the complete HSI data, which are then input into a histamine regression model. To reduce redundancy and extract useful spectral features, the inventors considered using a large-scale spectral foundation model. The inventors observed that current large-scale spectral foundation models, namely SpectralGPT (see reference
[0018] ), offer significantly powerful capabilities in fine-grained spectral feature representation. SpectralGPT is trained using one million spectral images and contains over 600 million parameters; its encoder has a deep 11-layer Transformer and feature-sharing mechanism, enabling efficient learning of spatial-spectral mixed token representations. Based on this, the inventors considered redirecting the output of the state-of-the-art spectral reconstruction model to the encoder of SpectralGPT. FSR can be achieved by applying the encoder of SpectralGPT to the target HSI.
[0103] However, the challenge here is that while SpectralGPT excels in spectral feature representation, it is trained on remote sensing images whose spectral properties differ significantly from those of sashimi and are unrelated to histamine. Therefore, directly using SpectralGPT's original encoder for spectral reconstruction may result in poor performance because the reconstructed features do not correspond to the histamine levels in sashimi, leading to inaccurate predictions.
[0104] To address this, the inventors considered using histamine-related spectral features in sashimi samples to fine-tune the SpectralGPT encoder. Specifically, they proposed a contrastive learning-based training scheme to enhance the feature distance between samples with significant histamine differences while reducing the distance between samples with similar histamine values. Subsequently, the inventors applied the frozen, fine-tuned SpectralGPT encoder to histamine-related spectral feature extraction in FSR, as shown in Figure 6.
[0105] To ensure that the SpectralGPT encoder can extract histamine-related embeddings, the inventors introduced the histamine value of the samples as prior knowledge and combined it with a contrastive learning scheme to guide the fine-tuning process. Specifically, as shown in Figure 6, positive and negative samples are selected based on the difference in histamine values. A histamine difference threshold is given. and The difference in histamine levels between the positive sample and the target sample is less than The difference in histamine levels between the negative sample and the target sample was greater than... Greater than These thresholds are determined based on the target prediction error range. Then, the HSI data (64×64×138) is input into the 3D convolutional layer of SpectralGPT for block embedding, followed by input into the 11-layer Transformer encoder of SpectralGPT to obtain features (294×768) in the latent space. Next, the Euclidean distance between every two sample features can be calculated. Finally, the features of each sample are reduced to their positive samples. The feature distance between them is increased, and each sample is increased in relation to its negative samples. The encoder is fine-tuned using the feature distance between the features. The goal is to minimize the following loss:
[0106] Among them, I, I - ,I + Let Θ(I) represent the HSI sample and its corresponding negative and positive samples, respectively. Θ(I) refers to the features extracted from the HSI sample by the SpectralGPT encoder. - ) refers to the features extracted by the SpectralGPT encoder from the negative sample corresponding to the HSI sample, Θ(I + ) refers to the features extracted by the SpectralGPT encoder from the positive sample corresponding to the HSI sample, ξ(Θ(I + ),Θ(I)) means Θ(I + The mean square distance between ξ(Θ(I) and Θ(I) - ),Θ(I)) means Θ(I - The mean square distance between Θ(I) and Θ(I) is used to enhance the model's ability to recognize these subtle differences, aiming to fine-tune the encoder of SpectralGPT to extract the most informative spectral features related to histamine.
[0107] After fine-tuning the encoder of SpectralGPT to train the feature-level SR model, the Transformer structure in the fine-tuned SpectralGPT is frozen and combined with the SOTA spectral reconstruction model (i.e., MST++, see reference [3]). Figure 6 shows the training process of the feature-level SR model. In this process, the spectral features are extracted using the fine-tuned SpectralGPT encoder, which ensures that the model can use its enhanced capabilities to capture important spectral information. In order to better adapt to the data in terms of resolution and number of channels, the block embedding layer of SpectralGPT before its encoder is not frozen, but is used as a soft constraint for joint retraining. Then, the mean square error between the spectral features extracted from the real HSI and the spectral features output by the feature-level SR model is used as the loss function, i.e.:
[0108] Among them, I hsi and I msi These refer to the hyperspectral image (HSI) data and multispectral image (MSI) data of sample I, respectively, Θ(I hsi ) refers to the features extracted from HSI data by the SpectralGPT encoder, Ψ(I msi The ) represents the features extracted from the MSI data by the feature-level SR model. Therefore, by minimizing the feature gap between the true HSI and MSI, the model can not only improve the quality of the reconstructed data, but also help to understand the inherent spectral characteristics in the data more deeply, paving the way for subsequent histamine regression.
[0109] 3.3 Unsupervised Continuous Histamine Regression
[0110] Following feature-level spectral reconstruction, a CNN-based histamine regression is designed to predict histamine levels based on the reconstructed features. To better preserve the integrity of the spatial-spectral information, the reconstructed 3D features in the latent space (rather than the 2D features in the last layer) are chosen as the input to the CNN model. Before being fed into the CNN model, some denoising operations are performed to avoid overfitting. First, average pooling and a sliding window are used for moving average to reduce the impact of feature noise. Then, the cleaned features (4×4×138) are input, passing sequentially through convolutional layers, batch normalization layers, and average pooling layers. This results in a 16-length vector. This vector is then used in a fully connected layer to obtain the histamine prediction result, as shown in Figure 7.
[0111] Initial Training of the Base Regression Model: The histamine regression model proposed in this application is first trained using supervised learning. During the initial training phase, the initial regression model is trained using labeled MSI data of known samples (e.g., from the training database), and this trained initial regression model is defined as the base regression model.
[0112] Despite the good performance of the CNN-based histamine regression model proposed in this application on known samples, a performance degradation was observed when the model was applied to unseen sashimi samples. As shown in Figure 8, the root mean square error on new samples was significantly higher for salmon, tuna, and sea bream than on known samples. This is due to differences in bacterial colonies and enzymes within individual sashimi samples, which may indicate differences in spectral characteristics between different samples. Therefore, it is necessary to provide an adaptive method to ensure the model's performance on unseen samples.
[0113] To address this challenge, this application proposes an unsupervised persistent histamine regression method that can adapt CNN regression models to new samples without labels. This method is based on two fundamentals: the monotonic cumulative properties of histamine and a large amount of unlabeled MSI data. Figure 9 shows the MSI images of salmon samples at different time points. Specifically, at each time point t... k MSI images of collected samples However, since histamine tags are unavailable in real-world scenarios, predicting histamine tags (e.g.) remains a challenge. For a sample taken at a longer time interval, the histamine level at previous time stamps (i.e., t0 and t1) is unknown. Therefore, supervised continuous learning to achieve model adaptation is not possible. Nevertheless, based on the monotonic cumulative properties of histamine, it can be ensured that… In this way, by restricting the model's predictions for consecutive timestamps of the same sample to follow an increasing property, unsupervised model calibration can be achieved on new samples. Specifically, if the pseudo-label predicted from the MSI of the current timestamp is smaller than the pseudo-label predicted from the MSI of the previous timestamp—which violates the histamine increasing property—the model will be adjusted to minimize this error difference to zero. That is:
[0114] Where Φ represents the current regression model that has not yet been adapted, used to obtain pseudo-labels from the MSI image. This represents the MSI image at the i-th timestamp. This represents the MSI image at the (i+1)th timestamp. This represents the features extracted from the MSI image at the i-th timestamp by the feature-level SR model. This represents the features extracted by the feature-level SR model from the MSI image at the (i+1)th timestamp. This represents the regression model prediction value corresponding to the features extracted by the feature-level SR model for the MSI image at the i-th time stamp. This represents the regression model prediction value corresponding to the feature extracted by the feature-level SR model for the MSI image at the (i+1)th timestamp.
[0115] The model is continuously adjusted for each sample at each time point. For example, for a sample at time point t2 to be predicted, it is first updated at time t1 based on the pseudo-labels generated using the base regression model, and then updated again at time t2 based on the pseudo-labels obtained at time t1 using the updated model, to obtain a regression model suitable for the sample at time point t2.
[0116] Furthermore, catastrophic forgetting is unavoidable during continuous learning, especially in unsupervised learning. Even if predicted values satisfy the histamine-increasing property, prediction bias can still increase. Therefore, in each iteration of the aforementioned model adaptation, the model is designed to be retrained using the training data from the base regression model. In this way, the convergence direction of the model can be controlled and prediction bias can be avoided.
[0117] Algorithm 1 describes the detailed steps of the unsupervised continuous histamine regression method proposed in this application, in which the regression model is continuously updated using the size relationship of pseudo-labels on the time series.
[0118] 4. Implementation
[0119] The inventors implemented a compact and low-cost prototype using an off-the-shelf multispectral camera (see reference
[0036] ) and lighting components. Figure 10A illustrates a real-world working scenario for FreshSpec, where the system is deployed on top of a sashimi display case, a distance from the sample. In the basic setup, this distance from the sample is approximately 20 cm, a common height for sashimi display cases. To minimize system size while avoiding direct light leakage to the camera, a two-layer prototype structure was designed. As shown in Figure 10B, an LED array is deployed on the top layer, while the multispectral camera is located on the bottom layer, with a 2 cm gap between the layers. Simultaneously, seven full-band halogen lamps (VCC7216-ND) are evenly placed on the top layer to ensure sufficient light intensity and a uniform light field. Therefore, the prototype measures only 9 cm × 9 cm × 2 cm, allowing for easy deployment in any commercial refrigerator in a sushi restaurant or fresh food store.
[0120] A SEETRUM SEE8820 MSI camera (see reference
[0036] ) was deployed in the center of the bottom layer to capture MSI images of sashimi at a cost of less than $100. The SEE8820 MSI camera covers the visible and near-infrared bands, with wavelengths ranging from 380 nm to 980 nm. The camera supports up to 31 different wavelength channels but has a relatively coarse spectral resolution of 50–60 nm. Furthermore, the images captured by this camera have a spatial resolution of 512 × 512 and a frame rate of 30 fps, which promises to provide rapid on-site histamine detection.
[0121] 5. Performance Evaluation
[0122] 5.1 Experimental Setup
[0123] 5.1.1 Dataset In the MSI-HSI reconstruction dataset, HSI images with 138 channels (400-1000 nm) and 1024 × 1024 pixels were acquired using a Cucumbert FireflEYE S185 camera, while MSI images were acquired using the aforementioned Seetrum SEE8820 camera. Salmon, tuna, and sea bream were placed in a laboratory environment for five hours to accumulate histamine, resulting in a total of 712 pairs of MSI-HSI samples.
[0124] In the MSI-histamine regression dataset, MSI images were acquired by a similar MSI camera deployed on top of a display case. There were 80 samples each of salmon, tuna, and sea bream. They were divided into 5 groups, corresponding to 5 time points for ground-based acquisition. Ground-based acquisition was achieved through destructive sampling and using a rapid histamine tester (see Figure 11), with an error of less than 0.05 mg / 100g. For continuous learning, MSI images were acquired at previous time points for each group of samples. That is, data from 80 samples were collected at the first time point, and data from 64 samples were collected at the second time point. Two MSI images were captured for each sample. Therefore, a total of 1440 pairs of MSI images were obtained for salmon, tuna, and sea bream. Both datasets will be publicly accessible (see reference
[0011] ).
[0125] 5.1.2 Training Scheme During reconstruction, both the aligned MSI and HSI images are segmented into 64×64 pixel blocks with a stride of 32. To fine-tune the SpectralGPT encoder, K is set... h1 =2 and K h2 =4, batch size is 32, learning rate is 10 -4 The encoder was trained for 10,000 iterations. Then, for feature-level spectral reconstruction, the learning rate was set to 5 × 10⁻⁶. -4 And use cosine annealing to reduce it to 10. -6 The reconstruction model was trained for a total of 50,000 iterations. For the regression model, the MSI image was first downsampled to 64×64 pixels using nearest neighbor interpolation, and then average pooling was performed to further reduce the image size to 4×4. Then, a batch size of 16 and a sum of 10 were used. -4 The initial learning rate is reduced to 10 using cosine decay on the Adam optimizer. -6 The model was trained for 3000 epochs to achieve convergence. Finally, this regression model was saved as the base regression model for subsequent continuous learning. For each time point, past MSI images were used to perform regressions on each sample at a rate of 10... -6The learning rate is adjusted to adapt the model for 8 epochs. Then, this new model is used to update the base regression model for samples at subsequent time points.
[0126] 5.1.3 Baselines The performance of FreshSpec is compared with the following approaches used as baselines: (i) MSI. The average spectrum of the MSI is input into a partial least squares (PLS) regressor, which is widely used in the field of spectral-based freshness detection (see reference
[0045] ). (ii) Recon (SOTA). The HSI is reconstructed using the existing SOTA spectral reconstruction method MST++, and the reconstructed HSI is input into a PLS and CNN regressor. Furthermore, dimensionality reduction is performed before regression using traditional feature extraction methods including principal component analysis, local linear embedding, and multi-order differencing. Among the above methods, the one with the best performance under the reconstructed spectrum is selected as its predictive performance.
[0127] 5.1.4 Evaluation Metrics Since histamine prediction is a regression problem, two main metrics are used for evaluation: the coefficient of determination (R²) and the root mean square error (RMSE). As mentioned in previous work (see references [6],
[0035] ), laboratory-grade HSI (cost > $10,000) can achieve an RMSE of approximately 2-3 mg / 100 g and an R² of 0.96-0.97. The system of this application (costing only $100) aims to achieve similar performance.
[0128] 5.2 Overall Performance
[0129] The performance of FreshSpec was evaluated using 5-fold cross-validation on 80 salmon, tuna, and sea bream samples, respectively. The data in the training and test sets came from completely different sashimi samples.
[0130] Table 1: R² comparison between FreshSpec and baseline
[0131] 5.2.1 Baseline Comparison Figures 12A and 12B show the overall prediction results for all three sashimi sample types. It is clear that FreshSpec can accurately predict histamine values for various sashimi types and histamine levels. Figure 13 and Table 1 show the RMSE and R² of the prediction results, revealing important findings. First, FreshSpec achieves significantly better prediction results compared to the baseline. Compared to the current state-of-the-art (SOTA) method, FreshSpec's average R² is improved by 0.1631, and the RMSE is reduced by 2.744. Second, it is noted that the solution using SOTA SR shows very limited improvement in histamine prediction compared to the MSI-based solution, with an average R² improvement of only 0.0941, which is different from previous tasks (see reference
[0038] ). This is because the high redundancy of HSI data wavelengths is very unfavorable for predicting trace amounts of histamine. In contrast, by using the feature-based spectral reconstruction proposed in this application, FreshSpec's prediction performance on histamine values is significantly improved, with an R² improvement of 0.2572, demonstrating the effectiveness of the system design.
[0132] 5.2.2 Feature-level Comparison To further explore the effectiveness of the FreshSpec design, a two-dimensional t-distributed stochastic neighbor embedding (t-SNE) projection (see reference
[0044] ) was performed to demonstrate the embedding representation of FreshSpec. According to the regulations of different countries on histamine safety values (see reference [8]), the samples were divided into five histamine level groups: 0-5 mg / 100g, 5-10 mg / 100g, 10-20 mg / 100g, 20-40 mg / 100g, and >40 mg / 100g. Figures 14A and 14B show the embedding representation of FreshSpec and the embedding of the same salmon samples after processing with the SOTA SR method. It can be found that FreshSpec's clustering of all five histamine levels is much clearer than the baseline, indicating its effectiveness in extracting histamine-related information during reconstruction.
[0133] 5.2.3 Model Overhead and Time Latency The total number of parameters in the reconstruction model of this application is 2.106M, and the number of floating-point operations per second (FLOPs) is 3,050.64M. This is due to the use of a large spectral base model encoder in AFSR. Furthermore, the inference time for reconstructing the spectral features of one MSI sample on an NVIDIA T4 Tensor Core GPU is 18.6 milliseconds, which is acceptable for practical applications of real-time spectral feature acquisition on the server side.
[0134] The regression model in this application has a total of 19.9K parameters and a FLOP of 318.48K, which can be easily deployed on any device, and can predict histamine in an MSI sample in just 0.4 milliseconds. Furthermore, for model updates in the HUCL section, the time delay at each new timestamp for a salmon sample is only 19.2 milliseconds. Overall, combining reconstruction and regression delays with image capture delay (33.3 milliseconds per image), the proposed system can achieve real-time histamine monitoring with a latency of less than 1 second.
[0135] 5.3 Ablation Study
[0136] Next, ablation studies were conducted to explore the various modules of FreshSpec in order to demonstrate the effectiveness of the system design.
[0137] Table 2: Ablation studies for AFSR and HUCL modules
[0138] 5.3.1 Effectiveness of the AFSR Module Table 2 presents the results of the ablation studies. Compared to the baseline using the SOTA super-resolution (SR) algorithm, the AFSR module achieved an improvement of approximately 0.1463 on R², exceeding the benefit of reconstructing HSI from MSI. This result highlights the importance of reconstructing useful features for trace substance prediction tasks. Furthermore, it underscores the effectiveness of the AFSR module proposed in this application. By employing a contrastive learning-based fine-tuning method, a large-scale spectral foundation model can extract information-rich histamine-related features to support subsequent regression tasks. Notably, the improvement of the AFSR module across the three sashimi samples is consistent, demonstrating its general effectiveness.
[0139] Furthermore, the importance of the fine-tuning step in AFSR was verified. In FSR, directly using the original SpectralGPT encoder trained on remote sensing images without fine-tuning—represented in Table 2 as Recon(SOTA+SFM)—achieved a limited performance increase of 0.0205 compared to the SOTA approach. The fine-tuning method in this application, using sashimi images and histamine-related feature distributions, significantly improved R² by 0.1258, demonstrating its necessity.
[0140] 5.3.2 Effectiveness of the HUCL Module A CNN-based model was designed to predict histamine levels in various sashimi samples, and an unsupervised continuous learning scheme, HUCL, was introduced to enhance its generalization ability to new samples. The effectiveness of the HUCL module is reflected in Table 2; Table 2 shows that the R² improved by 0.0537 compared to the baseline. Although the improvement of the HUCL module is relatively mild compared to the AFSR module, this is as expected. When using time-series MSI data, the HUCL module updates the loss function of the predicted samples without following the histamine increment principle, only fine-tuning it when the base model performs well. However, it is believed that HUCL will prove more beneficial in practical applications because it can continuously update the model using new, unlabeled data, allowing it to adapt and improve with the emergence of new samples during deployment.
[0141] 5.4 Robustness
[0142] Users may use FreshSpec in different scenarios, where ambient brightness, temperature, and even sample attributes such as size and location can affect system performance. Therefore, the robustness of FreshSpec under various experimental and environmental conditions was evaluated.
[0143] 5.4.1 Influence of Ambient Temperature The system operates in display cases within sushi restaurants or fresh food stores. Different stores have different temperature settings for their display cases, therefore it is necessary to ensure the system functions properly under varying ambient temperatures. The histamine prediction of FreshSpec on eight salmon samples at five different ambient temperature levels was evaluated, and the results are shown in Figure 15. As can be seen from the figure, the prediction RMSE of FreshSpec at all ambient temperatures was below 6 mg / 100g, superior to the best baseline result, demonstrating the stability of FreshSpec across various temperatures. Furthermore, it was observed that the prediction performance of FreshSpec improved with increasing ambient temperature. This may be due to the temperature drift of the multispectral camera used, and the fact that the training set samples were collected at room temperature, introducing some bias. This issue can be addressed by calibrating the hardware, but is not the focus of this application.
[0144] 5.4.2 Influence of Ambient Lighting Conditions Ambient light can interfere with MSI camera readings, potentially negatively impacting FreshSpec's performance. Therefore, the performance of FreshSpec under various realistic lighting conditions was investigated. Four typical lighting settings were considered: bright (approximately 550 lx), normal (approximately 250 lx), dim (approximately 50 lx), and dark (0 lx). As shown in Figure 16, FreshSpec's performance is highly stable under different lighting conditions. This is attributed to the background removal step in the data preprocessing workflow. Furthermore, FreshSpec consistently outperformed the baseline under all lighting conditions.
[0145] 5.4.3 Influence of Sample Location As shown in Figure 1, the sashimi samples were not placed directly below the camera in the central region. In most cases, FreshSpec captured sample images at different angles. Therefore, to evaluate the robustness of FreshSpec to different sample locations, the central region and four edge locations were considered, each at approximately a 30° angle to the camera center. Figure 17 shows the results, demonstrating that FreshSpec is robust to different locations. This is because it utilizes a circular light source that can uniformly illuminate the area. Furthermore, by employing a cropping step, only the smooth central region of the sample is extracted for analysis, thus avoiding problems caused by uneven illumination.
[0146] 5.4.4 Effect of Sample-to-Equipment Distance As shown in Figure 18, FreshSpec was deployed at the top of the display case. Considering the different sizes of the sashimi display cases, the distance between the sashimi samples and the camera varied. Therefore, the performance of FreshSpec was studied by placing samples at different distances from the camera (including 10 cm, 15 cm, and 20 cm). The results are shown in Figure 18. It can be found that the performance of FreshSpec was consistently better than the baseline. However, as the distance between the sample and the camera increased, the performance of FreshSpec decreased slightly. This is because the focal length of the selected MSI camera is greater than 20 cm. If the sample is too close to the camera, the captured image will be out of focus and contain unwanted noise, which may affect the spectral analysis. Fortunately, the height of the sashimi display case is usually above 20 cm, within the applicable working distance of the multispectral camera.
[0147] 5.4.5 Impact of Sample Size Fresh food stores may offer sashimi in various sizes for sale. Therefore, it is necessary to study the robustness of FreshSpec on sashimi samples of different sizes, especially on small samples. Three different sample sizes were selected: 3 × 1.5 cm², 6 × 3 cm², and 9 × 3 cm². Figure 19 shows the comparison results with the baseline. It is clear that the performance of both FreshSpec and the baseline decreases as the sample size decreases. Since small samples generate smaller MSI image patches containing less spatial information, it should be more difficult to obtain accurate predictions. Nevertheless, it is noted that even with very small sample sizes, the RMSE of FreshSpec is still below 6 mg / 100 g. Moreover, the performance decline of FreshSpec with decreasing sample size is much smaller than that of the baseline, indicating that FreshSpec has better robustness than the baseline.
[0148] 5.4.6 Effect of Sample Thickness The thickness of sashimi in the sales display case may vary. Different thicknesses of sashimi samples can affect the reflection and scattering of light within the sample. For example, thicker samples absorb a large amount of incident light, thus reducing the intensity of reflected light. Therefore, the robustness of FreshSpec at different sample thicknesses, including 2 cm, 4 cm, 6 cm, and 8 cm, was evaluated. Figure 20 shows the results, which can be seen that the RMSE of FreshSpec continuously increases with increasing sample thickness. This is because thicker sashimi samples absorb more light, thus reducing the quality of the collected MSI image and making histamine prediction more difficult. However, it is also noted that once the sample is sufficiently thick, the rate of increase in RMSE slows down, indicating a lower limit to FreshSpec performance. The performance degradation caused by sample thickness can be addressed by adding a stronger light source. Currently, only 7 LEDs are used for energy saving considerations, but this can be expanded to more LEDs in the future.
[0149] 6. Related work
[0150] This section briefly reviews work related to histamine detection in meat, hyperspectral reconstruction, and unsupervised continuous learning.
[0151] Histamine detection is crucial for determining histamine levels in food and is essential for food safety. Early detection methods primarily used colorimetric assays and chromatographic separation techniques (see references [4],
[0020] ,
[0041] , and
[0042] ). Newer techniques have emerged for histamine analysis (see references [6],
[0023] ,
[0024] , and
[0046] ), such as non-enzymatic biosensors, fluorescence, and spectral imaging. However, many of these methods involve time-consuming and expensive laboratory procedures to obtain reliable results. While some low-cost commercial biosensor test kits are available (see references
[0022] and
[0034] ), they still require multiple chemical operations, limiting their accessibility to the average user. In contrast, FreshSpec achieves passive detection without any chemical operations by utilizing non-invasive spectral imaging technology. More importantly, thanks to its carefully designed reconstruction and continuous learning algorithms, FreshSpec significantly reduces the cost of spectral imaging systems while maintaining good histamine detection accuracy.
[0152] Hyperspectral reconstruction is proposed to address the high cost of hyperspectral imaging systems. It can reconstruct hyperspectral information from limited spectral measurements (such as RGB images) (see reference
[0049] ). Generally, it is divided into two categories: prior-based methods and data-driven methods. Prior-based methods (see references [2],
[0012] , and
[0028] ) utilize statistical information from hyperspectral images, such as sparsity, spatial structure similarity, and spectral correlation, to identify possible solutions. In contrast, data-driven methods (see references [1], [3],
[0026] ,
[0039] ,
[0047] , and
[0051] ) utilize abstract features found in large RGB and hyperspectral image datasets to achieve more accurate results. However, these algorithms often reconstruct complete hyperspectral data without filtering out redundant or irrelevant information, leading to performance degradation in subsequent applications. In contrast, FreshSpec introduces a novel feature-based spectral reconstruction algorithm that focuses solely on application-relevant information, thereby enhancing the utility of the reconstructed data in practical applications.
[0153] Unsupervised continuous learning (UCL) focuses on incrementally learning new tasks without human labeling. Compared with supervised continuous learning, UCL faces several challenges besides catastrophic forgetting, including the representation of unlabeled data and the discovery of outliers. To address the problem of unlabeled data, existing UCL algorithms employ various self-supervised learning techniques, such as pseudo-labeling (see reference
[0015] ) and contrastive loss (see reference
[0043] ). Some UCL tasks are specifically designed for representation forgetting (see reference [7]) or output bias (see reference
[0016] ). While these general methods show good usability in tasks such as image classification, they are not suitable for histamine detection. In contrast, FreshSpec proposes a novel continuous learning framework based on the incremental properties of histamine and defines an appropriate loss function to mitigate the catastrophic forgetting problem.
[0154] 7. Discussion and Future Work
[0155] This section will discuss potential extensions to FreshSpec.
[0156] To enhance performance, this paper trained unique regression models for each of the three sashimi types (salmon, tuna, and sea bream) to address the inherent differences in their spectral characteristics. In practical applications, each sashimi sample can be easily classified based on its image appearance features (see reference
[0029] ), and its MSI data can then be fed into the corresponding model for accurate histamine prediction.
[0157] Real-world deployment allows FreshSpec to be further extended to achieve large-area shooting covering the entire display case by utilizing a pan-tilt to rotate the MSI camera, unaffected by the limitations of field of view (FOV) in a single direction. Evaluation results have confirmed that FreshSpec maintains satisfactory performance for each frame, regardless of where the salmon is placed within the FOV. Furthermore, by analyzing the MSI video, more information about the same sample can be obtained across different frames, potentially enhancing the system of this application, which will be explored in future work.
[0158] Extending to other trace substances, FreshSpec's core concepts can be extrapolated to the monitoring of other trace substances, such as total volatile basic nitrogen (see reference
[0050] ). Specifically, in feature-level spectral reconstruction, the encoder of the base model can be fine-tuned according to the characteristic distribution of the trace substance. For the unsupervised continuous learning part, the loss definition used for model updates can be modified according to the specific changing characteristics of each trace substance over time. Overall, FreshSpec provides new insights and solutions for advancing the field of trace substance monitoring in everyday food.
[0159] 8. Conclusion
[0160] FreshSpec, a low-cost spectral imaging system designed for precise histamine detection in sashimi, is introduced. This system operates autonomously without human intervention. To achieve this, FreshSpec employs a novel feature-level spectral reconstruction framework that, aided by a spectral base model encoder, minimizes irrelevant information and redundancy while preserving key histamine-related spectral features. Furthermore, FreshSpec introduces an unsupervised model improvement scheme that leverages the monotonic accumulation of histamine over time, enabling the model to continuously adapt and improve with new sashimi samples during real-world deployment. Experimental evaluations demonstrate that FreshSpec achieves high accuracy in histamine monitoring, achieving an R² value of 0.9319 across 240 samples from three sashimi types, significantly outperforming the baseline of 0.1631. FreshSpec shows great potential for direct deployment in sashimi shops and fresh food restaurants, ensuring food freshness and safety.
[0161] It should be understood that this application can be implemented as a computer device including a memory and a processor. The memory stores computer instructions executable by the processor, which, when executed by the processor, instruct the processor to perform the steps of the method of this application. The executable computer instructions can be embodied and implemented in the form of an application program. This computer device can be broadly defined as a server, a terminal, or any other electronic device with the necessary computing and / or processing capabilities. In one embodiment, the computer device may include a processor, memory, network interface, communication interface, etc., connected via a system bus. The processor of the computer device can be used to provide the necessary computing, processing, and / or control capabilities. The memory of the computer device may include a non-volatile storage medium and internal memory. The non-volatile storage medium may store an operating system, computer programs, etc. The internal memory can provide an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface and communication interface of the computer device can be used to connect and communicate with external devices via a network. When the computer program is executed by the processor, it performs the steps of the method of this application.
[0162] It should be understood that this application can be implemented as a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the steps of the method of this application to be performed. In one embodiment, the computer program is distributed across multiple network-coupled computer devices or processors, such that the computer program is stored, accessed, and executed in a distributed manner by one or more computer devices or processors. A single method step / operation, or two or more method steps / operations, may be executed by a single computer device or processor or by two or more computer devices or processors. One or more method steps / operations may be executed by one or more computer devices or processors, and one or more other method steps / operations may be executed by one or more other computer devices or processors. One or more computer devices or processors may execute a single method step / operation, or execute two or more method steps / operations. This application can also be implemented as a computer program product comprising a computer program that, when executed by a processor, causes the steps of the method of this application to be performed.
[0163] Those skilled in the art will understand that all or part of the steps of this application can be performed by a computer program instructing related hardware, such as a computer device or processor. The computer program may be stored in a non-transitory computer-readable storage medium, and when executed, it causes the steps of the method of this application to be performed. Depending on the context, any references herein to memory, storage, databases, or other media may include non-volatile and / or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.
[0164] Several references have been mentioned and cited above, and the contents of each of these references are incorporated into this paper in full by way of citation.
[0165] Although this application has been described in conjunction with embodiments, those skilled in the art will understand that the above description and drawings are exemplary and not restrictive, and this application is not limited to the disclosed embodiments. Various modifications and variations are possible without departing from the spirit of this application.
[0166] References:
Claims
1. A system for monitoring the content of a substance of interest in food, wherein the content of the substance of interest in the food indicates an indicator of the change of the food over time, the system comprising: The feature-level spectral reconstruction module is configured to: reconstruct a spectral feature representation related to the content of the substance of interest in the food based on the multispectral image data of the food, wherein the feature-level spectral reconstruction module includes a spectral base model, and the encoder of the spectral base model is trained to enhance the encoder's ability to extract spectral features related to the substance of interest; The estimation module is configured to estimate the content of the substance of interest in the food ingredient based on the reconstructed spectral feature representation.
2. The system according to claim 1, wherein, The encoder of the spectral base model is trained by iteratively performing a first training process using a first training set, the first training set including multiple hyperspectral images of multiple food samples of the food, the multiple food samples having different contents of the substance of interest, the first training process including: A set of hyperspectral images selected from the plurality of hyperspectral images is acquired. The set of hyperspectral images includes a target sample, a negative sample corresponding to the target sample, and a positive sample corresponding to the target sample. The content difference of the substance of interest between the positive sample and the target sample is less than a first content threshold, and the content difference of the substance of interest between the negative sample and the target sample is greater than a second content threshold. The second content threshold is greater than the first content threshold. The encoder of the spectral base model is trained using this set of hyperspectral images to reduce the feature distance between the target sample and the positive sample and increase the feature distance between the target sample and the negative sample.
3. The system according to claim 2, wherein, Training the encoder of the spectral base model using the set of hyperspectral images to reduce the feature distance between the target sample and the positive sample and increase the feature distance between the target sample and the negative sample includes: adjusting the encoder of the spectral base model using the set of hyperspectral images to minimize the following loss. Where I represents the target sample, I - I represents the negative sample corresponding to the target sample. + Let Θ(I) represent the positive sample corresponding to the target sample, and let Θ(I) represent the spectral features of the target sample extracted by the encoder. - ) represents the spectral features of the negative sample corresponding to the target sample extracted by the encoder, Θ(I + ) represents the spectral features of the positive sample corresponding to the target sample extracted by the encoder, ξ(Θ(I + ), Θ(I)) means Θ(I + The mean square distance between ξ(Θ(I) and Θ(I) - ),Θ(I)) means Θ(I - The mean square distance between ) and Θ(I).
4. The system according to claim 1, wherein, The feature-level spectral reconstruction module further includes a spectral reconstruction model. The output of the spectral reconstruction model is connected to the input of the spectral base model. The spectral reconstruction model is trained by iteratively executing a second training process using a second training set. The second training set includes multiple pairs of training images, each pair of training images including a hyperspectral image and a multispectral image of the same food sample corresponding to the same time point. The second training process includes: The hyperspectral image from a pair of training images is input into the trained encoder, and a first spectral feature representation is obtained at the output of the trained encoder; The multispectral image from the pair of training images is input into the spectral reconstruction model, and a second spectral feature representation is obtained at the output of the trained encoder; and The spectral reconstruction model is adjusted to minimize the difference between the first spectral feature representation and the second spectral feature representation.
5. The system according to claim 4, wherein, Adjusting the spectral reconstruction model to minimize the difference between the first spectral feature representation and the second spectral feature representation includes: adjusting the spectral reconstruction model such that the following loss... Minimize: Among them, I hsi and I msi Let Θ(I) represent the hyperspectral image and multispectral image in the pair of training images, respectively. hsi ) represents the first spectral feature representation, Ψ(I msi ) represents the second spectral feature.
6. The system according to claim 4, wherein, The spectral base model is SpectralGPT, and the second training process is performed with the Transformer of SpectralGPT, including the trained encoder, frozen.
7. The system according to claim 4, wherein, The spectral reconstruction model is MST++.
8. The system according to claim 1, wherein, The estimation module includes a regression model based on a convolutional neural network, which is configured to receive the reconstructed spectral feature representation as input and output the content of the substance of interest in the food ingredient.
9. The system according to claim 8, wherein, The content of the substance of interest in the food increases monotonically over time, and the regression model based on the convolutional neural network is obtained in the following way: An initial regression model is obtained, and the initial regression model is trained using a third training set to obtain a trained initial regression model as a base regression model. The third training set includes multiple labeled multispectral images of multiple food samples of the food. Unsupervised continuous learning is performed on the base regression model, and the unsupervised continuous learning iteratively performs an unsupervised adaptation process, which includes: Acquire multiple multispectral images of the currently monitored food ingredient captured at multiple consecutive time points; Using the current version of the basic regression model, a regression model prediction value is generated for each of the plurality of multispectral images. For each of the plurality of multispectral images, its spectral feature representation is reconstructed by the feature-level spectral reconstruction module based on the multispectral image, and its regression model prediction value represents the content of the substance of interest in the food at the time point at which the multispectral image was captured, predicted by the current version of the basic regression model based on the reconstructed spectral feature representation. Based on the monotonic cumulative characteristic of the content of the substance of interest in the food increasing monotonically over time, the current version of the basic regression model is updated using the magnitude relationship between the predicted values of the regression models of the multiple multispectral images captured at multiple consecutive time points. The updated base regression model obtained after performing unsupervised continuous learning is used as the regression model based on the convolutional neural network.
10. The system according to claim 9, wherein, Based on the monotonic cumulative characteristic of the content of the substance of interest in the food increasing monotonically over time, the current version of the basic regression model is updated using the magnitude relationship between the predicted values of the regression models of the multiple multispectral images captured at multiple consecutive time points, including: When at the later time point t i+1 The second multispectral image captured The regression model predicts a value no greater than the value at the preceding time point t, which is adjacent to and before the following time point. i The first multispectral image captured When calculating the predicted values of the regression model, adjust the current version of the base regression model to achieve the following error difference. Minimize: Where i is 0 or a positive integer, This represents the reconstructed spectral feature representation obtained by the feature-level spectral reconstruction module based on the first multispectral image. This represents the reconstructed spectral feature representation obtained by the feature-level spectral reconstruction module based on the second multispectral image. This represents the regression model prediction value of the first multispectral image. This represents the regression model prediction value of the second multispectral image.
11. The system according to claim 9, wherein, After performing the unsupervised adaptation process once or multiple times, the current version of the base regression model is retrained using the third training set.
12. The system according to claim 1, further comprising: A multispectral camera, through which the multispectral image data is acquired.
13. The system according to any one of the preceding claims, wherein, The substance of interest includes histamine, and the indicator is the freshness of the food ingredient.
14. The system according to any one of the preceding claims, wherein, The ingredients include sashimi.
15. A method for monitoring the content of a substance of interest in a food ingredient, wherein the content of the substance of interest in the food ingredient is an indicator of the change of the food ingredient over time, the method comprising: Acquire multispectral image data of the food ingredients; Based on the acquired multispectral image data, a feature-level spectral reconstruction module is used to reconstruct spectral feature representations related to the content of the substance of interest in the food. The feature-level spectral reconstruction module includes a spectral base model, and the encoder of the spectral base model is trained to enhance the encoder's ability to extract spectral features related to the substance of interest. Based on the reconstructed spectral characteristics, the content of the substance of interest in the food ingredient is estimated.
16. The method according to claim 15, wherein, The encoder of the spectral base model is trained by iteratively performing a first training process using a first training set, the first training set including multiple hyperspectral images of multiple food samples of the food, the multiple food samples having different contents of the substance of interest, the first training process including: A set of hyperspectral images selected from the plurality of hyperspectral images is acquired. The set of hyperspectral images includes a target sample, a negative sample corresponding to the target sample, and a positive sample corresponding to the target sample. The content difference of the substance of interest between the positive sample and the target sample is less than a first content threshold, and the content difference of the substance of interest between the negative sample and the target sample is greater than a second content threshold. The second content threshold is greater than the first content threshold. The encoder of the spectral base model is trained using this set of hyperspectral images to reduce the feature distance between the target sample and the positive sample and increase the feature distance between the target sample and the negative sample.
17. The method according to claim 16, wherein, Training the encoder of the spectral base model using the set of hyperspectral images to reduce the feature distance between the target sample and the positive sample and increase the feature distance between the target sample and the negative sample includes: adjusting the encoder of the spectral base model using the set of hyperspectral images to minimize the following loss. Where I represents the target sample, I - I represents the negative sample corresponding to the target sample. + Let Θ(I) represent the positive sample corresponding to the target sample, and let Θ(I) represent the spectral features of the target sample extracted by the encoder. - ) represents the spectral features of the negative sample corresponding to the target sample extracted by the encoder, Θ(I + ) represents the spectral features of the positive sample corresponding to the target sample extracted by the encoder, ξ(Θ(I + ), Θ(I)) means Θ(I - The mean square distance between ξ(Θ(I) and Θ(I) - ), Θ(I)) means Θ(I - The mean square distance between ) and Θ(I).
18. The method according to claim 15, wherein, The feature-level spectral reconstruction module further includes a spectral reconstruction model. The output of the spectral reconstruction model is connected to the input of the spectral base model. The spectral reconstruction model is trained by iteratively executing a second training process using a second training set. The second training set includes multiple pairs of training images, each pair of training images including a hyperspectral image and a multispectral image of the same food sample corresponding to the same time point. The second training process includes: The hyperspectral image from a pair of training images is input into the trained encoder, and a first spectral feature representation is obtained at the output of the trained encoder; The multispectral image from the pair of training images is input into the spectral reconstruction model, and a second spectral feature representation is obtained at the output of the trained encoder; and The spectral reconstruction model is adjusted to minimize the difference between the first spectral feature representation and the second spectral feature representation.
19. The method according to claim 18, wherein, Adjusting the spectral reconstruction model to minimize the difference between the first spectral feature representation and the second spectral feature representation includes: adjusting the spectral reconstruction model such that the following loss... Minimize: Among them, I hsi and I msi Let Θ(I) represent the hyperspectral image and multispectral image in the pair of training images, respectively. hsi ) represents the first spectral feature representation, Ψ(I msi ) represents the second spectral feature.
20. The method of claim 15, wherein, Based on the reconstructed spectral feature representation, the content of the substance of interest in the food ingredient is estimated, including: The reconstructed spectral feature representation is input into a regression model based on a convolutional neural network, so that the regression model outputs the content of the substance of interest in the food ingredient.
21. The method according to claim 20, wherein, The content of the substance of interest in the food increases monotonically over time, and the regression model based on the convolutional neural network is obtained in the following way: An initial regression model is obtained, and the initial regression model is trained using a third training set to obtain a trained initial regression model as a base regression model. The third training set includes multiple labeled multispectral images of multiple food samples of the food. Unsupervised continuous learning is performed on the base regression model, and the unsupervised continuous learning iteratively performs an unsupervised adaptation process, which includes: Acquire multiple multispectral images of the currently monitored food ingredient captured at multiple consecutive time points; Using the current version of the basic regression model, a regression model prediction value is generated for each of the plurality of multispectral images. For each of the plurality of multispectral images, its spectral feature representation is reconstructed by the feature-level spectral reconstruction module based on the multispectral image, and its regression model prediction value represents the content of the substance of interest in the food at the time point at which the multispectral image was captured, predicted by the current version of the basic regression model based on the reconstructed spectral feature representation. Based on the monotonic cumulative characteristic of the content of the substance of interest in the food increasing monotonically over time, the current version of the basic regression model is updated using the magnitude relationship between the predicted values of the regression models of the multiple multispectral images captured at multiple consecutive time points. The updated base regression model obtained after performing unsupervised continuous learning is used as the regression model based on the convolutional neural network.
22. The method according to claim 21, wherein, Based on the monotonic cumulative characteristic of the content of the substance of interest in the food increasing monotonically over time, the current version of the basic regression model is updated using the magnitude relationship between the predicted values of the regression models of the multiple multispectral images captured at multiple consecutive time points, including: When at the later time point t i+1 The second multispectral image captured The regression model predicts a value no greater than the value at the preceding time point t, which is adjacent to and before the following time point. i The first multispectral image captured When calculating the predicted values of the regression model, adjust the current version of the base regression model to achieve the following error difference. Minimize: Where i is 0 or a positive integer, This represents the reconstructed spectral feature representation obtained by the feature-level spectral reconstruction module based on the first multispectral image. This represents the reconstructed spectral feature representation obtained by the feature-level spectral reconstruction module based on the second multispectral image. This represents the regression model prediction value of the first multispectral image. This represents the regression model prediction value of the second multispectral image.
23. The method according to any one of claims 15-22, wherein, The substance of interest includes histamine, and the indicator is the freshness of the food ingredient.
24. The method according to any one of claims 15-22, wherein, The ingredients include sashimi.
25. A computer device comprising a memory and a processor, the memory storing computer instructions that, when executed by the processor, cause the method according to any one of claims 15-24 to be performed.
26. A non-transitory storage medium having stored thereon computer instructions, which, when executed by a processor, cause the method according to any one of claims 15-24 to be performed.
27. A computer program product comprising computer instructions that, when executed by a processor, cause the method according to any one of claims 15-24 to be performed.