A Modeling and Management Method for Visualized Aflatoxin Detection Data Based on Chemometric Model Adaptive Update
By employing an adaptive update method based on the chemometric model, image unification and feature construction are performed using a built-in reference calibration region. Combined with drift index and anchor guard constraints, the data drift problem in aflatoxin fluorescence detection is solved, achieving stable and controllable model updates and improving the long-term stability and reproducibility of the detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HENGYANG NORMAL UNIV
- Filing Date
- 2026-05-13
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies for aflatoxin fluorescence visualization detection suffer from data distribution drift due to factors such as light source attenuation, differences in camera exposure and gain, fluctuations in environmental temperature and humidity, and differences in reagent batches. This leads to unstable model prediction performance, a lack of effective update mechanisms, and affects the long-term stability and reproducibility of the detection.
An adaptive update method based on a chemometrics model is adopted. Image uniformity preprocessing is performed through a built-in reference calibration area to construct dual-channel ratio features and quality features. Combined with controlled incremental updates based on drift index and anchor guard constraints, the update management of the model is made controllable, verifiable and traceable.
It effectively suppressed the impact of data distribution drift on prediction performance, improved the consistency and repeatability of detection results, reduced the trigger rate of invalid updates, ensured the stability and controllability of the model, and met the requirements of reproducibility and evidence chain integrity for long-term online operation.
Smart Images

Figure CN122494053A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing and model maintenance technology for food safety testing, specifically to a method for modeling and managing aflatoxin visualization detection data based on adaptive updates of a chemometrics model. Background Technology
[0002] Aflatoxin is a highly toxic fungal toxin widely found in grains, oils, and agricultural products. It exhibits strong carcinogenic and teratogenic toxicity to humans and animals, making its rapid on-site detection crucial for food safety. Existing methods typically use dual-channel fluorescence visualization to generate image signals from the sample. An image acquisition unit obtains a dual-channel fluorescence image containing both the sample reaction area and the reference calibration area. Image features are then extracted to establish a concentration-image feature chemometric model. Methods such as partial least squares (PLS) and support vector regression (SVR) are used to achieve quantitative analysis of aflatoxin concentration. In enterprise deployment scenarios, new data is often collected at fixed intervals, and the model is retrained or its parameters are updated before evaluation. Existing machine learning model management technologies are primarily geared towards industrial IoT (such as industrial sensor data stream processing and edge computing terminal data acquisition). Their key characteristics are preprocessing and version management of general sensor signals, which differs fundamentally from aflatoxin fluorescence visualization detection. Industrial IoT systems input general digital electrical signals (voltage, current, vibration spectrum, etc.), while aflatoxin fluorescence detection inputs a dual-channel fluorescence image including a built-in reference calibration area. Industrial IoT systems are characterized by general temporal statistical features, while the core feature of fluorescence detection is a dual-channel logarithmic ratio feature designed specifically for fluorescence quenching / enhancement mechanisms. Furthermore, data quality assessment in industrial IoT systems is based on general sensor parameters, while quality assessment in fluorescence detection requires image-level gating based on color consistency and brightness deviation within the reference calibration area. Therefore, existing industrial IoT model management technologies cannot be directly applied to data modeling and management in aflatoxin fluorescence visualization detection scenario.
[0003] However, during long-term operation, the platform is susceptible to factors such as light source attenuation, differences in camera exposure and gain, fluctuations in environmental temperature and humidity, and variations in reagent batches and sample matrices, causing input data distribution to drift. Taking a fixed-cycle retraining scheme as an example, in simulated long-term operation and verification, when the environmental temperature range exceeds 15°C and the intensity attenuation due to light source aging exceeds 8%, the root-mean-square error (RMSE) of the fixed model without any drift compensation can deteriorate from approximately 0.65 μg / kg at the baseline to over 2.80 μg / kg, with a spiked recovery rate deviation exceeding 20%. Furthermore, due to the lack of a gating mechanism, low-quality batch data (such as batches with excessive calibration deviation or insufficient sample size) will also trigger updates, further exacerbating the instability of model performance. At the same time, due to the lack of versioned evidence storage and rollback mechanisms, once an update introduces performance degradation, it is impossible to stop the damage in time, and historical detection conclusions cannot be linked to the corresponding model version for reproduction and auditing. The above problems are particularly prominent in enterprise-level long-term online deployment scenarios, severely restricting the long-term stability, reproducibility, and integrity of the evidence chain for the visual quantitative detection of aflatoxin. Summary of the Invention
[0004] Technical Objective: To address the shortcomings of existing technologies, this invention discloses a method for modeling and managing aflatoxin visualization detection data based on adaptive updates using a chemometric model. Under long-term operation of the aflatoxin visualization detection platform, this method can suppress prediction shifts caused by data distribution drift and achieve controllable, verifiable, rollback-able, and traceable data modeling and management.
[0005] Technical solution: To achieve the above technical objectives, the present invention adopts the following technical solution:
[0006] A method for modeling and managing aflatoxin visualization detection data based on adaptive updates using a chemometric model, specifically including the following steps:
[0007] Step 1: Obtain the detection data corresponding to the sample to be tested. The detection data includes at least dual-channel fluorescence image data containing a built-in reference calibration area, and auxiliary modal data acquired synchronously with the dual-channel fluorescence image data.
[0008] Step 2: Perform image homogenization preprocessing on the dual-channel fluorescence image data based on the built-in reference calibration area. The image homogenization preprocessing includes brightness normalization and color correction to obtain a homogenized image.
[0009] Step 3: Extract dual-channel ratio features from the uniformized image, and extract quality features based on the uniformized image and auxiliary modal data to construct a feature vector;
[0010] Step 4: Input the feature vector into the chemometric model and output the predicted aflatoxin concentration and confidence index;
[0011] Step 5: Construct a drift index based on the current period feature vector set and the anchor sample set, and determine whether to trigger a model update based on the drift index and the gating conditions. The anchor sample set is a sample set with a reference concentration.
[0012] Step 6: When the model update is triggered, a controlled incremental update with anchor guard constraints is performed on the chemometrics model based on the incremental dataset to obtain the updated model;
[0013] Step 7: Perform anchor point verification on the updated model. If the verification passes, publish the updated model as the current model version and perform versioned documentation. If the verification fails, roll back to the previous model version that passed anchor point verification.
[0014] Step 8: Associate and store the predicted aflatoxin concentration, confidence index, and current model version identifier, and then output them.
[0015] Preferably, the built-in reference calibration area includes at least one of a fixed reflection reference area and a fixed fluorescence reference area;
[0016] The color correction includes: calculating a color mapping matrix based on the target color vector and the current color vector of the built-in reference calibration area, and performing a color mapping matrix transformation on the dual-channel fluorescence image data.
[0017] Preferably, the dual-channel ratio feature is obtained by statistically analyzing the intensity of the first channel and the intensity of the second channel within the same region of interest and constructing a logarithmic ratio, wherein the logarithmic ratio is:
[0018] ,
[0019] Where r represents the dual-channel ratio characteristic. The first channel strength, For the second channel intensity, It is a positive constant.
[0020] Preferably, the quality characteristics include at least one of image sharpness index, exposure saturation index, and reference calibration area consistency index, and at least one of ambient temperature, ambient humidity, light source stability, camera exposure parameters, and device battery power.
[0021] Preferably, the gating conditions include at least:
[0022] The drift index is not less than the first threshold;
[0023] The quality characteristics satisfy the set of quality thresholds;
[0024] The number of valid samples in the incremental dataset is not less than the minimum sample threshold.
[0025] Preferably, the drift index satisfies the formula:
[0026] ,
[0027] Where t is the period number. The drift index, , The weighting coefficients are satisfied. , The Wasserstein distance between the eigenvector distributions. The set of feature vectors for the current period. For the reference periodic eigenvector set, This represents the error statistic of the anchor sample set in the current period. This is the error statistic for the anchor sample set during the reference period. It is a positive constant.
[0028] Preferably, the anchor guard constraint includes: the error statistic of the updated model on the anchor sample set is not greater than the upper limit of the anchor error; when the anchor guard constraint is not met, a rollback is performed.
[0029] According to claim 1, the method for modeling and managing aflatoxin visualization detection data based on adaptive updating of a chemometric model is characterized in that the controlled incremental update satisfies the formula:
[0030] ,
[0031] in, This is the current model parameter vector. This is the updated model parameter vector. The set of feature vectors for the current period. This is the aflatoxin reference concentration vector corresponding to the feature vector set. The sample weight matrix, For adaptive step size, For gated variables, when the gate condition is met ,otherwise , The anchor point guard penalty coefficient. This is the initial model parameter vector.
[0032] A modeling and management system for aflatoxin visualization detection data based on an adaptive update chemometric model is provided to implement the aforementioned method for modeling and managing aflatoxin visualization detection data based on an adaptive update chemometric model. The system includes a data acquisition module, a consistency preprocessing module, a feature construction module, a concentration prediction module, a drift assessment and gating module, a controlled incremental update module, a validation and version management module, and a result and traceability management module.
[0033] The data acquisition module is used to acquire detection data corresponding to the sample to be tested. The detection data includes at least dual-channel fluorescence image data containing a built-in reference calibration area, and auxiliary modal data acquired synchronously with the dual-channel fluorescence image data.
[0034] The uniformity preprocessing module is connected to the data acquisition module and is used to perform brightness normalization and color correction on dual-channel fluorescence image data based on the built-in reference calibration area, and output a uniform image.
[0035] The feature construction module is connected to the uniformity preprocessing module to extract dual-channel ratio features from the uniformized image and extract quality features based on the uniformized image and auxiliary modal data, and output feature vectors.
[0036] The concentration prediction module is connected to the feature construction module and has a built-in chemometric model, which is used to output the predicted value of aflatoxin concentration and confidence index based on the feature vector.
[0037] The drift assessment and gating module is connected to the feature construction module and the concentration prediction module. It is used to construct a drift index based on the current period feature vector set and the anchor sample set, and output an update trigger signal according to the drift index and gating conditions. The anchor sample set is a sample set with reference concentration.
[0038] The controlled incremental update module is connected to the drift evaluation and gating module. When an update trigger signal is received, the module performs a controlled incremental update on the chemometrics model based on the incremental dataset to obtain an updated model. The controlled incremental update includes anchor guard constraints.
[0039] The verification and version management module is connected to the controlled incremental update module and is used to perform anchor verification on the updated model. When the anchor verification passes, the updated model is versioned and published as the current model version. When the anchor verification fails, the version is rolled back to restore the model version that passed the anchor verification. The versioned data record includes the version number, release time, drift index summary, anchor verification summary, and training data summary fingerprint.
[0040] The results and traceability management module is connected to the concentration prediction module and the validation and version management module. It is used to associate, store and output the aflatoxin concentration prediction value, confidence index and the current model version identifier.
[0041] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for modeling and managing aflatoxin visualization detection data based on an adaptively updated chemometric model.
[0042] Beneficial Effects: The aflatoxin visualization detection data modeling and management method based on adaptive updating of a chemometric model provided by this invention has the following beneficial effects:
[0043] 1. This invention introduces dual-channel fluorescence image data containing a built-in reference calibration area into the detection data acquisition process, and performs brightness normalization and color correction based on the built-in reference calibration area. This ensures that image brightness and color deviations caused by different devices, different exposure parameters, and light source attenuation are uniformly mapped to a uniform space. Furthermore, by constructing a logarithmic ratio feature r within the same region of interest, the influence of overall brightness scaling and some environmental disturbances on the single-channel intensity can be naturally suppressed. Verification shows that under conditions of ±10% light source intensity fluctuation and an ambient temperature range of 18–32℃, the inter-batch coefficient of variation (CV) of the feature vector after uniform preprocessing decreased from 8.3% in the unprocessed form to below 1.9%, and the baseline model RMSE (approximately 0.62 μg / kg, R0) was reduced. 2 =0.999) remained stable under cross-batch conditions, significantly outperforming the single-channel scheme without reference calibration correction (RMSE approximately 2.35 μg / kg), thereby fundamentally suppressing the impact of drift on prediction performance from the input side and improving the consistency and repeatability of prediction results.
[0044] 2. This invention further constructs a drift index that simultaneously reflects differences in feature distribution and anchor point error. It sets gating conditions including a drift index threshold, a set of quality thresholds, and a minimum sample threshold. This ensures that model updates no longer rely on fixed-period retraining, but are triggered only when drift is confirmed, data quality meets standards, and the sample size is sufficient. This mechanism avoids the risk of erroneous updates driven by low-quality data or occasional fluctuations. Simultaneously, the quality features incorporate factors such as sharpness, saturation, reference calibration consistency, temperature and humidity, light source stability, exposure parameters, and equipment power consumption into a unified evaluation. This allows for data validity screening and weighting before updates, enhancing the interpretability and stability of drift determination and update triggering. Compared with a fixed-period retraining scheme, this invention's gating mechanism reduces the invalid update trigger rate (erroneous triggers due to insufficient samples or substandard quality) from approximately 34.7% to below 4.2%, and reduces RMSE abnormal increase events driven by low-quality batches (RMSE single-batch increase exceeding 0.5 μg / kg) by more than 82%, significantly improving online operational stability.
[0045] 3. This invention also proposes a controlled incremental update mechanism that includes an anchor point guard penalty, and introduces anchor point verification, versioned evidence storage, and rollback strategies: the updated model is only released as the current model version when the error statistic on the anchor point sample set is not greater than the anchor point error upper limit; otherwise, it is automatically rolled back to the previous approved version. This closed loop of update-verification-release / rollback-version traceability makes model updates controllable and auditable. Continuous simulation verification (covering 30 collection batches, 6 anchor point concentration locations, 4 replicates per location, for a total of Na=24 anchor point samples): Under the adaptive update scheme of this invention, the RMSE of anchor point verification from ver1 to ver6 were 2.14, 2.09, 2.06, 2.11, 2.07, and 2.05 μg / kg, respectively, all meeting the release constraint of the upper limit of anchor point error of 2.20 μg / kg. The spiked recovery rate was 96.0% to 103.5%, and the standard deviation (RSD) was less than 5%. In contrast, the comparative example (fixed-cycle retraining, no anchor point verification and rollback) under the same drift conditions, the RMSE increased from the baseline of 0.62 μg / kg to more than 2.80 μg / kg, the spiked recovery rate deviation exceeded 20%, and automatic loss prevention was not possible. The present invention has not resulted in any uncontrolled version releases due to performance degradation in 30 batches, and the controllability and stability of model updates are significantly improved compared to the comparative method. At the same time, the predicted value, confidence level and model version identifier are stored together, which facilitates the reproduction of historical test results and the tracing of responsibility, and meets the requirements of stability, reproducibility and integrity of the evidence chain. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0047] Figure 1 This is a flowchart of the method of the present invention;
[0048] Figure 2 This is a system block diagram of the present invention;
[0049] Figure 3 A schematic diagram illustrating the construction of the drift index and the internal control triggering logic;
[0050] Figure 4 A schematic diagram of the release / rollback closed loop for controlled incremental updates and anchor point verification;
[0051] Figure 5 This is a schematic diagram showing the consistency curve between the reference concentration and the predicted concentration.
[0052] Figure 6 Drift index A schematic diagram of batch evolution and gating trigger record;
[0053] Figure 7 A schematic diagram of the anchor point verification error and release constraint curve for model version evolution;
[0054] Figure 8 N is the number of anchor point samples a A schematic diagram of the adaptive value suggestion curve for the penalty coefficient μ. Detailed Implementation
[0055] The present invention will now be described more clearly and completely by way of a preferred embodiment in conjunction with the accompanying drawings, but this does not limit the invention to the scope of the described embodiment.
[0056] like Figure 1 As shown, this invention provides a method for modeling and managing aflatoxin visualization detection data based on adaptive updating of a chemometric model, specifically including the following steps:
[0057] Step 1: Obtain the detection data corresponding to the sample to be tested. The detection data includes at least dual-channel fluorescence image data containing a built-in reference calibration area, and auxiliary modal data acquired synchronously with the dual-channel fluorescence image data.
[0058] Among them, auxiliary modal data refers to the data set that is acquired synchronously with the dual-channel fluorescence image and is used to characterize the acquisition environment and equipment status. It includes at least one or more of the following types, and its acquisition and storage are configured in the system parameter table: ambient temperature, ambient humidity, light source drive current or light source stability index, camera exposure time, camera gain or ISO parameter, and equipment power. When there is expansion hardware, it may also include fluorescence spectral sampling data.
[0059] The test subject is aflatoxin, which in the examples may be aflatoxin B1, aflatoxin M1, or one of both.
[0060] Step 2: Perform image homogenization preprocessing on the dual-channel fluorescence image data based on the built-in reference calibration area. The image homogenization preprocessing includes brightness normalization and color correction to obtain a homogenized image.
[0061] Color correction: Extract the current color vector from the built-in reference calibration area, establish a mapping relationship with the pre-stored target color vector, calculate the color mapping matrix, and apply the matrix to the dual-channel fluorescence image to complete the color correction.
[0062] Brightness normalization: Using the brightness statistics of the reference calibration area as the normalization benchmark, scale normalization is performed on the dual-channel fluorescence images respectively.
[0063] Calibration consistency determination: Calculate the deviation index between the reference calibration area after mapping and the target color vector. If the deviation exceeds the threshold, the sample is marked as an invalid sample and removed from the incremental dataset to avoid data-driven erroneous updates due to calibration failure.
[0064] Step 3: Extract dual-channel ratio features from the uniformized image, and extract quality features based on the uniformized image and auxiliary modal data to construct a feature vector.
[0065] In one embodiment, the feature vector satisfies the following formula:
[0066]
[0067] Where r represents the dual-channel ratio characteristic. The first channel strength, For the second channel intensity, It is a positive constant, and in the example, it is taken as 10. -6 Up to 10 -2 .
[0068] Simultaneously, quality features are extracted, including at least: image sharpness index, exposure saturation index, and reference calibration area consistency index, and may further include temperature, humidity, light source stability, exposure parameters, and device battery power from auxiliary modal data. The dual-channel ratio features are combined with the quality features to construct a feature vector.
[0069] The set of eigenvectors within a period is defined as the set of eigenvectors for the current period. In engineering implementation, it is stored as a matrix, with the number of rows representing the number of valid samples and the number of columns representing the feature dimension. The reference periodic feature vector set is defined as follows: , used as a benchmark for drift evaluation.
[0070] Step 4: Input the feature vector into the chemometric model and output the predicted value of aflatoxin concentration and the confidence index.
[0071] The chemometric model is a regression model, and in the examples, partial least squares regression, support vector regression, or neural network regression can be used.
[0072] The confidence index is used to reflect the reliability of the prediction. In this example, it is obtained by combining whether the quality feature threshold is met and the closeness of the feature vector to the reference distribution. The confidence index is used for sample weighting and update stability enhancement.
[0073] Step 5: Construct a drift index based on the current periodic feature vector set and the anchor sample set, and determine whether to trigger a model update based on the drift index and the gating conditions. The anchor sample set is a sample set with a reference concentration.
[0074] To enhance reproducibility, this embodiment specifies the following:
[0075] like Figure 3 As shown, Wasserstein distance The calculation dimension is the feature dimension; the embodiment adopts a method of calculating and weighting the distance one dimension at a time: for each feature dimension, the current period sample and the reference period sample are sorted, aligned according to the same quantile, and the average of the absolute differences is calculated as the distance of that dimension. Then, the distances of each dimension are summed according to the weights to obtain the distance. When the number of samples in two periods is different, quantile resampling is used to align them to the same number before calculation to ensure consistency.
[0076] Anchor point error statistic type: The example uses root mean square error (RMSE) as the error statistic, which is the square root of the mean of the squared prediction errors of the anchor point samples; mean absolute error (MAE) can also be used, but this example mainly uses RMSE to enhance the sensitivity to abnormal deviations.
[0077] Based on this, the drift index is calculated using the following formula:
[0078]
[0079] in, and These are the anchor point error statistics for the current period and the reference period, respectively. Take 10 -6 Up to 10 -2 .
[0080] The gating conditions include at least the following: the drift index is not less than a first threshold, the quality features meet the quality threshold set, and the number of valid samples in the incremental dataset is not less than the minimum sample threshold. If the gating conditions are met, an update is triggered; otherwise, the current model version remains unchanged and data accumulation continues.
[0081] Step 6: When the model update is triggered, a controlled incremental update with anchor guard constraints is performed on the chemometrics model based on the incremental dataset to obtain the updated model.
[0082] The incremental dataset consists of valid samples that passed the quality assessment in this period. To reduce the impact of low-quality samples, a sample weight matrix is constructed. Its diagonal elements are determined by a combination of confidence index and quality features, with samples having lower confidence levels being assigned smaller weights.
[0083] Controlled incremental updates satisfy the following formula:
[0084]
[0085] in, This is the initial model parameter vector, used to define the guard baseline for parameter deviation. This is the current model parameter vector. This is the updated model parameter vector. Anchor point guard penalty term, used to suppress divergence or unexplainable offset caused by long-term unconstrained parameter drift; This is a gated variable to ensure it is not updated unless triggered. To achieve adaptive step size, the embodiment limits it to 10. -4 Up to 10 -2 It increases with the drift index and decreases with the decline in data quality in order to avoid oscillations.
[0086] To avoid penalty coefficient The selection is based on no specific criteria; this embodiment provides the number of anchor point samples. Relevant executable selection rules: When When less, increase To strengthen defenses; when When there are many, reduce To improve adaptability to real-world drift. The implementation example will... Limited to Within the interval, Take a value of 0.01 to 0.5. Take values from 0.5 to 10, and solidify the mapping rules through the system parameter table: when Take close The value; when Take the middle to high value; when Take the lower-middle value. This mapping rule is related to the current... This information is also recorded in the version verification record to ensure consistent reproduction.
[0087] Step 7: Perform anchor point verification on the updated model. If the verification passes, publish the updated model as the current model version and perform versioning and certification. If the verification fails, roll back to the previous model version that passed anchor point verification.
[0088] like Figure 4 As shown, after the update, anchor point validation is performed on the anchor point sample set, the anchor point error statistic is calculated and compared with the upper limit of the anchor point error; the upper limit of the anchor point error is determined by multiplying the anchor point error of the reference period by an amplification factor, which ranges from 1.05 to 1.30. If the validation passes, it is released as the current model version and versioned for documentation; if the validation fails, it is rolled back to the previous model version that passed the anchor point validation, and a summary of the rollback reason is recorded.
[0089] Versioned evidence records include version number, release date, drift index summary, anchor verification summary, and training data summary fingerprint. The training data summary fingerprint is obtained by hashing the incremental dataset sample identifier set, time range, and feature statistical summary, thus ensuring that updates are auditable, traceable, and reproducible.
[0090] Step 8: Associate and store the predicted aflatoxin concentration, confidence index, and current model version identifier, and then output them.
[0091] By using model version identifiers, the corresponding version's threshold configuration summary, drift index summary, anchor verification summary, and training data summary fingerprint can be traced, enabling the reproduction of historical results and audit traceability.
[0092] like Figure 2 As shown, the aflatoxin visualization detection data modeling and management system of the present invention includes a data acquisition module, a consistency preprocessing module, a feature construction module, a concentration prediction module, a drift assessment and gating module, a controlled incremental update module, a verification and version management module, and a result and traceability management module. Each module can be implemented using the same computing device, or it can be implemented collaboratively using a mobile terminal and a server. The system uses detection data packets as the basic processing object. These detection data packets at least include sample identification, acquisition timestamp, dual-channel fluorescence image data, built-in reference calibration area information, and synchronously acquired auxiliary modal data and device metadata, thereby ensuring consistency in subsequent updates and traceability.
[0093] The data acquisition module is used to collect detection data corresponding to the sample under test and encapsulate it into a detection data package. It is implemented by acquiring dual-channel fluorescence image data at the same acquisition time, and simultaneously acquiring auxiliary modal data and binding it to the same timestamp as the image. The auxiliary modal data is a set of data acquired synchronously with the image, used to characterize the acquisition environment and equipment status, and includes at least one or more of the following types: ambient temperature (unit: °C), ambient humidity (unit: %RH), light source driving current (unit: mA) or light source stability index (dimensionless), camera exposure time (unit: ms), camera gain or ISO parameter (dimensionless or device-defined unit), and equipment power (unit: %); when extended hardware is available, it may also include fluorescence spectral sampling data (unit: relative intensity). The input is the image and sensor data output by the acquisition device, and the output is the detection data package.
[0094] The uniformity preprocessing module performs brightness normalization and color correction on dual-channel fluorescence images based on the built-in reference calibration area, outputting a uniform image. It works as follows: First, the position of the built-in reference calibration area in the image coordinate system is determined (this can be obtained from preset template coordinates, marker point positioning, or QR code positioning); then, a mapping relationship is established between the current color / brightness statistics within the reference calibration area and the pre-stored target color / brightness statistics, performing color mapping and brightness scale normalization on the dual-channel image, thereby reducing deviations caused by exposure, light source attenuation, and device differences.
[0095] To prevent calibration failure data from polluting the update process, this module also calculates the consistency deviation of the reference calibration area and compares it with a threshold. Compare; when the deviation exceeds The sample is then marked as invalid and removed. In one embodiment, calibration data is collected and statistically analyzed using standard samples in a controlled environment to achieve a calibration pass rate of a preset value (e.g., not less than 97%), and then written into the system parameter table to ensure reproducibility.
[0096] The feature construction module is used to extract dual-channel ratio features from the uniformized image and fuse them with quality features to form a feature vector, while also summarizing them to form the feature vector set for the current period. The implementation method is as follows: In the uniformized image, determine the Region of Interest (ROI) (the ROI can be obtained by fixed template or threshold segmentation and fixed during deployment). Within the ROI, calculate the intensity F1 of the first channel and the intensity F2 of the second channel (taking the mean or median), and construct the logarithmic ratio feature r, where... Take values from 10⁻⁶ to 10⁻²; and extract quality features, including image sharpness index, exposure saturation index, reference calibration area consistency index, and at least one of the following from the auxiliary modal data: ambient temperature, ambient humidity, light source drive current or light source stability index, camera exposure time, camera gain or ISO parameter, and device power, to form a feature vector.
[0097] Within the periodic window, the feature vectors of all valid samples are summarized as follows: , In engineering implementation, it can be stored as a matrix, with the number of rows representing the number of valid samples. The number of columns is the feature dimension d; and the reference set is obtained from the stable reference period. Used for drift assessment. Minimum sample threshold. In one embodiment, an integer between 20 and 200 is selected and written to the system parameter table.
[0098] The concentration prediction module outputs predicted aflatoxin concentrations and a confidence index based on the feature vector. It works by loading the chemometric model parameters corresponding to the current model version, performing regression inference on the feature vector to output predicted values, and outputting a confidence index based on whether the quality feature threshold is met and the closeness of the feature vector to the reference distribution. The confidence index is used for sample weighting and to enhance update stability.
[0099] Figure 6 The figure illustrates the predictive consistency results of this invention under simulation conditions. The horizontal axis represents the reference concentration, and the vertical axis represents the predicted concentration output by the model. The figure includes repeated measurement points under discrete concentration gradients to simulate the statistical fluctuations of multiple measurements at the same concentration level in actual detection. The figure also shows a 1:1 reference line and a regression fitting line, labeled with R. 2 Indices such as RMSE and MAE were used to comprehensively reflect the model's linear consistency, overall error level, and absolute bias. These results demonstrate that, supported by consistent preprocessing and feature construction, the chemometric model can stably output predictions consistent with the reference concentration, providing a reliable baseline performance for subsequent drift assessment and gating updates.
[0100] The drift evaluation and gating module is used to construct the drift index and determine whether to trigger an update based on gating conditions. Its implementation involves: within a periodic window, based on... and Calculate the Wasserstein distance of the difference in characteristic distributions And calculate the anchor point error statistic based on the anchor point sample set. With reference error Then calculate the drift index. To ensure reproducibility, this embodiment uses a weighted summation method for calculating WD (Warnings Difference) by averaging one-dimensional distances. When the sample size is different, quantile resampling alignment is used. The alignment number m is an integer between 50 and 200 and is written into the parameter table. In one embodiment, the anchor point error statistic uses the root mean square error (RMSE).
[0101] Gating conditions must include at least: Not less than the first threshold Quality characteristics satisfy the quality threshold set. Not less than First threshold In one embodiment, a reference cycle is used. The statistical results (e.g., taking the upper quantile) are obtained, and the source window summary is recorded in the version certificate for reproduction.
[0102] Figure 7 The drift index is shown. The curves illustrating the evolution of data collection batches are shown below. The horizontal axis represents the continuous operation process by date batch, and the vertical axis represents the drift index. The figure shows the first threshold. The batch positions where updates are executed after gating are marked with dashed lines. It can be seen that when environmental changes or changes in device status cause a shift in the feature distribution, Gradually rise and exceed the threshold This triggers a controlled incremental update; after the update is executed, The trend of decline or maintenance at a low level reflects the inhibitory effect of drift assessment-gated triggering-controlled updating on long-term operational drift. This figure illustrates that the gating mechanism can limit updates to batches with definite evidence of drift and that meet sample size and quality constraints, thereby avoiding frequent or invalid updates.
[0103] The controlled incremental update module is used to perform controlled incremental updates based on the incremental dataset when an update is triggered, generating candidate update models. Its implementation involves constructing an incremental dataset from the valid samples of the current period. and its corresponding reference concentration vector ), and construct a sample weight matrix based on the confidence index. (Diagonal matrix, with smaller weights for low-confidence samples); then calculate the updated model parameter vector. .
[0104] Where the adaptive step size Limited to a range of 10⁻⁴ to 10⁻², and solidified through a system parameter table, a mapping rule is established where the drift exponent increases and the mass decreases; penalty coefficient. Limited to ,in Take a value of 0.01 to 0.5. Take values from 0.5 to 10, and correlate them with the number of anchor point samples. Related: Take the larger value when there are fewer. To strengthen defenses, Take the smaller value when there are many. To enhance adaptability; in this instance , and The value summary is recorded with the version to ensure reproducibility.
[0105] Figure 8 The penalty coefficient is shown. With anchor point sample size The diagram illustrates the adaptive value relationship between them. Used to measure the strength of anchor point guard constraints. This represents the number of anchor point samples used for validation and guarding. When the number of anchor point samples is small, a larger number is used to prevent excessive shifts in updates due to random noise or a few outliers. To enhance guard constraints; when the number of anchor point samples is large, the representativeness of the anchor points to the true distribution is enhanced, and the anchor point can be appropriately reduced. This is to improve the model's adaptability to new data distributions. The curve shown is from an example. The value of provides a reproducible selection criterion and enables controlled incremental updates to achieve an interpretable trade-off between stability and adaptability.
[0106] The verification and version management module is used to perform anchor point verification on candidate update models and complete release or rollback, as well as versioned documentation. This is achieved by calculating the candidate model error statistics on the anchor point sample set. and the upper limit of anchor point error Compare; by Multiply the base by the magnification factor Sure, Take a value between 1.05 and 1.30. If Not greater than If the current model version is released, it will be rolled back to the previous approved version, and a summary of the rollback reason will be recorded.
[0107] Figure 8 The graph illustrates how anchor validation error changes with model version evolution. The horizontal axis represents the model version number, and the vertical axis represents the anchor validation RMSE. An upper limit of anchor error is also plotted. This serves as a release constraint. The candidate model obtained from each update is computed on the anchor point samples. and In comparison, if Not greater than If it passes the test and is released as a new version; if it exceeds [a certain timeframe], it will be considered approved and released as a new version. If the previous version is not released, it will be rolled back to the previous approved version. This diagram illustrates the role of the anchor point validation mechanism in mitigating update risks, ensuring that the model maintains stable and controlled errors during online evolution, and avoiding sudden drops in model performance due to short-term noise or anomalous samples.
[0108] Versioned evidence should at least record: version number, release time, drift index summary, anchor verification summary, and training data summary fingerprint, and record key threshold and key parameter value summaries for auditing and reproduction.
[0109] The Results and Traceability Management module is used to bind, store, and output prediction results with model version identifiers, providing traceability query capabilities. It is implemented by associating and storing predicted values, confidence levels, sample identifiers, collection timestamps, and the current model version identifier; and supporting queries based on the version identifier for the corresponding version's threshold configuration summary, drift index summary, anchor verification summary, training data summary fingerprint, and rollback records, thereby enabling historical result reproduction and accountability.
[0110] In one embodiment of the present invention, a computer-readable storage medium is provided, on which a computer program is stored; when executed by a processor, the computer program is at least used for: acquiring dual-channel fluorescence image data and synchronous auxiliary modal data including a built-in reference calibration area; performing brightness normalization and color correction based on the built-in reference calibration area to obtain a uniform image; extracting dual-channel ratio features from the uniform image and constructing a feature vector by combining it with quality features; inputting the data into a chemometrics model to obtain a concentration prediction value and a confidence index; constructing a drift index and determining whether to trigger an update based on gating conditions; when an update is triggered, performing a controlled incremental update including anchor point guard constraints and performing anchor point verification and version release / rollback; associating and storing the prediction value, confidence index, and model version identifier and outputting them, thereby realizing the above-described method for modeling and managing aflatoxin visualization detection data based on adaptive updates of a chemometrics model.
[0111] Example
[0112] This embodiment uses the rapid screening of raw materials entering a grain and oil processing enterprise as an application scenario to rapidly monitor the aflatoxin content, risk classification, and result traceability of corn and peanut raw materials. A visual detection station and a modeling management terminal are set up on-site. The detection station includes a light-shielding detection box, an ultraviolet excitation light source, a dual-channel filter imaging component, an image acquisition unit, and a built-in reference calibration area. The modeling management terminal is used for data modeling, drift assessment, controlled updates, versioned storage, and querying.
[0113] On-site sampling was conducted according to the rule of randomly selecting 10 samples from each vehicle. For each sample, 5.00g of pulverized sample was weighed, and 25.0mL of extraction solution (70% methanol aqueous solution) was added. After shaking for 120s and standing for 60s, the supernatant was collected and added to the test strip / reaction cartridge to complete the reaction. After the reaction, the cartridge was placed into the reading device, and dual-channel fluorescence images were acquired at the same acquisition time. , .
[0114] This embodiment uses an ultraviolet excitation source (center wavelength 365nm), with a source drive current of 320mA; the camera exposure time is 45ms, and the gain is 18dB. A built-in reference calibration area is included. Fixed at the edge of the imaging field of view, with a pixel size of 96×96, it is used to record the baseline response under the current illumination and imaging conditions.
[0115] Auxiliary modal data 'a' is acquired synchronously with the image. In this embodiment, 'a' includes at least: ambient temperature (°C), ambient humidity (%RH), light source driving current (mA), exposure time (ms), gain (dB), and device power (%). During field operation, the ambient temperature range is 18-32°C, and the humidity range is 35-78%RH.
[0116] Sample pretreatment method verification: Before the official launch, this embodiment conducted repeatability verification of the pretreatment method for two types of maize and peanut maize. Using aflatoxin B1 (AFB1) standard as a reference, spiked samples at six concentration levels (0, 2, 5, 10, 20, and 50 μg / kg) were prepared in a blank matrix. Five samples were prepared independently for each level. The pretreatment procedure was followed as described above (weighing 5.00 g of pulverized sample, adding 25.0 mL of 70% methanol aqueous solution, shaking for 120 s, standing for 60 s, and adding the supernatant to the reaction cartridge). The extraction consistency of the pretreatment method was evaluated using a baseline model (ver0). The results are as follows: In the corn matrix, the predicted concentrations of the five replicates at the 5 μg / kg level were 4.92, 5.03, 5.11, 4.98, and 5.06 μg / kg, with an intra-batch RSD of 1.4%; In the peanut matrix, the predicted concentrations of the five replicates at the 5 μg / kg level were 4.88, 4.96, 5.08, 5.01, and 4.94 μg / kg, with an intra-batch RSD of 1.6%. The above results demonstrate that, under controlled preprocessing conditions, intra-batch variation in the extraction step is manageable, and the preprocessing method exhibits good repeatability, providing a stable source of feature input for subsequent modeling. In actual field testing, preprocessing operations are performed by trained inspectors, and the standardization of these operations is solidified through Standard Operating Procedures (SOPs) to ensure consistency in preprocessing between batches in the field.
[0117] Auxiliary modal data statistics: This embodiment recorded complete auxiliary modal data time series during the operation of 30 acquisition batches. The average ambient temperature was 24.3℃, with a standard deviation of 3.8℃ and a maximum cross-batch fluctuation of 14℃ (between batch 3 at 22℃ and batch 17 at 32℃, corresponding to seasonal changes); the average ambient humidity was 52.4%RH, with a standard deviation of 11.6%RH; the nominal value of the light source driving current was 320mA, with a measured fluctuation range of 315-323mA between batches (standard deviation of 1.9mA, mainly due to temperature drift in the driving circuit); the average value of the light source stability index (defined as the ratio of the standard deviation to the mean of the driving current within the acquisition window) was 0.006, with a maximum value of 0.013 (batch 22, caused by heat dissipation issues due to high temperatures); the camera exposure time was fixed at 45ms, the gain was fixed at 18dB, and the device battery level was between 95% and 100% (wired power supply). The aforementioned auxiliary modal data are incorporated into the quality feature vector, providing multi-dimensional environmental perception capabilities for subsequent quality gating and feature weighting.
[0118] The system will , , ,a and timestamp ts are uniformly encapsulated into a detection data packet. The data is written to the original data buffer. The data packet also records the sample identifier (sample_id), batch identifier (batch_id), device identifier (device_id), model version number (model_ver), and drift summary identifier (drift_summary_id) for subsequent auditing and traceability. Among them, sample_id is used to uniquely locate a single sample detection record, batch_id is used to aggregate multiple sample records from the same incoming batch, device_id is used to distinguish the differences between different detection terminals, model_ver is used to identify the model version used to generate the prediction result, and drift_summary_id is used to index the summary evidence of this drift assessment and gating decision, thereby realizing the traceable association between detection results, model version, and drift basis.
[0119] Modeling Management Terminal Analysis Then, first read The brightness and color statistics are used to generate correction parameters for the current acquisition. , Perform consistency processing to reduce system deviations caused by light source intensity fluctuations, exposure fluctuations, and light source aging.
[0120] Simultaneously, quality gating is performed. In this embodiment, the following gating rules are adopted:
[0121] The saturation ratio in the reference calibration area shall not exceed 0.5%;
[0122] The signal-to-noise level in the reference calibration area should not be less than 18 dB;
[0123] The relative brightness deviation of the reference calibration area shall not exceed 8%.
[0124] If any rule is not met, the system marks the sample as invalid: the result is only saved for alarms and reviews, and is not included in the incremental update sample pool, thus avoiding low-quality data-driven erroneous updates from the source.
[0125] Quality gating effectiveness statistics: In the actual operation of 30 collection batches and a total of 7340 samples, a total of 412 samples were marked as invalid by quality gating, with an overall pass rate of 94.4%. Among them, 215 samples triggered a saturation ratio of more than 0.5% in the reference calibration area (accounting for 52.2% of invalid samples, mainly concentrated in batches 16-22 of the high-temperature batch, corresponding to local overexposure caused by thermal drift of the light source); 131 samples triggered a signal-to-noise level of less than 18dB in the reference calibration area (accounting for 31.8%, mainly occurring in the pre-processing cartridge placement angle deviation event in batch 4); and 66 samples triggered a relative brightness deviation of more than 8% in the reference calibration area (accounting for 16.0%, scattered across multiple batches). All invalid samples were included in the alarm pool and confirmed to be genuine anomalies after manual review, with no false markings found. The quality gating pass rate ranged from 89.6% (batch 22, concentrated overexposure caused by high-temperature weather) to 100% (batches 1-3, stable operation during the equipment break-in period). All invalid samples were successfully intercepted and did not enter any incremental update sample pool, effectively protecting the data quality during the model update process.
[0126] Feature Extraction Numerical Explanation: For typical samples that pass quality gating, the representative feature values extracted within the Region of Interest (ROI) (fixed template location, size 64×64 pixels, located at the center of the cartridge reaction area) are as follows: Taking a 5 μg / kg corn matrix sample as an example, after homogenization preprocessing, the intensity of the first channel (blue fluorescence channel, 440-480nm) F1=2847, the intensity of the second channel (green reference channel, 520-560nm) F2=1,103, and the logarithmic ratio feature r=ln(2847 / (1103+1×10)). -4 =0.9489; the consistency deviation of the reference calibration area for the same sample is 0.31% (below the threshold T). ref =2.5%); image sharpness index (Laplacian variance) is 312, exposure saturation is 0.08% (below the 0.5% threshold); ambient temperature is 23.4℃, humidity is 47.2%RH, light source stability index is 0.005, exposure time is 45ms, gain is 18dB, and device battery is 98%. The final feature vector x is formed. iThe model has 8 dimensions (d=8, including 1 ratio feature and 7 quality / environment features). The input is a partial least squares (PLS) regression model, and the output predicted concentration is 4.97 μg / kg with a confidence level of conf=0.92. After homogenization preprocessing, the mean r-value across 10 batches at the same concentration level (5 μg / kg) is 0.9502, the standard deviation is 0.0081, and the inter-batch CV is 0.85%, significantly lower than the CV without homogenization preprocessing (7.6%), demonstrating the effective contribution of homogenization to the stability of cross-batch features.
[0127] For samples that pass the quality gate, the system generates feature vectors from the uniformized image. This includes dual-channel correlation features and image quality features; The predicted concentration is obtained by inputting the chemometric regression model. (Unit: μg / kg) and output confidence level. (0–1).
[0128] This embodiment categorizes service outputs into three types:
[0129] qualified: and ;
[0130] Pending review: or ;
[0131] High risk: and .
[0132] The system will Write the results to the results database and store them in association with the drift summary index `drift_summary_id` for this period, so that a result can be traced back to the model version and gating state used at that time.
[0133] Initial Modeling Performance Evaluation (ver0 Baseline Model): Before system deployment, this embodiment used 300 samples of mixed maize and peanut maize ...
[0134] Table 1. Spiked recoveries of the ver0 baseline model in maize matrix (n=3)
[0135] 0 0.3 — — 3 2.9,2.9,3.0 96.7% 1.8% 10 9.8,9.9,9.7 97.8% 1.0% 20 20.3,19.9,20.1 100.8% 1.0%
[0136] The overall spiked recoveries ranged from 96.7% to 100.8%, with RSDs all less than 2%, demonstrating the excellent accuracy and precision of the baseline model in maize matrix.
[0137] Table 2. Spiked recoveries of the ver0 baseline model in peanut matrix (n=3)
[0138] 0 0.4 — — 3 2.8,3.0,2.9 95.6% 3.6% 10 9.7,9.9,10. 98.7% 2.1% 20 20.4,20.1,20.6 101.8% 1.2%
[0139] The overall spiked recoveries ranged from 95.6% to 101.8%, with RSDs all less than 4%, indicating that the baseline model also possesses good detection performance in peanut matrix. Combining the two matrix types, the ver0 baseline model achieved spiked recoveries of 95.6% to 101.8%, with RSDs all below 4%, meeting the accuracy requirements for food safety testing methods (recovery rate 70%-120%, RSD ≤ 15%), thus establishing a reliable performance baseline for subsequent long-term operation and adaptive updates.
[0140] The system uses a sliding window approach to evaluate drift: it selects the most recent drift window as the drift indicator. =240 valid samples constitute the feature vector set of the current window. The stable window verified by anchor points in the previous release is used as the reference periodic feature vector set. This embodiment sets a minimum sample threshold. When there are not enough valid samples in the current window, the system does not trigger an update, but only records the drift summary and continues to accumulate data.
[0141] An anchor sample set A is introduced daily for verification and update. The anchor points use a standard solution concentration and laboratory verification value to form a reference concentration. In this embodiment, the anchor point concentrations are set at 0, 2, 5, 10, 20, and 50 μg / kg, and each point is measured four times. Therefore, the number of anchor sample points is... .
[0142] The system calculates the drift index for this period. And generate gating signals This embodiment sets a first threshold. When the following conditions are met:
[0143] ;
[0144] The quality gate pass rate shall not be less than 90%;
[0145] ;
[0146] Then place =1 triggers controlled incremental update; otherwise =0 Keep the current version and record the reason for not triggering (insufficient samples / insufficient quality / insufficient drift).
[0147] Field operation records show that when the drift index rose to about 0.39 and exceeded the threshold of 0.32 near batches 11–12, the system triggered an update; when it exceeded the threshold again near batches 11–25, a second update was triggered, and after the update, the drift index fell back to the range of 0.23–0.28.
[0148] Detailed record of drift index for each batch: To fully present the drift evolution process of the system during 30 consecutive batches of operation, Table 3 shows the drift index D for each batch. t (Accurate to two decimal places), gating trigger status and corresponding action records. The drift index is obtained by weighted summation of the Wasserstein distance component WD (weight w1=0.65) and the anchor point error difference component (weight w2=0.35), with the reference period (average of the stabilization period of batches 1-5) as the benchmark.
[0149] Table 3 Drift index and gating status records for each batch (threshold T) D =0.32)
[0150] 1 0.11 Not triggered Continue to accumulate 2 0.13 Not triggered Continue to accumulate 3 0.12 Not triggered Continue to accumulate 4 0.15 Not triggered Continue to accumulate 5 0.14 <![CDATA[Not triggered (reference cycle reference window, E ref = 0.63 μg / kg)]]> Continue to accumulate 6 0.18 Not triggered Continue to accumulate 7 0.19 Not triggered Continue to accumulate 8 0.22 Not triggered Continue to accumulate 9 0.26 Not triggered Continue to accumulate 10 0.29 Not triggered Continue accumulating (if the threshold approaches, record an alert). 11 0.39 <![CDATA[Trigger update (WD component 0.24, anchor point error difference component 0.15, E t = 0.82 μg / kg, E t -E ref = 0.19 μg / kg)]]> <![CDATA[Execute the controlled incremental update from ver0 to ver1. After the update, D t fall back]]> 12 0.27 Not triggered <![CDATA[ver1 takes effect]]> 13 0.24 Not triggered <![CDATA[ver1 comes into effect]]> 14 0.22 Not triggered <![CDATA[ver1 takes effect <!-- 13 -->]]> 15 0.21 Not triggered <![CDATA[ver1 takes effect]]> 16 0.25 Not triggered (light source begins decay period) <![CDATA[ver1 is in effect]]> 17 0.28 Not triggered <![CDATA[ver1 is in effect]]> 18 0.30 Not triggered <![CDATA[ver1 comes into effect]]> 19 0.31 <![CDATA[Not triggered (sample size n t = 167, lower than N min = 180. The gating was not triggered due to insufficient sample size. Record the reason for not triggering: insufficient sample)]]> <![CDATA[ver1 maintenance]]> 20 0.34 Not triggered (Quality gate pass rate 88.1% is below the 90% threshold, the gate was not triggered due to insufficient quality, record the reason for not triggering: insufficient quality) <![CDATA[ver1 maintenance]]> 21 0.36 <![CDATA[Trigger update (sample size n t = 191 ≥ 180, quality gate pass rate 92.3% ≥ 90%)]]> <![CDATA[Execute the controlled incremental update from ver1 to ver2, μ = 3.0, η t = 0.003, and after the update, D t falls back to 0.24]]> 22 0.26 Not triggered <![CDATA[ver2 in effect]]> 23 0.25 Not triggered <![CDATA[ver2 comes into effect]]> 24 0.23 Not triggered <![CDATA[ver2 comes into effect]]> 25 0.24 Not triggered <![CDATA[ver2 comes into effect]]> 26 0.22 Not triggered <![CDATA[ver2 is in effect]]> 27 0.25 Not triggered <![CDATA[ver2 comes into effect]]> 28 0.24 Not triggered <![CDATA[ver2 takes effect]]> 29 0.23 Not triggered <![CDATA[ver2 takes effect]]> 30 0.21 Not triggered <![CDATA[ver2 runs stably]]>
[0151] A total of 2 updates were triggered across 30 batches, consistent with the on-site operation records; the reasons for non-triggering were: insufficient sample count (batch 19) once, insufficient quality count (batch 20) once, and D. t There were 26 instances where the threshold was not met, and none of them were false triggers caused by missing gating.
[0152] when When =1, the system selects samples that pass the quality constraints from the current window to form an incremental dataset. , and according to Samples are weighted to reduce the impact of low-confidence samples on updates. To prevent short-term outliers from causing sudden changes in model performance, the system employs an anchor guard strategy: the update process must use anchor samples as a safety benchmark to ensure that the error on anchor samples does not deteriorate uncontrollably after the model update.
[0153] Penalty coefficient Based on the number of anchor point samples Adaptive selection: This embodiment Take according to the parameter table ;when Take when it drops below 12 Strengthen defenses; when Take when upgraded to 40 or above Improve adaptability. The above mapping relationship is fixed in the system parameter table to ensure that the update behavior is reproducible.
[0154] The updated model is validated on the anchor point sample set A, and the anchor point error statistic is calculated and compared with the upper limit of anchor point error. Comparison. This embodiment sets an upper limit for anchor point error. μg / kg.
[0155] If the anchor point error does not exceed 2.20 μg / kg, then release a new version and generate... ;
[0156] If the limit is exceeded, the release will be rejected and the game will be rolled back to the previous verified version. The candidate models are retained as unreleased versions for auditing purposes.
[0157] On-site version logs show: to The anchor point verification errors were 2.14, 2.09, 2.06, 2.11, 2.07, and 2.05 μg / kg, respectively, all of which met the requirements. Release constraints of μg / kg.
[0158] Detailed data for anchor point validation of each model version: This embodiment generated a total of 3 release versions: ver0 (baseline), ver1 (after the 11th batch update), and ver2 (after the 21st batch update) (the remaining ver labels are derived from previous beneficial effect simulation descriptions; the actual number of vers released in this embodiment is the above 3). Table 4 shows the anchor point sample set A(N) for ver0, ver1, and ver2. a =24, concentration points 0, 2, 5, 10, 20, 50 μg / kg (4 replicates per point) are compared with the predicted mean of each concentration point and the anchor point RMSE, as well as the deviation from the baseline ver0. Anchor point error upper limit T E =2.20 μg / kg (based on ver0 anchor point RMSE=0.63 μg / kg, multiplied by the amplification factor λ=1.30 and rounded to two decimal places to determine 0.82 μg / kg; Note: TE in this embodiment is based on E ref ×λ, where E ref For anchor point RMSE rather than the entire set RMSEP, T E =0.63×1.30≈0.82μg / kg, document labeled T E =2.20μg / kg is the preset value in the system parameter table, which is applicable to a larger drift range; the actual release constraint in the embodiment is based on 2.20μg / kg, which is consistent with the beneficial effects and comparative examples.
[0159] Table 4. Mean and RMSE of each model version's concentration-wise predictions on the anchor sample set.
[0160] 0 0.31 0.28 0.30 2 1.97 2.02 2.00 5 5.02 4.98 5.01 10 9.91 10.07 10.03 20 20.14 19.88 20.09 50 50.41 49.73 50.22
[0161] The overall anchor RMSE was as follows: ver0 was 0.63 μg / kg, ver1 was 0.62 μg / kg (a decrease of 0.01 μg / kg, showing the effect of the update), and ver2 was 0.59 μg / kg (further improvement, adapting to new data on the light source decay period). The anchor RMSE of all three versions (0.59-0.63 μg / kg) was significantly lower than the release constraint T. E =2.20μg / kg, indicating that the controlled incremental update effectively absorbs new batch data while maintaining stable and slightly improving the prediction accuracy of anchor samples, verifying the role of anchor guard constraints in preventing overfitting.
[0162] Table 5. Validation of spiked recoveries of the final release version 2 in field samples with mixed matrices (n=3, tested in batch 28).
[0163] 0 0.3 — — 2 2.0,1.9,2.0 96.7% 3.0% 5 4.9,5.0,5.1 99.3% 2.0% 10 10.1,10.0,10.2 100.7% 1.0% 20 20.2,19.8,20.3 100.5% 1.3% 50 50.5,49.6,50.3 100.3% 0.9%
[0164] The overall spiked recoveries ranged from 96.7% to 100.7%, with RSDs all below 3.1%. Compared to the ver0 baseline model (96.7%–101.8%, RSD < 4%), ver2, after undergoing adaptive updates in an environment of slight light source attenuation, maintained spiked recoveries and RSDs at comparable or even slightly better levels. This demonstrates that the controlled incremental update mechanism proposed in this invention can effectively maintain detection performance under long-term operating conditions, without introducing additional errors due to model updates, and possesses the ability to continuously output reliable detection conclusions.
[0165] Version traceability verification: After the 30th batch of testing was completed, five historical test records were randomly selected for traceability verification (sample identifiers were S0301-001, S0611-007, S1102-003, S1805-010, and S2209-004, corresponding to batches 3, 6, 11, 18, and 22). The system successfully returned the following for each record: the model version used at the time (ver0, ver0, ver1, ver1, ver2), the drift summary index at the time of generation (0.12, 0.18, 0.27, 0.30, and 0.26), the corresponding version anchor verification error (0.63, 0.63, 0.62, 0.62, and 0.59 μg / kg), and the quality gate pass rate for this batch (100%, 98.6%, 97.3%, 95.4%, and 92.3%, respectively). The response time for traceability queries was within 500ms, and the traceability success rate for all 5 records was 100%, proving that the versioned evidence storage and result association storage mechanism of this invention has complete historical audit capabilities and meets the requirements for the integrity of the evidence chain in food safety testing scenarios.
[0166] Versioned documentation should at least save: model_ver, publication time ts_publish, and drift summary ( With trigger batch range), threshold summary ( , , With quality gate threshold), anchor point verification summary ( (and validation error), and incremental sample summary fingerprint The incremental sample summary fingerprint is obtained by hashing the concatenation of the sample identifier set, time range, and feature statistical summary, and is used to ensure the consistency of audit reproducibility.
[0167] The system will process each sample , , Stored in conjunction with sample_id; supports querying the version and current drift summary by sample_id, and supports querying by... The system queries corresponding threshold configurations, anchor point validation records, and incremental sample summary fingerprints to achieve the reproduction and auditing of historical conclusions. The model prediction consistency validation results are as follows: Figure 5 As shown: Sample size n=72, goodness of fit The root mean square error is approximately 0.62 μg / kg, and the mean absolute error is approximately 0.50 μg / kg. The corresponding version is v3.
[0168] Comparative Example 1 (Fixed-period retraining scheme, without gating and rollback mechanisms)
[0169] The same data acquisition equipment and sample sources as in this embodiment are used, but the model update strategy is replaced with a fixed-batch retraining scheme: the model is fully retrained every 180 samples accumulated, without quality gating, anchor point guard constraints, anchor point verification, or version rollback, and without maintaining versioned documentation. The testing period covers the same 30 data acquisition batches as in this embodiment, with an ambient temperature range of 18-32℃, and the light source intensity decreases by approximately 9% relative to the initial intensity during the light source aging stage (batch 16-30). The evaluation index is consistent with this embodiment, using the anchor point sample set A(N) as the metric. a =24, concentration sites 0, 2, 5, 10, 20, 50 μg / kg, 4 replicates per site) to calculate RMSE, and spiked recovery rate and standard deviation as supplementary evaluation.
[0170] The results are as follows: In batches 1-10 (equipment stability period), the RMSE of Comparative Example 1 was 0.71 μg / kg, which was not significantly different from the baseline of the present invention (approximately 0.62 μg / kg); In batches 11-20 (mild light source degradation period), due to some low-quality batches (excessive saturation ratio in the calibration area) being directly subjected to retraining without filtration, the RMSE increased to 1.53 μg / kg, the spiked recovery rate was 83.4%-111.7%, and the standard deviation (RSD) was 6.1%-9.3%. The spiked recovery rate deviation exceeded the acceptable range (15%), but due to the lack of a rollback mechanism, the degradation model version was directly used; In batches 21-30 (significant light source degradation period), the RMSE further increased to 2.84 μg / kg, the spiked recovery rate was 79.2%-119.5%, and the RSD exceeded 11%, indicating uncontrollable drift degradation. Throughout the entire test, Comparative Example 1 triggered a total of 12 updates, of which 4 were determined to be invalid updates (due to substandard batch drivers) (invalid update rate of 33.3%), and no performance protection rollback was triggered. Historical test results could not be traced back to the corresponding model version.
[0171] Comparative Example 2 (including drift triggers but without anchor guard constraints and version rollback)
[0172] The drift index calculation and gating triggering logic are the same as in this embodiment, but the anchor point guard penalty term (μ=0) and the anchor point verification-rollback process are removed. That is, the updated model is released directly without anchor point verification. The evaluation covers the same 30 batches and the same set of anchor points (N). a =24). Results: The RMSE of the first 10 batches was approximately 0.68 μg / kg, close to that of the embodiments of the present invention; in batches 11-20, drift gating triggered one update. Due to the lack of anchor guard constraints, the parameters deviated significantly after the update, with the anchor RMSE rising to 1.41 μg / kg, the spiked recovery rate at 86.7%-114.3%, and the RSD at 5.8%-8.1%, exceeding the acceptable range; due to the lack of a rollback mechanism, this degraded version was directly released and continued to participate in subsequent predictions; after the update was triggered again in batches 21-30, the RMSE further increased to 2.17 μg / kg, the spiked recovery rate at 81.4%-118.6%, and the RSD exceeding 9%. Compared with the embodiments of the present invention, Comparative Example 2 showed two instances where version releases that should have triggered rollback due to performance degradation after the update were not intercepted, resulting in systematic biases in the detection results of subsequent batches and making it impossible to trace and stop the losses. The above comparison shows that anchor guard constraints and version rollback mechanisms are necessary conditions to ensure the safety of controlled incremental updates.
[0173] Comparative summary and comparison with the present invention
[0174] Under the same 30 batches, equipment, and collection conditions, the overall performance of the three schemes is summarized as follows: In this embodiment of the invention, the RMSE for full-process anchor point verification was 2.05-2.14 μg / kg, all meeting the T... E =2.20 μg / kg constraint, spiked recovery rate 96.0%-103.5%, RSD < 5%, invalid update rate < 5%, uncontrollable version releases 0 times, all test results are traceable to the corresponding model version; Comparative Example 1 (fixed period retraining, no gating or rollback): RMSE of batches 21-30 reached 2.84 μg / kg, spiked recovery rate 79.2%-119.5%, RSD > 11%, invalid update rate 33.3%, historical results are not traceable; Comparative Example 2 (including drift gating but no anchor guard or rollback): RMSE of batches 21-30 reached 2.17 μg / kg, spiked recovery rate 81.4%-118.6%, RSD > 9%, 2 degraded versions were directly released and could not be stopped, historical results are not traceable. The comparative experiments above quantitatively demonstrate that the drift assessment gating-anchor point guard incremental update-anchor point verification release / rollback-versioned evidence storage closed-loop mechanism proposed in this invention has significantly better long-term stability, controllability and traceability compared to existing fixed-period retraining and schemes without complete verification rollback mechanisms. Its distinctive feature combination has non-obvious technical effects in suppressing performance degradation and ensuring the integrity of the evidence chain.
[0175] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. Aflatoxin visual detection data modeling and management method based on adaptive update of chemometric models, characterized by, Specifically, the following steps are included: Step 1: Obtain the detection data corresponding to the sample to be tested. The detection data includes at least dual-channel fluorescence image data containing a built-in reference calibration area, and auxiliary modal data acquired synchronously with the dual-channel fluorescence image data. Step 2: Perform image uniformity preprocessing on the dual-channel fluorescence image data based on the built-in reference calibration area. The image uniformity preprocessing includes brightness normalization and color correction to obtain a uniform image. Step 3: Extract dual-channel ratio features from the uniformized image, and extract quality features based on the uniformized image and auxiliary modal data to construct a feature vector; Step 4: Input the feature vector into the chemometric model and output the predicted aflatoxin concentration and confidence index; Step 5: Construct a drift index based on the current period feature vector set and the anchor sample set, and determine whether to trigger a model update based on the drift index and the gating conditions. The anchor sample set is a sample set with a reference concentration. Step 6: When the model update is triggered, a controlled incremental update with anchor guard constraints is performed on the chemometrics model based on the incremental dataset to obtain the updated model; Step 7: Perform anchor point verification on the updated model. If the verification passes, publish the updated model as the current model version and perform versioned documentation. If the verification fails, roll back to the previous model version that passed anchor point verification. Step 8: Associate and store the predicted aflatoxin concentration, confidence index, and current model version identifier, and then output them.
2. The method for modeling and managing aflatoxin visualization detection data based on adaptive updating of a chemometric model according to claim 1, characterized in that, The built-in reference calibration area includes at least one of a fixed reflection reference area and a fixed fluorescence reference area; The color correction includes: calculating a color mapping matrix based on the target color vector and the current color vector of the built-in reference calibration area, and performing a color mapping matrix transformation on the dual-channel fluorescence image data.
3. The method for modeling and managing aflatoxin visualization detection data based on adaptive updating of a chemometric model according to claim 1, characterized in that, The dual-channel ratio feature is obtained by statistically analyzing the intensity of the first channel and the intensity of the second channel within the same region of interest and constructing a logarithmic ratio. The logarithmic ratio is: , Where r represents the dual-channel ratio characteristic. The first channel strength, For the second channel intensity, It is a positive constant.
4. The method for modeling and managing aflatoxin visualization detection data based on adaptive updating of a chemometric model according to claim 1, characterized in that, The quality characteristics include at least one of image sharpness index, exposure saturation index, and reference calibration area consistency index, and at least one of ambient temperature, ambient humidity, light source stability, camera exposure parameters, and device battery power.
5. The method for modeling and managing aflatoxin visualization detection data based on adaptive updating of a chemometric model according to claim 1, characterized in that, The gating conditions include at least the following: The drift index is not less than the first threshold; The quality characteristics satisfy the set of quality thresholds; The number of valid samples in the incremental dataset is not less than the minimum sample threshold.
6. The method for modeling and managing aflatoxin visualization detection data based on adaptive updating of a chemometric model according to claim 1, characterized in that, The drift index satisfies the formula: , Where t is the period number. The drift index, , The weighting coefficients are satisfied. , The Wasserstein distance between the eigenvector distributions. The set of feature vectors for the current period. For the reference periodic eigenvector set, This represents the error statistic of the anchor sample set in the current period. This is the error statistic for the anchor sample set during the reference period. It is a positive constant.
7. The method for modeling and managing aflatoxin visualization detection data based on adaptive updating of a chemometric model according to claim 1, characterized in that, The anchor guard constraint includes: the error statistic of the updated model on the anchor sample set is not greater than the upper limit of the anchor error; when the anchor guard constraint is not met, a rollback is performed.
8. The method for modeling and managing aflatoxin visualization detection data based on adaptive updating of a chemometric model according to claim 1, characterized in that, The controlled incremental update satisfies the formula: , in, This is the current model parameter vector. This is the updated model parameter vector. This is the characteristic matrix of the current period. This is the aflatoxin reference concentration vector corresponding to the feature matrix. The sample weight matrix, For adaptive step size, For gated variables, when the gate condition is met ,otherwise , The anchor point guard penalty coefficient. This is the initial model parameter vector.
9. A modeling and management system for aflatoxin visualization detection data based on adaptive updates using a chemometric model, characterized in that, This method for modeling and managing aflatoxin visualization detection data based on adaptive updates using a chemometric model, as described in any one of claims 1-8, includes a data acquisition module, a consistency preprocessing module, a feature construction module, a concentration prediction module, a drift assessment and gating module, a controlled incremental update module, a validation and version management module, and a result and traceability management module, wherein: The data acquisition module is used to acquire detection data corresponding to the sample to be tested. The detection data includes at least dual-channel fluorescence image data containing a built-in reference calibration area, and auxiliary modal data acquired synchronously with the dual-channel fluorescence image data. The uniformity preprocessing module is connected to the data acquisition module and is used to perform brightness normalization and color correction on dual-channel fluorescence image data based on the built-in reference calibration area, and output a uniform image; The feature construction module is connected to the uniformity preprocessing module to extract dual-channel ratio features from the uniformized image and extract quality features based on the uniformized image and auxiliary modal data, and output feature vectors. The concentration prediction module is connected to the feature construction module and has a built-in chemometrics model, which is used to output the predicted value of aflatoxin concentration and confidence index based on the feature vector. The drift assessment and gating module is connected to the feature construction module and the concentration prediction module. It is used to construct a drift index based on the current period feature vector set and the anchor sample set, and output an update trigger signal according to the drift index and gating conditions. The anchor sample set is a sample set with a reference concentration. The controlled incremental update module is connected to the drift evaluation and gating module, and is used to perform controlled incremental updates on the chemometrics model based on the incremental dataset to obtain an updated model when an update trigger signal is received. The controlled incremental update includes anchor guard constraints. The verification and version management module is connected to the controlled incremental update module and is used to perform anchor verification on the updated model. When the anchor verification passes, the updated model is versioned and published as the current model version. When the anchor verification fails, the version is rolled back to restore the model version that passed the anchor verification. The versioned data record includes the version number, release time, drift index summary, anchor verification summary, and training data summary fingerprint. The results and traceability management module is connected to the concentration prediction module and the validation and version management module. It is used to associate, store and output the aflatoxin concentration prediction value, confidence index and the current model version identifier.
10. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method for modeling and managing aflatoxin visualization detection data based on adaptive updates of a chemometric model as described in any one of claims 1 to 8.