A method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials.

By constructing an intrinsic spectral fingerprint database and a graph neural network model, the influence of environmental disturbances is removed, which solves the problems of insufficient model generalization ability and high misjudgment rate in the risk monitoring of aflatoxin contamination in Chinese medicinal materials. It achieves dynamic monitoring with high applicability and low cost and supports full-cycle state analysis.

CN122409573APending Publication Date: 2026-07-17THE CHINESE MEDICINE RAW MATERIALO PROCESSING PLANT OF THE GUANGDONG TRADITIONAL CHINESE MEDICINE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE CHINESE MEDICINE RAW MATERIALO PROCESSING PLANT OF THE GUANGDONG TRADITIONAL CHINESE MEDICINE CORP
Filing Date
2026-06-18
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing methods for monitoring the risk of aflatoxin contamination in Chinese medicinal materials rely on threshold judgment models driven by a single perturbation factor. These models ignore the spectral response of the medicinal materials themselves and the ontological state characteristics of the toxin adsorption sites, resulting in insufficient generalization ability of the models. They are difficult to adapt to the dynamic changes in the storage of medicinal materials in multiple scenarios and cannot effectively isolate the influence of environmental disturbances, leading to a high misjudgment rate and high deployment costs.

Method used

By acquiring batch metadata and multimodal spectral data of Chinese medicinal materials, an intrinsic spectral fingerprint database is constructed. The influence of perturbation factors is calculated in real time, a perturbation feature mask matrix is ​​generated, spectral signal interference is removed, a graph neural network model is used to learn the toxin adsorption mode, and a heat map of the confidence level of pollution level and the contribution of key discrimination bands is generated to form a dynamic pollution judgment report.

Benefits of technology

It significantly improves the intrinsic characterization capability of spectral data, increases the accuracy of toxin identification, enhances the interpretability of judgment results, achieves high applicability and low cost monitoring in complex storage environments, supports full-cycle state evolution analysis, and has good engineering applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122409573A_ABST
    Figure CN122409573A_ABST
Patent Text Reader

Abstract

This invention relates to a method for monitoring and analyzing the risk of aflatoxin contamination in stored Chinese medicinal materials. The core scheme involves collecting batch-specific structured metadata of the medicinal materials, then using portable near-infrared-Raman dual-mode spectroscopy to obtain raw spectral data covering 400–2500 nm and 800–1800 nm. Through high-precision synchronization, environmental disturbance modeling, and multi-mode correction, the influence of environmental factors (such as water film, fluorescence background, and dust deposition) is eliminated. By integrating a medicinal material fingerprint database and microscopic CT structure, graph neural network analysis is used to remove perturbation spectral residuals, outputting the contamination level and interpretable discriminative bands, and generating a dynamic contamination assessment report. When contamination exceeds limits, warehouse early warning and storage location locking are implemented, with historical data providing real-time feedback to optimize the model. This invention improves the accuracy, traceability, and interpretability of contamination detection in stored medicinal materials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spectral detection and intelligent analysis of contamination risks of traditional Chinese medicine materials, and in particular to a method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicine materials. Background Technology

[0002] Existing methods for monitoring and analyzing the risk of aflatoxin contamination in traditional Chinese medicinal materials generally rely on threshold judgment models driven by a single perturbation factor and conventional spectral detection techniques. Mainstream approaches typically use sensor data on storage environment temperature and humidity, air cleanliness, or single-modal spectral analysis such as near-infrared or Raman spectroscopy, determining the contamination level of the medicinal materials through statistical thresholds or empirical formulas. The development trend of these approaches is to continuously increase the scale of environmental sensor deployment, expand the detection signal bands, and enhance the classification accuracy of the models; however, their core technical approach remains driven by environmental conditions, lacking in-depth characterization and dynamic modeling of the intrinsic response characteristics of the medicinal materials themselves.

[0003] Among existing publicly available technologies, representative achievements include rapid detection methods for medicinal toxins based on near-infrared or Raman spectroscopy, suitable for classifying contamination levels under single-batch, static conditions. These methods are limited by the type of hardware sensor and the data acquisition environment, and are primarily applicable to relatively stable or highly contaminated environments. They lack the dynamic response capability to varying storage conditions, different medicinal material matrices, and batch-to-batch microstructural differences. Furthermore, while some environment-driven multi-sensor fusion solutions can improve data dimensionality, they struggle to achieve accurate traceability of individual medicinal materials and ensure the interpretability and transparency of the contamination level determination process.

[0004] Current technologies suffer from the following main problems: First, they rely heavily on perturbation factors as the core basis for pollution level classification, neglecting the intrinsic state characteristics of the medicinal materials' spectral response and toxin adsorption sites. This results in insufficient model generalization ability and difficulty in adapting to the dynamic changes in medicinal material storage across various scenarios. Second, existing spectral detection schemes cannot effectively isolate the distortion effects of multi-source environmental disturbances such as temperature, humidity, water film thickness, light, and dust on the spectral signals of the medicinal material surface, leading to a high misjudgment rate in pollution risk assessment results. Third, the model's decision-making process lacks interpretability and cannot output decision elements related to batch traceability of medicinal materials, types of environmental interference, and toxin binding mechanisms, limiting the system's credibility and regulatory transparency. Fourth, some multimodal fusion solutions require a large number of additional hardware sensors, resulting in high implementation costs and significant practical deployment difficulties.

[0005] Therefore, existing technologies urgently need to break through the environmentally driven threshold judgment model and construct a dynamic pollution level determination mechanism based on the decoupling of the intrinsic spectral fingerprint of medicinal materials and environmental disturbances. This mechanism should not only achieve traceability of different medicinal materials and batches, but also possess the separability of environmental disturbances and the traceability of the decision-making process. Meeting the real-time and intelligent requirements of dynamic rotation storage scenarios for traditional Chinese medicine materials and promoting the development of pollution risk monitoring and analysis methods towards high interpretability, low cost, and broad adaptability are precisely the core technical challenges that the industry urgently needs to solve. Summary of the Invention

[0006] This application provides a method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicine materials, aiming to solve one of the problems or issues of the existing technology mentioned in the background.

[0007] This application provides a method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials, specifically including: S1: Obtain batch metadata and multimodal spectral raw data of Chinese medicinal materials to form an initial spectral dataset; S2: Based on the origin, harvesting season and primary processing method metadata contained in the initial spectral dataset, call the pre-constructed intrinsic spectral fingerprint database for matching and retrieval, and generate the standard intrinsic spectral reference vector of the corresponding medicinal material matrix. S3: Real-time acquisition and calculation of the nonlinear influence weight of the perturbation factor on the spectral signal, generating a perturbation feature mask matrix; S4: Apply the perturbation feature mask matrix to the initial spectral dataset to perform band-by-band stripping processing, and output the perturbation-free spectral residual sequence; S5: Based on the microscopic tissue texture topology obtained from previous micro-CT reconstruction, a graph neural network model is constructed. The perturbation-free spectral residual sequence is mapped as node attributes and input into the model. The response mode of adsorption sites of toxin molecules in the medicinal matrix is ​​learned through message passing mechanism, and a pollution level attribution confidence vector and a heat map of contribution of key discrimination bands are generated. S6: Determine the current aflatoxin contamination level based on the maximum probability value in the contamination level attribution confidence vector, and extract one or more key discriminant spectral intervals and their corresponding potential binding site inference texts by combining the key discriminant band contribution heatmap to form a dynamic contamination determination report; S7: If the current aflatoxin contamination level exceeds the preset safety threshold, a graded early warning command will be immediately triggered and the corresponding cargo location will be locked.

[0008] This application provides a method for monitoring and analyzing the risk of aflatoxin contamination during the storage of traditional Chinese medicinal materials, which has the following beneficial effects: (1) In response to the problem of spectral signal distortion and high misjudgment rate caused by environmental disturbance in traditional Chinese medicine storage quality detection, this solution constructs a "disturbance feature mask matrix" to model and remove common interferences such as temperature and humidity changes, light aging and dust deposition, which significantly improves the intrinsic characterization ability of spectral data. Compared with the conventional method of directly using the original spectrum for discrimination, this technology effectively overcomes the defects of poor model stability and weak generalization ability caused by the near-infrared and Raman signals being easily affected by water film absorption peak shift, fluorescence background rise and signal-to-noise ratio decrease in complex storage environment. The accuracy of toxin identification is significantly better than that without disturbance removal treatment. In particular, it has stronger resolution sensitivity in the medium and low concentration pollution range (5–20 μg / kg). It realizes the full utilization of existing environmental monitoring data to complete high robust spectral correction without adding additional sensors, which greatly improves the applicability and reliability of the system under dynamic storage conditions.

[0009] (2) By introducing a lightweight graph neural network model based on the microstructure and topology of medicinal materials and integrating dual-mode spectral node attributes for multi-channel joint learning, this scheme realizes a deep mechanistic analysis of aflatoxin adsorption behavior. It not only outputs the confidence level of pollution level, but also generates a heat map of key band contribution and inference of potential binding sites, which significantly enhances the interpretability of the judgment results. Compared with the traditional black box machine learning model that only provides classification labels, this technology links molecular interaction mechanism with spectral response mode for the first time, supporting the provision of scientific basis from the perspective of "why it is judged as high risk", which helps to guide the formulation of subsequent intervention measures. At the same time, the model adopts an edge computing architecture for deployment, and the time for a single analysis is controlled within 800 milliseconds, which meets the high-frequency detection requirements of Chinese medicine rotation storage, and does not rely on cloud computing power or complex parameter tuning process, which has good engineering implementation and promotion value.

[0010] (3) Furthermore, this solution establishes a closed-loop quality traceability system covering meta-information identification, standard calibration, historical response curve comparison and similarity ranking, realizing the leap from "single detection point" to "full-cycle state evolution analysis"; by establishing identity files for each batch of medicinal materials including origin, processing method, harvesting season and other dimensions, and combining the feedback of the main disturbance types and their impact intensity percentages extracted in this test, the system can dynamically evaluate the deterioration trend of medicinal materials under different storage conditions, providing data support for differentiated maintenance strategies; in addition, the three interpretable elements generated together constitute a transparent decision report, which can be used for regulatory audit traceability and can also assist professionals in understanding the model judgment logic, thereby enhancing user trust; the overall technical path is entirely based on the existing warehouse sensing infrastructure, without the need for additional image or gas detection equipment, and has the advantages of low cost, high compatibility and easy integration, and is particularly suitable for the intelligent upgrading and transformation of large-scale Chinese medicine warehouses.

[0011] The synergistic effect of the above-mentioned technologies has formed a new paradigm for the safety monitoring of traditional Chinese medicine storage, which integrates strong environmental adaptability, high discrimination accuracy, fast response speed and interpretable mechanism. It effectively solves the core pain points of existing detection methods, such as insufficient stability in real and complex environments, difficulty in interpreting results and high deployment costs, and provides reliable technical support for achieving accurate, efficient and reliable quality control of traditional Chinese medicine materials from storage to delivery. Attached Figure Description

[0012] Figure 1 This is the main flowchart of a method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials. Figure 2 This is a sub-flowchart of a method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicine materials; Figure 3 This is another sub-flowchart of a method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicine materials. Detailed Implementation

[0013] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0014] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0015] like Figure 1 As shown, this application provides a method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicine materials, specifically including: S1: Obtain batch metadata and multimodal spectral raw data of Chinese medicinal materials to form an initial spectral dataset.

[0016] Specifically, the original multimodal spectral data includes near-infrared reflectance spectral signals in the 400 to 2500 nanometer band and Raman scattering signals in the 800 to 1800 anti-centimeter range, and the initial spectral dataset is batch-identified.

[0017] S2: Based on the origin, harvesting season and primary processing method metadata contained in the initial spectral dataset, the pre-constructed intrinsic spectral fingerprint database is called for matching and retrieval to generate the standard intrinsic spectral reference vector of the corresponding medicinal material matrix, which serves as the benchmark reference object for subsequent perturbation decoupling.

[0018] S3: Real-time acquisition and calculation of the nonlinear influence weight of the perturbation factor on the spectral signal, generating a perturbation feature mask matrix.

[0019] Specifically, the disturbance factors include: temperature and humidity of the storage environment, light intensity, and air cleanliness parameters. Based on the real-time collected temperature and humidity, light intensity, and air cleanliness parameters of the storage environment, the physical effect simulation submodule is used to calculate the nonlinear influence weight of each disturbance factor on the spectral signal, generating a disturbance feature mask matrix characterizing the current environmental interference features.

[0020] S4: Apply the perturbation feature mask matrix to the initial spectral dataset to perform band-by-band stripping processing, outputting a deperturbed spectral residual sequence. The deperturbed spectral residual sequence only reflects the state of the medicinal material itself and can eliminate spectral distortions caused by changes in water film thickness, fluorescence background enhancement, and dust deposition.

[0021] S5: Based on the microstructure and topological structure of the medicinal material obtained by previous microscopic CT reconstruction, a graph neural network model is constructed. The perturbation-free spectral residual sequence is mapped as node attributes and input into the model. The adsorption site response mode of toxin molecules in the medicinal material matrix is ​​learned through message passing mechanism, and a pollution level attribution confidence vector and a heat map of contribution of key discrimination bands are generated.

[0022] S6: Determine the current aflatoxin contamination level based on the maximum probability value in the contamination level attribution confidence vector, and extract one or more key discriminant spectral intervals and their corresponding potential binding site inference texts by combining the key discriminant band contribution heatmap, forming a dynamic contamination determination report. The dynamic contamination determination report includes a level conclusion and interpretable elements.

[0023] S7: If the current aflatoxin contamination level exceeds the preset safety threshold, a graded early warning command will be immediately triggered and the corresponding cargo location will be locked.

[0024] S8: If the current aflatoxin contamination level does not exceed the preset safety threshold, the de-perturbation spectral residual sequence and perturbation factor are recorded in the historical contamination response curve database, and the standard intrinsic spectral reference vector is updated periodically.

[0025] Specifically, the periodically performed standard intrinsic spectral reference vector update includes: periodically comparing the similarity ranking of newly added samples in the historical pollution response curve database with the intrinsic spectral fingerprint database; and automatically updating the standard intrinsic spectral reference vector of a specific batch of medicinal materials if the spectral response mode of a particular batch of medicinal materials undergoes a significant drift.

[0026] Step S1: Obtain batch metadata and multimodal spectral raw data of Chinese medicinal materials to form an initial spectral dataset. Specifically, this includes: S1.1: The origin code, harvest season code, and initial processing method label of Chinese medicinal materials when they are put into storage are structured and parsed to generate a batch meta-information identifier with unique characteristics, which serves as the primary key index for subsequent multi-source data fusion.

[0027] The system performs structured parsing on the origin codes, harvest season codes, and initial processing method labels received when Chinese medicinal materials are put into storage. For the origin codes, segmented comparison and numerical conversion are performed according to a pre-defined geographic partitioning mapping table, outputting a regional category index value. The harvest season codes are grouped and parsed according to a seasonal cycle weight matrix, mapping seasonal indicator characters to cycle values ​​and performing cycle normalization to obtain a seasonal normalization coefficient. The initial processing method labels are classified and matched using a process flow classification dictionary, converting the processing method identifier into a process category code and generating a unique process attribute value. The regional category index value, seasonal normalization coefficient, and unique process attribute value are then concatenated according to primary key combination rules to form a composite batch attribute string to be encoded. A hash function is called on the composite batch attribute string to perform a hash operation, ensuring that the generated batch metadata identifier is unique and evenly distributed globally. The batch metadata identifier is stored in the batch index cache, providing a unique primary key index for subsequent multi-source data fusion module calls. By using hashing and primary key combination rules, the original origin, season, and processing method tags from the previous step are transformed into unique and easily retrieved batch metadata identifiers, enabling accurate indexing of batch data during multimodal spectral acquisition and fusion.

[0028] For example, when a batch of Chinese medicinal herbs enters the warehouse, the origin code is CN-HEB, the harvest season code is 2023-AUT, and the initial processing method label is DRY-SUN. The regional category index value corresponds to Hebei Province according to the mapping table, and the mapping value is 15; the harvest season code corresponds to autumn, and the cycle normalization coefficient is set to 3 / 4; the initial processing method label matches and identifies as "natural sun drying", and the process category code is set to 07. The multi-field concatenation rule is: origin index value + "-" + seasonal normalization coefficient text + "-" + process category code, and the concatenation result is "15-0.75-07". The MD5 hash function is called to calculate the 128-bit hash value of this string, and the batch meta-information identifier "4f7a2950b16dcb…" is obtained (example excerpt). The identifier is stored in the batch index cache. After retrieval testing, the indexing latency in a batch dataset of millions of data points is less than 2 milliseconds, and the global conflict probability of the batch metadata identifier is close to zero. This ensures that the scanning command of the corresponding batch can be accurately triggered and traceability binding can be achieved in the subsequent spectral acquisition stage.

[0029] S1.2: Based on the batch information identifier of the medicinal material, the synchronous acquisition command of the portable near-infrared-Raman dual-mode spectrometer is triggered to perform a non-contact scanning operation on the surface of the medicinal material, so as to obtain the near-infrared reflectance spectral signal sequence covering the 400 to 2,500 nanometer band and the Raman scattering signal sequence covering the 800 to 1,800 centimeter range, respectively.

[0030] Based on the generated batch information identifier of the medicinal material, this identifier is used as data input to trigger the acquisition control module of the portable near-infrared-Raman dual-mode spectrometer, realizing the parallel start signal generation of the dual-mode sensors. The acquisition control module reads the medicinal material category and batch characteristic parameters attached to the identifier, calls the device's preset band range configuration file, sets the sampling wavelength range of the near-infrared channel to 400 to 2500 nm, and sets the sampling wavenumber range of the Raman channel to 800 to 1800 inverse centimeters. The step position and spot coverage of the non-contact scanning are calculated through an optical scanning path planning algorithm to ensure that a uniform beam distribution is formed on the surface of the medicinal material during the scanning process. During the acquisition process, the near-infrared sensor records the reflected light intensity at each wavelength through a photodiode array, and the Raman sensor records the scattered spectral intensity at each wavenumber through a CCD detector, and caches them in real time to the dual-mode signal buffer. Band index binding processing is performed on the buffer data, and the wavelength or wavenumber of each sampling point is combined with the corresponding light intensity value to form the original signal pair sequence. The device's internal error detection module monitors the signal amplitude and noise levels of both channels. If these levels exceed a preset deviation threshold, automatic gain and integration time adjustments are immediately executed to ensure signal acquisition stability. The acquired near-infrared reflectance and Raman scattering spectral signal sequences are encapsulated into a dual-modal raw spectral data package, along with sampling parameters and batch identifier metadata, providing complete input for subsequent time-series alignment and traceability binding. Through the aforementioned acquisition control and optical signal processing methods, the batch-identified spectral scanning process is transformed into a high-confidence raw spectral signal sequence covering the specified wavelength band, achieving simultaneous acquisition of multimodal data and close correlation with medicinal material batch information.

[0031] For example, when performing near-infrared-Raman dual-mode acquisition on the Astragalus membranaceus sample with batch identifier "HY001-2024-08", the sampling wavelength range of the near-infrared channel was configured to be 400 to 2500 nm, with a resolution set to 2 nm. The sampling wavenumber range of the Raman channel was configured to be 800 to 1800 inverse centimeters, with a resolution set to 1 inverse centimeter. The spot step spacing of the scan path planning output was 1.5 mm, and the coverage area diameter was 50 mm. During the acquisition process, a significant moisture absorption peak was detected at 1450 nm in the near-infrared reflectance spectrum, and an enhanced protein-related peak was detected at 1650 inverse centimeters in the Raman scattering spectrum. The error detection module detected a high noise level in the Raman channel at the beginning of the acquisition and automatically adjusted the integration time from 0.5 seconds to 1 second, significantly improving the signal-to-noise ratio. At the end of the acquisition, the generated dual-mode raw spectral data package contained 1050 near-infrared sampling points and 1000 Raman sampling points. The signal amplitude was stable and the noise was low, providing high-quality input data for subsequent timestamp synchronization and perturbation stripping.

[0032] S1.3: The near-infrared reflectance spectral signal sequence and the Raman scattering signal sequence are time-aligned using a high-precision timestamp synchronization algorithm to eliminate the sampling delay difference between the two-mode sensors and generate a pair of multi-mode spectral raw data with strict spatiotemporal correspondence.

[0033] S1.4: Perform a deep binding mapping operation on the multimodal spectral raw data pairs and the medicinal material batch metadata identifier to construct a batch-identified spectral data unit containing complete traceability information, ensuring that each spectral data can be traced back to the specific medicinal material batch attribute.

[0034] The input objects are the original multimodal spectral data pairs after time-series alignment processing and the batch metadata identifiers of medicinal materials formed in the warehousing information parsing stage. Both of them have uniquely identifiable temporal and spatial correlation attributes.

[0035] A unique primary key binding is performed on the raw multimodal spectral data pairs, and the batch metadata identifier is written as the index key into the metadata structure of the data pair to ensure the consistency between the physical storage of the data and the logical identifier of the batch.

[0036] A bidirectional index mapping is implemented on the bound data structure, mapping the batch identifier to the spectral signal feature matrix, and at the same time mapping the spectral signal feature matrix back to the batch identifier, forming a bidirectional mapping table that can be quickly retrieved in any direction.

[0037] The integrity of the bidirectional mapping table is verified by using a hash verification function to calculate the summary value of the batch identifier and the spectral data matrix, and then checking against a preset verification list to detect whether the mapping data is missing or tampered with.

[0038] The traceability path is solidified for the mapping table that has passed the integrity verification. Batch information, spectral acquisition time, acquisition device number and operator identity information are embedded into the traceability field of the data unit to ensure that the data unit has a complete traceability chain throughout its life cycle.

[0039] Through deep binding mapping and bidirectional verification mechanisms, the results of the previous step are transformed into batch-identified spectral data units containing complete traceability information and capable of bidirectional retrieval, thereby meeting the input requirements for subsequent batch data aggregation and standardized encapsulation, and ensuring high reliability of system traceability.

[0040] For example, in the processing of Astragalus membranaceus batches entering storage, the origin code is set to 320101, the harvest season code is set to 3, and the initial processing method label is set to 01, resulting in a batch metadata identifier of "320101-3-01". The spectral data pairs acquired by a near-infrared and Raman dual-mode spectrometer and time-aligned have a near-infrared reflectance spectral matrix dimension of 2000×1 and a Raman scattering spectral matrix dimension of 520×1. The batch metadata identifier is written into the metadata structure of the spectral data pair, a bidirectional mapping index table is established for this data unit, and a summary value is calculated using a hash function. ; In this data unit, ID represents the batch identifier string, Spectra represents the serialized data of the spliced ​​near-infrared and Raman spectral matrices, and || is the concatenation operator. After verifying that the summary value matches the reference value, the acquisition time "2024-03-15T09:30:00Z", equipment number "NIR-RAMAN-05", and operator number "OP1023" are written into the traceability field. This data unit can be used to locate spectral data via the batch ID in subsequent searches, and can also be used to locate batch information in reverse via the spectral feature matrix. Verification results show that the retrieval latency is significantly reduced and the completeness of the traceability chain field is greatly improved.

[0041] S1.5: Perform batch aggregation and format standardization encapsulation on multiple batch-identified spectral data units to form an initial spectral dataset with batch identification that has a unified structure and is easy to retrieve and call later, thus completing the transformation from discrete sensor signals to systematic monitoring data.

[0042] Step S2: Based on the origin, harvesting season, and initial processing method metadata contained in the initial spectral dataset, a pre-constructed intrinsic spectral fingerprint database is invoked for matching and retrieval to generate a standard intrinsic spectral reference vector for the corresponding medicinal material matrix, which serves as the benchmark reference object for subsequent perturbation decoupling. Specifically, this includes: S2.1: Parse and process the batch identification information in the initial spectral dataset to extract the origin code, harvest season code, and initial processing method classification label to form a structured medicinal material matrix element information set, which serves as the index key value for subsequent fingerprint database retrieval.

[0043] S2.2: Based on the structured medicinal material matrix element information set, perform multi-dimensional joint matching retrieval in the intrinsic spectral fingerprint database to locate the historical calibration record cluster that is completely consistent with the current batch attributes of medicinal materials, and obtain the candidate intrinsic spectral prototype dataset containing multi-gradient toxin concentration response curves.

[0044] The intrinsic spectral fingerprint database is a structured spectral reference database constructed through a historical calibration experimental system. Its core data unit records the near-infrared reflectance and Raman scattering spectral response curves of specific medicinal material batches, uniquely identified by metadata such as origin, harvesting season, and initial processing method, collected under different aflatoxin concentration gradients. In application, the system uses the current medicinal material's metadata—origin, harvesting season, and initial processing method—as combined key values ​​to perform multi-dimensional joint matching searches within the database. This accurately locates historical record clusters with completely identical attributes and further filters out a set of spectral curves containing complete toxin concentration gradient responses, thereby outputting a highly reliable standard intrinsic spectral reference vector.

[0045] Based on the structured medicinal material matrix element information set, the intrinsic spectral fingerprint database retrieval interface is called to pass the combined key values ​​of the place of origin code, harvest season code and primary processing method classification label to the retrieval engine to form a multi-dimensional matching condition vector.

[0046] Attribute filtering operations are performed on the historical calibration record tuples in the intrinsic spectral fingerprint database to sequentially remove records that are inconsistent with any matching condition vector dimension, and retain the candidate record set that satisfies full dimension matching.

[0047] Within the candidate record set, the records are grouped according to the batch identifier. For each group of records, the concentration gradient index module is called to extract the spectral response curves containing multiple gradient points from zero to a set upper limit for aflatoxin B1 and B2 concentrations.

[0048] For each extracted spectral response curve, mode consistency verification is performed. Based on the matching degree of characteristic peak positions and morphology between near-infrared reflectance spectrum and Raman scattering signal, dual-mode consistency is determined, and curve samples with significant mismatch in either mode are removed.

[0049] The spectral response curves that have passed the pattern consistency verification are encapsulated into a candidate intrinsic spectrum prototype dataset. Each record in this dataset contains batch attribute labels, multi-gradient toxin concentration values ​​and corresponding bimodal spectral sequences. Through the above multi-dimensional joint matching retrieval and verification, the structured medicinal material matrix meta-information set of the previous step is transformed into candidate spectral data with complete traceability and concentration gradient coverage, so as to obtain the accurate input object for subsequent signal-to-noise ratio fusion.

[0050] For example, in one embodiment involving a batch of Astragalus membranaceus, the origin code was set to 4101, the harvest season code to Q3, and the initial processing method classification label to D2. The intrinsic spectral fingerprint database retrieval interface performed multidimensional matching using the combined key value [4101, Q3, D2], obtaining 54 historical calibration records that met the criteria. The concentration gradient index module extracted 324 corresponding spectral response curves based on aflatoxin B1 concentrations of 0, 10, 20, 30, 40, and 50 μg / kg and B2 concentrations of 0, 5, 10, 15, and 20 μg / kg. Using a dual-modal consistency judgment with a near-infrared reflectance spectrum and Raman scattering signal characteristic peak position matching degree threshold of 0.92, 17 mismatched curves were removed, ultimately forming a candidate intrinsic spectral prototype dataset containing 307 records. On this dataset, the average characteristic peak position shift of the dual-modal spectrum was controlled within 1.5 × 10⁻⁶. -3 Within the nm range, it meets the accuracy requirements for subsequent spectral fusion calculations.

[0051] S2.3: Perform signal-to-noise ratio weighted fusion calculation on the near-infrared reflectance spectral signal and Raman scattering signal in the candidate intrinsic spectral prototype dataset to eliminate single-mode noise interference and enhance the stability of characteristic peaks, generating a high-confidence dual-mode spectral fusion feature matrix.

[0052] S2.4: Principal component analysis algorithm is used to perform dimensionality reduction and reconstruction on the high-confidence dual-modal spectral fusion feature matrix to extract the low-dimensional manifold distribution that characterizes the intrinsic microstructure texture and component adsorption properties of medicinal materials, and obtain the intrinsic spectral feature basis vector of the medicinal material matrix.

[0053] The high-confidence dual-modal spectral fusion feature matrix obtained from the signal-to-noise ratio weighted fusion calculation in the candidate intrinsic spectral prototype dataset is centered by subtracting the mean of the corresponding full sample value from the reflection intensity value of each band to eliminate the offset based on data distribution. Covariance matrix calculation is then performed on the centered dual-modal spectral fusion feature matrix to obtain a multidimensional covariance structure describing the correlation and energy distribution characteristics between bands. Eigenvalue decomposition is performed on the covariance matrix to extract each eigenvector and its corresponding eigenvalue, and these eigenvalues ​​are sorted in descending order to determine the feature direction that contributes the most to the overall variance. Based on a preset cumulative contribution rate threshold, the first few principal components are selected to construct a dimensionality reduction transformation matrix, and the original high-dimensional dual-modal spectral fusion feature matrix is ​​projected onto the coordinate system of this transformation matrix to achieve low-dimensional reconstruction. Manifold distribution modeling is then performed on the reconstructed low-dimensional data to capture the relative positional relationship between the microstructure texture of the medicinal material surface and the adsorption characteristics of its components in the spectral space, outputting the intrinsic spectral feature basis vector of the medicinal material matrix. Principal component analysis is used to reduce and reconstruct the dimensionality, transforming the high-dimensional redundant information of the fused feature matrix into a low-dimensional basis vector with a compact structure and enhanced feature correlation, thus preparing the input for the subsequent zero-toxin concentration state mapping transformation.

[0054] For example, a dual-mode fusion feature matrix containing near-infrared reflectance signals from the 400 to 2500 nm band and anti-Centimeter Raman signals from the 800 to 1800 nm bands was centered to calculate a covariance matrix of 120 bands × 120 bands for a batch of Astragalus membranaceus samples, with values ​​ranging from 0.002 to 0.136. Eigenvalue decomposition was performed on this covariance matrix, yielding a maximum value of 0.136, corresponding to a eigenvector covering multiple key bands in both the near-infrared and Raman domains. Using a cumulative contribution rate threshold of 0.85, the first eight principal components were selected to construct a transformation matrix, projecting the original 120-dimensional feature vector into an 8-dimensional space. The sample point distribution in the low-dimensional space exhibited a clear clustering structure, with each cluster corresponding to a metal ion adsorption mode under different toxin concentration gradients. Based on this low-dimensional data, a manifold distribution model was further constructed to obtain the intrinsic spectral feature basis vector. The vector length is 8, and the element values ​​range from -0.45 to 0.37. The correlation coefficient between adjacent elements is significantly improved, which verifies its effectiveness in capturing the adsorption characteristics of medicinal materials.

[0055] S2.5: Based on the intrinsic spectral feature basis vector of the medicinal material matrix, perform zero toxin concentration state mapping transformation to remove the exogenous toxin response component introduced by the artificial inoculation experiment and restore the pure medicinal material's intrinsic state. Finally, output the standard intrinsic spectral reference vector as the benchmark reference object for perturbation decoupling.

[0056] like Figure 2 As shown, step S3: Real-time acquisition and calculation of the nonlinear influence weight of the perturbation factor on the spectral signal, generating a perturbation feature mask matrix. Specifically, this includes: S3.1: Acquire real-time temperature and humidity, light intensity and air cleanliness parameters of the storage environment, and perform nonlinear fitting processing on the temperature and humidity parameters of the storage environment based on the preset mapping relationship between the water film thickness on the surface of the medicinal material and the near-infrared absorption peak position shift, so as to generate a water film thickness interference coefficient vector that characterizes the degree of moisture interference.

[0057] Data collected from warehouse environment sensors, including temperature, relative humidity, light intensity, and air cleanliness values, serve as the input set for calculating water film thickness interference.

[0058] The temperature and relative humidity data are combined to form an environmental humidity state vector, and the nonlinear fitting module is called according to the preset mapping relationship between the water film thickness on the surface of the medicinal material and the near-infrared absorption peak position shift.

[0059] During the fitting process, the humidity state vector is input into the polynomial regression calculation unit, and the calibration coefficients in the mapping relationship are used to establish a functional expression of temperature and humidity versus peak position shift.

[0060] When using quadratic polynomial fitting, the peak offset Δλ is modeled as: ; Where RH represents relative humidity (%), T represents temperature (°C), and a, b, and c are coefficients obtained from calibration experiments.

[0061] The fitted peak offset Δλ is compared with the curve of water film thickness variation on the surface of the medicinal material, and the Δλ is converted into the water film thickness interference coefficient by the inversion algorithm.

[0062] The interference coefficients are segmented according to the spectral band index to form a water film thickness interference coefficient vector.

[0063] The above processing method transforms the environmental temperature and humidity data from the previous step into a water film thickness interference coefficient vector that characterizes the degree of moisture interference, thereby enabling a quantitative assessment of spectral distortion caused by temperature and humidity.

[0064] S3.2: Receive the water film thickness interference coefficient vector and the light intensity, and use the UV aging-induced fluorescence background rise quantization model to calculate the time accumulation effect of the light intensity to generate a fluorescence background rise weight sequence characterizing the intensity of photofluorescence interference.

[0065] The UV-aged fluorescence background rise quantification model is a mathematical-empirical model based on the photochemical aging mechanism, used to quantify the fluorescence background interference introduced by UV irradiation in the spectral detection of medicinal materials in the storage environment. The model takes real-time monitored light intensity (especially UV irradiance) time-series data as input, first integrating it over time to calculate the cumulative exposure, and then mapping this exposure to a series of time-varying fluorescence background rise coefficients through a pre-defined nonlinear saturation response function calibrated by prior aging experiments.

[0066] The system receives the water film thickness interference coefficient vector and light intensity as input, and calls the light response function in the UV aging-induced fluorescence background rise quantification model to map the light intensity parameter to the initial amplitude of the fluorescence background rise. During this mapping process, based on an empirical model showing that the excitation efficiency of fluorescent groups on the sample surface triggered by UV irradiation varies non-linearly with light intensity, a polynomial relationship is established between the intensity input and the background rise amplitude. After obtaining the initial amplitude of the background rise, combined with the illumination duration information, the cumulative rise value of the fluorescence background is calculated using a time-cumulative effect model. The cumulative model, based on an exponential integral form, integrates the instantaneous rise amplitude over the observation period to characterize the growth trend of the fluorescence signal over time. When using the cumulative model, the time-cumulative rise value is calculated using the following formula: ; Where B is the cumulative fluorescence background rise value, k is the model scaling factor, I is the illumination intensity, n is the intensity nonlinearity exponent, λ is the fluorescence attenuation coefficient, T is the illumination duration, and t is the integration variable. In the integration operation, the value of t continuously changes from 0 to T, traversing the entire illumination process, and is used to calculate the fluorescence contribution in each infinitesimal time interval.

[0067] The cumulative fluorescence background uplift value is multiplied by the water film thickness interference coefficient vector to introduce the modulation effect of the water film on the photoluminescence enhancement path. The fluorescence interference intensity in each band is adjusted by the product result so that the fluorescence background uplift weight distribution can reflect the dual effects of water film thickness and light intensity. The adjusted fluorescence background uplift value is then interpolated and extended across the spectral bands to generate a fluorescence background uplift weight sequence covering the entire near-infrared and Raman bands.

[0068] Through the above processing method, the water film thickness interference coefficient vector and light intensity in the previous step are transformed into a data sequence with both time accumulation effect and two-factor modulation characteristics, realizing the fluorescence background lifting weight sequence that characterizes the intensity of photoluminescence interference, and providing accurate light interference input for the subsequent construction of environmental disturbance feature tensors.

[0069] For example, in the medicinal material storage environment, the water film thickness interference coefficient vector is set to [1.05, 1.10, 1.15], the light intensity is 3200 Lux, the light duration T is 1800 seconds, the model scaling factor k is set to 0.002, the intensity nonlinearity exponent n is 1.3, and the fluorescence attenuation coefficient λ is 0.0005. Substituting the light intensity into the UV aging-induced fluorescence background rise quantification model, the initial rise amplitude calculation process is obtained: the light intensity I raised to the power of 1.3 is approximately 14177.3, multiplied by the coefficient k, the initial value is approximately 28.35. Calculated using the cumulative formula, the integral result is approximately 35400.7, and the cumulative fluorescence background rise value B is approximately 1004.2. Multiplying this value by the three elements of the water film thickness interference coefficient vector respectively, the adjusted rise value sequence [1054.41, 1104.62, 1154.83] is obtained. Linear interpolation was used to perform linear interpolation on the sequence in the 400–2500 nm near-infrared band and 800–1800 cm⁻¹ band. -1 Extending the Raman band to the full spectrum, a final fluorescence background enhancement weighted sequence was formed. This sequence was verified to be effective at wavelengths close to 1700 cm⁻¹. -1 The lifting intensity at the point was significantly increased, successfully reflecting the synergistic interference effect of water film thickness and photoluminescence under illumination.

[0070] S3.3: Call the water film thickness interference coefficient vector, the fluorescence background rise weight sequence, and the real-time air cleanliness parameters, and estimate the scattering loss of the real-time air cleanliness parameters according to the mapping table of standard dust deposition amount and spectral signal-to-noise ratio attenuation, so as to generate a dust deposition attenuation factor array characterizing the amplitude of dust adhesion interference.

[0071] The input conditions are the calculated water film thickness interference coefficient vector, the fluorescence background rise weight sequence, and the real-time air cleanliness parameters, where the air cleanliness parameters are represented by the particulate matter concentration values ​​output by the PM sensor in the storage environment.

[0072] Real-time air cleanliness parameters are mapped to a standard dust deposition data table, and the corresponding dust deposition values ​​are obtained through interpolation calculation, forming a dust deposition estimation vector with the same dimension as the water film thickness interference coefficient vector and the fluorescence background rise weight sequence.

[0073] The preset spectral signal-to-noise ratio attenuation mapping function is called to estimate the dust deposition amount vector, and the deposition amount value is converted into the scattering loss coefficient of the corresponding band. This mapping function is based on the coupling theory of particle size distribution and spectral scattering cross section.

[0074] Formula used: ; Calculate the attenuation factor for each band, where k is the dust scattering ratio coefficient, which is determined by the refractive index of the particles and the wavelength of the light source, and D is the particle mass density of the band corresponding to the dust deposition estimation vector.

[0075] The scattering loss coefficient vector, the water film thickness interference coefficient vector, and the fluorescence background lift weight sequence are dimensionally aligned to ensure that each interference component corresponds one-to-one with the band index.

[0076] The scattering loss coefficient vector is converted into a dust deposition attenuation factor array using an energy loss inversion algorithm based on a dust adsorption model. This array is used in the subsequent spectral signal compensation stage.

[0077] Through the above processing method, the disturbance factor in the previous step is transformed into a dust deposition attenuation factor array that can be directly used for spectral stripping processing, thereby achieving precise quantification of the dust adhesion interference intensity.

[0078] For example, in the Astragalus storage monitoring scenario, the real-time air cleanliness parameter output of the PM sensor is a particulate matter concentration of 150 micrograms per cubic meter. Mapping this to a standard dust deposition data table, the near-infrared band deposition mass density is obtained as 2.5 mg / cm². Using a dust scattering proportionality coefficient of 0.08 as input, the attenuation factor is calculated, yielding an amplitude attenuation value of approximately 0.833 for each band. This value is then transformed into a dust deposition attenuation factor array through band dimension alignment and energy loss inversion processing. This array is then combined with the water film thickness interference coefficient and the fluorescence background lift weight sequence to form a comprehensive tensor of the full-band interference distribution. Application results show that when the dust concentration is higher than the optimal air quality standard, the attenuation factor array significantly improves the recovery accuracy of the dust scattering correction algorithm, ensuring a significant improvement in the signal-to-noise ratio of the de-perturbed spectral residual sequence.

[0079] S3.4: Integrate the water film thickness interference coefficient vector, the fluorescence background lifting weight sequence, and the dust deposition attenuation factor array, and use a multi-source perturbation superposition algorithm to perform band alignment and weighted fusion processing on each interference component to generate a comprehensive environmental perturbation feature tensor containing full-band interference distribution information.

[0080] S3.5: Based on the comprehensive environmental perturbation feature tensor, perform normalization inversion and sparsification coding operations to convert continuous interference intensity values ​​into band-by-band suppression coefficients to generate the final perturbation feature mask matrix for spectral stripping processing.

[0081] like Figure 3 As shown, step S4: Apply the perturbation feature mask matrix to the initial spectral dataset to perform band-by-band stripping processing, eliminating spectral distortions caused by changes in water film thickness, fluorescence background enhancement, and dust deposition, and outputting a de-perturbed spectral residual sequence that only reflects the state of the medicinal material itself. Specifically, this includes: S4.1: Obtain the initial spectral dataset and the perturbation feature mask matrix. Based on the band alignment mechanism, map the water film thickness interference coefficient vector, fluorescence background lifting weight sequence and dust deposition attenuation factor array in the perturbation feature mask matrix to the near-infrared reflectance spectral signal band and Raman scattering signal band corresponding to the initial spectral dataset, respectively, to generate a band-level interference mapping tensor containing multi-dimensional environmental interference factors.

[0082] Obtain the initial spectral dataset and perturbation feature mask matrix, and call the band index matching module to read the full band index list of near-infrared reflectance spectral signals and the full band index list of Raman scattering signals.

[0083] Based on the band alignment mechanism, the water film thickness interference coefficient vector, the fluorescence background elevation weight sequence, and the dust deposition attenuation factor array are loaded respectively. The three types of interference coefficient vectors are mapped to the corresponding near-infrared signal bands and Raman signal bands in the initial spectral dataset according to their corresponding band index positions.

[0084] During the mapping process, a correlation matrix between the interference coefficient and the original signal amplitude is established for each band. Band-level interference matrices containing water film interference components, photoluminescence interference components, and dust interference components are generated using dual-mode signals.

[0085] Perform matrix stacking operations on the three types of band-level interference matrices according to interference type to generate a band-level interference mapping tensor containing multi-dimensional environmental interference factors.

[0086] An index check is performed on the generated band-level interference mapping tensor to ensure that the interference components of the near-infrared band and the Raman band occupy the correct positions in the tensor structure, thus realizing a one-to-one correspondence between the interference coefficient and the physical signal band.

[0087] Through the above mapping process, the comprehensive environmental disturbance characteristics from the previous step are solidified into the original spectral band structure in the form of precise band positions, and a band-level disturbance mapping tensor for nonlinear baseline fitting is output, realizing the structured input of environmental disturbance factors in the spectral stripping process.

[0088] For example, in a batch testing of traditional Chinese medicine stored in a warehouse, the initial spectral dataset contained near-infrared reflectance signals covering the 400–2500 nm band, divided into 1024 discrete sampling points, and Raman scattering signals covering the 800–1800 nm band, divided into 512 discrete sampling points. The length of the water film thickness interference coefficient vector in the perturbation feature mask matrix was consistent with that of the near-infrared sampling points. Each coefficient was calculated based on the relationship between relative humidity (RH) and absorption peak position shift, where RH values ​​were 65%, 75%, and 85%, corresponding to coefficients ranging from 0.02 to 0.05. The length of the fluorescence background enhancement weight sequence was consistent with that of the Raman sampling points, and the coefficients were calculated based on the cumulative time effect of light intensity, with light intensities of 500, 800, and 1200 lux, corresponding to coefficients ranging from 0.01 to 0.03. The dust deposition attenuation factor array covers the entire near-infrared and Raman bands. Its coefficients are calculated based on the relationship between air cleanliness level and PM2.5 deposition amount, with cleanliness levels of ISO 8 and ISO 9, and coefficient values ​​ranging from 0.005 to 0.02. The three types of interference coefficient vectors are mapped to their corresponding spectral bands according to their band indices. After performing a matrix stacking operation, a near-infrared interference mapping sub-matrix of size (3 × 1024) and a Raman interference mapping sub-matrix of size (3 × 512) are obtained. These are merged into a final band-level interference mapping tensor, and the mapping accuracy is verified. In the subsequent S4.2 step, this tensor is directly used as the input for nonlinear baseline fitting to achieve quantitative stripping of multi-source interference, significantly improving the accuracy of the de-perturbed spectral residuals.

[0089] S4.2: Perform nonlinear baseline fitting processing on the band-level interference mapping tensor, and use a polynomial regression algorithm to calculate the absorption peak shift caused by the change in water film thickness under each band and the fluorescence background rise curve caused by the duration of illumination, so as to generate an environmental interference synthetic substrate spectrum characterizing the superposition result of physical effects.

[0090] Nonlinear baseline fitting is performed on each spectral band signal in the band-level interference mapping tensor. The reference data for baseline fitting are set as the water film thickness influence weight and fluorescence background enhancement coefficient of the corresponding band in the perturbation feature mask matrix. The water film thickness influence weight is mapped to the band index position of the near-infrared reflectance spectrum signal to form the water film interference sub-vector for baseline fitting. The fluorescence background enhancement coefficient is mapped to the band index position of the Raman scattering signal to form the corresponding photoinduced background interference sub-vector. Polynomial regression modeling is performed on each band signal to fit the absorption peak shift caused by the change in water film thickness; the calculation process uses the following expression: ; Where Δλ is the peak displacement, a n R represents the polynomial coefficients. h Here, is the water film thickness interference coefficient, and n is the index of the polynomial order, starting from 0 (or 1) and increasing to the preset highest order. Based on illumination duration data and fluorescence background rise coefficient, a fluorescence background rise curve is fitted; its calculation process uses the following expression: ; Where F bg b m T represents the polynomial coefficients. uv Here, m represents the illumination duration, and m is the polynomial order index, starting from 0 (or 1) and increasing to the preset highest order. The water film absorption peak shift curve and the fluorescence background rise curve are weighted and superimposed within the band index space to form a synthetic environmental interference base spectrum characterizing the superposition of physical effects. Through nonlinear baseline fitting, the band-level interference mapping tensor from the previous step is transformed into a synthetic environmental interference base spectrum capable of comprehensively describing the coupling effect of water film interference and fluorescence interference, thus providing a precise baseline stripping reference for subsequent point-by-point subtraction operations.

[0091] For example, when performing nonlinear baseline fitting on the band-level interference mapping tensor of Astragalus membranaceus batches, the highest value of the water film thickness interference coefficient was 0.12 mm. Mapped to the near-infrared reflectance spectrum band of 800 to 1200 nm, the peak shift was fitted using third-order polynomial regression with coefficient vector a = [0.0012, -0.0004, 0.00005], and the maximum peak shift was 0.95 nm. The highest value of the fluorescence background lifting weight sequence was 0.034 under 720-hour illumination conditions, mapped to the Raman band of 1200 to 1400 cm⁻¹.-1 For the interval, a second-order polynomial regression was used to fit the background rise curve, with a coefficient vector b = [0.0025, -0.0009], and the maximum value of the background rise curve was 0.021. The peak shift curve and the background rise curve were then superimposed with a weighting factor of 0.6:0.4 to obtain the environmental interference synthetic base spectrum, whose main peak region corresponds to a composite interference intensity of approximately 0.78. Applying this base spectrum to the subsequent point-by-point subtraction process in step S4.3 significantly improves the accuracy of baseline stripping, significantly reduces the amplitude of interference residue in the restored spectral features, and ensures that the restored de-perturbed spectral residual sequence can accurately characterize the state of the medicinal material.

[0092] S4.3: Based on the environmental interference synthesized substrate spectrum, perform point-by-point subtraction on the initial spectral dataset to remove the water film absorption component from the original near-infrared reflectance spectral signal and subtract the fluorescence background component from the original Raman scattering signal to generate an intermediate denoised spectral sequence that has initially removed linear and nonlinear environmental noise.

[0093] Based on the environmental interference-synthesized substrate spectrum and the initial spectral dataset, a band-by-band correspondence was established. Water film absorption components were removed from the near-infrared reflectance spectral signal bands, and differential calculations were used for each data point to remove spectral intensity shifts caused by changes in water film thickness. Photofluorescence background subtraction was performed on the Raman scattering signal bands. Point-to-point subtraction was performed using the fluorescence background curve in the substrate spectrum according to the corresponding band values ​​to remove light-induced linear and nonlinear fluorescence uplift components. To ensure that high-frequency noise components do not accumulate during subtraction, the spectral difference sequence was smoothed using a bandpass filter, preserving the energy characteristics of the intrinsic peak regions of the medicinal material and suppressing noise in non-target bands. Amplitude normalization was performed on the smoothed near-infrared and Raman difference sequences to maintain consistency in the numerical range of different modal signals, providing a consistent scale benchmark for subsequent multimodal feature fusion. Intermodal amplitude synchronization adjustment was performed on the normalized dual-modal difference sequence to ensure precise correspondence between the two types of signals in time index, forming an intermediate denoised spectral sequence that initially removes environmental noise while preserving the response characteristics of the medicinal material. By employing band-by-band subtraction and filtering, the environmental interference synthesized substrate spectrum from the previous step is transformed into an intermediate denoised spectral sequence that eliminates water film absorption and fluorescence background components, thereby achieving a significant elimination effect on both linear and nonlinear environmental noise.

[0094] For example, in the batch spectral data processing of Astragalus membranaceus, the initial spectral dataset contains 1024 near-infrared bands and 256 Raman bands. The environmental interference synthesized substrate spectrum records a water film absorption peak position shift of 0.015 reflectance units and a fluorescence background rise of 4.2 count units. For the near-infrared bands, a point-by-point subtraction operation is performed, and the calculation formula is as follows: ; Where S is the original spectral signal, B is the water film component of the substrate spectrum, and S′ is the spectral signal after component removal. For the Raman band, a similar operation is performed to subtract the fluorescence background component. The fluorescence curve is obtained through polynomial fitting with a fitting order of 3. The calculation formula is as follows: ; Where R is the original Raman signal and F is the fitted background component. The subtracted difference sequence is processed through 0.5–1.8 μm and 850–1650 cm⁻¹. -1 The bandpass filter in the interval is smoothed, and the filter bandwidth is set to 0.2μm and 50cm. -1 This ensures that the intrinsic peak region signal of the medicinal material is preserved. Normalization processing unifies the amplitude range of the near-infrared and Raman difference sequences to the [-1,1] interval, and the maximum deviation of the inter-modal synchronization adjustment is controlled within 0.002 time units. The final output intermediate denoised spectral sequence maintains the complete peak positions while reducing the distortion amplitude caused by water film and fluorescence background to within 0.001 of the original amplitude, meeting the input requirements for subsequent dust scattering correction.

[0095] S4.4: Apply a dust scattering correction algorithm to the intermediate denoised spectral sequence, perform inverse gain compensation on the signal amplitude based on the dust deposition attenuation factor array, repair the spectral signal-to-noise ratio attenuation region caused by particulate matter occlusion, and generate a restored spectral feature vector after full-factor environmental perturbation correction.

[0096] The intermediate denoised spectral sequence is input into the dust scattering correction algorithm model. The dust deposition attenuation factor array is used as the basis for amplitude correction, establishing an inverse mapping relationship between particulate obstruction and signal-to-noise ratio attenuation across the entire spectral range. The dust deposition attenuation factor array is mapped to the corresponding spectral bands, forming a band-level attenuation compensation index matrix, ensuring a one-to-one correspondence between the dust interference weight and the amplitude of the denoised signal in each spectral band. Inverse gain calculation is performed on the band-level attenuation compensation index matrix, using a multiplicative correction method under energy conservation constraints. The compensated signal amplitude value is obtained using the following formula: ; Where A represents the original amplitude of the denoised signal, k is the dust deposition attenuation coefficient, and A′ is the compensated signal amplitude. During inverse gain calculation, the compensated band amplitude is smoothed and filtered to eliminate high-frequency noise components introduced by abrupt gain changes. The smoothed amplitude values ​​are reassembled back into the full-band spectral sequence to form a dust scattering-corrected spectral feature matrix. Normalization verification is performed on this feature matrix to ensure overall consistency between the energy distribution of each band and the reference intrinsic spectral characteristics, outputting a restored spectral feature vector corrected for full-factor environmental disturbances. Through dust scattering correction and inverse gain compensation, the intermediate denoised spectral sequence from the previous step is transformed into a restored spectral feature vector with particulate matter occlusion repair capabilities, achieving a significantly improved signal-to-noise ratio and characteristic peak stability.

[0097] For example, in the batch spectral detection scenario of Astragalus membranaceus, the amplitude of the intermediate denoised spectral sequence shows significant attenuation in the 1550 nm band, and the dust deposition attenuation factor array is set to 0.82 in this band. Mapping this coefficient to the 1550 nm band, the original amplitude A in the corresponding band is 0.56. The compensated amplitude A′ is calculated using the formula: ; That is, A′ is 0.6829. A Savitzky–Golay filter with a window length of 7 was used to smooth the compensation amplitude within a range of ±5 nm around 1550 nm, eliminating fluctuations caused by sharp gain. The resulting dust scattering correction spectral feature matrix showed an average signal-to-noise ratio increase of approximately 1.25 times across the entire band, with the main discrimination peak at 1650 cm⁻¹. -1 The stability of the interval is significantly improved, and the output restored spectral feature vectors show higher matching accuracy and interpretability in subsequent toxin adsorption site identification.

[0098] S4.5: Perform residual calculation processing on the restored spectral feature vector based on the standard intrinsic spectral reference vector, extract the residual fluctuation components in the restored spectral feature vector that are unrelated to the medicinal material matrix and filter them out, and finally output the perturbation-free spectral residual sequence that retains only the microstructure texture and toxin adsorption characteristics of the medicinal material.

[0099] Based on the obtained restored spectral feature vector and standard intrinsic spectral reference vector, the vector difference operation module is invoked to perform element-wise comparison between the two, generating a band-by-band difference value sequence. A signed amplitude discrimination strategy is applied to the difference value sequence, setting difference components with absolute amplitudes below the noise threshold to zero to eliminate weak fluctuations introduced by residual environmental disturbances. Spectral component analysis is performed on the retained difference components, mapping the difference value sequence to the spectral characteristic frequency domain, extracting high-frequency disturbance components unrelated to the medicinal material matrix, and labeling their band indices. The residual filtering function is invoked to locate the corresponding data points in the restored spectral feature vector according to the labeled band indices, and replace them with the corresponding values ​​in the standard intrinsic spectral reference vector. Overall smoothing filtering and energy normalization are performed on the data vector after residual replacement processing to ensure that the de-perturbed spectral residual sequence maintains continuity and amplitude consistency across the entire band range. Through the above residual calculation and filtering methods, the restored spectral feature vector is transformed into a de-perturbed spectral residual sequence that retains only the microstructure texture and toxin adsorption characteristics of the medicinal material, achieving an environmentally independent feature purification effect.

[0100] For example, a noise threshold of 0.002 reflectance units was set for the reconstructed spectral feature vector and the standard intrinsic spectral reference vector of Astragalus membranaceus batches. After differential calculation, the average amplitude of the absolute difference value detected in the 1450–1470 nm range was 0.006 reflectance units, significantly higher than the threshold, and determined to be a fluctuation not caused by matrix components. The difference value in this range was replaced with the corresponding value of the reference vector, and then a smoothing filter based on a five-point moving average was performed on the full-band data, and the maximum amplitude was normalized to 0.95 reflectance units. The formula for calculating the difference value is: ; Where V is the restored spectral feature vector, R is the standard intrinsic spectral reference vector, and Δ is the band-by-band difference sequence. The noise threshold discrimination formula is: ; Where M represents the discrimination result and T represents the noise threshold. In this embodiment, the peak shape of the processed, perturbation-free spectral residual sequence in the toxin binding site band is highly consistent with that of the reference sample, enabling the subsequent graph neural network model to significantly improve the accuracy of pollution level identification and enhance the interpretability of the judgment results during the feature input stage.

[0101] Step S5: Based on the microscopic tissue texture topology obtained from the previous microscopic CT reconstruction, a lightweight graph neural network model is constructed. The perturbation-free spectral residual sequence is mapped as node attributes and input into the model. Through a message passing mechanism, the model learns the adsorption site response patterns of toxin molecules in the medicinal matrix, generating a pollution level attribution confidence vector and a heatmap of contributions from key discrimination bands. Specifically, this includes: S5.1: Based on the microscopic tissue texture topology of medicinal materials obtained from previous microscopic CT reconstruction, the node set and edge connection relationship are initialized and defined. The data of each band in the perturbation-free spectral residual sequence are mapped to node attribute vectors, and a medicinal material spectral topology map data object containing a spatial adjacency matrix and a feature attribute matrix is ​​constructed.

[0102] Based on the microscopic tissue texture topology of medicinal materials obtained from previous micro-CT reconstruction, the perturbation-free spectral residual sequence was loaded as the input data source, and the corresponding microscopic three-dimensional structure file was matched with the batch identifier.

[0103] The microscopic tissue texture topology of medicinal materials obtained by the prior micro-CT reconstruction refers to the network skeleton formed by reconstructing and abstracting the high-resolution three-dimensional voxel data inside the medicinal materials obtained by micro-CT scanning technology. This skeleton is used to characterize the spatial adjacency relationships of its microscopic components, such as cell walls, vessels, and pores. The structure uses voxels or key spatial points as nodes and the spatial adjacency relationships between nodes as edges, thereby constructing a topological map that can encode the complex three-dimensional morphology and connection relationships inside the medicinal materials. This provides an analytical basis for subsequently mapping non-spatial attributes such as spectra to this spatial framework.

[0104] The voxel data from micro-CT are used to extract a set of nodes with actual physical connections according to the spatial distribution of the tissue texture of medicinal materials. Each node represents a scanned tissue unit.

[0105] The edge connection relationship initialization process is performed on the node set. A connectivity matrix is ​​established based on the distance threshold between voxels and the direction of tissue fibers, and the edge type is labeled to distinguish different tissue structure connection modes.

[0106] The perturbation-free spectral residual sequence of multi-band data is analyzed, the intensity value of each band is converted into the feature attribute vector corresponding to the node, and normalization is performed to eliminate the amplitude difference between different bands.

[0107] Construct a spectral topology map data object for medicinal materials containing a spatial adjacency matrix and a feature attribute matrix. The adjacency matrix records the spatial connection relationship between nodes, and the feature attribute matrix records the spectral response of each node.

[0108] This processing method transforms the perturbation-free spectral residual sequence from the previous step into structured topological data input acceptable to graph neural networks, ensuring that spectral information and the spatial relationship of medicinal tissues are synchronously integrated, thus achieving a high-fidelity data foundation for learning toxin adsorption locations.

[0109] For example, when processing micro-CT reconstruction data of Astragalus membranaceus batches, the voxel resolution was set to 5 μm, the number of nodes was 4096, and an Euclidean distance of less than 15 μm was used as the threshold for edge connection creation. Fiber connections in different directions were labeled as type1, and vertical crosslinks were labeled as type2. In the perturbation-free spectral residual sequence, there were 256 sampling points in the near-infrared band and 128 sampling points in the Raman band, which were mapped to the first 256 dimensions and the last 128 dimensions of node feature attributes, respectively. Z-score normalization was used to process each dimension of the features. When constructing the topology graph data object, the adjacency matrix was a 4096×4096 sparse matrix, and the feature attribute matrix was 4096×384 dimensions. The normalization formula is as follows: ; Where X is the original band intensity, μ is the average value of the band across all nodes, and σ is the standard deviation of the band across all nodes. After constructing this data object, the graph neural network can significantly improve the identification accuracy of aflatoxin adsorption sites during the training phase and exhibits a substantial improvement in the stability of pollution level determination on the test set.

[0110] S5.2: The graph convolution operator is used to perform multi-round neighborhood aggregation on the spectral topology graph data object of medicinal materials. The feature attribute matrices of adjacent nodes are weighted and fused to update the hidden state of the central node, and a high-order embedding representation vector of the node representing the coupling relationship between the local microenvironment and the spectral response is generated.

[0111] Based on the spatial adjacency matrix and feature attribute matrix of the medicinal herb spectral topology map data object as input, the predefined graph convolution operator kernel function is called to perform the first round of neighborhood feature aggregation operation on the node set. According to the connection relationship of the spatial adjacency matrix, the attribute vectors of the direct neighboring nodes of each central node are extracted, and the attribute vectors are weighted and summed with normalized weight coefficients to form the aggregated candidate feature vector of the central node.

[0112] The first round of aggregated candidate feature vectors are linearly combined with the original attribute vectors of the center node, and the ReLU nonlinear activation function is introduced to enhance the nonlinearity of feature representation, outputting the updated set of first-order hidden state vectors.

[0113] Using the updated set of first-order hidden state vectors as input, a second round of neighborhood feature aggregation operation is performed on the node set, extending the message passing range to two-hop neighbor nodes, adjusting the information contribution of different adjacent paths through the weight of adjacent edges, and merging the state vectors of two-hop neighbor nodes into the feature representation of the central node in a weighted manner.

[0114] Batch normalization is used to standardize the statistical properties of the second-round aggregation results, which reduces the bias of node feature distribution and improves the convergence stability of subsequent learning stages.

[0115] A high-order embedded representation vector set containing local microenvironment information and spectral response coupling characteristics is gradually generated through a multi-round iterative neighborhood aggregation process, ensuring that the feature representation of each node covers its local micro-tissue texture and toxin adsorption site pattern.

[0116] Through the above-mentioned multi-round neighborhood aggregation process, the perturbation spectral residual sequence is embedded into the high-dimensional feature space of the node, forming a high-order embedding representation of the node that characterizes the coupling relationship between the local environment and the spectral mode, providing a fine feature basis for the subsequent attention mechanism to screen key channels.

[0117] S5.3: Based on the node high-order embedded representation vector, an attention mechanism is introduced to calculate the edge weight coefficient, dynamically optimize the message passing path in the medicinal material spectral topology map data object, screen out the key feature channels that are strongly correlated with the adsorption characteristics of aflatoxin, and generate a sparse graph feature tensor focusing on the toxin binding site.

[0118] The input condition for the node high-order embedded representation vector is a set of node features formed after multiple rounds of neighborhood aggregation by the graph convolution operator, containing the coupling relationship between the local microenvironment and the spectral response. In this set, the relationship between each node and its adjacent edges has been fixed and mapped through a spatial adjacency matrix. An attention weight calculation operator is called on each node's high-order embedded representation vector to perform a correlation evaluation operation for each edge. The cosine similarity of the feature vectors between the central node and its adjacent nodes, the contribution of band features, and the peak intensity of the toxin adsorption response are fused according to a preset combination function to calculate the edge weight coefficient matrix. Based on the edge weight coefficient matrix, the message passing path in the medicinal herb spectral topology data object is dynamically reconstructed. Edges with weight coefficients below a significance threshold are pruned to reduce redundant propagation, while edges with weight coefficients above the threshold are retained and enhanced to concentrate on toxin-sensitive areas. Key feature channel screening is performed on the pruned graph structure. Based on the numerical distribution of the matching degree between the spectral band attributes of each node and the response mode of aflatoxin adsorption sites, band channels with a matching degree within a specified upper limit are selected as key node attributes. The selected key feature channels are bound to their corresponding node indices, and sparse encoding is performed. The feature vectors of non-key channels are set to zero, and only the feature values ​​of key channels are retained, generating a sparse graph feature tensor focused on toxin binding sites. Through the above attention mechanism and dynamic optimization processing, the high-order embedding representation vectors of nodes in the previous step are transformed into sparse graph feature tensors with the ability to focus on toxin binding sites, realizing the model's centralized modeling of high-discriminative feature regions in pollution level determination.

[0119] For example, in the high-order embedding representation vector of the node in the spectral topology map of Astragalus membranaceus batches, the input parameters include the 64-dimensional spectral features of each node and the spatial distance weights of the corresponding adjacent edges. A fusion function is used in the edge weight calculation: ; Where s is the band contribution score, c is the cosine similarity value, and p is the proportion of toxin peak intensity. After the weight coefficient matrix is ​​generated, the pruning threshold is set to 0.65; edges below this value are removed, and edges above this value are given a weight gain of 1.2 times during message passing. In the key feature channel screening stage, the matching degree upper limit is set to 0.9, and the selected bands include 1650cm. -1 ¹、1200cm -1 and 1450cm -1 The corresponding node attribute positions are channels 15, 28, and 42, respectively. In the sparse coding, other channels are set to zero, and only the values ​​of these three channels are retained to form a sparse graph feature tensor, whose non-zero elements account for 4.7% of the original feature tensor elements. This tensor significantly improves the concentration of toxin binding site localization in subsequent global readouts, and significantly expands the classification gap between high-risk and low-risk samples in the feature space in the pollution level determination.

[0120] S5.4: The pooling operation is performed on the sparse graph feature tensor by the global readout function, which aggregates the scattered node high-order embedding representation vectors into a fixed-dimensional graph-level global feature vector. This vector is then mapped to the pollution level classification space through a fully connected layer to generate a pollution level attribution confidence vector containing four states: safe, low-risk, medium-risk, and high-risk.

[0121] The sparse graph feature tensor optimized by the attention mechanism is input into the global readout function, and pooling operation is performed on the high-order embedding representation vector of each node to eliminate the differences in node position and form a unified graph-level feature structure.

[0122] The weighted average pooling method is adopted, and the weight coefficients of the node embedding representation are derived from the calculation results of the attention mechanism. The weight values ​​are proportional to the importance of the spectral channel corresponding to the node.

[0123] The weighted node vectors are summed within the graph structure to obtain a global feature vector of fixed dimensions, while maintaining the discriminative power of the original spectral features.

[0124] The global feature vector is input into the fully connected layer, and the high-dimensional embedding is mapped to the pollution level classification space through parameter matrix multiplication and bias term addition. The classification space includes four states: safe, low risk, medium risk and high risk.

[0125] The Softmax normalization function is used to probabilize the output of the classification space so that the output values ​​of each pollution level numerically satisfy the probability distribution constraints.

[0126] Through the above pooling, mapping and normalization processes, the sparse graph feature tensor in the previous step is transformed into a pollution level attribution confidence vector, thereby realizing a numerical expression of pollution level determination.

[0127] For example, in an embodiment where the number of nodes in the medicinal herb spectral topology graph is 128 and the embedding vector dimension of each node is 64, the sparse graph feature tensor is weighted average pooling to obtain a 64-dimensional global feature vector. The pooling weights are provided by the distribution vector output by the attention mechanism, with unit weights ranging from 0.0 to 1.0 and a mean of approximately 0.42. When the global feature vector is mapped to the 4-dimensional classification space through a fully connected layer, the parameter matrix dimension is 64×4 and the bias vector dimension is 4. After Softmax normalization, the output pollution level attribution confidence vector has an example value of [0.68, 0.21, 0.07, 0.04], indicating that the confidence of the safety level is significantly higher than other levels, reflecting that the actual batch of medicinal herbs presents a low pollution risk state under the input of the perturbation-free spectral residual sequence. In this embodiment, the confidence vector output in this step is used for interpretability analysis and risk decision-making in the subsequent dynamic pollution judgment report, significantly improving the judgment accuracy and system credibility.

[0128] S5.5: Based on the gradient backpropagation path in the confidence vector of pollution level attribution, extract the contribution weight values ​​of each band data to the pollution level classification results, map the contribution weight values ​​back to the original spectral band index and superimpose them onto the spatial coordinates of the microstructure texture topology of the medicinal material, and generate a key discrimination band contribution heatmap that intuitively displays the key discrimination spectral intervals and their potential binding sites.

[0129] The input conditions are the pollution level attribution confidence vector and the corresponding gradient backpropagation path information. The execution object is the gradient response of the de-perturbed spectral residual sequence in the graph neural network classification stage. Based on the gradient backpropagation path, the mapping interface from graph-level features to node-level features is called to distribute the backpropagation gradient of the category label corresponding to the confidence vector to the feature attribute matrix of each node. The sensitivity coefficient of each band is calculated by backpropagating through the node weights. The sensitivity coefficient sequence is numerically integrated using a band-specific difference operator to obtain the cumulative contribution of each band to the classification output, and a contribution weight value vector is generated through normalization. The contribution weight value vector is mapped back to the original spectral band index to establish a band-tissue spatial coordinate correspondence table. The coordinates are superimposed using the spatial adjacency matrix of the microscopic CT texture topology to form a joint encoding tensor of node spatial position and spectral band contribution weight. A high-threshold truncation operation is performed on the joint encoding tensor to retain the band node combinations whose contribution weight values ​​exceed the significance judgment threshold. These are marked as key discriminative feature units, and clustering is performed based on the spatial continuity of the node positions to extract the corresponding groups of key discriminative spectral intervals and potential binding sites. A heatmap generation algorithm is employed to overlay the contribution weights of each key discriminative feature unit onto a spatial layer representing the microstructure topology of the medicinal material using color gradients, outputting a heatmap of key discriminative band contributions. Through gradient weight backtracking and spatial mapping, the contamination level classification results from the previous step are transformed into visualized discriminative feature data with tissue location correlations, achieving the expected technical effect of significantly enhanced interpretability of the model.

[0130] For example, in the risk assessment of a batch of Astragalus membranaceus, the input pollution level assignment confidence vector is [0.05, 0.20, 0.35, 0.40], and the largest category shown in the gradient backpropagation path is high risk. The mapping interface is called to assign the backgradient of this category to 50 band nodes, with the sensitivity coefficient sequence ranging from 1450 to 1650 cm⁻¹. -1 Significant peaks were observed in the 2100-2200 nm range. The contribution weighting formula was used as follows: ; Where C is the contribution weight value, S is the sensitivity coefficient function, b is the band index, and max and min are the upper and lower bound indices of the interval, respectively. The normalized contribution weight value is at 1650cm. -1 The value reached 0.92 at 100 nm and 0.88 at 2100 nm. After establishing the correspondence between the band index and the micro-CT tissue coordinates, a joint coding tensor was generated, and the positions of the two key bands and their corresponding nodes were preserved at a threshold of 0.80. Weighted clustering was used to merge similar nodes into binding site clusters, and finally the location at 1650 cm⁻¹ was determined in the dermal cell structure of the medicinal material. -1The peak enhancement corresponds to protein secondary structure binding sites and polysaccharide group binding sites at 2200 nm. During heatmap generation, a color mapping range [blue (low) to red (high)] was used to overlay the 3D tissue topology. The output visualization results show significant red areas in the high-contribution bands. Evaluation shows that this heatmap can accurately indicate high-risk areas for toxin binding, significantly improving the interpretability and reliability of the model's judgment.

[0131] Step S6: Determine the current aflatoxin contamination level based on the maximum probability value in the contamination level attribution confidence vector, and extract one or more key discriminant spectral intervals with the highest discriminant power and their corresponding potential binding site inference text by combining the key discriminant band contribution heatmap, forming a dynamic contamination determination report containing the level conclusion and interpretable elements. In this embodiment, three key discriminant spectral intervals are preferably extracted. Specifically, these include: S6.1: Perform maximum value retrieval processing on the pollution level attribution confidence vector to extract the pollution level label with the highest probability value as the preliminary determination result of the current aflatoxin pollution level, providing core decision-making basis for subsequent interpretability analysis.

[0132] For the pollution level attribution confidence vector output by the lightweight graph neural network model, load the vector into the confidence resolution module as input.

[0133] The vector element index traversal operator is called to establish a location sequence, and a one-to-one mapping relationship is established between the confidence value of each pollution level label and the corresponding label.

[0134] The maximum value retrieval operator performs a global search on the set of confidence values, and uses an extreme value comparison strategy to continuously update the current highest confidence value and its corresponding label index during the traversal process.

[0135] Numerical stabilization measures are introduced to calculate the difference coefficient between the highest confidence value and the second highest value, ensuring that the maximum value retrieval result meets the significance judgment condition and avoiding label misjudgment caused by boundary value disturbance.

[0136] Decode the tag index, extract the pollution level tag text corresponding to the highest confidence value from the tag mapping table, and write it into the cache unit of the current aflatoxin pollution level preliminary determination result.

[0137] By using this maximum value retrieval and label parsing process, the pollution level attribution confidence vector generated in the previous step is transformed into a unique pollution level label, which serves as the core decision-making basis for subsequent interpretability analysis.

[0138] For example, after inputting the de-perturbed spectral residual sequence of a batch of medicinal materials into a graph neural network model, the resulting pollution level classification confidence vector is [0.12, 0.25, 0.48, 0.15], corresponding to the labels safe, low risk, medium risk, and high risk, respectively. This vector is loaded into the confidence resolution module, establishing a mapping relationship {safe: 0.12, low risk: 0.25, medium risk: 0.48, high risk: 0.15}. The set is traversed, and the maximum value retrieval operator updates the highest confidence value to 0.48, corresponding to the label medium risk. The difference coefficient calculation formula is: ; The numerator represents the difference between the highest and second-highest values, and the denominator is the sum of the two, used for significance determination. If the calculation result meets the preset significance condition, the decoded label index extracts "medium risk" as the determination result and writes it to the cache unit. In this embodiment, the output result is "medium risk," which will serve as the basis for subsequent key discrimination band extraction and binding site inference. The application effect shows that the system can complete the determination within 800 milliseconds and maintain high stability, avoiding result fluctuations caused by close confidence levels.

[0139] S6.2: Based on the preliminary determination result of the current aflatoxin contamination level, the contribution heatmap of the key discrimination band is called to perform spatial mapping analysis to identify three target spectral intervals in the heatmap whose response intensity exceeds the preset significance threshold, forming a set of high discrimination spectral features.

[0140] Based on the preliminary assessment of aflatoxin contamination levels and the contribution heatmap of key discrimination bands, a spatial mapping relationship is established. Each spectral band index in the heatmap is point-by-point bound to the spatial coordinates of the microstructure topology of the medicinal material, forming a band spatial mapping matrix with response intensity values. A threshold screening operation is performed on the band spatial mapping matrix. A preset significance threshold is used to compare the response intensity column element-by-element, marking all band indices greater than the threshold as candidate target bands. Adjacent band aggregation is performed on the candidate target bands. A continuous band merging algorithm is used to merge band sets with band index spacing less than the set merging window width into single spectral intervals, generating a set of redundancy-free target spectral intervals. Response intensity integral calculation is performed on the target spectral interval set, and a weighted summation method is used to evaluate the overall discrimination capability of each spectral interval, where the weight values ​​are taken from the contribution weight of the corresponding band position in the heatmap. The target spectral range set is sorted in descending order of overall discriminative power value. The top three spectral ranges in the sorting results are extracted to form a high-discriminative spectral feature set. This set effectively focuses on the optimal information band in pollution level judgment, achieving feature localization with enhanced interpretability.

[0141] S6.3: Using a pre-constructed medicinal material molecular structure-spectral response association knowledge base, semantic matching processing is performed on each target spectral interval in the high-discrimination spectral feature set to infer the potential binding sites of the medicinal material microstructure corresponding to each target spectral interval and generate a potential binding site prediction list.

[0142] For each target spectral region in the high-discrimination spectral feature set, a retrieval request corresponding to the molecular structure of the medicinal material is established, and a pre-constructed medicinal material molecular structure-spectral response association knowledge base is invoked to obtain the absorption or scattering response mode parameters of the spectral region in different medicinal material matrices.

[0143] The retrieved response mode parameters are matched with the peak position, peak intensity, and peak shape characteristics of the target spectral range at multiple scales. A similarity-based matching algorithm is used, and the matching degree is calculated using the cosine similarity formula. High-matching molecular structure candidates are then conformationally screened. The energy change curves of toxin molecules at the binding sites in the medicinal material matrix are calculated using a molecular dynamics simulation module, and candidate binding sites with binding energies below a preset threshold are extracted.

[0144] By projecting the candidate sites selected by binding energy onto the spatial coordinates of the microstructure of the medicinal material, the location of the site in the actual microstructure of the medicinal material and its surrounding chemical environment are confirmed.

[0145] By establishing a one-to-one correspondence between the mapped molecular structure information and spectral response characteristics, a potential binding site prediction list is formed. Through structure matching and energy screening, the high-discrimination spectral feature range of the previous step is transformed into molecular binding site data with precise spatial locations, thereby achieving a structured enhancement of the interpretability of pollution level determination.

[0146] For example, in the medium-risk assessment samples of dried tangerine peel batches, the target spectral range was located at 1650 cm⁻¹. -1 1450cm -1 and 1200cm -1The corresponding molecular structure response modes obtained from the knowledge base were protein amide I vibration, cellulose CH bending, and polyphenol CO stretching, respectively. Cosine similarity calculation was performed between the target spectral peak shape parameter vector and the knowledge base peak shape parameter vector, with a similarity threshold of 0.85. The results were 0.92, 0.88, and 0.90, respectively, all meeting the screening criteria. Molecular dynamics simulations were performed on the highly matched molecular structures, with binding energies of -35.4 kJ / mol, -28.5 kJ / mol, and -30.2 kJ / mol, respectively, all below the -25 kJ / mol threshold, thus entering the candidate set of binding sites. Combined with microscopic CT texture coordinate mapping, the protein amide I site was located in the cell wall protein region, the cellulose CH bending site in the cell wall fiber bundle, and the polyphenol CO stretching site in the intercellular matrix. The final generated potential binding site prediction list contained three spatially localizable structural records. Validation showed that this prediction list significantly improved the model's cross-medicinal material adaptability in contamination level interpretability analysis.

[0147] S6.4: Perform natural language synthesis based on the potential binding site prediction list and the standard toxicology description template to convert each potential binding site into a potential binding site prediction text containing a specific biochemical mechanism description, thereby constructing an interpretable text element set.

[0148] A matching retrieval index is established based on a potential binding site prediction list and a pre-set standard toxicological description template. The spatial coordinate labels of each binding site are mapped one-to-one with the corresponding biochemical mechanism entries in the template library, generating binding site data objects with mechanism codes. Feature retrieval is performed on the spectral interval index values ​​contained in the binding site data objects to extract toxicological keywords and reaction type parameters related to that interval, forming structured biochemical feature key-value pairs. These biochemical feature key-value pairs are input into a natural language generation algorithm module, which uses rule-based template filling logic to semantically render the mechanism codes, ensuring that the synthesized text contains three types of information: reaction type, molecular mechanism of action, and potential health effects. A multi-source feature weighting strategy is introduced during text synthesis. Multiplicative normalization is performed on the contribution weights of spectral peak positions and the significance factors of biochemical mechanisms to adjust the priority of each mechanism description in the text. The normalized weights are embedded in the syntactic structure selection module of the natural language generation process to achieve dynamic ordering of different mechanism descriptions. Terminology standardization is performed on the generated binding site texts, replacing non-standard terms and symbols using a standard toxicological terminology library to ensure that the output content conforms to professional expression standards. Through the above processing method, the binding site prediction results of the previous step are transformed into text elements containing descriptions of reaction type, molecular action mode and potential health effects, thereby realizing the construction of an interpretable text element set.

[0149] S6.5: Integrate the preliminary judgment results of the current aflatoxin contamination level, the high-discrimination spectral feature set, and the interpretable text element set into a structured encapsulation process to output a dynamic contamination judgment report containing the level conclusion, key discrimination bands, and biochemical mechanism description, thus completing the final decision output of the risk assessment module.

[0150] The preliminary determination of the current aflatoxin contamination level, the high-discrimination spectral feature set, and the interpretable text element set are subjected to multi-domain data structuring and encapsulation processing. Input conditions include contamination level label data, spectral interval index vectors, and a set of biochemical mechanism description texts. A hierarchical encapsulation algorithm is used to map the level labels to the core field areas of the risk assessment report structure and associate them with the band index parameters of the high-discrimination spectral feature set, constructing a feature metadata area containing an index-to-band mapping table. Natural language encoding and metadata binding are performed on the interpretable text element set, presenting the biochemical mechanism description as an independent interpretable annotation field in the report structure and associating it with the corresponding spectral interval through bidirectional reference pointers. Field validation rules are used to verify the consistency of multi-domain data elements in the structured report, including correlation verification between level labels and spectral intervals, and matching degree verification between spectral intervals and biochemical mechanism descriptions. The verification results form a publishable status flag.

[0151] By linking multi-domain data and encapsulating it in a structured manner, the analysis results from the previous step are transformed into dynamic pollution assessment report data objects, enabling the orderly presentation of level conclusions and interpretable elements in a unified report structure and enhancing traceability.

[0152] For example, in a traditional Chinese medicine (TCM) storage monitoring system, if the preliminary assessment of the current batch of medicinal materials as high-risk indicates a high level of contamination, the corresponding high-discrimination spectral feature set includes three band index values: 1450nm, 1680nm, and 1720nm. The interpretable text element set includes statements such as "the change in the 1450nm peak suggests an abnormal water film structure, which may enhance toxin adsorption," "the increased peak intensity at 1680nm represents the binding of lipid groups to toxin molecules," and "the peak shift at 1720nm corresponds to covalent modification of the protein backbone." During encapsulation, the high-risk level label is filled into the core field area of ​​the report, and the three band indexes are added to the feature metadata area in the form of key-value pairs <index:band>. In the process of associating the biochemical mechanism description, a bidirectional reference table is established, so that each band index corresponds to a specific text description, and the band index reference is embedded in the text description to facilitate location during report parsing. The verification rules confirmed a positive correlation between the high-risk level and the spectral range in the toxicology knowledge base, and the biochemical mechanism descriptions all exceeded the set threshold of 0.85 in content matching score. After verification, a publishable status flag was generated. This packaged report provides a complete level determination conclusion, clear spectral discrimination basis, and precise biochemical mechanism explanation when accessed by the system, effectively improving the interpretability and credibility of the pollution level determination.

[0153] Step S7: Determine whether the pollution level in the dynamic pollution assessment report exceeds a preset safety threshold. If it does, immediately trigger a graded early warning command and lock the corresponding cargo location. If it does not exceed the threshold, record the de-perturbed spectral residual sequence and perturbation factor of this assessment into the historical pollution response curve database for model iteration and optimization. Specifically, this includes: S7.1: Analyze the confidence vector of pollution level attribution in the dynamic pollution assessment report, extract the current aflatoxin pollution level label corresponding to the maximum probability value, and generate the real-time pollution status identifier to be determined.

[0154] S7.2: Based on the real-time pollution status identifier and the preset multi-level safety threshold matrix, perform numerical comparison calculation to calculate the difference coefficient of the pollution level deviating from the safety baseline, so as to generate a Boolean value of risk judgment result containing the over-limit flag.

[0155] A numerical comparison index is established based on the real-time pollution status identifier to be determined and the preset multi-level safety threshold matrix. The pollution level label in the real-time pollution status identifier is mapped to the corresponding level numerical code, forming a quantifiable risk assessment input parameter.

[0156] The threshold matrix retrieval operator is invoked to locate the security baseline value corresponding to the numerical code of the corresponding level in the multi-level security threshold matrix. The corresponding security baseline value is used as a reference parameter for risk assessment.

[0157] Execute the difference coefficient calculation module, inputting the pollution level numerical code and the safety baseline value into the difference calculation formula: ; Where L is the numerical code of the current pollution level, B is the safety baseline value, the numerator represents the deviation of the level, and the denominator represents the normalization factor of the total level.

[0158] The difference coefficient calculation result is compared with the corresponding difference threshold in the multi-level safety threshold matrix. The logical judgment operator is called to generate an over-limit flag. If the difference coefficient is greater than the difference threshold, the flag is set to true; otherwise, the flag is set to false.

[0159] The over-limit flag is encapsulated as a Boolean value representing the risk assessment result and output to the subsequent graded early warning triggering module or historical data recording module.

[0160] By calculating the difference coefficient and comparing the threshold, the real-time pollution status identifier from the previous step is transformed into a Boolean value containing the over-limit flag, thereby achieving accurate quantitative determination of the risk status.

[0161] For example, in a dynamic contamination assessment report for a batch of Codonopsis pilosula medicinal materials, the real-time contamination status is identified as "medium risk," with a mapped level numerical code of 3. The multi-level safety threshold matrix sets the safety baseline value for the medium risk level to 2.5 and the difference threshold to 0.05. The difference coefficient calculation formula is called, with input parameters L=3 and B=2.5. The calculation process is as follows: ; The numerator is 0.5, the denominator is 5.5, and the coefficient of variation is approximately 0.0909. Comparing this to the difference threshold of 0.05, the result is greater than the threshold, so the generated out-of-limit flag is set to true, the risk assessment result (Boolean value) is set to true, and the result is output to the tiered early warning trigger module. In this batch assessment, the system significantly improves the accuracy of the early warning response through numerical comparison calculations, and ensures that out-of-limit status can be determined using a unified standard in a dynamic warehousing environment, effectively avoiding misjudgments caused by drift in assessment rules between different batches.

[0162] S7.3: If the risk assessment result Boolean value indicates an over-limit state, a graded early warning instruction code of specific intensity is generated according to the difference coefficient mapping, and the warehouse control system interface is called to execute the electronic locking action of the corresponding storage location, so as to output the high-risk storage location locking status of physical isolation.

[0163] Step S8: If the current aflatoxin contamination level does not exceed the preset safety threshold, the de-perturbation spectral residual sequence and perturbation factor are recorded in the historical contamination response curve database, and the standard intrinsic spectral reference vector is periodically updated to maintain the adaptive accuracy of the dynamic contamination level determination mechanism under different storage conditions. Specifically, this includes: S8.1: If the risk judgment result Boolean indicates that the state is not exceeded, the generated de-perturbation spectral residual sequence and the synchronously collected storage perturbation factor are spatiotemporally aligned and packaged to construct a structured historical monitoring data tuple with timestamps, so as to form a standard training sample set for model iterative optimization.

[0164] S8.2: Write the standard training sample set into the incremental storage area of ​​the historical pollution response curve database and trigger the database version update mechanism to complete the persistent archiving of the historical pollution response curve dataset that supports adaptive accuracy maintenance.

[0165] S8.3: Obtain the newly added de-perturbed spectral residual sequence and the corresponding batch identification information from the historical pollution response curve database, and use the sliding time window algorithm to perform time-series aggregation processing on the newly added de-perturbed spectral residual sequence to generate a set of dynamic spectral feature vectors characterizing the recent state changes of the current batch of medicinal materials.

[0166] S8.4: Based on the dynamic spectral feature vector set, call the intrinsic spectral fingerprint database retrieval interface, extract the standard intrinsic spectral reference vector of the corresponding batch identifier as the benchmark comparison object, use the cosine similarity measurement algorithm to calculate the multidimensional spatial distance between the dynamic spectral feature vector set and the standard intrinsic spectral reference vector, and generate drift detection intermediate results including similarity numerical ranking.

[0167] S8.5: Set an adaptive drift judgment threshold based on the similarity value ranking in the intermediate drift detection results, perform abnormal pattern clustering analysis on dynamic spectral feature vectors that are lower than the adaptive drift judgment threshold, identify key band offset intervals that cause significant drift in spectral response modes, and generate a spectral drift diagnostic report containing drift type classification labels.

[0168] S8.6: Select the corresponding vector fusion strategy based on the drift type classification label in the spectral drift diagnosis report, and perform linear interpolation fusion operation on the set of dynamic spectral feature vectors that have undergone confidence weighting and the standard intrinsic spectral reference vector to generate an updated candidate standard intrinsic spectral reference vector.

[0169] S8.7: Perform smoothing filtering and normalization verification processing on the candidate standard intrinsic spectral reference vector to eliminate high-frequency noise interference introduced during the fusion process and ensure energy conservation constraints. Write the final vector that passes the verification into the intrinsic spectral fingerprint database to replace the original record and complete the adaptive update operation of the standard intrinsic spectral reference vector of a specific batch of medicinal materials.

[0170] For those skilled in the art, various other corresponding changes and modifications can be made based on the technical solutions and concepts described above, and all such changes and modifications should fall within the protection scope of the claims of this invention.

[0171] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains. The terms “first,” “second,” “third,” and similar terms used in this patent application specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “comprising” or “including” and similar terms mean that the elements or objects preceding “comprising” or “including” encompass the elements or objects listed following “comprising” or “including” and their equivalents, and do not exclude other elements or objects. The “multiple” mentioned in the embodiments of this application refers to two or more. A and / or B indicate three possibilities: A; B; and A and B.

[0172] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials, characterized in that, Specifically, it includes: S1: Obtain batch metadata and multimodal spectral raw data of Chinese medicinal materials to form an initial spectral dataset; S2: Based on the origin, harvesting season and primary processing method metadata contained in the initial spectral dataset, call the pre-constructed intrinsic spectral fingerprint database for matching and retrieval, and generate the standard intrinsic spectral reference vector of the corresponding medicinal material matrix. S3: Real-time acquisition and calculation of the nonlinear influence weight of the perturbation factor on the spectral signal, generating a perturbation feature mask matrix; S4: Apply the perturbation feature mask matrix to the initial spectral dataset to perform band-by-band stripping processing, and output the perturbation-free spectral residual sequence; S5: Based on the microscopic tissue texture topology obtained from previous micro-CT reconstruction, a graph neural network model is constructed. The perturbation-free spectral residual sequence is mapped as node attributes and input into the model. The response mode of adsorption sites of toxin molecules in the medicinal matrix is ​​learned through message passing mechanism, and a pollution level attribution confidence vector and a heat map of contribution of key discrimination bands are generated. S6: Determine the current aflatoxin contamination level based on the maximum probability value in the contamination level attribution confidence vector, and extract one or more key discriminant spectral intervals and their corresponding potential binding site inference texts by combining the key discriminant band contribution heatmap to form a dynamic contamination determination report; S7: If the current aflatoxin contamination level exceeds the preset safety threshold, a graded early warning command will be immediately triggered and the corresponding cargo location will be locked.

2. The method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials according to claim 1, characterized in that, After step S7, the following also includes: S8: If the current aflatoxin contamination level does not exceed the preset safety threshold, the de-perturbation spectral residual sequence and perturbation factor are recorded in the historical contamination response curve database, and the standard intrinsic spectral reference vector is updated periodically.

3. The method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials according to claim 2, characterized in that, The periodically performed standard intrinsic spectral reference vector update includes: periodically comparing the similarity ranking of newly added samples in the historical pollution response curve database with the intrinsic spectral fingerprint database; if the spectral response mode of a specific batch of medicinal materials drifts significantly, the standard intrinsic spectral reference vector of that batch will be automatically updated.

4. The method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials according to claim 1, characterized in that, The disturbance factors include: temperature and humidity of the storage environment, light intensity, and air cleanliness parameters.

5. The method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials according to claim 4, characterized in that, Step S3 specifically includes: The temperature, humidity, light intensity, and air cleanliness parameters of the storage environment are acquired in real time. Based on the preset mapping relationship between the water film thickness on the surface of the medicinal material and the near-infrared absorption peak position shift, the temperature and humidity parameters of the storage environment are subjected to nonlinear fitting processing to generate a water film thickness interference coefficient vector. The system receives the water film thickness interference coefficient vector and the light intensity, and uses the UV aging-induced fluorescence background rise quantification model to calculate the time accumulation effect of the light intensity, generating a fluorescence background rise weight sequence. The water film thickness interference coefficient vector, fluorescence background rise weight sequence, and real-time air cleanliness parameters are called. Based on the mapping table of standard dust deposition amount and spectral signal-to-noise ratio attenuation, the scattering loss of the real-time air cleanliness parameters is estimated, and a dust deposition attenuation factor array is generated. The water film thickness interference coefficient vector, fluorescence background rise weight sequence, and dust deposition attenuation factor array are integrated and subjected to band alignment and weighted fusion processing to generate a comprehensive environmental disturbance feature tensor containing full-band interference distribution information. Based on the comprehensive environmental perturbation feature tensor, normalization inversion and sparse encoding operations are performed to generate the perturbation feature mask matrix.

6. The method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials according to claim 5, characterized in that, The raw multimodal spectral data includes near-infrared reflectance spectral signals in the 400 to 2500 nanometer band and Raman scattering signals in the 800 to 1800 centimeter range.

7. The method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials according to claim 6, characterized in that, Step S4 specifically includes: The initial spectral dataset and perturbation feature mask matrix are obtained. Based on the band alignment mechanism, the water film thickness interference coefficient vector, fluorescence background lifting weight sequence and dust deposition attenuation factor array are mapped to the near-infrared reflectance spectral signal band and Raman scattering signal band corresponding to the initial spectral dataset, respectively, to generate a band-level interference mapping tensor. The nonlinear baseline fitting process is performed on the band-level interference mapping tensor, and the absorption peak shift caused by the change in water film thickness in each band and the fluorescence background rise curve caused by the duration of illumination are calculated using a polynomial regression algorithm to generate the environmental interference synthetic substrate spectrum. Based on the environmental interference synthesized substrate spectrum, a point-by-point subtraction operation is performed on the initial spectral dataset to remove the water film absorption component from the original near-infrared reflectance spectral signal and to subtract the fluorescence background component from the original Raman scattering signal, thereby generating an intermediate denoised spectral sequence. The dust scattering correction algorithm is applied to the intermediate denoised spectral sequence. Based on the dust deposition attenuation factor array, the signal amplitude is inversely compensated to repair the spectral signal-to-noise ratio attenuation region caused by particulate matter occlusion and generate the restored spectral feature vector. Based on the standard intrinsic spectral reference vector, residual calculation processing is performed on the restored spectral feature vector to extract and filter out the residual fluctuation components in the restored spectral feature vector that are unrelated to the medicinal material matrix, and output the perturbation-free spectral residual sequence that retains only the microstructure texture and toxin adsorption characteristics of the medicinal material.

8. The method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials according to claim 1, characterized in that, Step S5 specifically includes: Based on the microscopic tissue texture topology of medicinal materials obtained from previous micro-CT reconstruction, the node set and edge connection relationship are initialized and defined, and the data of each band in the perturbation-free spectral residual sequence are mapped to node attribute vectors to construct a medicinal material spectral topology map data object containing a spatial adjacency matrix and a feature attribute matrix. Perform multiple rounds of neighborhood aggregation processing on the spectral topology map data object of the medicinal materials to generate high-order embedding representation vectors for the nodes; Based on the node high-order embedded representation vector, an attention mechanism is introduced to calculate the edge weight coefficient, and the message passing path in the medicinal material spectral topology map data object is dynamically optimized to screen out key feature channels that are strongly correlated with aflatoxin adsorption characteristics and generate a sparse graph feature tensor. The sparse graph feature tensor is pooled by a global readout function to aggregate the scattered node high-order embedding representation vectors into a fixed-dimensional graph-level global feature vector, and then mapped to the pollution level classification space through a fully connected layer to generate a pollution level attribution confidence vector. Based on the gradient backpropagation path in the confidence vector of pollution level attribution, the contribution weight values ​​of each band data to the pollution level classification result are extracted, the contribution weight values ​​are mapped back to the original spectral band index and superimposed on the spatial coordinates of the microstructure texture topology of the medicinal material, and a heat map of the contribution of the key discriminant bands containing the key discriminant spectral intervals and their potential binding sites is generated.

9. The method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials according to claim 8, characterized in that, The step of performing multiple rounds of neighborhood aggregation processing on the medicinal herb spectral topology map data object to generate a high-order embedding representation vector for the node is as follows: using a graph convolution operator to perform multiple rounds of neighborhood aggregation processing on the medicinal herb spectral topology map data object, weighting and fusing the feature attribute matrices of adjacent nodes to update the hidden state of the central node, and generating the high-order embedding representation vector for the node.

10. The method for monitoring and analyzing the risk of aflatoxin contamination in the storage of traditional Chinese medicinal materials according to claim 1, characterized in that, The pollution level assignment confidence vector includes four states: safe, low risk, medium risk, and high risk; the dynamic pollution assessment report includes the level conclusion and interpretable elements.