Component detection method suitable for herbaceous plant essence

By combining a theoretical Raman spectroscopy database with a convolutional neural network model, the problem of isomer and stereoisomer identification in herbal plant extracts was solved, achieving high-precision component analysis.

CN121899111AInactive Publication Date: 2026-04-21SHAANXI UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHAANXI UNIV OF SCI & TECH
Filing Date
2026-03-19
Publication Date
2026-04-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively distinguish between structurally similar isomers and stereoisomers in the detection of herbal plant extracts, resulting in insufficient identification accuracy.

Method used

A theoretical Raman spectroscopy database was constructed, and the vibrational modes of candidate compounds were simulated using quantum chemical calculations. A convolutional neural network model was used to extract and compare deep features of the measured Raman spectra, and the accuracy of the identification results was ensured through dual verification conditions.

Benefits of technology

It enables high-precision identification of isomers and stereoisomers, improving the accuracy and reliability of herbal plant essence component analysis and avoiding the risk of misjudgment caused by a single verification method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121899111A_ABST
    Figure CN121899111A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of analytical chemistry, and discloses a component detection method suitable for herbaceous plant essence, and the method comprises the following steps: simulating a theoretical Raman spectrum of a candidate compound by using a quantum chemistry calculation module to construct a database; collecting the actually measured Raman spectrum of the sample to be detected; spectral deep features are extracted through a convolutional neural network model and matched and compared; and outputting an identification result based on dual verification of a matching degree threshold value and a feature similarity threshold value of the theoretical vibration peak position and the actually measured spectrum peak position. The method provided by the invention can solve the problem of insufficient discrimination of isomers and stereoisomers with similar structures in the prior art, and realizes accurate identification of structural analogues in a complex system without complete separation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of analytical chemistry, and more specifically to a method for detecting components in herbal plant extracts. Background Technology

[0002] Herbal plant extracts are an important subject of natural product research, and their component analysis has wide applications in drug development, quality control, and functional food development. Traditional chromatography-mass spectrometry (GC-MS) combines physical separation with mass-to-charge ratio detection for qualitative and quantitative analysis of compounds in complex matrices. Its typical workflow includes sample pretreatment, chromatographic separation, mass spectrometry ionization, and data analysis, and has become one of the standard methods for identifying plant chemical components. This technique establishes a correlation between retention time and mass spectrometric fragments to achieve the identification and characterization of target compounds.

[0003] However, existing technologies have limitations in distinguishing structurally similar isomers and stereoisomers, making it difficult to achieve high-precision identification in complex systems. Summary of the Invention

[0004] This invention provides a method for detecting components in herbal plant extracts, which solves the problem of insufficient differentiation between structurally similar isomers and stereoisomers using traditional techniques. To achieve the above objective, this invention provides the following technical solution:

[0005] This invention provides a method for detecting components in herbal plant extracts, comprising the following steps: S1. Theoretical Spectral Database Construction Steps: For the candidate compounds that may be contained in the target herbal plant extract, the quantum chemical calculation module is used to simulate the theoretical vibrational modes of each candidate compound under Raman spectroscopy, generate the corresponding theoretical Raman spectra, and construct the theoretical Raman spectral database. S2. Sample Spectrum Acquisition Steps: Raman spectroscopy was performed on the herbal plant essence sample to obtain the measured Raman spectrum; S3. Fusion, Comparison, and Component Identification Steps: The measured Raman spectra are input into a trained convolutional neural network model. The convolutional neural network model uses data from the theoretical Raman spectroscopy database as its training basis to extract deep features of the spectra. The model matches and compares the deep features of the measured Raman spectra with the features in the theoretical Raman spectroscopy database. S4. Dual Verification Result Output Steps: Based on the matching and comparison results, output the component identification results; among which, the validity of the identification results must simultaneously meet the matching degree threshold between the theoretically calculated and simulated vibration peak position and the measured spectral peak position, as well as the feature similarity threshold output by the convolutional neural network model, to complete the dual verification of calculation and measurement.

[0006] In an optional embodiment, the S1 theoretical spectral database construction step further includes: S1a. Dynamic Evolutionary Learning Module: Includes a meta-learning algorithm, enabling the quantum chemistry calculation module to automatically adjust calculation parameters and basis set settings based on the systematic deviations between measured and theoretical spectra in historical identification results, in order to iteratively optimize the accuracy of subsequent simulations; S1b. Synthetic Biology Auxiliary Calibration Module: For key active ingredients that are predicted to exist by quantum chemical calculations but lack physical standards and whose matching confidence level is higher than a preset threshold, the corresponding chemical structure information is input into the automated synthetic biology design platform. Based on gene circuit and metabolic pathway algorithms, this platform designs and guides microbial factories to rapidly synthesize trace amounts of target compounds, which serve as physical standards for the final verification and calibration of theoretical spectra.

[0007] In one optional embodiment, in the S2 sample spectral acquisition step: A nonlinear optical module integrating coherent anti-Stokes Raman scattering or stimulated Raman scattering is used for detection to completely suppress fluorescence background interference; Meanwhile, it includes an adaptive optics module, which uses a wavefront sensor to detect and compensate for optical aberrations caused by non-uniformity or micro-flow of the sample solution in real time. By controlling a deformable mirror to dynamically optimize the laser beam wavefront, it ensures the acquisition of a stable Raman signal with high signal-to-noise ratio and high spatial resolution under complex physical conditions.

[0008] In an optional embodiment, the convolutional neural network model in the S3 fusion alignment and component identification step is constructed as a collaborative system of a continuous learning framework and a causal discovery module: The continuous learning framework employs an elastic weight consolidation algorithm, enabling the model to learn the spectral characteristics of new compounds online without forgetting previously learned knowledge when encountering new herb samples not recorded in the original theoretical database, and to dynamically expand its recognition capabilities. The causal discovery module is based on the causal structure learning algorithm. It analyzes the statistical dependence between different spectral features and known component concentrations in a large number of mixed samples, infers and visualizes the potential causal networks between each spectral peak and between spectral peaks and components, and is used to verify and interpret the matching results of the neural network, reducing the risk of misjudgment caused by accidental spectral overlap.

[0009] In an optional embodiment, the S4 double verification result output step is further connected to a blockchain-enabled distributed verification and feedback network: Each successful computational and experimental double verification result, including the identified components, the corresponding theoretical / experimental spectral fingerprints, matching parameters, and sample source metadata, is generated into an encrypted hash value and stored in a permissioned blockchain network to ensure that the data is tamper-proof and traceable throughout the entire process. The network has a built-in global model optimization program based on federated learning: Under the coordination of the blockchain smart contract, each participating node only shares the encrypted model parameter updates, jointly trains a more powerful and general global spectral recognition model, and synchronously feeds the optimized model parameters back to the local convolutional neural network model of each node.

[0010] In an optional embodiment, a microfluidic chip preprocessing and spectral enhancement step is included before the S2 sample spectral acquisition step: The herbal extract sample to be tested was injected into an integrated microfluidic chip, which contained: Surface acoustic wave enrichment module: Utilizes an acoustic field to enrich the target active ingredient from a complex matrix into the optical detection cavity within the chip; Plasma resonance enhancement unit: The inner wall of the optical detection cavity is modified with an array of gold / silver nanoparticles optimized by a genetic algorithm to generate a local surface plasmonic resonance effect, thereby specifically enhancing the Raman scattering signal intensity of the target component.

[0011] This step achieves pre-screening and signal amplification of target components at the physical level, providing enhanced input signals for subsequent high-precision identification.

[0012] In an optional embodiment, the S3 fusion alignment and component identification step further includes a digital twin-driven virtual control experiment module: This module is activated when there is uncertainty or controversy regarding the identification results of a certain component; Based on the known components and physicochemical parameters of the sample being tested, this module constructs a digital twin of the sample in digital space. By using mechanistic models and Monte Carlo simulations, controversial components are virtually added or removed from the digital twin, and the corresponding expected Raman spectral changes are calculated. By comparing the expected spectral changes obtained from virtual experiments with the differences in measured spectra, third-party digital simulation evidence is provided for the existence and content of controversial components, forming a triple verification system of theory, experiment, and simulation.

[0013] In an optional embodiment, the method further includes an efficacy-oriented intelligent report generation step S5: S5a. Knowledge Graph Query: Input the list of components identified in step S4 into an interdisciplinary herbal medicine knowledge graph. This graph integrates pharmacology, phytochemistry, and clinical research databases, and automatically associates and extracts information on the known efficacy, target of action, synergistic combination, and potential side effects of each component. S5b. Natural Language Generation and Visualization: Based on graph neural networks, the query results are ranked by importance and relationships are sorted out, driving a natural language generation engine to automatically generate a structured test report containing key functional ingredient groups, synergistic network diagrams, quality evaluation and recommendations, directly transforming chemical composition data into biological and product insights that can be used for decision-making.

[0014] In one alternative embodiment, the adaptive optics module is coupled to a computational imaging algorithm library: This algorithm library includes compressed sensing, phase retrieval, and deep learning super-resolution reconstruction algorithms; While performing hardware-level aberration correction using adaptive optics, the algorithm library performs secondary calculations on the acquired raw light field information, further breaking through the optical diffraction limit at the software level. This enables sub-micron spatial resolution imaging of chemical composition microscopic distribution, allowing the method to not only identify components but also visually present the original spatial distribution of each component in plant cells or formulation particles.

[0015] In one alternative embodiment, a smart quality contract auto-execution system based on oracles is also deployed in the blockchain-enabled distributed verification and feedback network: The system predefines quality standards for different types of herbal extracts; Once new test data, along with its blockchain-based notarized data, is submitted to the network, the oracle automatically retrieves the data and the smart contract deployed on the chain compares it with preset quality standards. The comparison results automatically trigger corresponding on-chain operations: if the standard is met, a digital quality certification certificate with a timestamp and an immutable hash value is generated and broadcast; if the standard is not met, the relevant production nodes are automatically notified and the traceability process is initiated, realizing full automation and intelligence of testing, certification and quality control.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention provides a method for detecting components in herbal plant extracts. The method constructs a theoretical Raman spectral database by simulating the theoretical vibrational modes of candidate compounds under Raman spectroscopy using a quantum chemical calculation module. This enables precise characterization of the vibrational properties of isomers and stereoisomers, providing high-resolution spectral evidence for distinguishing structurally similar compounds. A convolutional neural network model trained based on the theoretical Raman spectral database extracts deep features from measured Raman spectra and performs matching comparisons. Utilizing the nonlinear feature extraction capabilities of deep learning, it effectively identifies subtle spectral differences, significantly improving the accuracy of identifying target components in complex mixtures. 2. By setting dual verification conditions—the matching threshold between the theoretically calculated and simulated vibration peak position and the measured spectral peak position, and the feature similarity threshold output by the convolutional neural network model—the identification results are ensured to meet the reliability requirements of both the physical model and the data-driven model, thus avoiding the risk of misjudgment caused by a single verification method.

[0017] 3. It improves the problem of insufficient differentiation between structurally similar isomers and stereoisomers in existing technologies, and ultimately achieves the technical effect of accurately identifying structurally similar substances in complex systems without complete separation.

[0018] 4. This invention establishes a computational-experimental dual verification mechanism by integrating quantum computing simulation and deep learning. While retaining the advantages of non-destructive testing of traditional spectroscopic technology, it breaks through the technical bottleneck of distinguishing structurally similar substances, providing an innovative solution for high-precision component analysis of herbal plant extracts. Attached Figure Description

[0019] Figure 1 This is an overall flowchart of the component detection method of the present invention; Figure 2 This is a flowchart illustrating the steps involved in constructing the S1 theoretical spectral database of the present invention. Figure 3 This is a flowchart of the S2 sample spectral acquisition steps of the present invention; Figure 4 This is a flowchart of the S3 fusion comparison and component identification steps of the present invention; Figure 5 This is a flowchart of the steps for outputting the S4 double verification results of the present invention; Figure 6 This is a flowchart of the S5 efficacy-oriented intelligent report generation steps of the present invention; Detailed Implementation Please refer to Figures 1 to 6 The present invention will now be described in further detail with reference to embodiments. It is to be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it.

[0020] Example 1: Traditional chromatography-mass spectrometry (GC-MS) technology has limited ability to distinguish between isomers and stereoisomers with highly similar structures in the detection of herbal plant extracts. At the same time, it relies on complex chemical pretreatment and standard comparison, resulting in long detection cycles, high operational thresholds, and difficulty in achieving rapid in-situ screening.

[0021] To address the above problems, this invention provides a method for detecting the components of herbal plant extracts, comprising the following steps: Step 1: Theoretical Spectral Database Construction Steps: For the candidate compounds that may be contained in the target herbal plant extract, the quantum chemical calculation module is used to simulate the theoretical vibrational modes of each candidate compound under Raman spectroscopy, generate the corresponding theoretical Raman spectra, and construct a theoretical Raman spectral database. The quantum chemical calculation module is a functional unit that numerically solves the electronic structure and nuclear coordinate response of a molecular system based on first-principles calculations (such as density functional theory, DFT). Its input is the three-dimensional chemical structure information of the candidate compound, and its output is the Raman activity, frequency position, and relative intensity of each vibrational mode of the molecule at a specific excitation wavelength. In this embodiment, this module is used to generate theoretical spectral data that is physically interpretable and independent of physical standards, serving as a benchmark for subsequent deep learning model training and comparison. Theoretical Raman spectroscopy is a digital representation of the intrinsic vibrational degrees of freedom of a molecule and its Raman scattering selection rules. It can include information such as peak position, full width at half maximum (FWHM), relative intensity, and polarization dependence, and can characterize the structural differences of candidate compounds at the atomic scale. The theoretical Raman spectroscopy database is a structured collection of theoretical Raman spectra corresponding to multiple candidate compounds. Its data format supports indexing and retrieval by compound ID, functional group type, molecular weight range, or structural similarity, providing an scalable and updatable reference base for the aforementioned feature matching.

[0022] In one alternative implementation, the theoretical spectrum generation method can be: based on the B3LYP functional and the 6-31G(d) basis set, perform geometric optimization and frequency analysis on the candidate compound under a gas phase or implicit solvent model, and output the theoretical spectrum after zero-point energy correction and Boltzmann weighting. In another alternative implementation, the theoretical spectrum generation method may include: employing a multi-scale modeling strategy, first using a semi-empirical method (such as PM6) to complete the initial conformation search for candidate compounds containing large conjugated systems or metal coordination structures, and then calling a higher-precision DFT method to complete the final spectrum calculation for low-energy conformation clusters; Furthermore, this theoretical spectrum generation method can also employ the following approach: introduce temperature and pressure parameters, calculate multiple sets of theoretical spectra in parallel under different thermodynamic conditions, and classify and store the results according to condition labels to enhance the database's adaptability to changes in the actual detection environment.

[0023] This invention generates theoretical Raman spectra through a quantum chemical calculation module, which not only avoids the difficulty of preparing physical standards, but also provides the database with clear physical mechanism support. The theoretical Raman spectrum database built on this basis provides a high-quality source of supervisory signals with causal traceability for subsequent deep learning models, thereby supporting the accurate differentiation of structurally similar objects.

[0024] Step 2: Spectral Acquisition of the Sample to be Tested: Raman spectroscopy was performed on the herbal plant essence sample to obtain the measured Raman spectrum; Among them, the measured Raman spectrum refers to the intensity of scattered light collected under actual experimental conditions after the sample is excited by a laser, as a function of the Raman shift (cm). -1 The digital spectral curves of the distribution can reflect the superposition response of various chemical bond vibration modes in the sample; the spectrum directly carries the composition and microstructure information of the sample under test, and is a key observational data connecting theoretical predictions and actual material states.

[0025] In one alternative implementation, the Raman spectroscopy detection method can be: using a 785nm near-infrared laser as the excitation source, combined with a fiber-coupled Raman probe and a high-pass filter to suppress Rayleigh scattering, and using a back-illuminated CCD detector to collect data at 500–2000 cm⁻¹. -1 The scattered signals within the range are preprocessed by dark current subtraction and smoothing / denoising to form the original spectrum; In another alternative implementation, the Raman spectroscopy detection method may include: integrating micro-area focusing and automatic platform scanning functions on the basis of conventional Raman detection, performing multi-point sampling of heterogeneous samples and taking the average spectrum to reduce local matrix effect interference; Furthermore, this Raman spectroscopy detection method can also employ the following approach: applying a controllable temperature control module to the sample slide and completing the acquisition under constant temperature (25℃±0.5℃) and light-protected conditions to reduce the impact of thermal drift and photodegradation on spectral stability.

[0026] This invention establishes a reliable mapping channel from real material systems to digital spectroscopic characterization by acquiring measured Raman spectra. The data obtained in this step serves as the input to the aforementioned model, and its signal-to-noise ratio and peak shape fidelity directly affect the accuracy and robustness of component identification.

[0027] Step 3: Fusion Comparison and Component Identification: The measured Raman spectrum is input into the trained convolutional neural network model. The convolutional neural network model uses data from the theoretical Raman spectroscopy database as its training basis to extract deep features of the spectrum. The model matches and compares the deep features of the measured Raman spectrum with the features in the theoretical Raman spectroscopy database. The convolutional neural network model is a deep learning architecture with hierarchical feature extraction capabilities. Its input is a one-dimensional spectral vector (the horizontal axis represents Raman shift, and the vertical axis represents intensity), and its output is the probability distribution or embedding space coordinates corresponding to each candidate compound. In this embodiment, the model does not directly fit peak shifts or intensity scaling, but instead learns higher-order statistical laws such as the global spectral morphology, multi-peak coupling relationships, and weak feature responses. Deep features refer to the low-dimensional dense representations obtained after multiple convolutions and nonlinear transformations of the model. These features can characterize combinations of patterns in the spectrum that are not easily identified manually but have component-specific characteristics, such as the joint vibrational fingerprints corresponding to the linkage mode of a certain type of flavonoid aglycone and glycosyl group. Matching and alignment refer to calculating the distance metric (such as cosine similarity or Euclidean distance) between the measured spectral feature vector and each theoretical spectral feature vector in the embedding space, and ranking the probability of candidate components accordingly.

[0028] In one alternative implementation, the matching and comparison method can be: inputting the measured spectrum and all theoretical spectra into a Siamese convolutional neural network with shared weights, outputting two embedding vectors, and then calculating the similarity score between the two through a learnable distance metric function; In another alternative implementation, the matching and comparison method may include: performing clustering compression on the theoretical spectral database in advance, constructing a hierarchical hash index structure, and performing coarse-grained retrieval followed by fine matching in the matching stage to improve the comparison efficiency under large-scale databases. Furthermore, this matching and comparison method can also employ the following approach: introducing an attention mechanism to dynamically weight the feature contributions of different Raman shift intervals at the end of the convolutional network, so that the model pays more attention to feature regions sensitive to structural differences (such as the C=O stretching vibration region and the aromatic ring breathing vibration region) during the comparison process.

[0029] This invention achieves collaborative analysis from physical modeling to data-driven analysis by fusing and comparing measured spectra with theoretical databases using a convolutional neural network model. With the help of deep feature extraction capabilities, the model can effectively suppress interference caused by fluorescence background, instrument noise and concentration fluctuations, and significantly enhance the ability to distinguish subtle structural differences such as isomers.

[0030] Step 4: Output of Dual Verification Results: Based on the matching and comparison results, output the component identification results; among which, the validity of the identification results must simultaneously meet the matching degree threshold between the vibration peak position calculated and simulated by theory and the measured spectral peak position, as well as the feature similarity threshold output by the convolutional neural network model, to complete the dual verification of calculation and measurement.

[0031] The matching degree between the theoretically calculated vibrational peak position and the measured spectral peak position refers to the comprehensive deviation index calculated after aligning the position of the main peak in the measured spectrum with the positions of the N strongest peaks in the corresponding theoretical spectrum one by one. This can be the weighted average absolute deviation (WMAE) or the dynamic time warping (DTW) distance, used to verify whether the spectral peak position conforms to the fundamental laws of quantum mechanics. The matching degree threshold is a preset upper tolerance limit (5cm). -1 If the threshold is exceeded, it indicates that there is a systematic deviation between the theoretical model of the candidate compound and the actual molecular vibrational behavior, and it should not be adopted; the feature similarity threshold refers to the lower limit of the cosine similarity (0.85) between the measured spectral embedding vector output by the convolutional neural network and the theoretical embedding vector of a candidate compound, which is used to verify whether the overall spectral morphology is consistent at the data-driven level; both thresholds must be met simultaneously for the component identification result to be considered valid.

[0032] In one alternative implementation, the dual verification method can be: first, the model outputs a Top-K candidate list, then calculates the peak position matching degree and feature similarity for each candidate, and only includes them in the final result set when both meet the criteria; In another alternative implementation, the dual verification method may include: weighting and fusing peak matching degree and feature similarity into a single confidence score, setting a joint threshold, and using an end-to-end differentiable design to make the model aware of the dual constraint requirements during the training phase. Furthermore, this dual verification method can also employ a hierarchical output mechanism: if peak position matching is satisfied but feature similarity is insufficient, it indicates that there may be impurity interference or the concentration is too low; if feature similarity is satisfied but peak position deviation exceeds the limit, it indicates that the theoretical model needs parameter optimization (such as basis set correction) and triggers the feedback path of the above-mentioned theoretical spectrum generation module.

[0033] This invention constructs a dual verification mechanism that combines interpretability and robustness by simultaneously applying physical constraints (peak matching) and data-driven constraints (feature similarity). This mechanism not only significantly reduces the risk of misjudgment caused by accidental spectral overlap or model overfitting, but also ensures that the identification results conform to the basic principles of quantum chemistry and are adapted to the complex spectral responses under actual detection conditions. Ultimately, it forms a computation-measurement closed-loop verification system that supports high-precision, separation-free, and traceable detection of structurally similar components in herbal plant extracts.

[0034] Example 2: In an optional embodiment, the present invention further provides the S1 theoretical spectral database construction step, which further includes: Step 1: S1a. Dynamic Evolutionary Learning Module: This module includes a meta-learning algorithm that enables the quantum chemistry calculation module to automatically adjust calculation parameters and basis set settings based on the systematic deviations between measured spectra and corresponding theoretical spectra in historical identification results, thereby iteratively optimizing the accuracy of subsequent simulations.

[0035] Meta-learning algorithms refer to a machine learning paradigm that learns how to learn. Their core capability lies in extracting transferable knowledge structures from multiple related tasks, enabling rapid adaptation to new tasks with a small sample size. In this embodiment, the algorithm does not directly participate in spectral prediction but models the systematic deviation patterns between measured spectra and corresponding theoretical spectra during historical identification processes. Examples include the overall redshift / blueshift trend of vibrational peaks of specific functional groups, peak width broadening consistency deviations, or intensity ratio mismatches. Its inherent function is to identify the distribution characteristics and potential causes of model errors and generate parameter correction strategies accordingly. In this embodiment, the algorithm feeds back the deviation analysis results to the quantum chemistry calculation module, driving it to dynamically update the exchange-correlation functional type, basis set shrinkage scheme (e.g., switching from 6-31G(d) to def2-TZVP), and solvation model parameters used in solving the Hartley-Fock equation. This improves the spectral fidelity of subsequent simulations for similar compounds while maintaining computational efficiency.

[0036] In one alternative implementation, the meta-learning algorithm can be a deviation perceptron built on a model-independent meta-learning (MAML) framework. It obtains initial weights sensitive to parameter space perturbations by performing inner-loop gradient updates and outer-loop aggregation optimizations on deviation subtasks of multiple validated compounds. In another alternative implementation, the meta-learning algorithm may include a lightweight graph neural network to encode the molecular topology, computational parameter combinations, and corresponding deviation vectors into a joint embedding representation, and to identify key atomic local environments that affect peak prediction through an attention mechanism. Furthermore, the meta-learning algorithm can also employ an online incremental update mechanism, triggering a mini-batch retraining whenever a newly added double-validated experimental-theoretical matching record is added, continuously strengthening the generalization recognition ability of novel deviation patterns.

[0037] This invention establishes a closed-loop learning path—from experimental feedback to bias modeling to parameter tuning to re-simulation—by embedding a meta-learning algorithm into the theoretical spectrum generation process. This enables the quantum chemical calculation module to no longer rely on fixed parameter configurations but instead has the ability to continuously calibrate its behavior based on real detection data, thereby ensuring that the theoretical Raman spectroscopy database maintains high-confidence prediction stability during long-term operation.

[0038] Step 2: S1b. Synthetic Biology Auxiliary Calibration Module: For key active ingredients predicted by quantum chemical calculations to exist but lacking physical standards and with a matching confidence level higher than a preset threshold, the corresponding chemical structure information is input into the automated synthetic biology design platform in the form of SMILES strings or InChI identifiers. This platform is an integrated software system that combines molecular modeling, pathway retrieval, gene circuit compilation, and fermentation control instruction generation. It consists of four core modules: compound syntheticity analysis module, metabolic pathway design module, gene circuit compilation module, and fermentation process generation module. Each module relies on the KEGG and MetaCyc natural product biosynthesis cloud databases to achieve data interoperability.

[0039] The platform hardware relies on an industrial-grade server with an Intel Xeon Gold 6330 CPU and an NVIDIA A100 GPU, running Linux CentOS 7.9 and employing a B / S architecture to support local and remote operation. Based on gene circuit and metabolic pathway algorithms, the platform uses biosyntheticity, host compatibility, and metabolic burden threshold as core constraints. First, a molecular fingerprint matching algorithm retrieves synthetic records of similar compounds and outputs a synthesability score. A score greater than or equal to 0.7 recommends *E. coli* BL21(DE3) or *Saccharomyces cerevisiae* CEN.PK2 as the optimal chassis strain. Then, a reverse pathway deduction algorithm breaks down the target compound into precursor metabolites, matching the endogenous metabolic pathway of the chassis strain and supplementing missing exogenous enzyme genes to construct a complete synthetic pathway, ensuring a carbon conversion rate greater than or equal to 30%. Subsequently, based on the synthetic pathway, the T7 / ADH1 constitutive promoter, a 0.5-2.0 strength ribosome binding site, and the rrnB / CYC1 terminator are selected to deliver the exogenous enzyme. Genes and regulatory elements were spliced ​​into a gene circuit and a GenBank-formatted DNA sequence file was generated. Codon bias was optimized to achieve a codon fitness index (CAI) greater than or equal to 0.8. Based on the growth characteristics of the chassis strain, standardized fermentation process parameters were generated, including LB / YPAD fermentation medium, initial pH 7.0-7.5, culture temperature 37℃ / 30℃, OD600 = 0.6-0.8 induction timing, 0.1-1.0 mmol / L IPTG inducer concentration, and fermentation duration of 24-72 h. Sampling, detection, and product separation and purification procedures were also designed and guided to complete the genetic engineering and fermentation culture of a microbial factory. High-performance liquid chromatography (HPLC) was used to separate and purify trace amounts of the target compound at a yield greater than or equal to 100 μg / L. This target compound was used as a physical standard for Raman spectroscopy detection, and peak positions and intensities were compared with theoretical Raman spectra to achieve final verification and calibration of the theoretical spectrum. The systematic deviation between the calibrated theoretical and measured spectra was greater than or equal to 5 cm⁻¹. -1 .

[0040] The automated synthetic biology design platform refers to an integrated software system that combines molecular modeling, pathway retrieval, gene circuit compilation, and fermentation control instruction generation. Its input is the SMILES string or InChI identifier of the compound, and its output is an executable set of microbial engineering operation instructions. The platform's inherent function is to retrieve feasible enzyme catalytic pathways from a natural product biosynthesis knowledge base under biosyntheticity constraints, and to design and combine them to regulate gene circuits to achieve the controllable expression of target molecules. In this embodiment, the platform receives high-confidence candidate component structure information output from the quantum chemical calculation module, while satisfying host compatibility and metabolic burden requirements. Under the premise of product toxicity threshold, the optimal chassis strain (such as Escherichia coli BL21(DE3) or Saccharomyces cerevisiae CEN.PK2) is automatically screened, a minimum gene editing set (including promoter, RBS, coding sequence and terminator) is planned, and corresponding fermentation process parameters (such as induction timing, feeding rate, pH and temperature program) are generated, which ultimately drives the microbial factory to complete the synthesis of target compounds at the microgram to milligram level; the matching confidence refers to the probability score output by the convolutional neural network model in step S3 for the feature matching between the theoretical spectrum and the measured spectrum of a candidate component. When the score is higher than the preset threshold of 0.92, the synthetic biology calibration process is triggered.

[0041] In one alternative implementation, the synthetic biology-assisted calibration module can be based on a rule-driven design process: first, it calls the KEGG and MetaCyc databases to match precursor metabolites, then uses PathwayTools for reverse pathway deduction, and combines CRISPR-Cas9 editing feasibility assessment to generate gene manipulation schemes. In another alternative implementation, the module may include a pre-trained Transformer model to directly predict the syntheticity score of molecular graph structures in typical microbial hosts and prioritize high-score pathways. Furthermore, the module can also be coupled with a digital twin fermentation simulator to conduct multiple rounds of virtual pilot production on candidate pathways based on kinetic models before actual synthesis initiation, screening out the implementation scheme with the highest expected yield and the fewest byproducts.

[0042] This invention connects quantum chemical prediction results with synthetic biology execution capabilities to construct a closed-loop verification chain of digital prediction → biomanufacturing → physical calibration. This not only solves the technical bottleneck of obtaining standards for rare active ingredients in herbal plants, but also enables theoretical Raman spectroscopy databases to have self-verification and authoritative endorsement capabilities, significantly enhancing the scientific rigor and regulatory acceptance of component identification conclusions.

[0043] This invention achieves bidirectional enhancement of the theoretical spectral database construction process through the synergistic effect of a dynamic evolutionary learning module and a synthetic biology-assisted calibration module: on the one hand, the meta-learning algorithm continuously improves the intrinsic accuracy of quantum chemical simulations by learning from historical biases; on the other hand, the synthetic biology platform fills the external verification gap of theoretical predictions through physical synthesis; together, they support the output of a theoretical Raman spectral database with high predictability and high reliability in step S1, providing a solid, reliable and continuously evolving data foundation for fusion comparison and component identification in step S3.

[0044] Example 3: In an optional embodiment, the present invention further provides the following step in the S2 sample spectral acquisition process: A nonlinear optical module integrating coherent anti-Stokes Raman scattering or stimulated Raman scattering is used for detection to completely suppress fluorescence background interference; Simultaneously, it includes an adaptive optics module that, based on a wavefront sensor, detects and compensates for optical aberrations caused by sample solution inhomogeneity or micro-flow in real time. By dynamically optimizing the laser beam wavefront through control of deformable mirrors, it ensures the acquisition of stable Raman signals with high signal-to-noise ratio and high spatial resolution under complex physical conditions. This includes: Step 1: Detection is performed using a nonlinear optical module that integrates coherent anti-Stokes Raman scattering or stimulated Raman scattering to completely suppress fluorescence background interference; in, Coherent anti-Stokes Raman scattering (CARS) is a four-wave mixing nonlinear optical process. Its signal generation has strict resonant selectivity and phase matching constraints, and it generates a strong coherent signal only at the vibrational frequency of the target molecule. In contrast, fluorescence emission is broadband, incoherent, and isotropic radiation, which does not satisfy the phase matching condition. Therefore, the CARS signal itself naturally excludes the fluorescence background. Stimulated Raman scattering (SRS) is a two-photon resonance enhancement process that relies on the precise frequency difference between the pump light and the Stokes light to match the molecular vibrational transitions. Its signal gain has narrow-band vibrational specificity, and fluorescence cannot be synchronously enhanced under this stimulated amplification mechanism, thus it also has an inherent fluorescence suppression capability. In this embodiment, the nonlinear optical module serves as the core of spectral excitation and acquisition in step S2. Its function is to directly obtain a pure, high-contrast Raman vibrational response signal from the original liquid / colloidal sample without applying chemical quenching or physical separation to the herbal plant essence sample. This avoids the problem of characteristic peak annihilation caused by fluorescence masking in traditional spontaneous Raman spectroscopy, thereby providing the convolutional neural network model in step S3 with a higher signal-to-noise ratio, more complete peak shape, and flatter baseline measured spectral input. In one alternative implementation, the nonlinear optical detection method can be: using a dual-color picosecond laser system to output pump light (1064nm) with a fixed wavelength and tunable Stokes light (720–850nm) respectively, which are collinearly focused into the detection cavity of a microfluidic chip to excite the SRS signal when the vibration resonance condition is met, and the intensity modulation component is extracted by a lock-in amplifier; In another alternative implementation, the nonlinear optical detection method may include: configuring a femtosecond optical parametric oscillator (OPO) to generate time-synchronized pump light and Stokes light, which are then incident on the sample after polarization combining; utilizing the forward coherence characteristics of the CARS signal, the anti-Stokes light (620–680 nm) in the phase-matching direction is collected through a high numerical aperture objective lens, and then the dispersion resolution and single-photon counting are performed by a grating spectrometer and an electron multiplier CCD. Furthermore, this nonlinear optical detection method can also be implemented by integrating the CARS and SRS channels onto the same optical path platform. By switching the optical parameter wavelength and detection mode, two nonlinear signals can be acquired alternately at the same detection site to form complementary verification. CARS provides high spatial resolution imaging capability, while SRS provides quantitative concentration response linearity, jointly supporting the joint analysis of component spatial distribution and relative content.

[0045] Step 2: This includes an adaptive optics module, which uses a wavefront sensor to detect and compensate for optical aberrations caused by non-uniformity or micro-flow of the sample solution in real time. By controlling a deformable mirror to dynamically optimize the laser beam wavefront, it ensures the acquisition of a stable Raman signal with high signal-to-noise ratio and high spatial resolution under complex physical conditions. in, A wavefront sensor is a spatially resolved detector used to measure the phase distortion of incident light waves. A typical type is the Shaker-Hartmann sensor, which divides the wavefront into multiple sub-apertures through a microlens array and then inverts the local wavefront slope based on the focal offset of each sub-aperture. A deformable mirror is an optical actuator with controllable surface deformation. It is usually composed of piezoelectric ceramic or electrostatically driven continuous surface mirrors. It can adjust the local curvature in real time according to the wavefront correction command and introduce equal amplitude reverse phase distortion to compensate for the aberration introduced by the sample. In this embodiment, the adaptive optics module functions as follows: to address the physical non-uniformities commonly found in herbal plant extract samples, such as multiphase states (e.g., suspended particles, lipid vesicles, viscous sols), concentration gradients, and microscale convection, by performing real-time closed-loop correction of low-order aberrations such as spherical aberration, coma, and defocus, as well as high-order wavefront distortion caused by local refractive index fluctuations. This maintains the diffraction-limited size and energy concentration of the laser focus within the optical detection cavity, ensuring the spatial positioning accuracy and signal strength stability of Raman signal acquisition. In one alternative implementation, the adaptive optics correction method can be: running a Shaker-Hartmann sensor at a sampling rate of 100Hz, acquiring wavefront slope data of ≥64 sub-apertures per frame, generating the first 20 order aberration coefficients after Zernike polynomial fitting, inputting them to a real-time controller, and driving a deformable mirror with a 37-unit piezoelectric actuator to complete millisecond-level wavefront reconstruction. In another alternative implementation, the adaptive optics correction method may include: adopting an indirect sensing strategy without wavefront sensors, that is, using the measured Raman signal intensity variance as the optimization objective function, and iteratively adjusting the deformable mirror surface parameters through the stochastic parallel gradient descent (SPGD) algorithm to achieve wavefront self-optimization without the need for additional optical detection paths; Furthermore, this adaptive optics correction method can also employ: coordinated control of wavefront correction and laser scanning trajectory—when the microscopic scanning system moves to a new region of the sample, the historical correction parameter matrix of the corresponding position is preloaded and fine-tuned in combination with the current sensor feedback, to achieve rapid aberration compensation across the field of view and support large-scale, high-throughput chemical imaging.

[0046] This invention achieves the dual goals of improving spectral purity and maintaining spatial fidelity simultaneously in a single detection process by physically coupling the intrinsic fluorescence suppression capability of the nonlinear optical module with the dynamic wavefront modulation capability of the adaptive optical module. Based on this, the obtained measured Raman spectra not only possess higher signal-to-noise ratios and clearer peak position identification, but also maintain a strict geometric correspondence between their spatial resolution information and spectral response intensity. This provides a reliable data foundation and physical guarantee for the convolutional neural network model in step S3 to accurately match the theoretical spectral database, especially for the distinguishable identification of compounds with highly similar structures (such as isomers and stereoisomers).

[0047] Example 4: In an optional embodiment, the present invention also provides a convolutional neural network model in the S3 fusion alignment and component identification step, which is constructed as a collaborative system of a continuous learning framework and a causal discovery module, including: Step 1: The continuous learning framework employs an elastic weight consolidation algorithm, enabling the model to learn the spectral characteristics of new compounds online without forgetting previously learned knowledge when encountering new herb samples not recorded in the original theoretical database, and to dynamically expand its recognition capabilities. Among them, the Elastic Weight Consolidation Algorithm is a typical parameter regularization-based continuous learning method. During the training of deep neural networks, this algorithm applies a secondary penalty term to the connection weights in the model that have a significant impact on the performance of existing tasks, thereby limiting their update magnitude in subsequent task training and protecting key knowledge representations from being overwritten. In this embodiment, the algorithm operates on the weight parameter space of the convolutional neural network model, enabling it to complete local gradient updates based on the current batch of data after receiving measured Raman spectra input from newly added herbal samples, while maintaining the ability to distinguish the features of existing compounds in the historical theoretical spectral database, thus achieving a gradual expansion of the recognition range without relying on repeated playback of the entire historical data or reconstruction of the model structure.

[0048] In one alternative implementation, the continuous learning method is based on the task increment scenario. Before introducing spectral data of new categories of compounds each time, the Fisher information matrix of the final training state of the model on the original theoretical database is estimated to lock the sensitivity of each weight parameter to the existing task loss function, and the corresponding L2 regularization constraint term of the weight is added to the optimization objective of the new task. In another alternative implementation, the continuous learning method uses a sliding window mechanism to maintain the gradient statistics of several recent batches of samples and dynamically adjusts the EWC penalty strength to adapt to the differences in the degree of spectral distribution drift between different batches. Furthermore, this continuous learning method employs an online Bayesian approximation strategy to model the weight uncertainty as a Gaussian distribution, thereby achieving a balance between knowledge retention and new knowledge absorption in a probabilistic manner.

[0049] Step 2: The causal discovery module, based on the causal structure learning algorithm, analyzes the statistical dependence between different spectral features and known component concentrations in a large number of mixed samples, infers and visualizes the potential causal networks between each spectral peak and between spectral peaks and components, and is used to verify and interpret the matching results of the neural network, reducing the risk of misjudgment caused by accidental spectral overlap. Among them, causal structure learning algorithms are a class of statistical learning methods used to infer the causal direction and structural relationship between variables from observational data. Common implementations include PC algorithms based on conditional independence tests, GES algorithms based on score search, and the NOTEARS framework that incorporates the assumption of nonlinear functions. These algorithms do not pre-determine the causal graph structure between variables. Instead, they identify directed acyclic graphs (DAGs) that satisfy Markov equivalence classes by analyzing the joint distribution characteristics between multidimensional spectral feature vectors and multidimensional component concentration vectors. This reveals which spectral peaks are more likely to be direct responses to vibrational modes of specific components, rather than collinear interference or random coupling. In this embodiment, the module takes the deep spectral embedding features extracted by the convolutional neural network and the corresponding laboratory-calibrated concentration data of the samples as input, constructs a feature-concentration joint variable set, performs causal graph structure search, and outputs a visualized causal network containing nodes (spectral peaks / components), edges (causal orientation), and edge weights (causal strength), which serves as an independent logical verification basis for the matching results of the neural network output.

[0050] In one alternative implementation, the causal discovery method uses the principal components of spectral features and the concentrations of known components as the variable set, and employs a PC algorithm based on conditional mutual information to gradually eliminate spurious correlation edges, eventually converging to the minimum I-MAP equivalence class causal graph. In another alternative implementation, the causal discovery method maps spectral features to a latent space, fits a linear causal structure with sparse regularization using the continuously optimized NOTEARS framework, and then quantifies the marginal causal contribution of each component to the key spectral peak through Shapley value decomposition. Furthermore, this causal discovery method employs an integrated causal discovery strategy, jointly running multiple basic algorithms (such as PC, GES, and LiNGAM) to assign higher confidence to consistently identified causal edges, and displays them as a heatmap superimposed on the original Raman spectrum to intuitively identify the peak positions of strong causal associations.

[0051] This invention integrates a continuous learning framework and a causal discovery module into the same convolutional neural network model architecture. This enables the model to possess both online adaptability to unknown compounds and traceable interpretability for matching results. Specifically, the continuous learning framework ensures that the model's accuracy in identifying known compounds in the original theoretical database does not significantly decrease as it continuously receives spectral data from new herbal samples, thus supporting the model's robustness under long-term deployment. The causal discovery module automatically calls upon the historical mixed sample data pool after each component matching to generate a causal evidence chain between the spectral peaks involved in the current matching and the components. If a key peak on which a matching result depends is not identified as a strong causal sub-node of the component by the causal network, a confidence reduction or manual review prompt is triggered. The synergistic effect of these two modules constitutes a technical extension of the computation-experimental dual verification mechanism at the model layer, not only improving the breadth and depth of identification but also fundamentally enhancing the scientific credibility and regulatory acceptability of the identification conclusions.

[0052] Example 5: In an optional embodiment, the present invention further provides an S4 dual verification result output step that is further connected to a blockchain-enabled distributed verification and feedback network, including: Step 1: Each successful calculation and experimental double verification result, including the identified components, the corresponding theoretical / experimental spectral fingerprints, matching parameters, and sample source metadata, is generated into an encrypted hash value and stored in a permissioned blockchain network to ensure that the data is tamper-proof and traceable throughout the entire process. Among them, the encrypted hash value refers to the unique fixed-length digest string generated by performing a one-way mapping operation on the structured detection result data using a cryptographic hash algorithm (such as SHA-256). Technically, it possesses input sensitivity and collision resistance. This technical feature is used in the field of information security to achieve data integrity verification and identity binding. In this embodiment, its function is to encapsulate all verifiable elements output from step S4—namely, component name, characteristic peak positions and intensity information of the corresponding theoretical and measured spectra, the threshold judgment results of vibration peak position matching degree and feature similarity, and metadata such as sample collection time, origin, batch number, and testing equipment ID—into an irreversible and unforgeable data fingerprint, serving as the basic unit for on-chain evidence storage. A permissioned blockchain network refers to a distributed ledger system employing an access mechanism, where only authorized testing institutions, regulatory nodes, or certification platforms can participate in ledger recording and verification as consensus nodes. Its consensus mechanism can be... Practical Byzantine Fault Tolerance (PBFT) or Raft; this technical feature is used in the field of distributed systems to balance security, scalability, and governance controllability; in this embodiment, its role is to build a trusted collaborative infrastructure, enabling all participants to jointly maintain an authoritative, consistent, and auditable historical record of test results without sharing the original spectral data and model weights, supporting cross-institutional result mutual recognition and quality traceability; the immutability and full traceability of data are technical effects achieved by the collaboration of permissioned blockchain networks and encrypted hash values, rather than independent technical features; its achievement depends on the chronological linking of blocks, forward constraint of hash pointers, and a multi-node redundant storage mechanism; in this embodiment, this feature enables any test report to have timestamp anchoring, operation traceability, and path backtracking capabilities from the time of its generation, providing underlying data trust support for regulatory authorities to conduct spot checks, enterprises to implement quality backtracking, and third parties to conduct methodological verification.

[0053] In one alternative implementation, the hash generation and on-chain process can be as follows: after the local detection terminal confirms the double verification in step S4, it automatically calls the embedded cryptographic module to perform SHA-256 hash calculation on the standardized JSON format detection result package, generate a 32-byte digest, and package the digest together with the digital signature, timestamp and node public key certificate into a transaction through the lightweight blockchain SDK, and submit it to the permissioned blockchain gateway. In another alternative implementation, the hash generation and on-chain process may include: receiving the verification result streams from multiple local detection nodes in a unified manner by the edge computing gateway, constructing a Merkle tree according to a preset aggregation strategy (such as packaging 10 results into a Merkle root), generating a single root hash, and then on-chaining it in batches to reduce on-chain overhead and improve throughput efficiency. Furthermore, the hash generation and on-chain process can also be achieved by integrating a blockchain coprocessor into the firmware layer of the detection instrument, synchronously triggering hash calculation and zero-knowledge proof generation the instant the spectral data completes the S4 determination, and uploading only the proof rather than the original data digest, thereby further compressing the on-chain storage and communication load while ensuring the validity of the verification.

[0054] Step 2: The network has a built-in global model optimization program based on federated learning: Under the coordination of the blockchain smart contract, each participating node only shares the encrypted model parameter updates, jointly trains a more powerful and general global spectral recognition model, and synchronously feeds the optimized model parameters back to the local convolutional neural network model of each node. Among them, the global model optimization procedure of federated learning refers to a distributed machine learning paradigm, the core of which lies in the aggregation of model parameters on a central server and the retention of original training data locally; this technical feature is used in the field of artificial intelligence to solve collaborative modeling problems under the constraints of data privacy, data silos, and compliant transmission; in this embodiment, its role is to enable different testing institutions (such as traditional Chinese medicine processing plants, third-party testing laboratories, and research institutes) to drive the continuous evolution of the global model by uploading only encrypted local model gradients or weight difference updates without exchanging the original measured spectral data or exposing the composition of their local theoretical databases; blockchain smart contracts refer to deterministic program code deployed on permissioned blockchains, which can automatically execute predefined logic, such as verifying node qualifications, verifying parameter update signatures, triggering aggregation tasks, and distributing update packages; this technical feature is used in the field of blockchain applications to realize The decentralized process is automated and rules are rigidly enforced. In this embodiment, its role is to replace the traditional centralized coordination server, undertaking the scheduling, verification, aggregation, and distribution functions in federated learning, ensuring that each node contributes authentically, behaves compliantly, and updates reliably, eliminating single points of failure and the risk of human intervention. Sharing only encrypted model parameter updates means applying homomorphic encryption or secure multi-party computation protection to the parameter changes (such as ΔW) generated by the local model in this round of training, so that it is always in a ciphertext state during transmission and aggregation. This technical feature is used in the field of privacy computing to ensure the confidentiality and integrity of the model update process. In this embodiment, its role is to prevent malicious nodes from reconstructing the local sample distribution characteristics or theoretical spectral library structure from the plaintext parameter updates through reverse inference, meeting the compliance requirements of the Personal Information Protection Law and the Regulations on the Management of Human Genetic Resources for biological sample-related data.

[0055] In one alternative implementation, the federated learning optimization procedure can be as follows: Each node locally runs the convolutional neural network model used in step S3. After completing the component identification task of a batch of test samples, based on the newly added reliable samples (including measured spectra and corresponding component labels) accumulated locally and confirmed by double verification in step S4, it performs a round of local gradient descent update to generate an encrypted gradient ΔW_enc. After the smart contract collects ΔW_enc from at least N valid signatures, it calls the on-chain aggregation function to perform encrypted domain averaging, generates a global update ΔW_global_enc, and broadcasts it to the entire network. Each node decrypts the updated data and applies it to its local model. In another alternative implementation, the federated learning optimization procedure may include: introducing an Elastic Weight Consolidation (EWC) constraint term into the local loss function, so that nodes actively protect key weights that are sensitive to the spectral features of learned compounds during updates, avoiding degradation of the recognition performance of classic herbal ingredients (such as flavonoids and saponins) due to new sample bias; the coefficient of this constraint term is dynamically issued by the smart contract based on the historical model drift monitoring results; Furthermore, the federated learning optimization process can also adopt the following: Introduce a robust weighting mechanism in the aggregation phase, where smart contracts automatically allocate aggregation weights based on the hash evidence credibility level of each node's updated data (such as the node's double verification pass rate in the past 30 days and on-chain audit score) and local data diversity indicators (such as spectral peak coverage width entropy value), thereby suppressing the interference of low-quality or abnormal updates on the global model.

[0056] This invention achieves a qualitative leap in detection data, transforming it from recordable to authoritative, auditable, and traceable, by generating encrypted hashes from S4 dual-verification results and storing them in a permissioned blockchain network. Building upon this, it leverages blockchain smart contracts for automated scheduling and trusted verification of the entire federated learning process. This allows participating nodes to continuously contribute encrypted parameter updates while strictly adhering to data sovereignty boundaries, collectively cultivating a global spectral recognition model that covers more herbaceous species, has stronger anti-interference capabilities, and a higher generalization level. Ultimately, the model parameters, after secure transmission back, enhance the component discrimination robustness of each node's local S3 model, forming a closed-loop enhancement mechanism of individual detection → on-chain evidence storage → co-evolution → capability feedback. This fundamentally breaks through the bottlenecks of data scale and compound coverage of a single institution, supporting the evolution of the herbal plant essence quality evaluation system towards standardization, intelligence, and ecological sustainability.

[0057] Example 6: In an optional embodiment, the method further includes: Step 1: Inject the herbal plant extract sample to be tested into an integrated microfluidic chip, which is equipped with a surface acoustic wave enrichment module: using an acoustic field to enrich the target active ingredient from the complex matrix into the optical detection cavity inside the chip; Among them, the surface acoustic wave enrichment module refers to the microscale force field control unit based on the surface acoustic waves excited by the piezoelectric substrate; This module typically possesses the capability to non-contactly manipulate, migrate, and focus micro- or nano-scale particles or molecules through acoustic radiation force within its technical field. In this embodiment, the module is configured in the transition region between the sample inlet channel and the optical detection cavity of the microfluidic chip. Its function is to selectively drive the target component to gather along the acoustic pressure node towards the central region of the optical detection cavity based on the difference in acoustic contrast between the target active component and the matrix component, while repelling macromolecular impurities and suspended particles. Thus, the physical enrichment and spatial localization of the target component can be completed without introducing chemical reagents or centrifugation. The enrichment result is directly used as the target for subsequent Raman spectroscopy acquisition, significantly improving the local concentration and spatial distribution uniformity of the target molecules in the detection cavity.

[0058] Step 2: Plasma Resonance Enhancement Unit: The inner wall of the optical detection cavity is modified with an array of gold / silver nanoparticles optimized by a genetic algorithm to generate a localized surface plasmonic resonance effect, thereby specifically enhancing the Raman scattering signal intensity of the target component; Among them, the plasma resonance enhancement unit refers to a nanostructure functional layer with controllable electromagnetic field localization capability constructed at the solid-liquid interface of the optical detection cavity. This unit typically possesses the capability to significantly enhance the Raman scattering cross section of neighboring molecules by exciting localized surface plasmon resonance (LSPR) in its respective technical field; In this embodiment, the unit uses a gold / silver nanoparticle array as the physical carrier. Its geometry, size distribution, and spatial arrangement are iteratively optimized using a genetic algorithm to match the LSPR response peak position with the excitation laser wavelength and scattering spectrum range corresponding to the main Raman vibrational modes of the target active ingredient. This matching results in a strong electromagnetic field enhancement during the Raman scattering process of the target ingredient in the microenvironment of the optical detection cavity after enrichment. This enhances the measured Raman signal intensity by 2–3 orders of magnitude without altering the intrinsic spectral characteristics of the molecule. This enhancement effect directly affects the raw spectral data acquired in step S2, providing a higher signal-to-noise ratio input basis for the convolutional neural network model in step S3 to identify and compare weak feature peaks.

[0059] Step 3: This step achieves pre-screening and signal amplification of the target components at the physical level, providing an enhanced input signal for subsequent high-precision identification; Among them, the physical level of pre-screening and signal amplification of target components means that the entire microfluidic chip pretreatment process does not rely on biochemical or physicochemical means such as chemical derivatization, enzymatic hydrolysis or chromatographic separation, but instead completes two types of physical operations, namely component enrichment and spectral enhancement, simultaneously through the spatial selective manipulation of acoustic force field and the localized energy coupling of plasma electromagnetic field. This technical feature is typically manifested in the field as a sample preprocessing paradigm with multi-physics field coordinated control at the micro-nano scale. In this embodiment, the specific function of this feature is manifested as follows: the high-density target molecule clusters formed by the surface acoustic wave enrichment module become the most effective enhancement targets of the plasma resonance enhancement unit; and the Raman response of the target molecules in the electromagnetic hotspot region of the cluster is enhanced by the LSPR effect; the two overlap in space, are coupled in time, and are complementary in function, forming an enrichment-enhancement closed loop; the enhanced input signal output by this closed loop not only shows an overall improvement in signal-to-noise ratio, but also shows the stable maintenance of the peak shape fidelity and relative intensity relationship of key discrimination peaks (such as C=O stretching vibration, aromatic ring breathing vibration, etc.), thereby ensuring that the model's ability to distinguish structurally similar isomers in step S3 is not affected by low signal-to-noise ratio interference.

[0060] In one alternative implementation, the acoustic field generation method of the surface acoustic wave enrichment module is as follows: an interdigital transducer (IDT) is integrated on a piezoelectric material substrate, and a radio frequency alternating voltage is applied to generate surface acoustic waves with a frequency of 10–100 MHz. The direction and intensity of the acoustic radiation force are controlled by adjusting the amplitude and phase of the input signal, thereby dynamically guiding the migration path of the target active ingredient. In another alternative implementation, the sound field generation method of the surface acoustic wave enrichment module includes: using a multi-mode IDT array to construct a two-dimensional acoustic potential well within the chip, so that the target component forms a stable accumulation at the intersection of multiple sound pressure nodes; Furthermore, this sound field control method also adopts a feedback closed-loop control strategy, that is, real-time acquisition of fluorescent markers or refractive index change signals in the optical detection cavity, which are used as inputs for adaptive adjustment of sound field parameters to cope with changes in sound propagation characteristics caused by fluctuations in viscosity and ionic strength of different batches of herbal extracts.

[0061] In one alternative implementation, the modification method of the gold / silver nanoparticle array in the plasma resonance enhancement unit is as follows: by using electron beam lithography combined with metal evaporation process, a periodic nanopore template is defined on the inner wall of the optical detection cavity, and then nanoparticles with uniform size and controllable lattice orientation are grown by chemical replacement or in-situ reduction method. In another alternative implementation, the modification method of the gold / silver nanoparticle array in the plasma resonance enhancement unit includes: using a drop-coating-self-assembly method, injecting a pre-synthesized solution of gold / silver nanospheres modified with thiol ligands into the chip, so that they are oriented and anchored to the cavity wall surface treated with thiol silanization through Au–S bonds to form a monolayer close-packed array; Furthermore, this modification method also employs a vapor-phase deposition-assisted thermally induced rearrangement process. Based on the initial disordered particle layer, the particles undergo controllable fusion and morphological evolution through gradient heating, ultimately obtaining a heterogeneous nanostar array with a multi-level tip structure, thereby broadening the LSPR response bandwidth and enhancing the hotspot density.

[0062] This invention achieves physical pre-screening and in-situ signal amplification of trace active ingredients in herbal plant extracts through the synergistic effect of a surface acoustic wave enrichment module and a plasmonic resonance enhancement unit. By utilizing the spatial enrichment of target components through an acoustic force field, the background interference of complex matrices on Raman detection is reduced, and the effective concentration of target molecules within the detection volume is increased. Furthermore, the localized surface plasmonic resonance effect excited by a nanoparticle array optimized by a genetic algorithm specifically enhances the Raman scattering efficiency of target components within the enrichment region. Together, these two mechanisms constitute a dual-gain mechanism of concentration enhancement and signal amplification, substantially improving the measured Raman spectra obtained in step S2 in terms of peak signal-to-noise ratio, peak position stability, and weak peak identifiability. This improvement directly supports the accurate matching of highly similar vibrational modes in the theoretical spectral database by the convolutional neural network model in step S3, especially ensuring reliable differentiation of components with subtle structural differences such as isomers and stereoisomers, thus fully realizing the computational-measurement dual verification objective defined by this invention.

[0063] Example 7: In an optional embodiment, the present invention further provides a digital twin-driven virtual control experiment module in the S3 fusion alignment and component identification step: This module is activated when there is uncertainty or controversy regarding the identification results of a certain component; Based on the known components and physicochemical parameters of the sample being tested, this module constructs a digital twin of the sample in digital space. By using mechanistic models and Monte Carlo simulations, controversial components are virtually added or removed from the digital twin, and the corresponding expected Raman spectral changes are calculated. By comparing the expected Raman spectral changes obtained from virtual experiments with the differences in measured spectra, third-party digital simulation evidence is provided to determine the presence and content of controversial components, forming a triple verification system of theory, experiment, and simulation, including: Step 1: This module is activated when there is uncertainty or controversy regarding the identification results of a certain component; Uncertainty or controversy refers to situations where the feature similarity output by the convolutional neural network model is within a preset critical range, or the matching degree between the theoretically calculated vibration peak position and the measured peak position does not reach the double verification threshold but is close to the lower limit of the threshold, or the difference in matching scores of multiple candidate components is less than the set tolerance range. This state is jointly determined by the confidence label generated in the S4 double verification result output step and the deviation distribution statistics, serving as the logical condition for triggering the virtual control experiment module driven by digital twins. This condition does not rely on manual intervention and can be automatically identified by the system to activate the subsequent simulation process.

[0064] Step 2: Based on the known components and physicochemical parameters of the sample being tested, this module constructs a digital twin of the sample in digital space; The known components refer to the set of components that have been preliminarily confirmed through the S3 fusion comparison and component identification steps and meet the dual verification threshold requirements. Their chemical structure, relative concentration range, solvent environment information, pH value, ionic strength, temperature and other physicochemical parameters are all used as input data. The digital twin is a dynamic virtual mapping that characterizes the multi-physics coupling characteristics of the sample under test in digital space. Its core function is to reproduce the molecular vibrational response behavior of the target sample under Raman excitation conditions. The digital twin does not include complete microstructure modeling, but is constructed by co-constructing a coarse-grained molecular dynamics model and a quantum chemical semi-empirical parameterization method. It can reflect the conformational distribution and vibrational coupling effect of key active ingredients in complex matrices while maintaining computational efficiency. The construction process of the digital twin does not require the participation of physical standards, but only relies on the structural parameters of the components already included in the S1 theoretical spectral database and the macroscopic spectral background features obtained by S2.

[0065] Step 3: Using mechanistic models and Monte Carlo simulations, the controversial components are virtually added or removed from the digital twin, and the corresponding expected Raman spectral changes are calculated. The mechanistic model refers to a causal model of the spectral response established based on the molecular vibrational selection rule and the resonance Raman enhancement mechanism, used to describe the relative intensity distribution of each vibrational mode of a specific chemical structure under a given laser wavelength excitation. Monte Carlo simulation, under the constraints of this mechanistic model, involves random sampling and statistical weighting of the conformational sampling space, local microenvironmental perturbation amplitude, and concentration gradient distribution of the controversial component to generate a representative set of virtual spectral responses. This simulation process does not solve the Schrödinger equation but utilizes the theoretical Raman spectral database already generated in S1 as a priori knowledge base. The structural characteristics of the controversial component are mapped to its possible vibrational peak position shift range and linear broadening range; the virtual addition or removal of the controversial component refers to updating the mass fraction and interaction term weight of the component in the state variables of the digital twin according to the set concentration gradient increment or decrement, thereby driving the mechanism model to recalculate the vibrational response function of the overall system; the output expected Raman spectral changes are represented as a set of differential spectral curves with wavenumber as the horizontal axis and normalized intensity change as the vertical axis, the shape of which can reflect typical interference modes such as displacement pulling of adjacent peak positions, peak intensity suppression or new peak generation after the introduction of the controversial component.

[0066] Step 4: Compare the differences between the expected Raman spectral changes obtained from the virtual experiment and the measured spectra to provide third-party digital simulation evidence for the presence and content of the disputed components, thus forming a triple verification system of theory, measurement, and simulation. The comparison between the expected Raman spectral changes and the measured spectra involves quantitatively evaluating the spatial correlation, peak consistency, and statistical significance of the differential signals under the same wavenumber grid and signal-to-noise ratio normalization conditions. This evaluation does not employ the absolute error minimization criterion but instead uses a joint criterion of Wasserstein distance and Kolmogorov–Smirnov test to determine whether the observed changes in the measured spectra fall within the 95% confidence interval of the expected change distribution generated by the Monte Carlo simulation. If they do, the existence of the disputed component is considered statistically supported, and its concentration estimate is taken as the median of the simulated distribution; if not, the component is deemed to have no statistical significance. The component is absent or below the detection limit in the current sample; this comparison result is independent of the black-box matching result of the convolutional neural network in S3 and the hard decision logic of the double threshold in S4, forming the third type of orthogonal verification path; the theoretical, experimental and simulation triple verification system is thus reflected as follows: the first layer provided by S1-S2 (consistency between theoretical calculation and experimental data), the second layer provided by S4 (dual constraints of model feature similarity and peak position matching degree), and the third layer provided by this step (causal interpretability verification under the intervention of controllable variables in digital space). The three complement each other in the verification dimension and jointly improve the robustness of identification of highly similar structural substances, trace active components and unknown coexisting interference substances.

[0067] This invention, by embedding a digital twin-driven virtual control experiment module in the S3 fusion comparison and component identification steps, enables the system to proactively initiate interpretability verification in uncertain scenarios. Using digital twins constructed from known components, combined with mechanism-guided Monte Carlo simulations, it achieves isolated analysis and quantitative inversion of the effects of controversial components. Furthermore, statistical comparison of virtual experimental results with measured spectral differences not only avoids the risks of relying on a single algorithm output but also enhances the traceability and reproducibility of conclusions from a causal modeling perspective. The resulting theoretical-experimental-simulation triple verification system effectively supports high-confidence identification of isomers, stereoisomers, and low-abundance active ingredients in herbal plant extracts, and is particularly suitable for quality judgment scenarios involving new resource plants or batches with process variations that lack standard references.

[0068] Example 8: In an optional embodiment, the method further includes: Step 1: S5a. Knowledge Graph Query: Input the list of ingredients identified in step S4 into an interdisciplinary herbal medicine knowledge graph. This graph integrates pharmacology, phytochemistry, and clinical research databases, and automatically associates and extracts information on the known efficacy, target of action, synergistic combination, and potential side effects of each ingredient. The herbal medicine knowledge graph refers to a structured semantic network organized in the form of triplets (entity-relationship-entity). Its nodes include multi-source heterogeneous medical and botanical entities such as compounds, target proteins, pathways, diseases, efficacy descriptions, and clinical trial numbers. Edges represent verified or high-confidence inferred relationships between entities. This knowledge graph has semantic reasoning capabilities and can support entity alignment and relationship completion across databases. In this embodiment, its role is to serve as a semantic bridge connecting the chemical component identification results and biological meaning. When the component list output by S4 (e.g., chlorogenic acid, rutin, tanshinone IIA) is used as the query subject input, the graph automatically performs multi-hop retrieval along predefined relationship paths such as component-pharmacological effects, component-metabolic pathways, component-synergistic compatibility, and component-adverse reactions, outputting a set of structured association information to provide semantically complete original knowledge materials for subsequent report generation.

[0069] Step 2: S5b. Natural Language Generation and Visualization: Based on graph neural networks, the query results are sorted by importance and relationships, driving a natural language generation engine to automatically generate a structured test report containing key functional ingredient groups, synergistic network diagrams, quality evaluation and recommendations, directly transforming chemical component data into biological and product insights that can be used for decision-making; Among them, Graph Neural Networks (GNNs) are deep learning models specifically designed for processing graph-structured data. Their core capability lies in aggregating neighbor node information through message passing mechanisms, thereby learning node representations and capturing complex dependencies in the graph. In this embodiment, the GNN uses the knowledge subgraph output by S5a as input, encoding each component node, target node, pathway node, and their connecting edges into embedding vectors. It then calculates the comprehensive influence score of each component in the efficacy network based on multi-dimensional indicators such as node centrality, path weight, and clinical evidence level, achieving importance ranking. The natural language generation engine refers to a text generation module based on sequence-to-sequence modeling. Its input is a structured knowledge tuple weighted and sorted by the GNN, and its output is a Chinese detection report paragraph conforming to professional standards. In this example, the engine employs a combination of template guidance and content filling: First, it loads the corresponding semantic template framework based on the report type (e.g., R&D evaluation version, quality inspection compliance version, market promotion version). Then, it fills the template slots with key component groups extracted after GNN sorting, synergistic pairs (e.g., baicalin and wogonin exhibit synergistic inhibitory effects in the TLR4 / NF-κB pathway), and risk warning items (e.g., samples containing aristolochic acid trigger nephrotoxicity warnings) according to logical hierarchy. Finally, it generates semantically coherent, terminologically accurate, and hierarchically clear natural language text. The synergistic network diagram is an interactive topology diagram in SVG format, where the node size reflects the component influence score, the edge thickness represents the synergistic strength, the color distinguishes the direction of action (activation / inhibition), and it supports clicking to expand literature source links.

[0070] This invention uses the component list output by S4 as the entry point for knowledge graph queries, initiating a cross-disciplinary semantic mapping process to obtain comprehensive biomedical annotations covering efficacy, targets, synergies, and risks. Furthermore, it leverages graph neural networks to perform important modeling and relation compression of these annotations based on graph structure awareness, forming a concise, task-oriented knowledge representation. Finally, a natural language generation engine decodes this representation into a human-readable, scenario-adaptable, and decision-making-usable structured report, fully constructing a closed-loop transformation link from chemical components to biological functions to product value. This link does not rely on human expert interpretation, avoiding experience bias and significantly improving the usability and transformation efficiency of test results in downstream processes such as traditional Chinese medicine research and development, quality control, and health product design.

[0071] Example 9: In an optional embodiment, the present invention further provides an adaptive optics module coupled with a computational imaging algorithm library: This algorithm library includes compressed sensing, phase retrieval, and deep learning super-resolution reconstruction algorithms; While performing hardware-level aberration correction using adaptive optics, the algorithm library performs secondary calculations on the acquired raw light field information, further breaking through the optical diffraction limit at the software level. This enables sub-micron spatial resolution imaging of chemical composition microscopic distribution, allowing the method to not only identify components but also visually present the original spatial distribution of each component in plant cells or formulation particles, including: Step 1: The adaptive optics module is coupled with a computational imaging algorithm library; The adaptive optics module refers to a real-time optical aberration compensation system based on wavefront sensing and deformable mirror control. Its inherent function is to dynamically correct wavefront distortion caused by sample solution inhomogeneity, micro-flow, or interface refractive index changes during laser confocal or nonlinear Raman imaging, thereby improving beam focusing quality and signal stability. In this embodiment, the module serves as a hardware-level preprocessing unit, outputting aberration-corrected raw light field data to provide a high-quality input foundation for subsequent algorithm processing.

[0072] The computational imaging algorithm library refers to a set of algorithms for post-processing light field data. Its inherent function is to recover spatial structure information that exceeds the physical limits of traditional optical systems from limited, noisy, or undersampled raw optical measurement data through mathematical modeling and information reconstruction. In this embodiment, the algorithm library does not replace the hardware correction function of the adaptive optics module, but performs collaborative software enhancement on its output results, forming a two-level imaging optimization path of hardware correction + software restoration.

[0073] Step 2: This algorithm library includes compressed sensing, phase retrieval, and deep learning super-resolution reconstruction algorithms; Each algorithm operates collaboratively according to the logic of "parallel preprocessing - joint computation - feature matching - spatial mapping". While adaptive optics completes hardware-level aberration correction based on wavefront sensors and deformable mirrors, the algorithm library performs staged secondary calculations on the original light field information containing multi-dimensional data such as two-dimensional light intensity distribution, three-dimensional light intensity distribution, Raman shift channel, and spatial coordinates. This breaks through the optical diffraction limit at the software level and ultimately achieves microscopic distribution imaging of chemical composition with a spatial resolution of less than or equal to 800 nanometers. The specific implementation logic and operation process are as follows: an iterative reconstruction method based on total variational regularization is adopted with a sampling rate of 30%, 50 iterations, and a regularization coefficient of 0.01. The original light field intensity distribution after correction is sparsely encoded and iteratively reconstructed using a compressed sensing algorithm. While maintaining the integrity of spectral features, the data dimension is reduced and redundant noise is suppressed, and the denoised light field intensity matrix is ​​output. Using the wavefront sensor measurements from the adaptive optics module as the initial phase estimate, a hybrid input-output iterative algorithm is employed, with the support domain constraint set as the physical boundary of the optical detection cavity and an intensity matching error threshold of 1×10⁻⁶. -6The complex amplitude wavefront is inverted from the light field intensity matrix using a phase retrieval algorithm to complete the missing phase information of the original light field and output high-fidelity complex amplitude light field data. Based on an improved U-Net architecture multi-input multi-output network, the intensity matrix obtained from compressed sensing and the complex amplitude light field obtained from phase retrieval are used as dual inputs. After 4 layers of downsampling, 4 layers of upsampling, and an attention mechanism, the light field region corresponding to the Raman feature peak is focused. The light field data is then super-reconstructed using a deep learning super-resolution reconstruction algorithm, outputting light field distribution data at the level of less than or equal to 800 nanometers, completing the secondary calculation at the software level. The super-reconstructed light field distribution data is then feature-matched with a theoretical Raman spectral database, and the weighted average absolute deviation between the theoretical vibration peak position and the measured light field peak position is less than or equal to 5%. -1 Centimeter-level verification completes the qualitative identification of components. At the same time, the characteristic light field signals of each component are mapped to the grid according to spatial coordinates to generate a pseudo-color heat map that uses color depth to represent the relative abundance of components. This achieves sub-micron level spatial resolution imaging of chemical component micro-distribution, intuitively presenting the original spatial distribution of each component in plant cells or formulation particles.

[0074] Compressed sensing (CS) is a mathematical framework that uses sparse priors to reconstruct complete information from data at a rate far below the Nyquist sampling rate. Its inherent function is to reduce the dimensionality of data acquisition, shorten imaging time, and suppress redundant noise. In this embodiment, the algorithm is used to perform sparse coding and iterative reconstruction of the Raman light field intensity distribution after preliminary adaptive optics correction, thereby reducing the number of pixels required for a single scan and improving imaging throughput while maintaining the integrity of spectral features.

[0075] Phase retrieval is a technique for inverting the complex wavefront phase from a diffraction pattern containing only intensity information. Its inherent function is to compensate for the missing phase information in traditional Raman imaging and restore the subtle interference and diffraction details during the propagation of the light field. In this embodiment, the algorithm uses the wavefront sensor feedback data provided by the adaptive optics module as the initial constraint, combined with the measured light intensity distribution, and reconstructs a high-fidelity complex amplitude light field through iterative strategies such as error subtraction or mixed input and output, providing a phase-sensitive basis for subsequent spatial positioning.

[0076] Deep learning super-resolution reconstruction refers to the technical approach of using a trained convolutional neural network model to map low-resolution or blurred optical images into high-resolution, high-fidelity images. Its inherent function is to learn the nonlinear mapping relationship between image degradation and restoration through a data-driven approach in the absence of a clear physical model. In this embodiment, the algorithm uses the complex light field of compressed sensing reconstruction result and phase recovery output as joint input. Guided by multi-scale feature fusion and attention mechanism, it generates a Raman spectral spatial distribution map with sub-micron level detail resolution, so that each pixel not only corresponds to a Raman shift value, but also characterizes the relative abundance and spatial orientation difference of a specific compound at that location.

[0077] Step 3: While performing aberration correction at the hardware level using adaptive optics, the algorithm library performs secondary calculations on the acquired raw light field information, further breaking through the optical diffraction limit at the software level to achieve sub-micron level spatial resolution imaging of chemical composition micro-distribution. Among them, hardware-level aberration correction refers to the closed-loop control process that uses a wavefront sensor to detect beam distortion in real time and drives a deformable mirror to apply conjugate phase compensation, which belongs to the category of physical optics. In this embodiment, the process ensures that the excitation light entering the detection cavity is always in the optimal focused state, providing the prerequisite for obtaining a raw Raman light field with a signal-to-noise ratio better than 50dB.

[0078] The raw light field information refers to the two-dimensional / three-dimensional light intensity distribution dataset directly recorded by the detector without any image interpolation, filtering or enhancement processing. It may contain multi-dimensional information such as spatial coordinates, wavelength channels and time frame indexes. In this embodiment, the information includes both intensity amplitude and residual phase perturbation after adaptive optics modulation, which constitutes the raw processing object of the computational imaging algorithm library.

[0079] Secondary computation processing refers to the process after completing a hardware correction, where no additional optical elements or higher numerical aperture objectives are required. Instead, multiple complementary algorithms are applied to the same set of raw data to expand the information dimension and increase the resolution. In this embodiment, the processing is not executed serially, but rather adopts a staged parallel scheduling strategy: compressed sensing prioritizes data dimensionality reduction and denoising, phase recovery synchronously constructs the complex amplitude field, and the deep learning model finally integrates the outputs of both to generate a unified super-resolution spatial distribution map.

[0080] Step 4: This method not only identifies components but also visually presents the original spatial distribution of each component in plant cells or formulation particles. The identification of components refers to determining which candidate compounds and their relative contents are present in the sample to be tested, based on the fusion comparison and double verification mechanism defined in Examples 1–3. In this example, this function has been completed by steps S3 and S4, and its technical implementation will not be explained again in this step.

[0081] The intuitive presentation of the original spatial distribution of each component in plant cells or formulation particles refers to superimposing the spatial location information of each identified component onto the microstructure image in the form of a pseudo-color heatmap. The color depth reflects the Raman characteristic peak intensity integral value of the component at the corresponding pixel position, and the spatial coordinate accuracy reaches the sub-micrometer level (≤800nm). In this embodiment, the presentation result is derived from the spatially resolved spectral cube output by the above-mentioned deep learning super-resolution reconstruction. After principal component analysis and non-negative matrix decomposition, the specific spectral fingerprint of each component is extracted and mapped to the corresponding coordinate grid through spatial matching, thereby constructing a component-location correlation matrix.

[0082] In one alternative implementation, the compressed sensing algorithm is an iterative reconstruction method based on total variation (TV) regularization, which significantly reduces the sampling rate while preserving the edges by applying sparse constraints to the gradient domain of the original light field intensity image. In another alternative implementation, the compressed sensing algorithm includes a random undersampling reconstruction process based on wavelet domain sparse representation, uses the Haar wavelet basis to decompose the light field data at multiple scales, and reconstructs the dominant coefficients through a greedy search strategy, which is suitable for real-time imaging requirements in rapid screening scenarios. Furthermore, this compressed sensing method combines a structured random measurement matrix with a joint sparse model to incorporate the correlation between different Raman shift channels into the prior modeling, thereby improving the quality of cross-wavelength consistent reconstruction.

[0083] In one alternative implementation, the phase retrieval algorithm is an iterative algorithm based on hybrid input / output (HIO), which uses wavefront sensor measurements provided by the adaptive optics module as the initial phase estimate, and alternately applies support domain constraints and intensity matching constraints in each iteration, gradually converging to a physically realizable complex amplitude solution; In another alternative implementation, the phase retrieval algorithm includes a depth-expanded differentiable phase solving network that embeds the traditional HIO iterative process into the neural network layer structure, enabling it to have end-to-end training capabilities and adapt to dynamic adjustments based on the scattering characteristics of different samples. Furthermore, this phase retrieval method employs a multifocal illumination-assisted phase retrieval strategy. By acquiring multiple sets of intensity images at different focal planes, three-dimensional light field constraints are constructed, thereby improving the robustness of phase reconstruction in deep tissue regions.

[0084] In one alternative implementation, the deep learning super-resolution reconstruction algorithm is a multi-input multi-output (MIMO) network based on the U-Net architecture. It receives the real part and imaginary part images of the compressed sensing reconstruction map and the phase recovery output as three-channel inputs, and outputs a super-resolution Raman intensity distribution map with spatial coordinate alignment. In another alternative implementation, the deep learning super-resolution reconstruction algorithm includes a generative adversarial network that introduces a physically guided loss function. Its discriminator not only evaluates the realism of the image, but also embeds a Rayleigh diffraction limit constraint term, which forces the generator to output a spatial spectral distribution that conforms to the laws of optical propagation. Furthermore, this deep learning super-resolution reconstruction method employs an online fine-tuning mechanism. During each new sample detection process, the backbone network is updated with lightweight parameters using a small amount of high-quality reference data from the current batch to adapt to the scattering characteristic drift caused by different herbal essence matrices.

[0085] This invention uses the high-quality raw light field information output by the adaptive optics module as a unified input source, organically linking three algorithms—compressed sensing, phase retrieval, and deep learning super-resolution reconstruction—at the data level. Compressed sensing provides a compact and denoised intensity representation for subsequent processing; phase retrieval supplements crucial wavefront phase information, supporting spatial coherence modeling; and the deep learning model integrates the outputs of the two aforementioned algorithms, completing cross-modal information alignment and super-resolution mapping in an abstract feature space. Based on this, and utilizing the spatially resolved spectral cube output by the three algorithms, the system can simultaneously perform qualitative and quantitative identification of components and subcellular-scale spatial localization without the need for fluorescent labeling, section staining, or chemical separation. Ultimately, it constructs a three-in-one paradigm for analyzing herbal plant essences, encompassing components, structure, and function, providing traceable, reproducible, and scalable technical support for spatial metabolomics research and the quality evaluation of micro / nano formulations.

[0086] Example 10: In an optional embodiment, the present invention also provides a blockchain-enabled distributed verification and feedback network, in which an oracle-based smart quality contract automatic execution system is deployed, including: Step 1: The system predefines quality standards for different types of herbal extracts; The quality standard refers to a set of compliance thresholds for ingredients strongly correlated with herbal plant essence categories, stored in the form of structured data in a permissioned blockchain network. These thresholds include, but are not limited to, the minimum detection limit of the target active ingredient, the maximum allowable content of impurity components, the lower limit of characteristic peak matching, the lower limit of characteristic similarity, and spectral signal-to-noise ratio requirements. The quality standard is classified and coded according to herbal species and mapped to the metadata identification field of the corresponding sample. This allows the appropriate standard template to be automatically retrieved based on the source category of the sample to be tested during subsequent comparisons. Its purpose is to provide clear, executable, and tamper-proof judgment criteria for on-chain automated decision-making, avoiding inconsistencies in scale or response delays caused by human experience.

[0087] Step 2: When new detection data, along with its blockchain evidence, is submitted to the network, the oracle automatically obtains the data and the smart contract deployed on the chain compares it with the preset quality standards. Here, the oracle refers to a trusted data relay module connecting the external world of the blockchain with on-chain smart contracts. Its function is to securely and deterministically transform off-chain real-world detection result data (including component identification results, theoretical / measured spectral fingerprint hash values, matching parameters, and sample source metadata output from step S4) into on-chain readable structured input. The smart contract refers to an automated program written with deterministic logic and deployed on permissioned blockchain nodes, embedding pre-defined quality standard parsing rules and comparison logic. The technical action of this step is to achieve trusted migration and instant semantic parsing of detection results from the physical experimental domain to the digital governance domain. In one optional implementation method... In one implementation, the comparison method involves: after confirming data integrity through hash verification, extracting the component list and corresponding confidence parameters from the detection data, sequentially matching each threshold condition in the preset standard, and performing Boolean logic judgment. In another optional implementation, the comparison method involves: mapping the detection data into a standardized quality assessment vector, calculating the deviation modulus with the preset standard vector in Euclidean distance space, and using whether the deviation modulus is lower than the tolerance threshold as the overall compliance criterion. Furthermore, this comparison method adopts a hierarchical triggering mechanism: first, performing hard threshold screening of key components, and only initiating a joint soft assessment of secondary components and spectral quality parameters if the screening passes, thereby balancing efficiency and robustness.

[0088] Step 3: The comparison results automatically trigger the corresponding on-chain operations: If the standard is met, a digital quality certification certificate with a timestamp and an immutable hash value is generated and broadcast; if the standard is not met, the relevant production nodes are automatically notified and the traceability process is initiated, realizing full automation and intelligence of testing, certification and quality control. The digital quality certification certificate is an on-chain credential natively supported by blockchain, possessing cryptographic signatures and globally unique hash identifiers. Its content includes at least: certificate number, issuance timestamp, corresponding test report hash, the version number of the quality standard on which it is based, certification status identifier, and digital signature of the issuing node. Automatic notification to relevant production nodes refers to the targeted push of structured alarm messages to upstream production units (such as extraction workshops, formulation production lines, or raw material suppliers) registered in the traceability path of the sample through a preset node communication protocol. The messages carry anomaly location information (e.g., the detection amount of flavonoid component A is lower than the requirements of Article 4.2 of the Q / YY-2023 standard). Initiating the traceability process refers to calling the on-chain stored full lifecycle metadata (covering harvest batch, solvent type, extraction temperature, storage time, etc.) to automatically generate a multi-level causal graph and mark high-risk links, allowing quality management personnel to quickly locate the source of deviation. The role of this step is to elevate a single testing behavior to a closed-loop quality governance event, so that quality judgment no longer stops at static conclusion output, but becomes a dynamic starting point driving subsequent actions.

[0089] This invention achieves a trusted bridge between off-chain testing data and on-chain governance logic through oracles. It uses smart contracts to transform preset quality standards into executable, verifiable, and unbypassable digital rules, automatically forking and executing two types of on-chain transactions—certification and tracing—based on comparison results. Furthermore, the generation and broadcasting of digital quality certification certificates ensure that quality credit is verifiable, transferable, and accumulative, while targeted notifications and causal tracing in abnormal scenarios guarantee that quality risks are perceptible, locatable, and interventionable. Ultimately, this forms an autonomous operating mechanism covering the entire chain from testing input to rule comparison, status determination, action execution, and credit accumulation, completely eliminating response gaps caused by manual review, cross-departmental coordination, and delayed paper records in traditional quality management systems, truly achieving millisecond-level transformation from sampling results to quality decisions.

[0090] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting components in herbal plant extracts, characterized in that, Includes the following steps: S1. Theoretical Spectral Database Construction Steps: For the candidate compounds that may be contained in the target herbal plant extract, the quantum chemical calculation module is used to simulate the theoretical vibrational modes of each candidate compound under Raman spectroscopy, generate the corresponding theoretical Raman spectra, and construct the theoretical Raman spectral database. S2. Sample Spectrum Acquisition Steps: Raman spectroscopy was performed on the herbal plant essence sample to obtain the measured Raman spectrum; S3. Fusion, comparison, and component identification steps: The measured Raman spectrum is input into a trained convolutional neural network model. The convolutional neural network model uses data from the theoretical Raman spectroscopy database as its training basis to extract deep features of the spectrum. The model matches and compares the deep features of the measured Raman spectrum with the features in the theoretical Raman spectroscopy database. S4. Dual verification result output step: Based on the matching and comparison results, output the component identification results; wherein, the validity of the identification results must simultaneously meet the matching degree threshold between the theoretically calculated and simulated vibration peak position and the measured spectral peak position, as well as the feature similarity threshold output by the convolutional neural network model, to complete the calculation and measurement dual verification.

2. The method for detecting components in herbal plant extracts according to claim 1, characterized in that, The steps for constructing the S1 theoretical spectral database further include: S1a. Dynamic Evolutionary Learning Module: Includes a meta-learning algorithm, enabling the quantum chemical calculation module to automatically adjust calculation parameters and basis set settings based on the systematic deviation between measured and theoretical spectra in historical identification results, so as to iteratively optimize the accuracy of subsequent simulations; S1b. Synthetic Biology Auxiliary Calibration Module: For key active ingredients that are predicted to exist by the quantum chemical calculations but lack physical standards and whose matching confidence level is higher than a preset threshold, the corresponding chemical structure information is input into the automated synthetic biology design platform; this platform, based on gene circuit and metabolic pathway algorithms, designs and guides microbial factories to rapidly synthesize trace amounts of target compounds, which serve as physical standards for the final verification and calibration of theoretical spectra.

3. The method for detecting components in herbal plant extracts according to claim 1, characterized in that, In the S2 sample spectral acquisition step: A nonlinear optical module integrating coherent anti-Stokes Raman scattering or stimulated Raman scattering is used for detection to completely suppress fluorescence background interference; Meanwhile, it includes an adaptive optics module, which uses a wavefront sensor to detect and compensate for optical aberrations caused by non-uniformity or micro-flow of the sample solution in real time. By controlling a deformable mirror to dynamically optimize the laser beam wavefront, it ensures the acquisition of a stable Raman signal with high signal-to-noise ratio and high spatial resolution under complex physical conditions.

4. The method for detecting components in herbal plant extracts according to claim 1, characterized in that, The convolutional neural network model in the S3 fusion comparison and component identification step is constructed as a collaborative system of a continuous learning framework and a causal discovery module: The continuous learning framework employs an elastic weight consolidation algorithm, enabling the model to learn the spectral characteristics of new compounds online without forgetting previously learned knowledge when encountering new herbaceous samples not recorded in the original theoretical database, and to dynamically expand its recognition capabilities. The causal discovery module is based on a causal structure learning algorithm. It analyzes the statistical dependence between different spectral features and known component concentrations in a large number of mixed samples, infers and visualizes the potential causal networks between each spectral peak and between spectral peaks and components, and is used to verify and interpret the matching results of the neural network, thereby reducing the risk of misjudgment caused by accidental spectral overlap.

5. The method for detecting components in herbal plant extracts according to claim 1, characterized in that, The S4 dual verification result output step is further connected to a blockchain-enabled distributed verification and feedback network: Each successful computational and experimental double verification result, including the identified components, the corresponding theoretical / experimental spectral fingerprints, matching parameters, and sample source metadata, is generated into an encrypted hash value and stored in a permissioned blockchain network to ensure that the data is tamper-proof and traceable throughout the entire process. The network incorporates a global model optimization program based on federated learning: Under the coordination of the blockchain smart contract, each participating node only shares encrypted model parameter updates, jointly trains a more powerful and general global spectral recognition model, and synchronously feeds the optimized model parameters back to the local convolutional neural network model of each node.

6. The method for detecting components in herbal plant extracts according to claim 1, characterized in that, Prior to the S2 sample spectral acquisition step, a microfluidic chip preprocessing and spectral enhancement step is also included: The herbal extract sample to be tested was injected into an integrated microfluidic chip, which contained: Surface acoustic wave enrichment module: Utilizes an acoustic field to enrich the target active ingredient from a complex matrix into the optical detection cavity within the chip; Plasma resonance enhancement unit: The inner wall of the optical detection cavity is modified with an array of gold / silver nanoparticles optimized by a genetic algorithm to generate a local surface plasmonic resonance effect, thereby specifically enhancing the Raman scattering signal intensity of the target component; This step achieves pre-screening and signal amplification of target components at the physical level, providing enhanced input signals for subsequent high-precision identification.

7. The method for detecting components of herbal plant extracts according to claim 1, characterized in that, The S3 fusion comparison and component identification step also includes a digital twin-driven virtual control experiment module: This module is activated when there is uncertainty or controversy regarding the identification results of a certain component; Based on the known components and physicochemical parameters of the sample being tested, this module constructs a digital twin of the sample in digital space. By using mechanistic models and Monte Carlo simulations, controversial components are virtually added or removed from the digital twin, and the corresponding expected Raman spectral changes are calculated. By comparing the expected spectral changes obtained from virtual experiments with the differences in measured spectra, third-party digital simulation evidence is provided for the existence and content of controversial components, forming a triple verification system of theory, experiment, and simulation.

8. The method for detecting components in herbal plant extracts according to claim 1, characterized in that, The method also includes a performance-oriented intelligent report generation step S5: S5a. Knowledge Graph Query: Input the list of components identified in step S4 into an interdisciplinary herbal medicine knowledge graph. This graph integrates pharmacology, phytochemistry, and clinical research databases, and automatically associates and extracts information on the known efficacy, target of action, synergistic combination, and potential side effects of each component. S5b. Natural Language Generation and Visualization: Based on graph neural networks, the query results are ranked by importance and relationships are sorted out, driving a natural language generation engine to automatically generate a structured test report containing key functional ingredient groups, synergistic network diagrams, quality evaluation and recommendations, directly transforming chemical composition data into biological and product insights that can be used for decision-making.

9. The method for detecting components in herbal plant extracts according to claim 3, characterized in that, The adaptive optics module is coupled with a computational imaging algorithm library: This algorithm library includes compressed sensing, phase retrieval, and deep learning super-resolution reconstruction algorithms; While performing hardware-level aberration correction using adaptive optics, the algorithm library performs secondary calculations on the acquired raw light field information, further breaking through the optical diffraction limit at the software level. This enables sub-micron spatial resolution imaging of chemical composition microscopic distribution, allowing the method to not only identify components but also visually present the original spatial distribution of each component in plant cells or formulation particles.

10. The method for detecting components in herbal plant extracts according to claim 5, characterized in that, The blockchain-enabled distributed verification and feedback network also includes an oracle-based smart quality contract automatic execution system. The system predefines quality standards for different types of herbal extracts; Once new test data, along with its blockchain-based notarized data, is submitted to the network, the oracle automatically retrieves the data and the smart contract deployed on the chain compares it with preset quality standards. The comparison results automatically trigger corresponding on-chain operations: if the standard is met, a digital quality certification certificate with a timestamp and an immutable hash value is generated and broadcast. If the requirements are not met, the relevant production nodes will be automatically notified and the traceability process will be initiated, achieving full automation and intelligence in testing, certification, and quality control.