Detection method for determining content of each component of ganoderma lucidum polysaccharide
By standardizing the spectral data of Ganoderma lucidum polysaccharide sample solutions and analyzing the characteristics of microbial interference, combined with DNA identification and compensation correction, the problem of distortion in spectral quantitative results caused by microbial contamination was solved, and stable and comparable detection of component content was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHANG ZHOU HALTH VOCATIONAL COLLEGE
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies for detecting Ganoderma lucidum polysaccharide samples, microbial contamination or metabolic byproducts cause a systematic shift between the spectral response and the actual component content, resulting in distorted quantitative results and reduced reliability of cross-batch comparisons and quality traceability.
By standardizing the spectral data of Ganoderma lucidum polysaccharide sample solution, calculating the microbial interference characteristics, and performing compensation correction based on DNA identification and microbial category identification, a set of component number templates was constructed to achieve the initial calculation and correction of component content.
Even in the presence of microbial contamination or metabolic byproducts, it can stably output component content results, reduce systematic bias, improve the stability and comparability of component content, and ensure the accuracy and traceability of quantitative conclusions.
Smart Images

Figure CN121877796A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of spectral detection and microbial gene detection technology, and more specifically, to a method for determining the content of various components of Ganoderma lucidum polysaccharides. Background Technology
[0002] As a commonly used edible and medicinal fungus, Ganoderma lucidum's polysaccharide content and composition ratio are frequently used for raw material quality evaluation, process stability monitoring, and product consistency control. Current testing practices include quantitative analysis based on chemical colorimetry and chromatographic separation, as well as methods using spectroscopic methods for rapid detection of sample solutions combined with models to estimate component content. Compared to traditional wet chemical detection, spectroscopic methods offer advantages such as smaller sample volumes, faster detection speed, and suitability for online or batch screening, thus possessing high application value in the rapid quantification and process monitoring of polysaccharides.
[0003] Microbial contamination or the generation of metabolic byproducts during the storage, transportation, or preparation of Ganoderma lucidum polysaccharide sample solutions can occur. These microbial factors alter the background absorption, scattering characteristics, and baseline morphology of the sample solution, causing a systematic shift between the spectral response and the actual component content. Because existing methods largely rely on general spectral preprocessing and single prediction models, and lack mechanisms to incorporate interference conditions such as microbial type and load intensity into the detection process for tiered confirmation and parameterized compensation correction, actual detection often requires manual retesting or changes in handling procedures. This makes it difficult to identify the source of distortion and to standardize correction criteria, ultimately leading to inaccurate quantitative results for the content of each component of Ganoderma lucidum polysaccharide and reducing the reliability of cross-batch comparisons and quality traceability.
[0004] In view of this, the present invention proposes a detection method for determining the content of various components of Ganoderma lucidum polysaccharides to solve the above problems. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art and achieve the above objectives, the present invention provides the following technical solution: a method for determining the content of various components of Ganoderma lucidum polysaccharides, comprising:
[0006] S1. Obtain the Ganoderma lucidum polysaccharide sample solution and generate a sample number; obtain the spectral acquisition parameters and acquire the raw spectrum; perform standardization processing on the raw spectrum to obtain the standardized spectral vector.
[0007] S2. Calculate the microbial interference characteristics, including scattering interference index and residual structure index, based on the standardized spectral vector; output the microbial distortion identifier based on the microbial interference characteristics and the preset interference threshold, and output the DNA recognition trigger identifier when the microbial distortion identifier meets the DNA confirmation trigger condition.
[0008] S3. When a DNA recognition trigger is present, obtain homologous sample processing solution based on sample number and perform nucleic acid extraction to obtain DNA feature data; generate DNA sequence feature data based on DNA feature data and compare it with the microbial reference sequence library to obtain microbial category identifier; calculate microbial load index based on DNA feature data and output microbial load level;
[0009] S4. Obtain component caliber parameters and generate a set of component number templates and a set of effective component numbers for the sample. The component caliber parameters include a molecular weight range threshold sequence and acidity property determination rules. Call the component content prediction model to output the initial calculation result of component content from the standardized spectral vector. Construct an initial calculation vector of component content according to the set of component number templates and write the corresponding initial calculation value of component content into the corresponding position of the initial calculation vector of component content according to the set of effective component numbers for the sample. When the microbial distortion label is non-distortion label, output the component content result according to the initial calculation vector of component content. When the microbial distortion label is distortion label, perform compensation correction on the initial calculation vector of component content based on the microbial category label and microbial load level to obtain the corrected component content result.
[0010] Furthermore, the method for obtaining the component number template set and the sample valid component number set includes:
[0011] Based on the molecular weight interval threshold sequence, the molecular weight of polysaccharides is mapped to a set of molecular weight interval numbers, which includes the first interval to the Gth interval. The acidity property determination rule is to output an acidity determination result identifier based on the acidity characteristic index and the preset acidity threshold. The acidity determination result identifier includes an acidity identifier and a non-acidity identifier. The acidity characteristic index is defined as the ratio of the acidity peak integral value to the skeleton peak integral value. The acidity peak integral value is the integral result of the standardized spectral vector in the acidity peak interval, and the skeleton peak integral value is the integral result of the standardized spectral vector in the skeleton peak interval. Based on the Cartesian combination relationship between the first interval to the Gth interval and the acidity determination result identifier, a set of component number templates is generated.
[0012] For each sample number, the polysaccharide molecular weight corresponding to the sample number is mapped to a set of sample molecular weight interval numbers based on the molecular weight interval threshold sequence; the acidity determination result identifier corresponding to the sample number is calculated based on the acidity property determination rule; and each molecular weight interval number in the sample molecular weight interval number set is combined with the corresponding acidity determination result identifier to generate a set of effective component numbers for the sample.
[0013] Furthermore, the method of calling the component content prediction model to output the preliminary calculation results of component content from the standardized spectral vector, constructing the preliminary calculation vector of component content according to the component number template set, and writing the corresponding preliminary calculation value of component content into the corresponding position of the preliminary calculation vector of component content by referring to the set of valid component numbers of the sample includes:
[0014] Based on the sample number, the standardized spectral vector is read and the component content prediction model is called to obtain a mapping table of the initial calculated component content values corresponding to each component number in the set of valid component numbers of the sample. Based on the component number template set, an initial calculated component content vector is constructed. Each component number in the component number template set is traversed. When the current component number does not belong to the set of valid component numbers of the sample, a null value mark is written in the sequence position corresponding to the component number in the initial calculated component content vector. When the current component number belongs to the set of valid component numbers of the sample, the initial calculated component content value is written in the sequence position of the component number in the initial calculated component content vector.
[0015] Furthermore, the method for performing compensation correction on the initial component content vector based on microbial category identifiers and microbial load levels to obtain the corrected component content results includes:
[0016] Read the microbial category identifier and microbial load level corresponding to the sample number, and use the microbial category identifier, microbial load level and component number as indexes to read the multiplicative compensation coefficient vector and additive compensation bias vector from the pre-constructed compensation parameter table;
[0017] Based on the multiplicative compensation coefficient vector, the additive compensation bias vector, and the initial calculation vector of component content, a component content correction formula is set. Each valid component number in the set of valid sample component numbers is traversed, and the initial calculation value of the component content corresponding to that valid sample component number is obtained from the initial calculation vector of component content. The corrected value of the component content corresponding to that valid sample component number is then calculated according to the component content correction formula. The corrected value of the component content corresponding to each valid sample component number is written into the corresponding index position in the corrected component content vector. The index positions corresponding to non-valid sample component numbers remain unchanged, resulting in the corrected component content vector. The corrected component content result is then output according to the component number template set.
[0018] Furthermore, the method for obtaining the DNA feature data includes:
[0019] Based on the sample number, the homologous sample processing solution is retrieved from the sample management record; the homologous sample processing solution is subjected to microbial enrichment pretreatment to obtain an enriched precipitate, the microbial enrichment pretreatment including centrifugation, discarding the supernatant and retaining the precipitate, adding buffer for resuspension, and repeated washing; lysis buffer and cell wall lysis enzyme are added to the enriched precipitate to form a lysis mixture, and the lysis mixture is incubated, mechanically disrupted, and centrifuged to obtain a lysis supernatant; magnetic bead adsorption solution is added to the lysis supernatant and mixed well to adsorb DNA onto the surface of the magnetic beads; the magnetic beads are placed on a magnetic rack for separation and the supernatant is discarded, and washing is performed; a preset elution volume of elution solution is added to the magnetic beads for elution to obtain a DNA extract; the DNA concentration and DNA purity index of the DNA extract are measured; the DNA concentration, DNA purity index, sample retention volume, and elution volume are packaged into DNA characteristic data; the sample retention volume is the volume of homologous sample processing solution used for DNA extraction in this case.
[0020] Furthermore, methods for generating DNA sequence feature data based on DNA feature data include:
[0021] The process involves: reading the version number of the microbial reference sequence library, which contains labeled region reference sequences and corresponding microbial category identifiers; constructing amplification feed parameters based on DNA feature data, including amplification feed volume and dilution factor; performing labeled region amplification on the DNA extract based on the amplification feed parameters to obtain amplification products, and acquiring reads from the amplification products to form a sequence read set; performing quality and length screening on the sequence read set, retaining reads with a length between 250 and 550 bases to obtain retained read sequences; and encapsulating the retained read sequence set and its read count information into DNA sequence feature data.
[0022] Furthermore, the method for obtaining the microbial category identifier includes:
[0023] Each read in the DNA sequence feature data is aligned with a microbial reference sequence library to calculate sequence similarity. The alignment results of all reads are aggregated by microbial category identifier, and a category score is calculated for each microbial category identifier. The category score is calculated as a weighted sum of the highest similarity of the category and the proportion of reads in the category. The microbial category identifier with the highest category score is selected as the candidate microbial category identifier. A similarity threshold and an advantage difference threshold are preset, and a first microbial category discrimination condition is preset. The first microbial category discrimination condition is as follows: when the highest similarity corresponding to the candidate microbial category identifier is not less than the similarity threshold and the difference between the category score of the candidate microbial category identifier and the second highest category score is not less than the advantage difference threshold, the microbial category identifier is output as the candidate microbial category identifier; when the first microbial category discrimination condition is not met, the microbial category identifier is output as the superior category identifier.
[0024] Furthermore, methods for calculating microbial load indicators and outputting microbial load levels based on DNA characteristic data include:
[0025] Quantitative amplification detection was performed on the DNA extract to obtain the cycle threshold, and the version number of the standard curve parameters was read. The standard curve parameters include the standard curve slope and the standard curve intercept, and a preset quantitative feed volume was used. Based on the cycle threshold and the standard curve parameters, the quantitative reaction copy number was calculated. Using the quantitative reaction copy number as a conversion benchmark, and combined with the quantitative feed volume and the elution volume and retention volume in the DNA characteristic data, a volume conversion was performed to obtain the copy number corresponding to each milliliter of sample solution as a microbial load index. The microbial load index was mapped to a microbial load level, which includes low load level, medium load level and high load level.
[0026] Furthermore, the implementation methods of S2 include:
[0027] Read the normalized spectral vector based on the sample number and obtain the corresponding reference normalized spectral vector;
[0028] Linear fitting is performed on the normalized spectral vector and the reference normalized spectral vector to obtain the second multiplicative coefficient and the second additive coefficient; the multiplicative deviation is calculated based on the second multiplicative coefficient; the additive deviation is calculated based on the second additive coefficient; the additive scaling factor is obtained; the multiplicative deviation and the additive deviation are divided by the additive scaling factor and added to obtain the scattering interference index;
[0029] The spectral reconstruction model is invoked to perform reconstruction inference on the standardized spectral vector to obtain the reconstructed spectral vector; the standardized spectral vector and the reconstructed spectral vector are subtracted point by point to obtain the residual sequence; the residual energy is calculated based on the residual sequence; the residual difference mean square is calculated based on the residual sequence to obtain the stability term parameter; the residual difference mean square is divided by the sum of the residual energy and the stability term parameter to obtain the residual structure index.
[0030] A comprehensive interference score is calculated based on the microbial interference characteristics. The comprehensive interference score is compared with a preset interference threshold to obtain a microbial distortion label, which includes both distortion and non-distortion labels. When the microbial distortion label is a distortion label and the comprehensive interference score is greater than the preset DNA trigger threshold, a DNA recognition trigger label is output.
[0031] Furthermore, the implementation method of S1 includes:
[0032] Obtain Ganoderma lucidum polysaccharide sample solutions and generate a sample number for each Ganoderma lucidum polysaccharide sample solution;
[0033] Spectral acquisition parameters are obtained, including wavelength start point, wavelength end point, wavelength step size, integration time, and number of scans. A wavelength sequence is generated based on the wavelength start point to the wavelength end point and according to the wavelength step size. The spectral acquisition unit is controlled to scan and acquire the Ganoderma lucidum polysaccharide sample solution according to the wavelength sequence to obtain the original spectrum. The original spectrum is then subjected to standardization processing to obtain a standardized spectral vector. The standardization processing includes dark current subtraction, baseline correction, noise reduction, scattering correction, and normalization.
[0034] Compared with existing technologies, the detection method for determining the content of various components of Ganoderma lucidum polysaccharides proposed in this invention has the following technical effects and advantages:
[0035] This invention addresses the problem that Ganoderma lucidum polysaccharide samples are prone to distortion of spectral quantitative results and insufficient consistency of component content output under microbial contamination or interference from microbial metabolic byproducts. By incorporating microbial category identification and microbial load level into the detection link, a closed-loop mechanism is achieved to first obtain the initial calculation result of component content and then perform directional compensation correction according to the interference conditions. This enables the output of stable corrected component content results even under complex sample conditions, reducing the impact of systematic deviations caused by microbial factors on quantitative conclusions.
[0036] This invention uses a molecular weight range threshold sequence and acidity property determination rules to fix the component scope. It combines the acidity determination result identifier to construct a component number template set and generate a set of valid sample component numbers. This ensures consistent indexing of component content results and a fixed data structure dimension, thereby improving comparability and traceability across samples and batches. In acidity determination, this invention sets acidity peak ranges and framework peak ranges based on mid-infrared spectroscopy, and constructs an acidity characteristic index using the integral result of the standardized spectral vector within these two ranges. This index is used to output acidity and non-acidity identifiers. This index solidifies the acidity difference into the component number with a calculable scope, reducing the risk of component classification drifting with background changes.
[0037] Furthermore, this invention obtains microbial category identifiers and microbial load levels based on DNA feature data, and establishes a compensation parameter table using microbial category identifiers, microbial load levels, and component numbers as indexes. This table is used to perform linear compensation correction on the initial component content vector bit by bit, thereby correcting the proportional distortion and background rise-type offset caused by microorganisms and improving the stability and accuracy of component content output.
[0038] In summary, this invention uses mid-infrared spectroscopy for quantitative analysis as the main approach, combined with microbial nucleic acid identification and parameterized compensation correction mechanisms, to ensure that the content of each component of Ganoderma lucidum polysaccharides can be output stably, comparablely, and traceably, and effectively reduces the risk of deviation in the content of Ganoderma lucidum polysaccharides caused by quantitative distortion due to microbial contamination or metabolic byproducts. Attached Figure Description
[0039] Figure 1 This is a flowchart of a detection method for determining the content of various components of Ganoderma lucidum polysaccharides according to Embodiment 1 of the present invention;
[0040] Figure 2 This is a schematic diagram of a detection system for determining the content of various components of Ganoderma lucidum polysaccharides according to Embodiment 2 of the present invention;
[0041] Figure 3 This is a flowchart of the method for obtaining the corrected component content results in Embodiment 1 of the present invention;
[0042] Figure 4 This is a flowchart of the method for obtaining DNA feature data according to Embodiment 1 of the present invention. Detailed Implementation
[0043] The technical solutions of the embodiments of the present invention will be described in detail, clearly, and completely below with reference to the accompanying drawings. It should be particularly noted that the specific embodiments described below are only for better illustrating and explaining the technical solutions of the present invention, and are intended to enable those skilled in the art to better understand and implement the present invention, and should not be construed as limiting the scope of protection of the present invention. Without departing from the spirit and substance of the present invention, those skilled in the art can modify, adjust, or make equivalent substitutions based on the content disclosed in the present invention, and these should all be considered within the scope of protection of the present invention.
[0044] Example 1
[0045] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.
[0046] Please see Figure 1 As shown in the figure, this embodiment discloses a method for determining the content of various components of Ganoderma lucidum polysaccharides, including:
[0047] S1. Obtain the Ganoderma lucidum polysaccharide sample solution and generate a sample number. Obtain the spectral acquisition parameters and acquire the raw spectrum. Perform standardization processing on the raw spectrum to obtain the standardized spectral vector.
[0048] S2. Calculate the microbial interference characteristics, including scattering interference index and residual structure index, based on the standardized spectral vector; output the microbial distortion identifier based on the microbial interference characteristics and the preset interference threshold, and output the DNA recognition trigger identifier when the microbial distortion identifier meets the DNA confirmation trigger condition.
[0049] S3. When a DNA recognition trigger is present, obtain homologous sample processing solution based on sample number and perform nucleic acid extraction to obtain DNA feature data; generate DNA sequence feature data based on DNA feature data and compare it with the microbial reference sequence library to obtain microbial category identifier; calculate microbial load index based on DNA feature data and output microbial load level.
[0050] S4. Obtain component caliber parameters and generate a set of component number templates and a set of effective component numbers for the sample. The component caliber parameters include a molecular weight range threshold sequence and acidity property determination rules. Call the component content prediction model to output the initial calculation result of component content from the standardized spectral vector. Construct an initial calculation vector of component content according to the set of component number templates and write the corresponding initial calculation value of component content into the corresponding position of the initial calculation vector of component content according to the set of effective component numbers for the sample. When the microbial distortion label is a non-distortion label, output the component content result according to the initial calculation vector of component content. When the microbial distortion label is a distortion label, perform compensation correction on the initial calculation vector of component content based on the microbial category label and microbial load level to obtain the corrected component content result.
[0051] Specifically, a method for determining the content of various components of Ganoderma lucidum polysaccharides is implemented as follows:
[0052] S1. Obtain the Ganoderma lucidum polysaccharide sample solution and generate a sample number. Obtain the spectral acquisition parameters and acquire the raw spectrum. Perform standardization processing on the raw spectrum to obtain the standardized spectral vector.
[0053] Specifically, the implementation method of S1 includes:
[0054] A sample solution of Ganoderma lucidum polysaccharide is obtained, which is a detection solution formed by dissolving Ganoderma lucidum polysaccharide samples. A sample number is generated for each sample solution, which is used to uniquely identify the current detection object. When generating the sample number, the collection date and timestamp and the daily serial number are obtained, and the sample number is obtained by concatenating the collection date and timestamp and the daily serial number. The sample number is then bound to the container identifier of the Ganoderma lucidum polysaccharide sample solution and written into the storage unit.
[0055] The spectral acquisition parameters are obtained, including wavelength start point, wavelength end point, wavelength step size, integration time, and number of scans. A wavelength sequence is generated based on the wavelength start point to the wavelength end point and according to the wavelength step size. The spectral acquisition unit is then controlled to scan and acquire the Ganoderma lucidum polysaccharide sample solution according to the wavelength sequence. The integration time controls the signal integration duration for a single acquisition, and the number of scans controls the number of repeated acquisitions. The raw spectrum is obtained, including the wavelength sequence and the spectral response intensity sequence. The raw spectrum is bound to the sample number and saved. The wavelength start point, wavelength end point, wavelength step size, integration time, and number of scans are written into the acquisition metadata corresponding to the sample number for subsequent parameter replay and result traceability in standardized processing.
[0056] It should be noted that the original spectrum includes the wavelength sequence and the spectral response intensity sequence; dark current subtraction, baseline correction, denoising, scattering correction and normalization are all applied to the spectral response intensity sequence, while the wavelength sequence remains unchanged as an index; the normalized spectral vector is the normalized spectral response intensity sequence, and the normalized spectral vector and the wavelength sequence are bound and saved through the sample number.
[0057] The original spectrum is normalized to obtain a normalized spectral vector. The normalization process includes dark current subtraction, baseline correction, denoising, scattering correction and normalization. The normalized spectral vector is then bound to the sample number and saved.
[0058] During dark current subtraction, the dark current spectrum is acquired. This spectrum is a sequence of dark current response intensity obtained by the spectral acquisition unit under shading conditions, using spectral acquisition parameters consistent with the current sample number. The dark current spectrum is written to the storage unit of the smart sensor, generating a dark current spectrum version number. The smart sensor integrates the spectral acquisition unit, processing unit, and local storage unit. The dark current spectrum version number identifies the acquisition date and timestamp and spectral acquisition parameter aperture corresponding to the dark current spectrum. Based on the sample number, the spectral response intensity sequence of the original spectrum is read. Under the condition of one-to-one correspondence of wavelength sequences, the dark current response intensity sequence of the dark current spectrum is subtracted point by point from the spectral response intensity sequence of the original spectrum to obtain the dark current subtracted spectrum. The dark current subtracted spectrum is bound and saved with the sample number, and the dark current spectrum version number is recorded in the smart sensor. This allows the dark current spectrum used to be traced and reused when the standardized spectral vector of the sample number is replayed or verified later.
[0059] The role of dark current subtraction in this invention is as follows: Dark current response biases the spectral response intensity under conditions of no incident light or low incident light, and this bias changes with device temperature drift, integration time, and sensor status. When the Ganoderma lucidum polysaccharide sample solution contains microbial contamination or microbial metabolic byproducts, the scattering background and weak signal segments of the sample are more easily amplified by the dark current bias, leading to spurious changes in the scattering interference index and residual structure index of microbial interference characteristics due to non-microbial factors, thus affecting the accuracy of microbial distortion identification. By performing dark current subtraction on the original spectrum, the device background bias is removed from the spectral response intensity, making the weak signal segments and baseline segments of the standardized spectral vector more stable, thereby reducing the false trigger probability of microbial distortion diagnosis and improving the repeatability and traceability of subsequent initial component content calculation results and component-level correction results.
[0060] When performing baseline correction, the baseline fitting order and baseline fitting constraint parameters are obtained. The baseline fitting constraint parameters include the baseline point selection rules and the upper limit of the number of iterations. Based on the dark current subtraction spectrum, the baseline point set is extracted according to the baseline point selection rules, and the baseline curve is obtained by fitting the baseline point set based on the baseline fitting order. The baseline correction spectrum is obtained by subtracting the baseline curve point by point from the dark current subtraction spectrum, and the baseline fitting order and the number of iterations are written into the processing metadata corresponding to the sample number.
[0061] During denoising, a noise level index is calculated to determine whether denoising should be performed. The noise level index is calculated based on the baseline-corrected spectrum, which is the mean square difference between adjacent points. When the noise level index is greater than the noise threshold, denoising is performed on the baseline-corrected spectrum. The denoising process uses a preset smoothing window length and a preset smoothing order to generate a denoised spectrum. When the noise level index is not greater than the noise threshold, the baseline-corrected spectrum is directly used as the denoised spectrum, and a marker indicating that denoising was not triggered is recorded.
[0062] When performing scattering correction, a reference spectral vector is obtained. The reference spectral vector is a spectral vector obtained by the same standardization process for a preset standard sample or a historical stable sample. The scattering correction coefficient is calculated based on the denoised spectrum and the reference spectral vector. The scattering correction coefficient includes a first multiplicative coefficient and a first additive coefficient. The denoised spectrum is corrected based on the first multiplicative coefficient and the first additive coefficient to obtain the scattering correction spectrum. The scattering correction coefficient is then bound to the sample number and saved.
[0063] During normalization, the normalization method parameters are obtained, including the normalization type and the normalization scale parameter. Normalization is then performed on the scattering correction spectrum to obtain a standardized spectral vector. In this embodiment, vector norm normalization is used: the vector norm of the scattering correction spectrum is calculated, and the standardized spectral vector is obtained by dividing the scattering correction spectrum point by point using the vector norm.
[0064] The standardized spectral vector is bound to the sample number and saved, and then written into the storage unit of the smart sensor.
[0065] In one specific embodiment of the present invention, spectral acquisition parameters are used to determine the acquisition range and sampling density of the spectral acquisition unit. For the spectral quantification of Ganoderma lucidum polysaccharide sample solution, the wavelength start point is set at [900nm, 1100nm], and the wavelength end point is set at [1600nm, 1800nm], ensuring that the wavelength end point is greater than the wavelength start point; the wavelength step size is set at [1nm, 5nm] to guarantee the spectral resolution required for spectral feature extraction and modeling. The integration time is set at [50ms, 500ms] to control the integration duration of a single acquisition; the number of scans is set at [3, 10] to improve acquisition stability and reduce occasional fluctuations. During acquisition, a wavelength sequence is generated based on the wavelength start point, wavelength end point, and wavelength step size. The spectral acquisition unit is then controlled to output a spectral response intensity sequence to the sample solution according to the wavelength sequence, thereby forming the original spectrum. The wavelength start point, wavelength end point, wavelength step size, integration time, and number of scans are written into the acquisition metadata to support subsequent standardized processing, playback, and result traceability.
[0066] Dark current subtraction is used to eliminate the background bias of the spectral acquisition unit. The dark current spectrum is acquired under shading conditions using the same wavelength sequence and integration time as the current acquisition, and written to a storage unit to generate a dark current spectrum version number. The storage unit can be configured as a local storage unit for the smart sensor, allowing the end-user to record the dark current spectrum and version number. During dark current subtraction, the dark current response intensity sequence is subtracted point-by-point from the spectral response intensity sequence of the original spectrum at the same wavelength position to obtain the subtracted dark current spectrum. Simultaneously, the dark current spectrum version number is recorded in the processing metadata corresponding to the sample number, enabling subsequent verification to accurately trace the source of the dark current spectrum upon which this subtraction was based.
[0067] Baseline correction is used to eliminate baseline shifts caused by optical path drift, container background, and slowly varying scattering. When using polynomial baseline fitting for baseline correction, the baseline fitting order is [1, 3], and the upper limit of the number of iterations is recommended to be [10, 30]. The baseline points are selected using the lower envelope rule, which divides the wavelength sequence into segments with a fixed number of points, with a segment length of [30, 80] sampling points. The response intensity of each segment at the [5%, 15%] quantile is taken as the baseline point. The baseline curve is obtained by fitting the baseline point set, and the baseline-corrected spectrum is obtained by subtracting the baseline curve point by point from the dark current subtraction spectrum. The baseline fitting order, upper limit of the number of iterations, segment length, and quantile value are written as baseline parameters into the processing metadata to ensure that the baseline correction is reproducible.
[0068] Denoising is used to suppress the disturbances caused by high-frequency random noise to the accurate diagnosis and component content prediction of microorganisms. The noise level index can be calculated using the mean square difference between adjacent points, and the noise threshold is set at... When the noise level exceeds the noise threshold, a smoothing denoising process is performed. The smoothing window length is set in [7,21], and the smoothing order is set in [2,4]. The noise threshold, smoothing window length, smoothing order, and denoising trigger flag are written into the processing metadata to make the denoising behavior traceable.
[0069] Scattering correction is used to reduce the multiplicative or additive scattering background caused by turbidity, particles, and microbial cells, making the subsequent microbial interference characteristic quantity and component content prediction inputs more stable. Scattering correction can adopt a multiplicative scattering correction approach: obtain a reference spectral vector, which is obtained by subtracting the same dark current and correcting the baseline from a preset standard sample or a historical stable sample; perform linear fitting between the denoised spectrum and the reference spectral vector to obtain the first multiplicative coefficient and the first additive coefficient as scattering correction coefficients, and correct the denoised spectrum based on the scattering correction coefficients to obtain the scattering correction spectrum. To avoid abnormal amplification, the first multiplicative coefficient is set in [0.7, 1.3].
[0070] Normalization is used to unify the spectral amplitude scale and reduce scale differences caused by varying acquisition intensities and light source attenuation. When using vector norm normalization, the vector norm is calculated for the scattering-corrected spectrum, and the normalized spectral vector is obtained by dividing the scattering-corrected spectrum point by point using the vector norm. To avoid normalization amplifying noise and causing numerical instability, a minimum norm threshold is set... When the vector norm falls below the minimum norm threshold, an unavailable flag is written and output is terminated. The normalization type, minimum norm threshold, and vector norm value are written to the processing metadata.
[0071] To ensure traceability across time points and support subsequent model maintenance, the spectral acquisition parameters and standardized processing parameters are versioned and encapsulated: the wavelength start point, wavelength end point, wavelength step size, integration time, number of scans, dark current spectral version number, baseline fitting order, upper limit of iterations, segment length, quantile value, noise threshold, smoothing window length, smoothing order, reference spectral vector number, scattering correction coefficient, normalization type, and minimum norm threshold are written into the parameter version record and a parameter version number is generated; the parameter version number is bound to the sample number for storage, so that the parameter version record can be read based on the sample number and the standardization process can be repeated.
[0072] More specifically, taking the Ganoderma lucidum polysaccharide sample solution with sample number "sample 20250801-0007" as an example, the wavelength start point was 950 nm, the wavelength end point was 1700 nm, the wavelength step size was 2 nm, the integration time was 200 ms, and the number of scans was 5; the generated wavelength sequence length was 376 points. Dark current spectroscopy was acquired under shading conditions using the same wavelength sequence and integration time, generating the dark current spectrum version number "dark current 20250801-01". Baseline correction used second-order polynomial fitting, with an upper limit of 20 iterations, a segment length of 60 points, and a quantile of 10%. The noise level index calculated from the baseline-corrected spectrum was... The noise threshold is set to The smoothing window length is 11, and the smoothing order is 3. The noise level of this sample exceeds the noise threshold, triggering denoising. Multiplicative scattering correction is used for scattering correction. The reference spectral vector is numbered "reference spectrum 20250208-03". Linear fitting yields a first multiplicative coefficient of 1.08 and a first additive coefficient of -0.003. The first multiplicative coefficient is within the range [0.7, 1.3] and does not trigger scattering anomaly markers. Normalization uses vector norm normalization, with a minimum norm threshold of [missing value]. The vector norm of this sample is 0.42, resulting in a normalized spectral vector of length 376.
[0073] S2. Calculate the microbial interference characteristics, including scattering interference index and residual structure index, based on the standardized spectral vector; output the microbial distortion identifier based on the microbial interference characteristics and the preset interference threshold, and output the DNA recognition trigger identifier when the microbial distortion identifier meets the DNA confirmation trigger condition.
[0074] The implementation methods of S2 include:
[0075] The standardized spectral vector is read based on the sample number, and the corresponding reference standardized spectral vector is obtained. The reference standardized spectral vector is generated from a historical stable sample set, which is the set of standardized spectral vectors corresponding to Ganoderma lucidum polysaccharide sample solutions that have been confirmed to be free of microbial contamination and have stable measurements. The amplitude average of the historical stable sample set is calculated point by point according to the wavelength sequence to obtain the reference standardized spectral vector, and a reference spectral number is generated for the reference standardized spectral vector. The reference spectral number and the reference standardized spectral vector are written into the storage unit so that the source of the reference spectrum can be traced for the calculation of the microbial interference characteristics of each subsequent sample number.
[0076] The scattering interference index is calculated. A linear fit is performed between the normalized spectral vector and the reference normalized spectral vector to obtain the second multiplicative coefficient and the second additive coefficient. The linear fit satisfies that the normalized spectral vector is approximately equal to the second multiplicative coefficient multiplied by the reference normalized spectral vector plus the second additive coefficient. Based on the second multiplicative coefficient, the multiplicative deviation is calculated; the multiplicative deviation is the absolute value of the difference between the second multiplicative coefficient and 1. Based on the second additive coefficient, the additive deviation is calculated; the additive deviation is the absolute value of the second additive coefficient. An additive scaling factor is obtained, used to normalize the additive deviation, and is set in the range [0.005, 0.02]. The multiplicative deviation and the additive deviation are divided by the additive scaling factor and added to obtain the scattering interference index. The scattering interference index, the second multiplicative coefficient, the second additive coefficient, the additive scaling factor, and the reference spectral number are bound and written into the processing metadata. The scattering interference index is used to characterize the combined perturbation intensity of the spectral shape caused by multiplicative scattering and additive background induced by turbidity, particles, or microbial cells. Multiplicative scattering refers to the scattering caused by particles such as microbial cells, which causes a proportional scaling or slope change in spectral intensity.
[0077] The calculation method for the scattering interference index includes:
[0078] ;
[0079] in, As a scattering interference index, This is the multiplicative deviation. It is an additive deviation. It is an additive scaling factor.
[0080] It should be noted that the scattering interference index is set using a combination of multiplicative and additive deviations. This is because microbial contamination or microbial metabolic byproducts entering the Ganoderma lucidum polysaccharide sample solution typically introduce two types of typical spectral perturbations: one type manifests as an overall amplitude scaling due to particle scattering and turbidity changes, resulting in an approximate proportional amplification or reduction based on the reference spectral shape, which can be characterized by the deviation of the second multiplicative coefficient from the ideal value; the other type manifests as an overall baseline rise or fall due to background scattering, impurity fluorescence, or stray light in the optical path, which can be characterized by the absolute value of the second additive coefficient. Based on this, the normalized spectral vector is fitted to the reference normalized spectral vector as "the normalized spectral vector is approximately equal to the second multiplicative coefficient multiplied by the reference normalized spectral vector plus the second additive coefficient," with the multiplicative deviation characterizing the multiplicative scattering perturbation and the additive deviation characterizing the background rise perturbation. Since additive and multiplicative deviations may differ in magnitude, an additive scaling factor is introduced to normalize the additive deviations, enabling the two types of disturbances to be superimposed and compared using the same scoring caliber, and facilitating the setting of a unified threshold for distortion determination.
[0081] The connection between the scattering interference index and the technical problem this invention aims to solve lies in the fact that this invention addresses scenarios where microbial contamination or microbial metabolic byproducts cause distortion in spectral quantitative results. This distortion is not merely manifested as increased random noise, but more often as a systematic shift in the spectral shape. The most representative form of this shift is multiplicative scaling and baseline shift caused by scattering. If this type of systematic shift is not quantified, the component content prediction model will misinterpret the spectral shape changes caused by the scattering background as differences in Ganoderma lucidum polysaccharide components during the inference stage, leading to systematic biases in the initial calculation results of component content, and subsequently triggering unnecessary DNA confirmation or outputting incorrect correction strategies.
[0082] By constructing a scattering interference index and incorporating it as a component of microbial interference characteristics, scattering distortion can be pre-identified before quantitative spectral output. This distinguishes between changes in scattering background caused by microorganisms and spectral shape changes caused by differences in Ganoderma lucidum polysaccharide components, thereby reducing the probability of misdiagnosis of microbial distortion and providing a quantitative basis for subsequent compensation and correction based on microbial load level and microbial category identification. Ultimately, this improves the stability and traceability of the output of the content of each component of Ganoderma lucidum polysaccharide.
[0083] Calculate the residual structure index. Invoke the spectral reconstruction model to perform reconstruction inference on the normalized spectral vector, obtaining the reconstructed spectral vector; subtract the normalized spectral vector from the reconstructed spectral vector point by point to obtain the residual sequence; calculate the residual energy based on the residual sequence, where the residual energy is the average of the squares at each point in the residual sequence; calculate the residual difference mean square based on the residual sequence, where the residual difference mean square is the average of the squares at each point in the difference sequence of adjacent points; obtain the stability term parameters, which are set as follows: The residual structure index is obtained by dividing the residual difference mean square by the sum of the residual energy and the stability term parameters. The residual structure index, residual energy, residual difference mean square, and spectral reconstruction model version number are then bound and written into the processing metadata. The residual structure index is used to characterize residual features in the normalized spectral vector that cannot be interpreted by the reconstructed spectral vector and exhibit a structural morphology, thereby reducing the probability of false triggers caused solely by random noise.
[0084] It should be noted that the spectral reconstruction model is used to learn the typical structure of the standardized spectral vector of Ganoderma lucidum polysaccharide sample solution under conditions without microbial interference. This allows for structural reconstruction of the input spectrum during the inference stage, and the parts that cannot be explained by the typical structure are represented as residual sequences to support the diagnosis of microbial distortion.
[0085] Specifically, the construction process of the spectral reconstruction model involves first establishing a training sample set, generated from a historically stable sample set. This historically stable sample set consists of Ganoderma lucidum polysaccharide sample solutions that have been confirmed free of microbial contamination and have stable component content through DNA identification or culture methods. For each sample number in the historically stable sample set, a standardized spectral vector is read and used to form a training input matrix. Each row of the training input matrix corresponds to a standardized spectral vector for a sample number. To avoid including abnormal samples in the "normal structure," anomaly removal is performed on the training input matrix: the Euclidean distance between each standardized spectral vector and the mean vector of the training input matrix is calculated. Sample numbers whose distance exceeds a preset anomaly distance threshold are marked for removal and removed from the training sample set, resulting in a purified training sample set. The spectral reconstruction model is then trained based on this purified training sample set. The spectral reconstruction model employs an autoencoder structure, including an encoder and a decoder. The encoder compresses the standardized spectral vectors into low-dimensional latent feature vectors, and the decoder restores the low-dimensional latent feature vectors to reconstructed spectral vectors. The encoder consists of several fully connected layers, with layer widths set according to a "layer-by-layer dimensionality reduction from the input dimension," and the latent dimension set at 5%–20% of the input dimension. The decoder uses a layer width setting symmetrical to the encoder, progressively increasing the latent dimension to the input dimension. During training, standardized spectral vectors are used as input and supervision targets, and mean squared error is used as the reconstruction loss function. The cleaned training sample set is iteratively updated according to a preset number of training rounds and batch size. Training stops when the reconstruction loss on the validation set no longer decreases. The structural parameters, training parameters, and loss convergence information of the spectral reconstruction model are encapsulated into a spectral reconstruction model version number, which is bound to the sample number range of the cleaned training sample set for storage to support subsequent model traceability and incremental updates.
[0086] The process by which the spectral reconstruction model performs reconstruction inference on the normalized spectral vector is as follows:
[0087] The standardized spectral vector is read based on the sample number of the sample to be tested, and is used as the input of the encoder. The encoder performs layer-by-layer linear transformation and nonlinear activation on the standardized spectral vector, and outputs a latent feature vector. The latent feature vector is used to characterize the main components of the spectral structure change under the condition of no microbial interference. The latent feature vector is used as the input of the decoder. The decoder performs layer-by-layer linear transformation and nonlinear activation on the latent feature vector, and outputs a reconstructed spectral vector with the same dimension as the standardized spectral vector. The reconstructed spectral vector is subtracted from the standardized spectral vector point by point to obtain the residual sequence, and the residual energy and residual structure index are calculated based on the residual sequence.
[0088] It should be noted that, since the spectral reconstruction model only fits the typical spectral structure of the historical stable sample set during the training phase, when there are atypical components introduced by microbial cell scattering, metabolic byproduct background, or related structural perturbations in the standardized spectral vector, these atypical components are difficult to be effectively expressed by the latent feature vector. As a result, they are weakened in the reconstructed spectral vector and highlighted in the residual sequence, making the residual structure index more sensitive to the measurement distortion caused by microorganisms, thereby improving the reliability of the microbial distortion label output.
[0089] Microbial interference characteristics are constructed and microbial distortion identifiers are output. Weighting parameters are obtained, including scattering weight and residual weight, both set in the range [0.3, 0.7], and their sum is 1. Scattering interference indicators are weighted based on the scattering weight, and residual structure indicators are weighted based on the residual weight. The weighted results are summed to obtain a comprehensive interference score. The comprehensive interference score is compared with a distortion threshold set in the range [0.6, 1.2]. When the comprehensive interference score is greater than the distortion threshold, a microbial distortion identifier is output as a distortion identifier; when the comprehensive interference score is not greater than the distortion threshold, a microbial distortion identifier is output as a non-distortion identifier. The comprehensive interference score, distortion threshold, and microbial distortion identifier are bound to the sample number and saved. The microbial distortion identifier is used to characterize whether the spectral quantitative results corresponding to the current sample number are significantly interfered with by microbial contamination or microbial metabolic byproducts.
[0090] When the microbial distortion marker meets the DNA confirmation trigger condition, a DNA identification trigger marker is output. A DNA trigger threshold is obtained and set in the range [1.0, 1.8], and this threshold must be greater than the distortion judgment threshold. When the microbial distortion marker is valid and the overall interference score is greater than the DNA trigger threshold, a DNA identification trigger marker is output and bound to the sample number. When the microbial distortion marker is valid but the overall interference score is not greater than the DNA trigger threshold, a verification marker is output to trigger re-sampling and re-judgment, and this marker is bound to the sample number and saved. This tiered threshold system achieves a closed-loop logic of first judging distortion and then triggering DNA confirmation, ensuring that DNA identification resources are used for high-risk sample numbering and reducing unnecessary confirmation overhead.
[0091] More specifically, taking the standardized spectral vector of sample number "sample 20250801-0007" as an example, with reference spectrum number "reference spectrum 20250208-03" and additive scaling factor of 0.01; linear fitting yields a second multiplicative coefficient of 1.12 and a second additive coefficient of 0.004, resulting in a multiplicative deviation of 0.12 and an additive deviation of 0.004. The scattering interference index is... The residual energy was calculated using the spectral reconstruction model. The residual mean square difference is Stability term parameters are taken The residual structure index is The scattering weight is set to 0.45, the residual weight to 0.55, and the overall interference score is... If the distortion threshold is set to 0.65, the output microbial distortion identifier will be the distortion identifier. If the DNA trigger threshold is set to 1.20, since the overall interference score of 0.67 does not exceed the DNA trigger threshold, a verification marker will be output instead of triggering DNA recognition. For example, for the standardized spectral vector of sample number "sample 20250801-0008", the scattering interference index is 0.90 and the residual structure index is 1.05, and the overall interference score is 0.98 under the same weights. When the distortion threshold is 0.65 and the DNA trigger threshold is 0.85, the overall interference score exceeds the DNA trigger threshold, and a DNA recognition trigger identifier will be output and bound to the sample number for storage.
[0092] S3. When a DNA recognition trigger is present, obtain homologous sample processing solution based on sample number and perform nucleic acid extraction to obtain DNA feature data; generate DNA sequence feature data based on DNA feature data and compare it with the microbial reference sequence library to obtain microbial category identifier; calculate microbial load index based on DNA feature data and output microbial load level.
[0093] The implementation methods of S3 include:
[0094] The homologous sample processing solution is read from the sample management record based on the sample number; the homologous sample processing solution is the retained and dispensed solution of the Ganoderma lucidum polysaccharide sample solution with the same source as the standardized spectral vector.
[0095] Microbial enrichment pretreatment was performed on homologous sample solutions to reduce the inhibitory effect of Ganoderma lucidum polysaccharide matrix. The homologous sample solutions were placed in centrifuge tubes and centrifuged at 9000–12000 rpm for 8–15 minutes. The supernatant was discarded, and the precipitate was retained. Buffer was added to the precipitate for resuspending, where the volume of buffer added was defined as the resuspending volume, and the resuspending volume was set to 0.3–1.0 mL. This resuspending volume was used to uniformly disperse the precipitate and reduce soluble polysaccharide residues, thereby reducing the residue of polysaccharides and small molecule interfering substances in the solution. The resuspending was centrifuged again, and the supernatant was discarded. Washing was repeated 1–2 times to obtain the enriched precipitate. Centrifugation speed, centrifugation time, number of washes, and resuspending volume were used as enrichment parameters and linked to the sample number for storage.
[0096] The enriched precipitate was lysed to release nucleic acids and form a lysis supernatant. Specifically, lysis buffer and cell wall lysin were added to the enriched precipitate, and the mixture was stirred to form a lysis mixture. The total volume of the lysis mixture was set to 0.4–1.2 mL to ensure thorough lysis and subsequent purification. The lysis mixture was incubated at 55–65 °C for 10–30 minutes, followed by mechanical shaking for 1–3 minutes. After lysis, the mixture was centrifuged at 9000–12000 rpm for 3–8 minutes, and the supernatant was collected as the lysis supernatant. The incubation temperature, incubation time, shaking time, and total volume of the lysis mixture were recorded as lysis parameters and linked to the sample number for storage.
[0097] like Figure 4 As shown, DNA purification was performed on the lysate supernatant to obtain DNA extract and generate DNA characteristic data. Specifically, magnetic bead adsorption solution was added to the lysate supernatant and mixed to adsorb DNA onto the surface of the magnetic beads; the magnetic beads were placed on a magnetic rack for separation and the supernatant was discarded, followed by washing; a preset elution volume of elution solution (40–80 μL) was added to the magnetic beads for elution to obtain the DNA extract. The DNA concentration and purity index of the DNA extract were measured. The DNA concentration was determined using quantitative fluorescence method, and the DNA purity index was determined using absorbance ratio. The DNA concentration, DNA purity index, sample retention volume, and elution volume were encapsulated as DNA characteristic data and bound to the sample number, then written into the intelligent sensor storage unit. The sample retention volume was defined as the volume of homologous sample processing solution used for microbial DNA identification.
[0098] DNA sequence feature data is generated based on DNA feature data to support microbial category identification. Specifically, the version number of the microbial reference sequence library is read. The microbial reference sequence library contains reference sequences of labeled regions and corresponding microbial category identifiers. Amplification feed parameters are constructed based on the DNA feature data, including amplification feed volume and dilution factor. Based on the amplification feed parameters, the labeled regions of the DNA extract are amplified to obtain amplification products. Reads are then obtained from the amplification products to form a sequence read set. Quality and length screening are performed on the sequence read set, retaining reads with a length between 250 and 550 bases to obtain retained read sequences. The retained read sequence set and its read count information are encapsulated into DNA sequence feature data. The version number of the microbial reference sequence library, the number of retained reads, and the length distribution are written into the smart sensor storage unit to ensure the traceability of sequence features.
[0099] Methods for constructing amplification feed parameters based on DNA feature data include:
[0100] Read DNA concentration and purity values; when the DNA purity value is below the lower limit or above the upper limit of the DNA purity threshold, output a verification marker and trigger repurification or resampling until the DNA purity value falls within the range defined by the lower and upper limits of the DNA purity threshold. Set the target feed mass range and the operable volume range for amplification, and set the target feed mass within the target feed mass range; calculate the theoretical feed volume based on the DNA concentration and target feed mass, which is equal to the ratio of the target feed mass to the DNA concentration.
[0101] It should be noted that, for example, in the technical solution of this invention, the target amplification mass range can be set to [2, 20] nanograms. This target amplification mass range is used to constrain the DNA mass entering the amplification reaction system, so as to avoid amplification failure or excessive randomness due to too low a mass, and inhibition or non-specific amplification due to too high a mass. The operable volume range is set to [1, 10] microliters to ensure the accuracy of pipetting and the stability of the reaction system ratio.
[0102] The theoretical feed volume and the operable volume range are compared to determine the dilution factor and amplification feed volume: when the theoretical feed volume falls within the operable volume range, the dilution factor is set to 1, and the theoretical feed volume is determined as the amplification feed volume; when the theoretical feed volume is less than the lower limit of the operable volume range, it indicates a high DNA concentration. To reduce pipetting errors and maintain the volume ratio of the amplification reaction system, a preset dilution factor greater than 1 is set, and the DNA extract is pre-diluted. The theoretical feed volume is then recalculated based on the diluted DNA concentration. The theoretical feed volume, which is recalculated and falls within the operable volume range, is determined as the amplification feed volume, with the dilution factor set to 2-10. When the theoretical feed volume is greater than the upper limit of the operable volume range, it indicates that the DNA concentration is low. To avoid the feed volume being too large and crowding out the reaction system volume or introducing more inhibitors, the target feed mass is adjusted to the lower limit of the amplification target feed mass range, and the theoretical feed volume is recalculated. The theoretical feed volume, which is recalculated and falls within the operable volume range, is determined as the amplification feed volume, and the dilution factor is set to 1.
[0103] It should be noted that the method for setting the target feed mass within the amplification target feed mass range includes: determining the target feed mass according to the rule of first satisfying volume operability, then satisfying the feed mass range, and finally fixing it to a reproducible single value, so that the calculation of amplification feed volume and dilution factor is clear. After reading the DNA concentration value, the operable volume range is first converted into the feed mass range that can be achieved at the current concentration: calculate "the mass corresponding to the lower limit of volume = DNA concentration value × lower limit of operable volume range", and "the mass corresponding to the upper limit of volume = DNA concentration value × upper limit of operable volume range"; then, the intersection with the amplification target feed mass range is used to obtain the candidate feed mass range. The lower limit of the candidate feed mass is the larger of the mass corresponding to the lower limit of the amplification target feed mass range and the mass corresponding to the lower limit of volume, and the upper limit of the candidate feed mass is the smaller of the mass corresponding to the upper limit of the amplification target mass range and the mass corresponding to the upper limit of volume. If the lower limit of the candidate feed quality is less than or equal to the upper limit of the candidate feed quality, it indicates that there exists a target feed quality that simultaneously satisfies the feed quality constraint and volume operability. In this case, the target volume is set as the median of the operable volume range, and "median quality = DNA concentration value × target volume" is calculated. The target feed quality is determined as the result after interval truncation of the median quality. That is, when the median quality is less than the lower limit of the candidate feed quality, the lower limit of the candidate feed quality is taken; when the median quality is greater than the upper limit of the candidate feed quality, the upper limit of the candidate feed quality is taken; otherwise, the median quality is taken. Thus, a single and reproducible target feed quality is obtained.
[0104] If the candidate feed mass range is empty, it means that two types of constraints cannot be satisfied simultaneously at the current DNA concentration: When the mass corresponding to the lower volume limit is greater than the upper limit of the amplification target feed mass range, the target feed mass is selected as the upper limit of the amplification target feed mass range. This allows for maximizing the theoretical feed volume without exceeding the upper limit, followed by pre-dilution to bring the final amplification feed volume into the operable volume range. When the mass corresponding to the upper volume limit is less than the lower limit of the amplification target feed mass range, the target feed mass is selected as the lower limit of the amplification target feed mass range. This aims to minimize the theoretical feed volume without falling below the lower limit. If the theoretical feed volume still exceeds the upper limit of the operable volume range, the target feed mass remains unchanged, and a verification mark is output to indicate feed limitations caused by low sample concentration. Through these rules, the target feed mass can be uniquely determined at any DNA concentration, maintaining consistency with the process for determining the amplification feed volume and dilution factor.
[0105] The process of amplifying labeled regions of DNA extract based on amplification feed volume and dilution factor to obtain amplification products and sequence read sets can be disclosed as follows: First, read the amplification feed volume and dilution factor corresponding to the sample number: when the dilution factor is equal to 1, directly take a sample from the DNA extract according to the amplification feed volume as the amplification template; when the dilution factor is greater than 1, mix the DNA extract with dilution buffer according to the dilution factor to obtain diluted DNA extract, and then take a sample from the diluted DNA extract according to the amplification feed volume as the amplification template. The dilution status, dilution factor, amplification feed volume, and sample number are then linked and saved.
[0106] During the labeled region amplification stage, a labeled region relevant to the identification of the target microorganism is identified, and a pair of primers is designed for this region. Each primer includes at least an amplification guide sequence for complementary pairing with sequences flanking the labeled region. A sample barcode sequence and a sequencing adapter sequence are linked before the amplification guide sequence, ensuring that the amplification product carries both the sample barcode sequence and the sequencing adapter sequence. The amplification template, the primer pair, the polymerase reaction system, and the buffer system are mixed to form the amplification reaction solution, and thermal cycling amplification is performed to obtain the amplification product. Thermal cycling amplification includes at least denaturation, annealing, and extension cycles, with the number of cycles determined based on the amount of amplification product and the risk of non-specific amplification. After amplification, the amplification product is purified to remove free primers and short fragments. The concentration of the amplification product is quantified, and the fragment length is verified. The quality record of the amplification product is linked to the sample number and stored to support the identification of the sample barcode and the reading of the sequencing adapter in the subsequent read acquisition stage.
[0107] During the read acquisition stage, amplified products are pooled according to sample barcodes for library construction, and read acquisition is performed on the sequencer to obtain a set of reads. After cluster generation or equivalent immobilization of the amplified products, the sequencer reads the sequence information of the amplified fragments and outputs reads containing base sequences and quality values. All reads corresponding to the same sample barcode are summarized to form a set of reads corresponding to that sample number. The read set is then subjected to quality filtering and adapter removal to obtain a set of effective reads that can be used for subsequent DNA sequence feature data generation and alignment with microbial reference sequence libraries.
[0108] DNA sequence feature data is compared with a microbial reference sequence library to perform DNA alignment and output microbial category identifiers. Specifically, each read sequence in the DNA sequence feature data is compared with the microbial reference sequence library, and sequence similarity is calculated. The similarity is obtained as "number of matching bases / number of effective aligned bases". The alignment results of all reads are aggregated by microbial category identifier, and a category score is calculated for each microbial category identifier. The category score is calculated as a weighted sum of the highest similarity of the category and the proportion of reads in the category. The microbial category identifier with the highest category score is selected as the candidate microbial category identifier. The similarity threshold is set to 0.95-0.98, and the dominance difference threshold is set to 0.02-0.05. A first microbial category discrimination condition is preset. The first microbial category discrimination condition is: when the highest similarity corresponding to the candidate microbial category identifier is not less than the similarity threshold and the difference between the category score of the candidate microbial category identifier and the second highest category score is not less than the dominance difference threshold, the microbial category identifier is output as the candidate microbial category identifier; when the first microbial category discrimination condition is not met, the microbial category identifier is output as the superior category identifier.
[0109] The method for calculating the category score includes:
[0110] ;
[0111] in, This represents the category score for category h. This represents the sequence similarity corresponding to category h. This indicates the percentage of the segment read corresponding to category h. The corresponding weighting coefficient is set to 0.5 to 0.8. For example, in the technical solution of the present invention, the weighting coefficient can be set to 0.55.
[0112] Microbial load indices are calculated and microbial load levels are output based on DNA characteristic data. Specifically, quantitative amplification is performed on the DNA extract to obtain the cycle threshold, and the version number of the standard curve parameters is read. The standard curve parameters include the standard curve slope and intercept, and the quantitative feed volume is set to 2–5 μL. The quantitative reaction copy number is calculated based on the cycle threshold and standard curve parameters. Using the quantitative reaction copy number as a conversion benchmark, combined with the quantitative feed volume and the elution and retention volumes from the DNA characteristic data, a volume conversion is performed to obtain the copy number per milliliter of sample solution as the microbial load index. The microbial load index is mapped to a microbial load level, which is divided into low load, medium load, and high load levels. The low load level corresponds to a copy number less than [missing value]. Copies / mL, medium load level corresponds to copy number at copies / ml, high load level corresponds to a copy number greater than Copy / mL. The cycle threshold, standard curve parameter version number, quantitative feed volume, microbial load index, microbial load level and sample number are bound and saved, and written to the smart sensor storage unit for subsequent selection and verification decisions of calibration strategy identifiers.
[0113] It should be noted that the cycle threshold for quantitative amplification detection of DNA extract can be obtained as follows: Read the DNA extract based on the sample number, add a sample to the quantitative amplification reaction system according to the preset amplification feed volume, the reaction system including primer pairs targeting the marker region of the target microorganism, polymerase, buffer, and fluorescence signal generating components; place the reaction system in a quantitative amplification instrument to perform thermal cycling amplification, and collect fluorescence signals at the end of each cycle to form a fluorescence curve; set a baseline interval and determine the fluorescence threshold line based on the fluorescence curve; record the cycle number corresponding to the first time the fluorescence signal exceeds the fluorescence threshold line as the cycle threshold, and bind the cycle threshold to the sample number for storage.
[0114] Table 1 shows examples of data obtained from the cycle threshold by quantitative amplification detection.
[0115] Table 1. Example of cycle threshold data obtained from quantitative amplification detection.
[0116]
[0117] It should be noted that at the end of each cycle, fluorescence signals are collected to form a fluorescence curve, the baseline is calculated according to the baseline interval, and a threshold fluorescence intensity is set; when the fluorescence signal first exceeds the threshold fluorescence intensity, the corresponding number of cycles is recorded as the cycle threshold, and the average value of the replicate wells is used as the cycle threshold output value for that sample number.
[0118] The method for calculating the quantitative reaction copy number includes:
[0119] The standard curve for the amplification of the set quantity is: ;
[0120] The equivalent can be written as: ;
[0121] Where C is the quantitative reaction copy number, representing the number of target nucleic acid copies participating in the amplification in this quantitative amplification reaction system, in units of "copy / reaction". T is the cycle threshold, i.e., the cycle threshold value output by the quantitative amplification detection, representing the number of cycles corresponding to when the amplification signal reaches the preset threshold. b is the standard curve intercept, obtained by fitting the standard curve, representing the... The theoretical value of the cycle threshold is given by k. k is the slope of the standard curve, obtained by fitting the standard curve; it is usually negative and used to characterize the strength of the linear relationship between the cycle threshold and the logarithm of the copy number, and the amplification efficiency. In this invention, the slope of the standard curve is constrained to [-3.8, -3.1] to ensure that the amplification efficiency is within a usable range.
[0122] It should be noted that in quantitative amplification (QA), the amplification signal follows the basic rule during the exponential growth phase: "the higher the initial template copy number, the fewer cycles are required to reach the preset fluorescence threshold." Using the initial copy number of the target nucleic acid in the reaction system as the QA copy number, when the amplification efficiency is within the usable range and the threshold setting remains consistent, the cycle threshold and the logarithm of the QA copy number have an approximately linear relationship; that is, the cycle threshold decreases linearly as the decimal logarithm of the copy number increases. Based on this linear relationship, a standard curve is established using a set of standards with known copy numbers, yielding the slope and intercept of the standard curve. In actual sample detection, the measured cycle threshold is substituted into the standard curve to deduce the QA copy number, thus forming... The conversion model uses the slope of the standard curve to characterize the response strength of the cycle threshold to changes in copy number, and the intercept of the standard curve to characterize the overall shift of the cycle threshold due to the threshold setting and system background, thereby achieving a reproducible quantitative estimate of the target nucleic acid copy number in the reaction system.
[0123] This formula is directly related to the problem of distorted spectral quantitative results caused by microbial contamination or microbial metabolic byproducts that this invention aims to solve. This invention does not rely solely on spectral characteristics to determine distortion, but rather uses DNA identification to confirm microbial interference and further outputs a microbial load level to drive subsequent correction strategy selection and review decisions. The microbial load level needs to be obtained as a comparable quantitative indicator from the quantitative amplification detection results. The cycle threshold itself is only the number of cycles and cannot directly reflect the microbial quantity; moreover, cycle thresholds between different batches and different reaction systems are not directly comparable. By converting the cycle threshold into the quantitative reaction copy number using a standard curve, the microbial quantity is transformed from device readings into the interpretable and thresholdable dimension of copy number. This allows for stable classification of microbial interference intensity in this invention and links the microbial load level with the microbial distortion marker on the spectral side, forming a closed-loop path of distortion identification, DNA confirmation, load classification, and strategy selection. This reduces misjudgments and false triggers caused by relying solely on spectral data and improves the reliability and traceability of the output of the content of each component of Ganoderma lucidum polysaccharide.
[0124] The formula for calculating the microbial load index is as follows:
[0125] ;
[0126] Where L is the microbial load index, and the unit is copies / mL. The quantitative feed volume represents the volume of DNA extract added to the quantitative amplification reaction system, expressed in microliters. Elution volume represents the volume of elution buffer added during the purification step to obtain the final DNA extract, in microliters (µL). U represents the sample retention volume, representing the volume of homologous sample processing solution used for DNA extraction, in milliliters (mL). 1000 is a unit conversion factor used to convert the µL / mL volume ratio to per milliliter (mL) caliber.
[0127] It should be noted that the formula for calculating the microbial load index establishes a backward relationship from the volume of the reaction system to the volume of the original sample, based on the quantitative reaction copy number. The quantitative reaction copy number C obtained by quantitative amplification detection represents the number of target nucleic acid copies carried by the added DNA extraction solution in a single quantitative amplification reaction system. Therefore, the quantitative reaction copy number C is related to the volume of DNA extraction solution added to the reaction system. Corresponding. Through The copy number density (copies / µL) per unit volume of the DNA extract can be obtained, which reflects the actual concentration level of the purified and eluted DNA extract. This is because the DNA extract originates from a homologous sample treatment solution of sample volume U, which has been enriched, lysed, and purified, and then eluted to a volume... The copy number density is obtained by elution and therefore multiplied by the elution volume. This allows us to calculate the estimated total copy number corresponding to the sample volume. Further dividing the total copy number by the sample volume U yields the copy number per milliliter of sample solution, i.e., the microbial load index L. The coefficient 1000 in the formula is used to... and The microliter unit and the milliliter unit of U are unified to make the load index output the standard caliber of copies / ml, thereby ensuring the comparability of load results under different sample numbers and different experimental batches.
[0128] This formula is directly related to the problem of distortion in spectral quantitative results caused by microbial contamination or microbial metabolic byproducts, which this invention aims to address. This invention not only needs to determine whether microorganisms cause distortion but also needs to quantify the intensity of microbial interference to drive the selection of subsequent compensation and correction strategies. Using only the cycle threshold or the copy number C of a single reaction cannot reflect the microbial load level in the original Ganoderma lucidum polysaccharide sample solution because the retention volume, elution volume, and quantitative feed volume of different samples may vary, resulting in C having only intra-reaction significance and lacking cross-sample comparability. By uniformly converting C to the copy number per milliliter of sample solution using the above volume back-calculation formula, the DNA confirmation result can be further extended from the presence or absence of microorganisms to the microbial load level, thus forming a closed-loop linkage with the microbial distortion labeling on the spectral side: when the microbial load level is high, the priority of retesting or strong correction strategies is increased; when the microbial load level is low, unnecessary treatment intensity is reduced and over-correction is avoided. Therefore, the microbial load index becomes a key dimension connecting DNA confirmation and spectral quantitative correction, enabling the present invention to more reliably identify and suppress measurement distortions caused by microorganisms, and ultimately improve the stability and traceability of the output of the content of each component of Ganoderma lucidum polysaccharide.
[0129] The following is an example of the data from this invention: The sample volume of sample number "sample 20250801-0008" was 2.0 mL. It was centrifuged at 11,000 rpm for 10 minutes and washed once with 0.5 mL of resuspension to obtain an enriched precipitate. The lysis parameters were as follows: 0.8 mL of lysis buffer and lysis enzyme were added at the preset concentration, the total volume of the lysis mixture was 0.9 mL, the incubation temperature was 60°C, the incubation time was 20 minutes, and the shaking time was 2 minutes. The magnetic bead purification elution volume was 50 μL, the DNA concentration was measured to be 4.1 ng / μL, the DNA purity index was 1.86, the lower limit of the DNA purity threshold was set to 1.70, the upper limit of the DNA purity threshold was set to 2.10, and the purity abnormality marker was not triggered. The target amplification feed mass was set at 10 nanograms, resulting in a calculated amplification feed volume of 2.4 μL and a dilution factor of 1. After read acquisition, 320 reads were retained, with read lengths concentrated at 420 ± 30 bases. The microbial reference sequence library version number was "Reference Library 202501-V2". Alignment showed a highest similarity of 0.97 for the candidate microbial category identifier, with a category score advantage difference of 0.04. Under similarity thresholds of 0.96 and advantage difference thresholds of 0.03, the output microbial category identifier was "Bacteria". Quantitative amplification detection yielded a cycle threshold of 28.9, with the standard curve parameter version number being "Standard Curve 202501-Q1". The standard curve slope was -3.35, and the standard curve intercept was 38.2. The quantitative feed volume was 3 μL, and the calculated microbial load index was... Copy / mL and map to high load level; bind and save microbial category identifier, microbial load level, DNA characteristic data, reference library version number, standard curve parameter version number and sample 20250801-0008, and write to the smart sensor storage unit.
[0130] Examples of microbial category identification data are shown in Table 2.
[0131] Table 2. Examples of Microbial Category Identification
[0132]
[0133] S4. Obtain component caliber parameters and generate a set of component number templates and a set of effective component numbers for the sample. The component caliber parameters include a molecular weight range threshold sequence and acidity property determination rules. Call the component content prediction model to output the initial calculation result of component content from the standardized spectral vector. Construct an initial calculation vector of component content according to the set of component number templates and write the corresponding initial calculation value of component content into the corresponding position of the initial calculation vector of component content according to the set of effective component numbers for the sample. When the microbial distortion label is a non-distortion label, output the component content result according to the initial calculation vector of component content. When the microbial distortion label is a distortion label, perform compensation correction on the initial calculation vector of component content based on the microbial category label and microbial load level to obtain the corrected component content result.
[0134] The implementation methods of S4 include:
[0135] A molecular weight interval threshold sequence and acidity property determination rules are obtained, and a component number template set is constructed to fix the scope. The molecular weight interval threshold sequence is used to map the polysaccharide molecular weight to a set of molecular weight interval numbers, which includes the first interval to the Gth interval. The acidity property determination rules output an acidity determination result identifier based on acidity characteristic indicators and preset acidity thresholds. The acidity determination result identifier includes an acidity identifier and a non-acidity identifier. Specifically, the acidity characteristic indicator is compared with the preset acidity threshold. If the acidity characteristic indicator is greater than or equal to the preset acidity threshold, the acidity determination result identifier is an acidity identifier; if the acidity characteristic indicator is less than the preset acidity threshold, the acidity determination result identifier is a non-acidity identifier. Based on the Cartesian combination relationship between the first interval to the Gth interval and the acidity identifier and the non-acidity identifier, a component number template set is pre-generated. Each component number in the component number template set consists of a molecular weight interval number and an acidity determination result identifier.
[0136] For example, in the technical solution of the present invention, the number of molecular weight intervals is set to G=6, and the molecular weight interval threshold sequence is denoted as {10,30,100,300,800}kDa, where kDa is the unit of molecular weight. The first interval is [0,10)kDa, the second interval is [10,30)kDa, the third interval is [30,100)kDa, the fourth interval is [100,300)kDa, the fifth interval is [300,800)kDa, and the sixth interval is [800,+∞)kDa.
[0137] It should be noted that the component number template set covers all combinations from "first interval and acidic label, first interval and non-acidic label" to "G interval and acidic label, G interval and non-acidic label"; the component number template set is used to fix the data structure dimension and index scope of compensation parameters for subsequent component content results.
[0138] A set of effective component numbers for each sample is generated based on the sample number, ensuring that only one acidity category corresponds to the same molecular weight range. For each sample number, the polysaccharide molecular weight corresponding to that sample number is mapped to a set of sample molecular weight range numbers based on the molecular weight range threshold sequence. The acidity characteristic index corresponding to that sample number is calculated based on the acidity property determination rules, and a unique acidity determination result identifier is output by comparing the acidity characteristic index with the acidity threshold. The acidity characteristic index is defined as the ratio of the acidity peak integral value to the skeleton peak integral value. The acidity peak integral value is the integral result of the standardized spectral vector within the acidity peak range, and the skeleton peak integral value is the integral result of the standardized spectral vector within the skeleton peak range. After outputting the acidity determination result identifier, each molecular weight range number in the sample molecular weight range number set is combined with this unique acidity determination result identifier to generate a set of effective component numbers for the sample.
[0139] For example, in the technical solution of this invention, the unit for peak intervals in the mid-infrared spectrum is usually written as cm. -1 The acid peak range can be set to [1400, 1450] cm. -1 The skeletal peak range can be set to [900, 1200] cm. -1 .
[0140] For example, assuming an acidity threshold of 0.20, the acidity peak integral value of sample one is 0.84, the skeleton peak integral value is 6.00, and the acidity characteristic index is 0.84 / 6.00≈0.14, resulting in a non-acidity label. Sample two has an acidity peak integral value of 1.05, a skeleton peak integral value of 5.50, and an acidity characteristic index of 1.05 / 5.50≈0.1909, also resulting in a non-acidity label. Sample three has an acidity peak integral value of 1.22, a skeleton peak integral value of 5.40, and an acidity characteristic index of 1.22 / 5.40≈0.2259, resulting in an acidity label. Taking sample three as an example, when the molecular weight interval number obtained by mapping the polysaccharide molecular weight interval threshold sequence of sample three is the second interval, a component number is generated based on "second interval + acidity label".
[0141] It should be noted that the integral value of the acid peak is calculated by discretely summing the response intensities of each sampling point within the acid peak interval according to the sampling step size, and the integral value of the skeleton peak is calculated by discretely summing the response intensities of each sampling point within the skeleton peak interval according to the sampling step size. The number of component numbers in the effective component number set of the sample is consistent with the number of intervals in the sample molecular weight interval number set, so that only one component number is generated for each molecular weight interval. Component numbers corresponding to other acid categories in the component number template set that are not selected by the sample number are not included in the effective component number set of the sample.
[0142] Based on the sample number, a standardized spectral vector is read and the component content prediction model is invoked to obtain a mapping table of the initial calculated component content values corresponding to each component number in the set of valid component numbers for the sample. An initial calculated component content vector is constructed based on the component number template set. Each component number in the template set is traversed. When the current component number does not belong to the set of valid component numbers for the sample, a null value is written at the corresponding index position in the initial calculated component content vector. When the current component number belongs to the set of valid component numbers for the sample, the initial calculated component content value is written at the index position of the current component number in the initial calculated component content vector. This results in an initial calculated component content vector that corresponds one-to-one with the set of component number templates and has a fixed scope.
[0143] An example of the mapping table for the initial calculated values of component content is shown in Table 3.
[0144] Table 3. Examples of mapping representations of initial calculated component content values
[0145]
[0146] It should be noted that the initial vector for calculating component content is denoted as... Where J is the number of component numbers in the component number template set, and j ranges from 1 to J. This represents the initial calculated value of the component content at the j-th position in the initial component content calculation vector. The j-th position corresponds one-to-one with the j-th component number in the component number template set. This represents the initial calculated value of the component content corresponding to the j-th component number.
[0147] When the microbial distortion label is set to a non-distortion label, the component content results are output based on the initial vector calculated from the component content.
[0148] like Figure 3 As shown, when the microbial distortion label is a distortion label, the method for performing compensation correction on the initial calculated vector of component content based on the microbial category label and microbial load level to obtain the corrected component content result includes:
[0149] Read the microbial category identifier and microbial load level corresponding to the sample number, and use the microbial category identifier, microbial load level and component number as indexes to read the multiplicative compensation coefficient vector and additive compensation bias vector from the pre-constructed compensation parameter table; the multiplicative compensation coefficient vector and additive compensation bias vector are arranged according to the sequence position of the component number template set, and correspond one-to-one with the component content initial calculation vector.
[0150] It should be noted that the multiplicative compensation coefficient vector is denoted as... The additive compensation bias vector is denoted as .in, Let be the multiplicative compensation coefficient at the j-th position in the multiplicative compensation coefficient vector. This is the additive compensation bias at the j-th position in the additive compensation bias vector.
[0151] Based on the multiplicative compensation coefficient vector, the additive compensation bias vector, and the initial calculation vector of component content, the component content correction formula is set as follows: the corrected component content value equals the initial calculation value of component content multiplied by the corresponding multiplicative compensation coefficient plus the corresponding additive compensation bias. For each valid component number in the set of valid sample component numbers, the initial calculation value of the component content corresponding to that valid component number is obtained from the initial calculation vector of component content, and the corrected component content value corresponding to that valid component number is calculated according to the component content correction formula. Optionally, when the corrected component content value is less than zero, it is truncated to zero and marked with a truncation flag. The corrected component content value corresponding to each valid sample component number is written into the corresponding index position in the corrected component content vector. The index positions corresponding to non-valid sample component numbers remain unchanged with a null value flag, resulting in the corrected component content vector. The corrected component content result is then output according to the component number template set.
[0152] It should be noted that the construction of the compensation parameter table is driven by the technical problem that microbial contamination or microbial metabolic byproducts cause distortion of spectral quantitative results, which this invention aims to address. Its core objective is not to fit the component content of arbitrary samples, but rather to establish a reproducible mapping relationship between the systematic deviation of spectral prediction results relative to the true content under specific microbial interference conditions and the compensation correction parameters. To this end, this invention solidifies microbial interference into two quantifiable and gradable fields as table construction conditions: microbial category identifier and microbial load level. The microbial category identifier is used to distinguish the types of metabolic byproducts and spectral background interference morphologies produced by different microbial communities in the sample solution, while the microbial load level is used to characterize the magnitude difference in interference intensity. When constructing the table, the microbial category identifier and microbial load level are used as primary conditions, enabling the compensation parameter table to provide differentiated compensation for different interference sources and intensities, rather than using a single parameter to cover all scenarios, thus avoiding over- or under-correction.
[0153] In constructing the data foundation, this invention collects Ganoderma lucidum polysaccharide sample solutions covering multiple batches and multiple load levels as a calibration sample set, and simultaneously acquires three types of data for each calibration sample: First, a standardized spectral vector is obtained by mid-infrared spectral acquisition and standardization processing according to this invention, and an initial calculation vector of component content is output based on the component number template set, which is used to characterize the perturbation prediction result of the spectral model under microbial interference conditions; Second, microbial category identifiers and microbial load indicators are obtained by DNA identification, and the microbial load indicators are discretely mapped to microbial load levels, which are used to characterize the interference category and interference intensity of the sample; Third, the reference component content results of the same calibration sample under the same component number caliber are obtained by reference detection method, which is used as the true content benchmark. The above three types of data are bound by sample number to form a complete calibration record, so that each component number can simultaneously obtain the ternary relationship of "initial calculation content value - reference content value - microbial interference conditions", which directly corresponds to the technical problem of this invention in terms of data structure: the distortion caused by microorganisms is manifested as the deviation of the initial calculation content value from the reference content value.
[0154] In the parameter solving stage, this invention groups calibration records by microbial category identifier and microbial load level, and estimates compensation parameters one by one according to the sequence position of the component number template set within each group. For any component number, the initial calculated content value sequence and reference content value sequence of all calibration samples within that group are collected, and the systematic deviation between the two is modeled as a linear compensation relationship. That is, proportional distortion is corrected by multiplicative compensation coefficient, and systematic offset caused by background rise is corrected by additive compensation bias. Based on this, a pair of parameters is solved under that group and that component number, and the parameters corresponding to all component numbers are arranged in the order of the template set to form a multiplicative compensation coefficient vector and an additive compensation bias vector. The compensation parameter table constructed in this way realizes an explicit mapping from microbial interference conditions to full component compensation parameters. This allows the matching set of compensation parameters to be directly indexed after reading the microbial category identifier and microbial load level in actual detection, and the initial calculated vector of component content is corrected bit by bit, thereby specifically eliminating spectral distortion caused by microbial contamination or metabolic byproducts, and improving the stability, comparability, and traceability of the output of the content of each component of Ganoderma lucidum polysaccharide.
[0155] Example 2
[0156] Please see Figure 2 As shown in the figure, this embodiment provides a detection system for determining the content of various components of Ganoderma lucidum polysaccharides, including a spectral acquisition standardization module, a microbial distortion discrimination module, a nucleic acid typing module, and a component prediction and correction module. Each module is connected by wires and / or wirelessly to realize data transmission.
[0157] The spectral acquisition and standardization module is used to acquire Ganoderma lucidum polysaccharide sample solutions and generate sample numbers, acquire spectral acquisition parameters and acquire raw spectra; and perform standardization processing on the raw spectra to obtain standardized spectral vectors.
[0158] The microbial distortion discrimination module calculates microbial interference features, including scattering interference index and residual structure index, based on standardized spectral vectors; it outputs microbial distortion identifiers based on microbial interference features and preset interference thresholds, and outputs DNA recognition trigger identifiers when the microbial distortion identifiers meet the DNA confirmation trigger conditions.
[0159] The nucleic acid typing module, when a DNA recognition trigger is present, obtains homologous sample processing solution based on the sample number and performs nucleic acid extraction to obtain DNA feature data; generates DNA sequence feature data based on the DNA feature data and compares it with the microbial reference sequence library to obtain microbial category identification; calculates microbial load index based on DNA feature data and outputs microbial load level.
[0160] The component prediction and correction module is used to acquire component caliber parameters and generate a set of component number templates and a set of effective component numbers for the sample. The component caliber parameters include a molecular weight range threshold sequence and acidity property determination rules. It calls the component content prediction model to output the initial calculation result of the component content from the standardized spectral vector, constructs an initial calculation vector of component content according to the set of component number templates, and writes the corresponding initial calculation value of the component content into the corresponding position of the initial calculation vector of component content by referring to the set of effective component numbers for the sample. When the microbial distortion identifier is non-distortion, the component content result is output according to the initial calculation vector of component content. When the microbial distortion identifier is distortion, compensation correction is performed on the initial calculation vector of component content based on the microbial category identifier and microbial load level to obtain the corrected component content result.
[0161] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0162] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for determining the content of various components of Ganoderma lucidum polysaccharides, characterized in that, include: S1. Obtain the Ganoderma lucidum polysaccharide sample solution and generate a sample number, obtain the spectral acquisition parameters and acquire the raw spectrum; The original spectrum is normalized to obtain a normalized spectral vector; S2. Microbial interference characteristic quantities, including scattering interference index and residual structure index, are calculated based on standardized spectral vectors. Microbial distortion identifiers are output based on microbial interference characteristics and preset interference thresholds, and DNA recognition trigger identifiers are output when the microbial distortion identifiers meet the DNA confirmation trigger conditions. S3. When a DNA recognition trigger is present, obtain homologous sample processing solution based on sample number and perform nucleic acid extraction to obtain DNA feature data; generate DNA sequence feature data based on DNA feature data and compare it with the microbial reference sequence library to obtain microbial category identifier; calculate microbial load index based on DNA feature data and output microbial load level; S4. Obtain component caliber parameters and generate a set of component number templates and a set of effective component numbers for the sample. The component caliber parameters include a molecular weight range threshold sequence and acidity property determination rules. Call the component content prediction model to output the initial calculation result of component content from the standardized spectral vector. Construct an initial calculation vector of component content according to the set of component number templates and write the corresponding initial calculation value of component content into the corresponding position of the initial calculation vector of component content according to the set of effective component numbers for the sample. When the microbial distortion label is non-distortion label, output the component content result according to the initial calculation vector of component content. When the microbial distortion label is distortion label, perform compensation correction on the initial calculation vector of component content based on the microbial category label and microbial load level to obtain the corrected component content result.
2. The method for determining the content of various components of Ganoderma lucidum polysaccharides according to claim 1, characterized in that, The methods for obtaining the component number template set and the sample valid component number set include: Based on the molecular weight interval threshold sequence, the molecular weight of polysaccharides is mapped to a set of molecular weight interval numbers, which includes the first interval to the Gth interval. The acidity property determination rule is to output an acidity determination result identifier based on the acidity characteristic index and the preset acidity threshold. The acidity determination result identifier includes an acidity identifier and a non-acidity identifier. The acidity characteristic index is defined as the ratio of the acidity peak integral value to the skeleton peak integral value. The acidity peak integral value is the integral result of the standardized spectral vector in the acidity peak interval, and the skeleton peak integral value is the integral result of the standardized spectral vector in the skeleton peak interval. Based on the Cartesian combination relationship between the first interval to the Gth interval and the acidity determination result identifier, a set of component number templates is generated. For each sample number, the polysaccharide molecular weight corresponding to the sample number is mapped to a set of sample molecular weight interval numbers based on the molecular weight interval threshold sequence; the acidity determination result identifier corresponding to the sample number is calculated based on the acidity property determination rule; and each molecular weight interval number in the sample molecular weight interval number set is combined with the corresponding acidity determination result identifier to generate a set of effective component numbers for the sample.
3. The method for determining the content of various components of Ganoderma lucidum polysaccharides according to claim 1, characterized in that, The method of calling the component content prediction model to output the preliminary calculation results of component content from the standardized spectral vector, constructing the preliminary calculation vector of component content according to the component number template set, and writing the corresponding preliminary calculation values of component content into the corresponding positions of the preliminary calculation vector of component content by referring to the set of valid component numbers of the sample includes: Based on the sample number, the standardized spectral vector is read and the component content prediction model is called to obtain a mapping table of the initial calculated component content values corresponding to each component number in the set of valid component numbers of the sample. Based on the component number template set, an initial calculated component content vector is constructed. Each component number in the component number template set is traversed. When the current component number does not belong to the set of valid component numbers of the sample, a null value mark is written in the sequence position corresponding to the component number in the initial calculated component content vector. When the current component number belongs to the set of valid component numbers of the sample, the initial calculated component content value is written in the sequence position of the component number in the initial calculated component content vector.
4. The method for determining the content of various components of Ganoderma lucidum polysaccharides according to claim 1, characterized in that, Methods for compensating and correcting the initial component content vector based on microbial category identification and microbial load level to obtain the corrected component content results include: Read the microbial category identifier and microbial load level corresponding to the sample number, and use the microbial category identifier, microbial load level and component number as indexes to read the multiplicative compensation coefficient vector and additive compensation bias vector from the pre-constructed compensation parameter table; Based on the multiplicative compensation coefficient vector, the additive compensation bias vector, and the initial calculation vector of component content, a component content correction formula is set. Each valid component number in the set of valid sample component numbers is traversed, and the initial calculation value of the component content corresponding to that valid sample component number is obtained from the initial calculation vector of component content. The corrected value of the component content corresponding to that valid sample component number is then calculated according to the component content correction formula. The corrected value of the component content corresponding to each valid sample component number is written into the corresponding index position in the corrected component content vector. The index positions corresponding to non-valid sample component numbers remain unchanged, resulting in the corrected component content vector. The corrected component content result is then output according to the component number template set.
5. The method for determining the content of various components of Ganoderma lucidum polysaccharides according to claim 1, characterized in that, The method for obtaining the DNA feature data includes: Based on the sample number, the homologous sample processing solution is retrieved from the sample management record; the homologous sample processing solution is subjected to microbial enrichment pretreatment to obtain an enriched precipitate, the microbial enrichment pretreatment including centrifugation, discarding the supernatant and retaining the precipitate, adding buffer for resuspension, and repeated washing; lysis buffer and cell wall lysis enzyme are added to the enriched precipitate to form a lysis mixture, and the lysis mixture is incubated, mechanically disrupted, and centrifuged to obtain a lysis supernatant; magnetic bead adsorption solution is added to the lysis supernatant and mixed well to adsorb DNA onto the surface of the magnetic beads; the magnetic beads are placed on a magnetic rack for separation and the supernatant is discarded, and washing is performed; a preset elution volume of elution solution is added to the magnetic beads for elution to obtain a DNA extract; the DNA concentration and DNA purity index of the DNA extract are measured; the DNA concentration, DNA purity index, sample retention volume, and elution volume are packaged into DNA characteristic data; the sample retention volume is the volume of homologous sample processing solution used for DNA extraction in this case.
6. The method for determining the content of various components of Ganoderma lucidum polysaccharides according to claim 1, characterized in that, Methods for generating DNA sequence feature data based on DNA feature data include: The process involves: reading the version number of the microbial reference sequence library, which contains labeled region reference sequences and corresponding microbial category identifiers; constructing amplification feed parameters based on DNA feature data, including amplification feed volume and dilution factor; performing labeled region amplification on the DNA extract based on the amplification feed parameters to obtain amplification products, and acquiring reads from the amplification products to form a sequence read set; performing quality and length screening on the sequence read set, retaining reads with a length between 250 and 550 bases to obtain retained read sequences; and encapsulating the retained read sequence set and its read count information into DNA sequence feature data.
7. The method for determining the content of various components of Ganoderma lucidum polysaccharides according to claim 6, characterized in that, The method for obtaining the microbial category identifier includes: Each read in the DNA sequence feature data is aligned with a microbial reference sequence library to calculate sequence similarity. The alignment results of all reads are aggregated by microbial category identifier, and a category score is calculated for each microbial category identifier. The category score is calculated as a weighted sum of the highest similarity of the category and the proportion of reads in the category. The microbial category identifier with the highest category score is selected as the candidate microbial category identifier. A similarity threshold and an advantage difference threshold are preset, and a first microbial category discrimination condition is preset. The first microbial category discrimination condition is as follows: when the highest similarity corresponding to the candidate microbial category identifier is not less than the similarity threshold and the difference between the category score of the candidate microbial category identifier and the second highest category score is not less than the advantage difference threshold, the microbial category identifier is output as the candidate microbial category identifier; when the first microbial category discrimination condition is not met, the microbial category identifier is output as the superior category identifier.
8. The method for determining the content of various components of Ganoderma lucidum polysaccharides according to claim 1, characterized in that, Methods for calculating microbial load indices and outputting microbial load levels based on DNA feature data include: Quantitative amplification detection was performed on the DNA extract to obtain the cycle threshold, and the version number of the standard curve parameters was read. The standard curve parameters include the standard curve slope and the standard curve intercept, and a preset quantitative feed volume was used. Based on the cycle threshold and the standard curve parameters, the quantitative reaction copy number was calculated. Using the quantitative reaction copy number as a conversion benchmark, and combined with the quantitative feed volume and the elution volume and retention volume in the DNA characteristic data, a volume conversion was performed to obtain the copy number corresponding to each milliliter of sample solution as a microbial load index. The microbial load index was mapped to a microbial load level, which includes low load level, medium load level and high load level.
9. The method for determining the content of various components of Ganoderma lucidum polysaccharides according to claim 1, characterized in that, The implementation methods of S2 include: Read the normalized spectral vector based on the sample number and obtain the corresponding reference normalized spectral vector; Linear fitting is performed on the normalized spectral vector and the reference normalized spectral vector to obtain the second multiplicative coefficient and the second additive coefficient; the multiplicative deviation is calculated based on the second multiplicative coefficient; the additive deviation is calculated based on the second additive coefficient; the additive scaling factor is obtained; the multiplicative deviation and the additive deviation are divided by the additive scaling factor and added to obtain the scattering interference index; The spectral reconstruction model is invoked to perform reconstruction inference on the standardized spectral vector to obtain the reconstructed spectral vector; the standardized spectral vector and the reconstructed spectral vector are subtracted point by point to obtain the residual sequence; the residual energy is calculated based on the residual sequence; the residual difference mean square is calculated based on the residual sequence to obtain the stability term parameter; the residual difference mean square is divided by the sum of the residual energy and the stability term parameter to obtain the residual structure index. A comprehensive interference score is calculated based on the microbial interference characteristics. The comprehensive interference score is compared with a preset interference threshold to obtain a microbial distortion label, which includes both distortion and non-distortion labels. When the microbial distortion label is a distortion label and the comprehensive interference score is greater than the preset DNA trigger threshold, a DNA recognition trigger label is output.
10. The method for determining the content of various components of Ganoderma lucidum polysaccharides according to claim 1, characterized in that, The implementation methods of S1 include: Obtain Ganoderma lucidum polysaccharide sample solutions and generate a sample number for each Ganoderma lucidum polysaccharide sample solution; Spectral acquisition parameters are obtained, including wavelength start point, wavelength end point, wavelength step size, integration time, and number of scans. A wavelength sequence is generated based on the wavelength start point to the wavelength end point and according to the wavelength step size. The spectral acquisition unit is controlled to scan and acquire the Ganoderma lucidum polysaccharide sample solution according to the wavelength sequence to obtain the original spectrum. The original spectrum is then subjected to standardization processing to obtain a standardized spectral vector. The standardization processing includes dark current subtraction, baseline correction, noise reduction, scattering correction, and normalization.