Microbial colony quantitative analysis method based on fusion of conductance change and TOC (Total Organic Carbon) model
By fusing conductivity changes with the TOC model, and combining multi-source data with machine learning algorithms, the problem of long processing time and low accuracy in traditional microbial colony quantification methods has been solved, achieving high-precision, real-time microbial colony quantification analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-04-21
AI Technical Summary
Existing methods for quantifying microbial colonies are time-consuming, susceptible to human error and environmental factors, and fail to effectively handle the variability in carbon dioxide production rates, resulting in inaccurate quantitative results.
A method combining conductivity change and TOC model was adopted. By collecting conductivity difference and carbon dioxide concentration data in real time, and combining them with environmental temperature and humidity parameters, a multilayer perceptron neural network and state space model were used for feature extraction and smoothing. A Gaussian process regression model was constructed to calculate the colony equivalent and its confidence interval, and the accuracy was verified by sterilized sucrose.
It achieves high-precision, real-time quantitative analysis of microbial colonies, reduces the impact of carbon dioxide production rate variability, improves the accuracy and efficiency of quantitative results, and adapts to the metabolic characteristics of different microbial species and sample matrices.
Smart Images

Figure CN121901557A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, and in particular to a method for quantitative analysis of microbial colonies based on the fusion of conductivity changes and TOC models. Background Technology
[0002] Quantitative analysis of microbial colonies is a core component in ensuring food safety, environmental quality monitoring, and pharmaceutical quality control. In the food industry, quantitative analysis is needed to control the levels of microbial contamination in raw materials and finished products. In environmental monitoring, the number of microorganisms is used to quantify the degree of pollution in water and soil. Therefore, efficient and accurate methods for quantitative analysis of microbial colonies play an irreplaceable role in quality control across various fields.
[0003] Current mainstream methods for quantifying microbial colonies have several technical limitations. Traditional methods, such as plate counting, require inoculating microbial samples into a culture medium and culturing for 24-48 hours until colonies form before manual counting. This method is not only time-consuming and cannot meet real-time monitoring needs, but it is also susceptible to variations in inoculation volume, culture conditions, and subjective errors in manual counting, resulting in poor repeatability of quantitative results. To improve efficiency, existing automated methods, such as the vision-assisted carbon-based difference colony calculation method proposed in patent CN119169619A, achieve automated counting by extracting colony features through image processing. However, such methods are highly dependent on visual data and are easily affected by changes in light intensity and background interference, leading to distorted feature extraction. Furthermore, these methods do not consider the dynamic changes in carbon dioxide production during microbial metabolism. For example, different microbial species and growth stages can lead to significant differences in carbon dioxide production rates. Ignoring this dynamic can cause bias in the estimation of the mean colony count. In addition, existing automated algorithms lack the ability to adaptively handle variability in carbon dioxide production rates. In scenarios with small samples or non-uniform distribution, it is difficult to accurately quantify the colony count with limited data, resulting in a significant decrease in accuracy.
[0004] Therefore, there is an urgent need for a quantitative analysis method for microbial colonies based on the fusion of conductivity changes and the TOC model to solve the above problems. Summary of the Invention
[0005] The purpose of this invention is to provide a method for quantitative analysis of microbial colonies based on the fusion of conductivity changes and the TOC model, comprising the following steps: Real-time acquisition of conductivity difference data and carbon dioxide concentration data of microbial samples, and simultaneous recording of environmental temperature and humidity parameters to form a multi-source input vector; Based on the conductivity difference data, the conductivity difference data is mapped to carbon-based values and combined with carbon dioxide concentration data to extract microbial quantity characteristics; The extracted microbial quantity characteristics are subjected to time series analysis, smoothing and state estimation, and the smoothed microbial quantity estimates and their statistical covariance data are output. The probability distribution of microbial quantity is obtained based on the smoothed microbial quantity estimate and the statistical covariance data value, and the colony equivalent and its confidence interval are calculated and output based on the probability distribution of microbial quantity. The accuracy of the colony equivalent is verified, and the verified final colony quantification result is output.
[0006] Furthermore, the step of real-time acquisition of conductivity difference data and carbon dioxide concentration data of microbial samples, and simultaneous recording of environmental temperature and humidity parameters to form a multi-source input vector includes: The initial conductivity value of the microbial sample before complete decomposition of organic carbon and the final conductivity value after complete decomposition were obtained respectively. The conductivity difference data was obtained by calculating the difference between the final conductivity value and the initial conductivity value. Continuously collect data on the concentration of carbon dioxide produced by microbial metabolism; Simultaneously monitor the environmental temperature and humidity parameters of the microbial samples; Data preprocessing was performed on conductivity difference data, carbon dioxide concentration data, ambient temperature parameters, and ambient humidity parameters. The preprocessed conductivity difference data, carbon dioxide concentration data, ambient temperature parameters, and ambient humidity parameters are combined according to time series to form a multi-source input vector, and a timestamp is added to each data point.
[0007] Furthermore, the step of mapping the conductivity difference data to carbon-based values and combining it with carbon dioxide concentration data to extract microbial abundance characteristics includes: Map the conductivity difference data to carbon-based values and output the first and second carbon-based values; A multilayer perceptron neural network is constructed as a machine learning regression model. The input layer of the neural network receives a multi-source input vector containing carbon dioxide concentration data, conductivity difference data, ambient temperature parameters, and ambient humidity parameters. The neural network structure is configured with a linear rectified unit as the activation function and the output layer is set as a single neuron structure to output the preliminary estimate of the number of microorganisms. The neural network model was trained using a historical calibration dataset, which contained sample data with known numbers of microorganisms. The first carbon-based value, the second carbon-based value, and the preliminary microbial quantity estimate output by the neural network are fused at the feature level to form microbial quantity features.
[0008] Furthermore, the steps of performing time series analysis on the extracted microbial quantity characteristics, smoothing the data, estimating the state, and outputting the smoothed microbial quantity estimate and its statistical covariance data include: The microbial quantity characteristics are reconstructed over time to form a microbial quantity time series; The microbial population time series was subjected to stationarity testing and detrending processing to extract the steady-state growth component and short-term fluctuation component from the sequence. Establish a state-space model, define state variables including microbial quantity state and microbial growth rate state, and define observation variables as microbial quantity characteristics after time series analysis. The state-space model is solved, and the microbial quantity characteristics are smoothed and the optimal state is estimated through iterative calculation. The smoothed estimate of the number of microorganisms and its corresponding statistical covariance data are output synchronously.
[0009] Furthermore, the step of obtaining the probability distribution of microbial quantity based on the smoothed microbial quantity estimate and the statistical covariance data value, and calculating and outputting the colony equivalent and its confidence interval based on the probability distribution of microbial quantity includes: Using a Gaussian process regression framework, the smoothed microbial quantity estimate is used as the mean of the training samples, and the statistical covariance data value is used as the noise variance of the training samples to construct the Gaussian process prior distribution. Radial basis functions are selected as kernel functions, and the similarity between training samples is calculated through the kernel functions to construct the covariance matrix; Calculate the posterior distribution of the Gaussian process to obtain the complete probability distribution of the number of microorganisms; The mean of the microbial population is extracted from the probability distribution as a baseline value for colony equivalent, and the confidence interval is calculated based on the standard deviation of the probability distribution. A variance threshold monitoring mechanism is set up so that when the variance of the output probability distribution exceeds a preset threshold, the data quality check process is automatically triggered.
[0010] Furthermore, the step of verifying the accuracy of the colony equivalent and outputting the verified final colony quantification result includes: Establish a sterile sucrose validation database, which stores the colony equivalent baseline values and their corresponding validation status identifiers obtained in historical sterile sucrose validation experiments. Obtain the colony equivalent and its confidence interval of the current batch of microbial samples, and assess the reliability level of the current measurement results based on the width of the confidence interval; The similarity between the current colony equivalent and the benchmark value in the sterile sucrose validation database is calculated, and it is determined whether the current colony equivalent passes the validation. When the similarity value is within the preset range, a verification pass mark is generated, and the current colony equivalent is marked as the verified final colony quantification result; When the similarity value is outside the preset range, a verification failure flag is generated, triggering a model parameter calibration command and outputting a warning message suggesting that data be collected again.
[0011] Furthermore, this invention also discloses a microbial colony quantitative analysis system based on the fusion of conductivity changes and the TOC model, comprising: The acquisition module is used to collect the conductivity difference data and carbon dioxide concentration data of microbial samples in real time, and simultaneously record the ambient temperature and humidity parameters to form a multi-source input vector; The extraction module is used to map the conductivity difference data into carbon-based values based on the conductivity difference data, and combine it with carbon dioxide concentration data to extract microbial quantity characteristics. The analysis module is used to perform time series analysis on the extracted microbial quantity characteristics, and to perform smoothing and state estimation, outputting the smoothed microbial quantity estimate and its statistical covariance data value. The calculation module is used to obtain the probability distribution of the number of microorganisms based on the smoothed microbial quantity estimate and the statistical covariance data value, and to calculate and output the colony equivalent and its confidence interval based on the probability distribution of the number of microorganisms. The verification module is used to verify the accuracy of the colony equivalent and output the verified final colony quantification result.
[0012] Furthermore, the analysis module includes: The reconstruction unit is used to reconstruct the microbial quantity characteristics over time to form a microbial quantity time series. The extraction unit is used to perform stationarity testing and detrending processing on the microbial quantity time series, and to extract the steady-state growth component and short-term fluctuation component from the sequence. Establish a unit to build a state-space model, define state variables including microbial quantity state and microbial growth rate state, and define observation variables as microbial quantity characteristics after time series analysis; The computing unit is used to solve the state-space model and to smooth the microbial quantity characteristics and estimate the optimal state through iterative calculation. The output unit is used to synchronously output the smoothed estimate of the number of microorganisms and its corresponding statistical covariance data.
[0013] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method for quantitative analysis of microbial colonies based on the fusion of conductivity changes and TOC model.
[0014] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method for quantitative analysis of microbial colonies based on the fusion of conductivity changes and a TOC model.
[0015] The beneficial effects of this application are as follows: Firstly, this invention employs a full-process hybrid model, extracting multi-dimensional features of conductivity difference, carbon dioxide concentration, and environmental parameters through machine learning regression algorithms. It combines time series analysis to model microbial growth dynamics to compensate for short-term fluctuations in carbon dioxide. Finally, it uses Gaussian process regression to regress the probability distribution of the number of microorganisms generated. This design comprehensively covers the impact of variability in carbon dioxide production rates. Based on simulation data verification, the mean estimation error is reduced compared to existing image recognition methods. At the same time, the closed-loop mechanism verified by sterilized sucrose further ensures the accuracy of quantitative results, effectively solving the bias problem caused by neglecting the dynamics of carbon dioxide in existing methods.
[0016] Secondly, this invention adopts a modular design. The data acquisition layer is based on a high-precision conductivity meter and a non-dispersive infrared carbon dioxide sensor to achieve real-time data acquisition. The feature extraction and dynamic processing layer algorithms are lightweight and can be deployed on edge devices, which greatly improves efficiency compared with the traditional flat plate counting method and meets the needs of online monitoring. At the same time, the algorithm supports online updates. Through continuous training with historical calibration datasets and new scene samples, it can adapt to the metabolic characteristics of different microbial species and different sample matrices, solving the problem of limited adaptability of existing methods. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of a method flow proposed in an embodiment of this application.
[0018] Figure 2 This is a schematic diagram of the system structure proposed in an embodiment of the present invention.
[0019] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0020] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0021] like Figure 1As shown, this application provides a method for quantitative analysis of microbial colonies based on the fusion of conductivity changes and the TOC model, including the following steps: S1 collects real-time conductivity difference data and carbon dioxide concentration data of microbial samples, and simultaneously records environmental temperature and humidity parameters to form a multi-source input vector; S2, Based on the conductivity difference data, the conductivity difference data is mapped to carbon-based values using the total organic carbon algorithm, and combined with carbon dioxide concentration data, a machine learning regression model is used to extract microbial quantity characteristics; S3. Time series analysis is performed on the extracted microbial quantity characteristics. A state-space model is applied for smoothing and state estimation. The smoothed microbial quantity estimate and its statistical covariance data are output to compensate for short-term fluctuations in carbon dioxide concentration. S4. Obtain the probability distribution of the number of microorganisms based on the smoothed microbial quantity estimate and the statistical covariance data value, and calculate and output the colony equivalent and its confidence interval based on the probability distribution of the number of microorganisms. S5, verify the accuracy of the colony equivalent using the sterile sucrose verification method, and output the verified final colony quantification result.
[0022] As described in steps S1-S5 above, microorganisms continuously consume organic carbon in the environment and produce carbon dioxide during their growth and metabolism. Their metabolic activity is directly related to the number of colonies, and the conductivity difference can reflect the degree of change in organic carbon content. Changes in carbon dioxide concentration can reflect the dynamic process of microbial metabolism. However, differences in microbial species and different growth stages can lead to significant variability in the rate of carbon dioxide production. A single data source is difficult to comprehensively and accurately quantify the number of colonies. In addition, environmental temperature and humidity directly affect the growth status of microorganisms, further interfering with the accuracy of quantitative results.
[0023] In traditional methods for quantifying microbial colonies, plate counting is cumbersome, time-consuming, and highly susceptible to human error. Existing automated detection methods often rely on a single data source or algorithm, making them vulnerable to interference from environmental factors such as light and background. They also struggle to effectively address the variability in carbon dioxide production rates, resulting in significant deviations in quantitative results. Therefore, it is necessary to develop a technical solution that integrates multi-source data and processes it using scientific algorithms to eliminate the adverse effects of variability and environmental interference, thereby achieving accurate quantification of microbial colony counts.
[0024] This invention specifically employs a multi-source data fusion and multi-algorithm collaborative processing technical solution. By integrating conductivity difference, carbon dioxide concentration, and environmental temperature and humidity parameters, and combining steps such as feature extraction, time series optimization, and probability modeling, it solves the problem of quantitative inaccuracy caused by the limitations and variability of a single data source. Through a complete technical process of real-time acquisition of multi-source related data, mapping carbon-based numerical values to extract microbial quantitative characteristics, optimizing the estimation results through time series analysis, constructing probability distribution to calculate core parameters, and completing accuracy verification, it achieves high-precision, real-time quantitative analysis of microbial colonies while ensuring the reliability of the analysis results.
[0025] In one embodiment, the steps of real-time acquisition of conductivity difference data and carbon dioxide concentration data of microbial samples, and simultaneous recording of environmental temperature and humidity parameters, to form a multi-source input vector include: S11. A high-precision conductivity meter is used to obtain the initial conductivity value of the microbial sample before the complete decomposition of organic carbon and the final conductivity value after the complete decomposition. The conductivity difference data is obtained by calculating the difference between the final conductivity value and the initial conductivity value. S12 uses a non-dispersive infrared carbon dioxide sensor to continuously collect carbon dioxide concentration data produced by microbial metabolism at a sampling frequency of not less than 1 Hz, ensuring the capture of the dynamic changes in microbial metabolism. S13 synchronously monitors the environmental temperature and humidity parameters of microbial samples through temperature and humidity sensors. The environmental temperature measurement range is 0-50 degrees Celsius, and the environmental humidity measurement range is 0-100% relative humidity. S14. Perform data preprocessing on conductivity difference data, carbon dioxide concentration data, ambient temperature parameters and ambient humidity parameters, including using moving average filtering to remove high-frequency noise interference and performing data standardization to eliminate dimensional differences between different parameters. S15 combines the preprocessed conductivity difference data, carbon dioxide concentration data, ambient temperature parameters, and ambient humidity parameters into a multi-source input vector according to the time series, and adds a timestamp identifier to each data point.
[0026] As described in steps S11-S15 above, by collecting microbial metabolic correlation data and environmental parameters, and after standardized preprocessing, integrating them in chronological order, a multi-source input vector with unified structure and reliable quality is formed.
[0027] During microbial growth and metabolism, the decomposition of organic carbon alters the electrical conductivity of the surrounding environment. Metabolic activity directly produces carbon dioxide, while environmental temperature and humidity affect microbial enzyme activity, changing their growth rate and metabolic intensity. These parameters collectively constitute the core information dimensions reflecting the quantity of microorganisms. To achieve accurate colony quantification, it is essential to acquire multi-dimensional data that truly reflects the metabolic state and growth environment of microorganisms. Insufficient data acquisition precision, asynchronous timing, or the presence of noise interference will directly lead to distortion in subsequent feature extraction, ultimately affecting the accuracy of the quantitative results. Therefore, it is necessary to establish a standardized multi-source data acquisition and preprocessing workflow.
[0028] Traditional data acquisition methods for microbial detection have significant drawbacks. They often use ordinary conductivity meters to measure conductivity values, which are not accurate enough and result in large errors in the calculation of conductivity differences. The carbon dioxide sensors used are susceptible to interference and have slow response speeds. The temperature and humidity monitoring range is narrow and cannot be adapted to the diverse growth environments of microorganisms. At the same time, there is a lack of systematic data preprocessing. Differences in the dimensions of different parameters and high-frequency noise directly affect the usability of data. Asynchronous acquisition by different devices also disrupts the temporal correlation between data.
[0029] The core of using a high-precision conductivity meter to obtain conductivity difference data is to reflect the degree of carbon consumption by accurately measuring the change in conductivity before and after the decomposition of organic carbon. Before the complete decomposition of organic carbon in a microbial sample, the ion concentration in the solution is relatively stable. At this time, the initial conductivity value is measured using a high-precision conductivity meter. After the complete decomposition of organic carbon, the ions produced by microbial metabolism increase the solution conductivity, and the final conductivity value is measured. The difference between the two values is then calculated to obtain the conductivity difference data. The measurement accuracy of the high-precision conductivity meter directly determines the accuracy of the difference, avoiding deviations in carbon-based numerical mapping due to insufficient equipment precision.
[0030] A non-dispersive infrared carbon dioxide sensor is used to collect data at a sampling frequency of at least 1 Hz because this sensor has fast response and anti-interference characteristics. It detects concentration through the principle of infrared radiation absorption, and the light source and detector are protected from direct contact with the gas being measured, avoiding the poisoning problem of traditional catalytic combustion sensors. The concentration of carbon dioxide produced by microbial metabolism changes dynamically with the growth stage. For example, the concentration rises rapidly during the growth period. A sampling frequency of at least 1 Hz can record at least one data point per second, fully capturing this dynamic process. If the frequency is too low, key concentration change nodes will be missed, leading to inaccurate judgment of metabolic intensity.
[0031] Parameters are monitored synchronously using temperature and humidity sensors. A temperature range of 0-50 degrees Celsius and a relative humidity range of 0-100% are set because the optimal growth temperature for most common microorganisms is between 0-50 degrees Celsius. Humidity changes affect the water activity of the culture medium, and these two ranges comprehensively cover the typical growth environment for microorganisms. Synchronous monitoring ensures that temperature and humidity parameters are perfectly matched with conductivity and carbon dioxide data in time. When the carbon dioxide concentration suddenly increases at a certain moment, it can be determined, in conjunction with the concurrent temperature and humidity, whether it is due to the microorganisms entering a rapid growth phase or environmental interference, providing an environmental reference for subsequent data interpretation.
[0032] During data preprocessing, the moving average filter removes noise by calculating the average value of data within a fixed window. The window length is experimentally determined to balance smoothing effect and data response speed. For example, for carbon dioxide concentration data, a window length of 5 data points is selected to filter out high-frequency noise introduced by the sensor circuit without lagging behind the actual concentration changes. Data standardization uses the Z-score method to convert parameters with different dimensions, such as conductivity difference and carbon dioxide concentration, into dimensionless data with a mean of 0 and a standard deviation of 1. This eliminates the magnitude difference between conductivity difference and carbon dioxide concentration, ensuring balanced weights for each parameter during subsequent neural network model training.
[0033] The preprocessed data is combined by time series and timestamped because microbial metabolism exhibits temporal continuity. For example, the conductivity difference at a certain moment is correlated with the carbon dioxide concentration change at the previous moment, and combining data by time series preserves this dynamic correlation. Timestamping ensures that each data point is traceable to a specific collection time. When subsequent analysis reveals anomalies in a data segment, the timestamp can be used to pinpoint environmental changes or equipment status during the collection period, providing a basis for data quality checks. The resulting multi-source input vector achieves structured integration of multi-dimensional data, directly meeting the input requirements of subsequent multilayer perceptron neural networks.
[0034] In one embodiment, the step of mapping the conductivity difference data to carbon-based values using a total organic carbon algorithm based on the conductivity difference data, and then combining this with carbon dioxide concentration data to extract microbial abundance features using a machine learning regression model includes: S21, The total organic carbon conversion algorithm is applied to map the conductivity difference data into carbon-based values. The total organic carbon conversion algorithm is based on the linear relationship between conductivity difference and total organic carbon content, and outputs the first carbon-based value and the second carbon-based value. S22, construct a multilayer perceptron neural network as a machine learning regression model. The input layer of the neural network receives a multi-source input vector containing carbon dioxide concentration data, conductivity difference data, ambient temperature parameters, and ambient humidity parameters. S23, Configure the neural network structure, including setting two hidden layers, each containing multiple neurons, using a linear rectified unit as the activation function, and setting the output layer as a single neuron structure to output a preliminary estimate of the number of microorganisms; S24, a neural network model is trained using a historical calibration dataset, which contains sample data with known microbial quantities. During training, mean squared error is used as the loss function, and the model parameters are optimized through backpropagation algorithm. S25, the first carbon-based value, the second carbon-based value, and the preliminary microbial quantity estimate output by the neural network are fused at the feature level to form microbial quantity features, which are used as input for subsequent time series analysis.
[0035] As described in steps S21-S25 above, the conductivity difference is mapped to a carbon-based value through the total organic carbon conversion algorithm, and multi-layer perceptron neural network is used to process multi-source data and output a preliminary estimate of the number of microorganisms. Finally, the core features that can comprehensively reflect the number of microorganisms are extracted through feature-level fusion.
[0036] The essence of microbial growth and metabolism is the consumption of organic carbon in the environment and the production of metabolic products. Conductivity difference data can indirectly reflect the degree of organic carbon decomposition. The more thorough the decomposition of organic carbon, the greater the change in ion concentration in the solution, and the larger the conductivity difference. Carbon dioxide concentration data directly reflects metabolic intensity. Environmental temperature and humidity parameters affect the metabolic rate and organic carbon consumption efficiency by influencing microbial enzyme activity. These data together constitute multi-dimensional information strongly correlated with microbial abundance. To accurately extract microbial abundance characteristics, conductivity differences must first be converted into carbon-based values with clear physical meaning, i.e., values related to organic carbon consumption. Then, multi-source data such as carbon dioxide, temperature, and humidity should be integrated to process nonlinear correlations. Otherwise, single data or unconverted raw data cannot comprehensively characterize microbial abundance and are prone to feature distortion due to the lack of data correlation.
[0037] Among them, the total organic carbon (TOC) conversion algorithm is applied to map the conductivity difference data into carbon-based values. This algorithm is based on the linear relationship between conductivity difference and total organic carbon content. The TOC conversion formula is as follows: TOC= ; Wherein, TOC represents the carbon-based value, that is, the total organic carbon content. The data represents the conductivity difference. a and b are the linear coefficients of the total organic carbon conversion algorithm, which are derived from the historical calibration process established based on the linear relationship between conductivity difference and total organic carbon content. The first carbon-based value and the second carbon-based value calculated by this formula correspond to the organic carbon consumption of microorganisms at different growth stages. For example, the first carbon-based value reflects the logarithmic growth phase, which corresponds to the total carbon consumption when metabolism is vigorous and organic carbon is consumed rapidly. The second carbon-based value reflects the total carbon consumption during the stationary phase, which corresponds to the slowdown of metabolism and the stabilization of organic carbon consumption. The two values together constitute the dynamic characteristics of organic carbon consumption.
[0038] A multilayer perceptron neural network was constructed as a machine learning regression model. Its input layer receives the multi-source input vector, which includes carbon dioxide concentration data, conductivity difference data, ambient temperature parameters, and ambient humidity parameters. The number of nodes in the input layer is consistent with the vector dimension, totaling four nodes, ensuring that the multi-source data can be completely input into the model. Two hidden layers were configured in the neural network structure, each containing 16 neurons. This number has been experimentally verified to ensure computational efficiency while fully capturing the nonlinear correlations of the multi-source data. Too few neurons would lead to insufficient model fitting ability and an inability to handle the cross-influence of temperature and humidity on carbon dioxide and conductivity; too many neurons would increase redundant computation. A linear rectified unit (RCU) was used as the activation function. This function effectively solves the gradient vanishing problem of the traditional sigmoid function in deep networks by setting negative input values to 0 and keeping positive input values unchanged, ensuring that the model can deeply learn the complex relationships between multi-source data. The output layer is set as a single-neuron structure, specifically used to output preliminary estimates of microbial quantity, realizing the mapping from multi-source data to quantitative estimation.
[0039] The neural network model was trained using a historically calibrated dataset, constructed through prior experiments. This dataset contains multiple sets of sample data with known microbial counts, such as standard microbial samples prepared with colony counts of 100 CFU / mL, 500 CFU / mL, and 1000 CFU / mL. For each sample, corresponding data on carbon dioxide concentration, conductivity difference, ambient temperature, and humidity were collected, forming multiple input-output pairs. The inputs consisted of multi-source data, and the outputs represented known colony counts. The mean squared error was used as the loss function during training, calculated using the following formula: ; Where M represents the mean squared error, and n represents the number of samples in the historical calibration dataset. Indicates the actual number of colonies. This function represents the model's predicted value and quantifies the deviation between the predicted value and the actual colony count. The backpropagation algorithm is used to optimize the model parameters. For example, by calculating the partial derivative of the loss function with respect to the weights of each layer, the connection weights between the hidden layer and the input layer, and between the hidden layer and the output layer, are adjusted backward from the output layer to the input layer. This gradually reduces the loss function value, ultimately ensuring that the deviation between the model's initial estimate of the microbial count and the actual colony count is controlled within a preset range, guaranteeing the accuracy of the initial estimate.
[0040] The first and second carbon-based values are fused with the preliminary microbial quantity estimate output by the neural network at the feature level. The fusion process uses a weighted summation method; for example, the microbial quantity feature = 0.4 × first carbon-based value + 0.3 × second carbon-based value + 0.3 × preliminary microbial quantity estimate. The weight allocation is determined based on the correlation coefficient between each feature and the microbial quantity. The carbon-based value directly reflects organic carbon consumption and has a higher correlation with microbial quantity, therefore it is assigned a higher weight. The preliminary estimate integrates multi-source data and assigns corresponding weights to compensate for the influence of environmental factors. This fusion method combines the direct characteristics of organic carbon consumption with the comprehensive estimation characteristics of multi-source data to form a more dimensional and accurate microbial quantity feature, avoiding bias caused by incomplete information from a single feature and reducing errors in subsequent smoothing and state estimation.
[0041] In one embodiment, the steps of performing time-series analysis on the extracted microbial quantity characteristics, applying a state-space model for smoothing and state estimation, and outputting the smoothed microbial quantity estimate and its statistical covariance data include: S31, The microbial quantity characteristics are reconstructed over time to form a microbial quantity time series; S32, Perform stationarity test and detrending processing on the microbial quantity time series, and extract the steady-state growth component and short-term fluctuation component from the sequence; S33, Establish a state-space model, define state variables including microbial quantity state and microbial growth rate state, and define the observed variable as the microbial quantity characteristics after time series analysis; S34. The Kalman filter algorithm is applied to solve the state-space model. Through iterative calculation of the prediction and update steps, the microbial quantity characteristics are smoothed and the optimal state is estimated. S35, In each iteration of the Kalman filter, the smoothed estimate of the number of microorganisms and its corresponding statistical covariance data value are output synchronously. The statistical covariance data value is used to quantify the uncertainty of the estimation result.
[0042] As described in steps S31-S35 above, the present invention achieves smooth estimation and state quantification of microbial quantity by reconstructing the time series of microbial quantity characteristics, processing stationarity, modeling state space, and optimizing Kalman filtering. At the same time, it outputs statistical covariance data values to characterize the estimation uncertainty.
[0043] Because microbial growth exhibits significant temporal dynamics, its population changes systematically with growth stages such as the lag phase, logarithmic phase, and stationary phase. However, the extracted microbial population characteristics are affected by short-term metabolic fluctuations, such as temporary changes in enzyme activity leading to fluctuations in carbon dioxide release. These fluctuations, coupled with sensor measurement noise, result in random fluctuations in characteristic values. These fluctuations do not reflect the true changes in microbial populations and will introduce estimation bias if directly used in subsequent analyses. Furthermore, microbial population changes are directly regulated by the growth rate; focusing solely on the population itself cannot fully describe its dynamic changes or achieve optimal state estimation. Therefore, time series analysis is needed to separate the true growth trend from disturbance fluctuations, and a model should be constructed based on the growth rate to achieve accurate estimation and quantify uncertainty.
[0044] The core of the time-series reconstruction of microbial abundance features is to arrange the extracted microbial abundance features sequentially according to the timestamps of the multi-source input vectors, forming a microbial abundance time series. Since the data acquisition sampling frequency is no less than 1 Hz, and each data point is timestamped, the reconstruction is performed by sorting the data from earliest to latest timestamp. For example, the microbial abundance features collected per second are arranged sequentially as [F1, F2, ..., Ft], where Ft is the abundance feature at second t, forming a continuous time series. This reconstruction method can completely preserve the dynamic change trajectory of microbial abundance over time, providing a temporal basis for subsequent analysis of its growth patterns and fluctuation characteristics, and avoiding misjudgments of dynamic patterns due to disordered data order.
[0045] Stationarity testing and detrending processing were performed on the microbial population time series. The Augmented Dickey-Fuller (ADF) test was used to determine if the series was stationary. If a trend was observed, such as a sustained increase in population characteristics during the logarithmic phase, detrending processing was required. For example, a trend term was fitted to the series using linear regression, and the trend term was subtracted from the original series to obtain a detrended series. This detrended series retains the steady-state growth component, including the exponential growth portion after excluding the linear trend during the logarithmic phase, as well as short-term fluctuations (i.e., random fluctuations). This processing effectively separates fluctuations that do not reflect true population changes, focusing on the true growth trend of microorganisms. For example, it removes feature value jumps caused by instantaneous sensor noise, ensuring that subsequent modeling is based on true dynamic patterns.
[0046] When establishing the state-space model, the state variables are defined as the microbial population state (N) and the microbial growth rate state (N). Because the growth rate directly determines the rate of change in quantity, the higher the growth rate, the greater the increase in the number of microorganisms per unit time. These two factors together constitute the core state describing the dynamics of microbial growth. The observed variable is defined as the microbial quantity characteristic after stationarity treatment and detrending. This characteristic has eliminated fluctuation interference and can accurately reflect the actual situation of the state variable. The model is specifically constructed as follows: Wherein, the state equation is set as: ; in, This represents the state of microbial quantity at time t, i.e., the true state value of microbial quantity at time t. This represents the state of microbial population at time t-1, and is the optimal estimate from the previous iteration. This represents the microbial growth rate at time t-1, which is the optimal estimate of the growth rate from the previous iteration. This is represented as state noise and follows a normal distribution. The value represents the sampling interval, which is 1 second. This formula can describe the dynamic relationship between the quantity at time t and the quantity at time t-1, and the growth rate.
[0047] The observation equation is set as follows: = + ; in, The observed variable at time t represents the microbial quantity characteristics after time series analysis. It represents the observation noise at time t, which follows a normal distribution. It is used to quantify the error in the observation process, establish the correlation between the observed value and the quantity state, and fully characterize the mapping relationship between the microbial growth state and the observation characteristics.
[0048] The Kalman filter algorithm is applied to solve the state-space model. This process is achieved through two iterative steps: prediction and update. In the prediction step, based on the state estimate and state equation at time t-1, the prior state estimate at time t is calculated, along with the prior covariance matrix. In the update step, using the observed variable Zt at time t and the observation equation, the Kalman gain is calculated, and the prior estimate is corrected with the observed value to obtain the posterior state estimate at time t, which is the smoothed estimate of microbial abundance and the posterior covariance matrix. Through iterative calculation once per second, the estimation results can be corrected in real time with new observation data. For example, when the observed value is too high due to short-term fluctuations, the filter will correct the estimate towards the true trend based on historical state and covariance, achieving smooth processing of microbial abundance characteristics and optimal state estimation, significantly reducing estimation bias caused by fluctuation interference.
[0049] In each iteration of the Kalman filter, the smoothed estimate of microbial population and its corresponding statistical covariance are output synchronously. The statistical covariance value is derived from the diagonal elements of the posterior covariance matrix during the filtering process, directly quantifying the uncertainty of the microbial population estimate at that moment. A smaller covariance value indicates a smaller deviation between the observed and estimated values, and a more reliable estimation result. A larger covariance value indicates higher estimation uncertainty. This data provides a crucial noise variance parameter for Gaussian process regression, ensuring that subsequent probability distribution construction accurately incorporates estimation uncertainty information and avoids probability distribution bias caused by ignoring uncertainty.
[0050] In one embodiment, the step of obtaining the probability distribution of microbial quantity based on the smoothed microbial quantity estimate and the statistical covariance data value, and calculating and outputting the colony equivalent and its confidence interval based on the probability distribution of microbial quantity includes: S41. Using the Gaussian process regression framework, the smoothed microbial quantity estimate is used as the mean of the training samples, and the statistical covariance data value is used as the noise variance of the training samples to construct the Gaussian process prior distribution. S42, select the radial basis function as the kernel function, calculate the similarity between training samples through the kernel function, and construct the covariance matrix; S43, based on the Bayesian inference principle, calculate the posterior distribution of the Gaussian process to obtain the complete probability distribution of the number of microorganisms; S44, extract the distribution mean from the probability distribution of microbial quantity as the benchmark value of colony equivalent, and calculate the confidence interval based on the standard deviation of the probability distribution; S45, set a variance threshold monitoring mechanism. When the variance of the output probability distribution exceeds the preset threshold, the data quality check process will be automatically triggered.
[0051] As described in steps S41-S45 above, the smoothed microbial quantity estimate and statistical covariance data are integrated through the Gaussian process regression framework to construct the probability distribution of microbial quantity, and then the colony equivalent benchmark value and confidence interval are extracted. At the same time, the data quality is ensured by monitoring the variance threshold, so as to realize the probabilistic characterization and reliability control of the microbial colony quantification results.
[0052] In quantitative analysis of microbial colonies, relying solely on a single estimate cannot fully reflect the possible range of the true quantity. While the smoothed estimate reduces fluctuations, uncertainties remain due to individual differences in microbial metabolism and residual minor sensor noise. Statistical covariance data quantifies this uncertainty. Directly using the estimate as the final result makes it impossible for users to assess its reliability; for example, an estimate might fall within the margin of the true quantity without a clear confidence boundary. Furthermore, without monitoring the variance of the probability distribution, poor data quality, such as temporary sensor malfunctions causing data bias, cannot be promptly identified and addressed, leading to invalidation of subsequent quantitative results. Therefore, it is necessary to expand point estimates to interval estimates through probability distribution construction, combine uncertainty quantification to achieve a comprehensive and reliable representation of the results, and establish a variance monitoring mechanism to ensure data quality.
[0053] When constructing the Gaussian process prior distribution using the Gaussian process regression framework, the core principle is to use the smoothed microbial quantity estimate as the mean of the training samples and the output statistical covariance as the noise variance of the training samples. Since Kalman filtering has already achieved the optimal smoothed estimate of the microbial quantity, and the statistical covariance accurately quantifies the uncertainty of this estimate, constructing the prior distribution based on this ensures that the initial assumptions of the Gaussian process closely match the true dynamic characteristics and uncertainty level of the microbial quantity. This avoids modeling bias caused by using unfounded default priors, such as a prior with a mean of 0 and a fixed noise variance. For example, if the output smoothed estimate at a certain moment is 500 CFU / mL and the statistical covariance is 0.02, then the mean of the training samples at that moment should be set to 500 CFU / mL and the noise variance to 0.02, ensuring the rationality of the prior distribution.
[0054] The radial basis function is chosen as the kernel function to calculate the similarity between training samples, and the covariance matrix is constructed. The expression for the radial basis function is: ; in, The kernel function value quantifies the similarity between the i-th and j-th training samples (range: (0,1]). The higher the similarity, the closer the kernel function value is to 1. Let represent the i-th training sample vector (with a dimension of 1, i.e., scalar form), representing the core feature of the number of microorganisms at time i. Let represent the j-th training sample vector (with a dimension of 1, i.e., scalar form), representing the core feature of microbial quantity at time j. express and The Euclidean distance is used to quantify the numerical difference between training samples at two different time points. This represents the kernel width parameter, used to adjust the kernel function's sensitivity to sample differences. The smaller the value, the more sensitive the function is to differences. The larger the value, the smoother the difference between the two samples. This kernel function can effectively capture the nonlinear similarity between samples. For example, when two time points are close (e.g., 1 second apart) and the microbial population changes slowly, the similarity value is close to 1, and the corresponding element value in the covariance matrix is large, indicating a strong correlation between the two. When the time interval is large (e.g., 10 seconds apart) and the microbial population enters a stable period with small changes, the similarity value can still maintain a reasonable level, accurately reflecting the temporal correlation characteristics of the microbial population. The covariance matrix constructed through this similarity calculation can accurately describe the statistical correlation between training samples, providing core matrix support for subsequent probability distribution calculations.
[0055] The calculation of the posterior distribution of a Gaussian process based on Bayesian inference integrates the prior distribution with the observational information of the training samples, i.e., the smoothed estimate of the microbial population output. Specifically, it first obtains the initial probability characteristics of the microbial population through the prior distribution, then combines the sample association represented by the covariance matrix, and uses the Bayesian formula to calculate the posterior distribution, ultimately obtaining the complete probability distribution of the microbial population. This process can fully utilize the information accumulated in the previous steps, preserving the dynamic trend of the estimate while incorporating the uncertainty information of the statistical covariance data. At the same time, it corrects the bias in the prior assumptions through Bayesian inference, making the final probability distribution more closely match the true statistical characteristics of the microbial population, avoiding distribution distortion caused by relying solely on priors or observations.
[0056] The mean of the probability distribution of microbial quantity is extracted as the benchmark value for colony equivalent because the mean is the central tendency of the probability distribution, which best reflects the possible true value of microbial quantity, consistent with the definition logic of colony equivalent. Confidence intervals are calculated based on the standard deviation of the probability distribution, using the industry-standard 95% confidence interval calculation method: Confidence interval = mean ± 1.96 × standard deviation. For example, when the mean is 500 CFU / mL and the standard deviation is 5 CFU / mL, the confidence interval is [490.2, 509.8] CFU / mL. This interval representation allows users to clearly understand the possible range of the true microbial quantity. Compared to traditional single-point estimation, it better supports subsequent quality control decisions. For example, in food hygiene testing, if the upper limit of the confidence interval does not exceed the safety threshold, the product can be judged as qualified, avoiding misjudgments due to accidental biases in single-point estimation.
[0057] When setting up the variance threshold monitoring mechanism, first preset the variance threshold, which can be set to 0.05. This value corresponds to a relative error of no more than 5% in the estimation of microbial quantity. The variance of the output probability distribution is monitored in real time. When the variance exceeds the preset threshold, it indicates that the uncertainty of the current training sample is too high. This may be due to incomplete elimination of fluctuations by the Kalman filter, or sensor anomalies during the data acquisition phase, such as a temporary reduction in the sampling frequency of the carbon dioxide sensor. In this case, a data quality check process is automatically triggered. The check includes whether there are abnormal jumps in the preprocessed raw data and whether the iteration parameters of the Kalman filter are reasonable. This mechanism can promptly identify unreliable probability distributions caused by low-quality data, avoid outputting invalid colony equivalents and confidence intervals, and ensure the reliability of quantitative results.
[0058] In one embodiment, the step of verifying the accuracy of the colony equivalent using the sterile sucrose validation method and outputting the validated final colony quantification result includes: S51, Establish a sterilized sucrose validation database, wherein the database stores the colony equivalent baseline values and their corresponding validation status identifiers obtained in historical sterilized sucrose validation experiments. S52, obtain the colony equivalent and its confidence interval of the current batch of microbial samples, and evaluate the reliability level of the current measurement results based on the width of the confidence interval; S53, calculate the similarity between the current colony equivalent and the benchmark value in the sterile sucrose verification database, and use a statistical hypothesis testing method based on confidence intervals to determine whether the current colony equivalent passes the verification. S54, when the similarity value is within the preset range, generate a verification pass mark and mark the current colony equivalent as the verified final colony quantification result; S55: When the similarity value is not within the preset range, a verification failure flag is generated, a model parameter calibration command is triggered, and a warning message suggesting re-collecting data is output.
[0059] As described in steps S51-S55 above, this invention provides a benchmark reference by establishing a sterile sucrose validation database, assesses reliability by combining the confidence interval of the current batch colony equivalent, completes accuracy verification through similarity calculation and statistical hypothesis testing, and finally outputs qualified quantitative results or triggers abnormal handling based on the validation results, forming a closed-loop quality control for colony quantification analysis, ensuring that the final output microbial colony quantification results are accurate and reliable.
[0060] The colony equivalent and its confidence interval have been output above, but this result may still be affected by various factors and thus biased. For example, long-term use of the sensor may cause accuracy drift, resulting in slight deviations in the collected data, which will then be passed on to the colony equivalent in subsequent steps. Alternatively, differences in microbial sample batches, such as slight variations in the metabolic intensity of different batches of microorganisms, may lead to decreased adaptability of the neural network model, resulting in feature extraction biases that ultimately affect the colony equivalent. If the colony equivalent is output directly without verification, it may cause misjudgments in critical scenarios such as food hygiene and pharmaceutical quality control, for example, classifying unqualified products as qualified, thus posing a safety risk. Therefore, a systematic verification process is needed to confirm the accuracy of the colony equivalent, and an anomaly handling mechanism should be established to ensure the reliability of the final output result.
[0061] The core of establishing a sterile sucrose validation database lies in accumulating baseline data through multiple sterile sucrose validation experiments. The principle of this experiment is to add high-temperature sterilized sucrose to a microbial sample with a known initial colony equivalent. After sterilization, the sucrose contains no microorganisms and does not introduce additional colonies. Based on the patterns of microbial sucrose metabolism, the theoretical range of colony equivalent changes after addition can be predetermined. The colony equivalent under this scenario is obtained through actual testing and used as a baseline value. The status of each validation is then marked, such as "validation passed" or "validation failed." The passing standard is that the deviation between the actual detected value and the theoretical range is less than 5%. This experiment is repeated multiple times; for example, 50 sterile sucrose validation experiments are conducted on E. coli samples. The colony equivalent baseline value and the corresponding validation status are recorded each time, forming a sterile sucrose validation database. This provides a stable and traceable baseline reference for subsequent validation of the current batch.
[0062] Obtain the colony equivalent and its confidence interval for the current batch of microbial samples. For example, if the colony equivalent of a batch of E. coli samples is 500 CFU / mL, the confidence interval is [490, 510] CFU / mL. Assess the reliability level based on the width of the confidence interval. The smaller the confidence interval, the less dispersion of the probability distribution, the lower the uncertainty of the colony equivalent, and the higher the reliability. For example, if the confidence interval width of the above batch is 20 CFU / mL, and another batch has a confidence interval width of 40 CFU / mL, the former's reliability is significantly higher than the latter. This pre-assessment allows for more rigorous subsequent validation of batches with lower reliability, avoiding wasted resources.
[0063] The current colony equivalent is compared with the baseline value in the sterile sucrose validation database. The similarity calculation uses the relative error method, and the formula is as follows:
[0064] Where P represents the similarity value and Y represents the current colony equivalent. Denote the reference value. For example, if the current colony equivalent is 500 CFU / mL and the matched reference value in the database is 505 CFU / mL, then the similarity value = |500 - 505| / 505 × 100% ≈ 0.99%. Subsequently, a statistical hypothesis testing method based on the confidence interval is adopted, specifically the t-test. The current colony equivalent, its corresponding standard deviation of the confidence interval, the reference value, and its historical standard deviation are used as inputs. The significance level is set to 0.05 to determine whether there is a statistical difference between the current colony equivalent and the reference value. If there is no difference, it is determined to pass the verification. This process uses statistical tests rather than simple numerical comparisons, which can effectively exclude misjudgments caused by accidentally close numerical values. For example, if a current value is close to the reference value but the confidence interval is too wide, the statistical test will identify its insufficient reliability and determine that it fails to pass the verification, thus improving the scientific nature of the verification.
[0065] When the similarity value is within the preset range, the preset range is determined according to the similarity distribution of the historical verified data in the database and is usually set to 0 - 5%. For example, the above similarity value of 0.99% is within this range, and a verification passed flag is generated. The current colony equivalent (500 CFU / mL) is marked as the final colony quantification result after verification, and this result can be directly used for subsequent quality determination. For example, in food hygiene testing, if this result is lower than the safety threshold, it is determined that the food microbial index is qualified. When the similarity value is not within the preset range, for example, the current colony equivalent is 530 CFU / mL, the reference value is 500 CFU / mL, and the similarity value = 6% exceeds the preset range. A verification failed flag is generated, triggering a model parameter calibration instruction, specifically to re-optimize the parameters of the multi-layer perceptron neural network model. For example, the model is supplemented with training based on the data that failed to pass the verification this time. At the same time, a warning message suggesting to re-collect data is output, prompting the operator to check whether the sensor is normal, such as whether the non-dispersive infrared carbon dioxide sensor needs to be calibrated and whether the sample collection is standardized, to ensure the data quality of subsequent re-detection and form a closed loop for exception handling.
[0066] As Figure 2 shown, the present invention also discloses a microbial colony quantitative analysis system based on the fusion of conductivity change and TOC model, including: A collection module 1 for real-time collecting the conductivity difference data and carbon dioxide concentration data of the microbial sample, and synchronously recording the environmental temperature and humidity parameters to form a multi-source input vector; An extraction module 2 for mapping the conductivity difference data to carbon-based values based on the conductivity difference data and combining with the carbon dioxide concentration data to extract microbial quantity characteristics; An analysis module 3 for performing time series analysis on the extracted microbial quantity characteristics, and performing smoothing processing and state estimation, and outputting the smoothed microbial quantity estimated value and its statistical covariance data value; Calculation module 4 is used to obtain the probability distribution of microbial quantity based on the smoothed microbial quantity estimate and the statistical covariance data value, and to calculate and output the colony equivalent and its confidence interval based on the probability distribution of microbial quantity. The verification module 5 is used to verify the accuracy of the colony equivalent and output the verified final colony quantification result.
[0067] In one embodiment, the analysis module includes: The reconstruction unit is used to reconstruct the microbial quantity characteristics over time to form a microbial quantity time series. The extraction unit is used to perform stationarity testing and detrending processing on the microbial quantity time series, and to extract the steady-state growth component and short-term fluctuation component from the sequence. Establish a unit to build a state-space model, define state variables including microbial quantity state and microbial growth rate state, and define observation variables as microbial quantity characteristics after time series analysis; The computing unit is used to solve the state-space model and to smooth the microbial quantity characteristics and estimate the optimal state through iterative calculation. The output unit is used to synchronously output the smoothed estimate of the number of microorganisms and its corresponding statistical covariance data.
[0068] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method for quantitative analysis of microbial colonies based on the fusion of conductivity changes and TOC model.
[0069] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method for quantitative analysis of microbial colonies based on the fusion of conductivity changes and a TOC model.
[0070] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0071] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0072] The above description is merely a preferred embodiment of the present invention and does not limit the scope of this application. Any equivalent results or equivalent process transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of protection of this application.
Claims
1. A method for quantitative analysis of microbial colonies based on the fusion of conductivity changes and the TOC model, characterized in that, Includes the following steps: Real-time acquisition of conductivity difference data and carbon dioxide concentration data of microbial samples, and simultaneous recording of environmental temperature and humidity parameters to form a multi-source input vector; Based on the conductivity difference data, the conductivity difference data is mapped to carbon-based values and combined with carbon dioxide concentration data to extract microbial quantity characteristics; The extracted microbial quantity characteristics are subjected to time series analysis, smoothing and state estimation, and the smoothed microbial quantity estimates and their statistical covariance data are output. The probability distribution of microbial quantity is obtained based on the smoothed microbial quantity estimate and the statistical covariance data value, and the colony equivalent and its confidence interval are calculated and output based on the probability distribution of microbial quantity. The accuracy of the colony equivalent is verified, and the verified final colony quantification result is output.
2. The method for quantitative analysis of microbial colonies based on the fusion of conductivity change and TOC model according to claim 1, characterized in that, The steps of real-time acquisition of conductivity difference data and carbon dioxide concentration data of microbial samples, and simultaneous recording of environmental temperature and humidity parameters to form a multi-source input vector include: The initial conductivity value of the microbial sample before complete decomposition of organic carbon and the final conductivity value after complete decomposition were obtained respectively. The conductivity difference data was obtained by calculating the difference between the final conductivity value and the initial conductivity value. Continuously collect data on the concentration of carbon dioxide produced by microbial metabolism; Simultaneously monitor the environmental temperature and humidity parameters of microbial samples; Data preprocessing was performed on conductivity difference data, carbon dioxide concentration data, ambient temperature parameters, and ambient humidity parameters. The preprocessed conductivity difference data, carbon dioxide concentration data, ambient temperature parameters, and ambient humidity parameters are combined according to time series to form a multi-source input vector, and a timestamp is added to each data point.
3. The method for quantitative analysis of microbial colonies based on the fusion of conductivity change and TOC model according to claim 1, characterized in that, The step of mapping the conductivity difference data to carbon-based values and combining it with carbon dioxide concentration data to extract microbial abundance characteristics includes: Map the conductivity difference data to carbon-based values and output the first and second carbon-based values; A multilayer perceptron neural network is constructed as a machine learning regression model. The input layer of the neural network receives a multi-source input vector containing carbon dioxide concentration data, conductivity difference data, ambient temperature parameters, and ambient humidity parameters. The neural network structure is configured with a linear rectified unit as the activation function and the output layer is set as a single neuron structure to output the preliminary estimate of the number of microorganisms. The neural network model was trained using a historical calibration dataset, which contained sample data with known numbers of microorganisms. The first carbon-based value, the second carbon-based value, and the preliminary microbial quantity estimate output by the neural network are fused at the feature level to form microbial quantity features.
4. The method for quantitative analysis of microbial colonies based on the fusion of conductivity change and TOC model according to claim 1, characterized in that, The steps of performing time-series analysis on the extracted microbial quantity characteristics, smoothing and state estimation, and outputting the smoothed microbial quantity estimate and its statistical covariance data include: The microbial quantity characteristics are reconstructed over time to form a microbial quantity time series; The microbial population time series was subjected to stationarity testing and detrending processing to extract the steady-state growth component and short-term fluctuation component from the sequence. Establish a state-space model, define state variables including microbial quantity state and microbial growth rate state, and define observation variables as microbial quantity characteristics after time series analysis. The state-space model is solved, and the microbial quantity characteristics are smoothed and the optimal state is estimated through iterative calculation. The smoothed estimate of the number of microorganisms and its corresponding statistical covariance data are output synchronously.
5. The method for quantitative analysis of microbial colonies based on the fusion of conductivity change and TOC model according to claim 1, characterized in that, The steps of obtaining the probability distribution of microbial quantity based on the smoothed microbial quantity estimate and the statistical covariance data value, and calculating and outputting the colony equivalent and its confidence interval based on the probability distribution of microbial quantity include: Using a Gaussian process regression framework, the smoothed microbial quantity estimate is used as the mean of the training samples, and the statistical covariance data value is used as the noise variance of the training samples to construct the Gaussian process prior distribution. Radial basis functions are selected as kernel functions, and the similarity between training samples is calculated through the kernel functions to construct the covariance matrix; Calculate the posterior distribution of the Gaussian process to obtain the complete probability distribution of the number of microorganisms; The mean of the microbial population is extracted from the probability distribution as a baseline value for colony equivalent, and the confidence interval is calculated based on the standard deviation of the probability distribution. A variance threshold monitoring mechanism is set up so that when the variance of the output probability distribution exceeds a preset threshold, the data quality check process is automatically triggered.
6. The method for quantitative analysis of microbial colonies based on the fusion of conductivity change and TOC model according to claim 1, characterized in that, The steps of verifying the accuracy of the colony equivalent and outputting the verified final colony quantification result include: Establish a sterile sucrose validation database, which stores the colony equivalent baseline values and their corresponding validation status identifiers obtained in historical sterile sucrose validation experiments. Obtain the colony equivalent and its confidence interval of the current batch of microbial samples, and assess the reliability level of the current measurement results based on the width of the confidence interval; The similarity between the current colony equivalent and the benchmark value in the sterile sucrose validation database is calculated, and it is determined whether the current colony equivalent passes the validation. When the similarity value is within the preset range, a verification pass mark is generated, and the current colony equivalent is marked as the verified final colony quantification result; When the similarity value is outside the preset range, a verification failure flag is generated, triggering a model parameter calibration command and outputting a warning message suggesting that data be collected again.
7. A microbial colony quantitative analysis system based on the fusion of conductivity change and TOC model, characterized in that, include: The acquisition module is used to collect the conductivity difference data and carbon dioxide concentration data of microbial samples in real time, and simultaneously record the ambient temperature and humidity parameters to form a multi-source input vector; The extraction module is used to map the conductivity difference data into carbon-based values based on the conductivity difference data, and combine it with carbon dioxide concentration data to extract microbial quantity characteristics. The analysis module is used to perform time series analysis on the extracted microbial quantity characteristics, and to perform smoothing and state estimation, outputting the smoothed microbial quantity estimate and its statistical covariance data value. The calculation module is used to obtain the probability distribution of the number of microorganisms based on the smoothed microbial quantity estimate and the statistical covariance data value, and to calculate and output the colony equivalent and its confidence interval based on the probability distribution of the number of microorganisms. The verification module is used to verify the accuracy of the colony equivalent and output the verified final colony quantification result.
8. The microbial colony quantitative analysis system based on the fusion of conductivity change and TOC model according to claim 7, characterized in that, The analysis module includes: The reconstruction unit is used to reconstruct the microbial quantity characteristics over time to form a microbial quantity time series. The extraction unit is used to perform stationarity testing and detrending processing on the microbial quantity time series, and to extract the steady-state growth component and short-term fluctuation component from the sequence. Establish a unit to build a state-space model, define state variables including microbial quantity state and microbial growth rate state, and define observation variables as microbial quantity characteristics after time series analysis; The computing unit is used to solve the state-space model and to smooth the microbial quantity characteristics and estimate the optimal state through iterative calculation. The output unit is used to synchronously output the smoothed estimate of the number of microorganisms and its corresponding statistical covariance data.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Sterile and sterile carbon-based difference bacterial colony calculation method based on visual auxiliary detection
CN119169619A