A method for constructing an air microorganism real-time monitoring model
By constructing a real-time airborne microbial monitoring model using Gaussian mixture models and piecewise linear functions, the problems of long detection cycles, poor timeliness, and large errors in existing technologies are solved, achieving low-cost and high-accuracy real-time airborne microbial monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHONGKE AIXIN MICRO-ENVIRONMENTAL INFORMATION TECH (SHENZHEN) CO LTD
- Filing Date
- 2022-08-05
- Publication Date
- 2026-04-21
AI Technical Summary
Existing airborne microbial detection technologies suffer from problems such as long detection cycles, poor timeliness, large errors, and low accuracy. Furthermore, real-time airborne microbial particle detection instruments are expensive and cannot be widely adopted.
A Gaussian mixture model combined with piecewise linear functions was adopted to establish a real-time monitoring model using data such as airborne microbial concentration, particulate matter concentration, CO2 concentration, temperature and humidity. The data was wrapped in an ellipse by the Gaussian mixture model and piecewise linear regression was performed to construct a real-time airborne microbial monitoring model.
It enables low-cost, easy-to-implement and promote real-time online monitoring of airborne microorganisms, improving the accuracy and timeliness of monitoring and reducing errors.
Smart Images

Figure CN115116547B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of microbial monitoring technology. Specifically, it relates to a method for constructing a real-time airborne microbial monitoring model. Background Technology
[0002] Airborne microorganisms refer to microorganisms that exist in the air. The presence and harmfulness of airborne microorganisms must be confirmed through on-site sampling and testing, as well as epidemiological investigations. In public place hygiene monitoring, the total number of bacteria in the air is often used to characterize the degree of pollution. While airborne microorganism monitoring technologies have diversified, the following are the main ones with practical application value:
[0003] 1. Culture and Counting Techniques: Sampling microorganisms in the air or on object surfaces is a technique with over a century of history and remains the most reliable method for identifying pathogens. The natural sedimentation method is the most traditional: sampling plates are exposed to the environment for 10 minutes and then incubated for 12–48 hours. The resulting colonies can be identified and counted visually, and the degree of environmental microbial contamination can be evaluated by referring to relevant assessment standards. This method is simple, but it is time-consuming, has low sampling efficiency, and cannot accurately reflect the level of contamination. The use of automated sampling systems and multi-purpose microbial samplers overcomes this drawback; representative instruments include the Anderson 6-stage sampler. The disadvantage of air sampling and culture techniques is that test results take several days to report, resulting in poor timeliness. The advantage is that some common airborne microorganisms can be collected simultaneously, which helps establish a record of indoor air quality or airborne microbial concentrations, facilitating timely protective and remedial measures.
[0004] 2. Laser Scattering Particle Counting Technology: This method utilizes photoelectric technology to identify the physical parameters of airborne particles, and a particle counter has been developed based on this technology. Microbial aerosols typically exhibit a skewed distribution in the air, with most environmental particles concentrated in the 5–10 micrometer range. Using the principle of laser scattering, the particle counter can identify particles in the 0.3–20 micrometer range; this particle size range includes most bacteria and all spores, making it a feasible alternative method for detecting airborne pathogens. The particle counter is a low-cost aerosol monitoring device that can be used for monitoring and alarming microbial aerosols in special environments, such as when environmental particle concentration suddenly changes significantly during a biological warfare attack. The drawback of this technology is that it can only identify the total number of airborne particles and lacks specificity.
[0005] 3. Laser-Induced Fluorescence Biological Particle Counting Technology: This technology can characteristically identify the biological characteristics of particles. It identifies the physical parameters of particle size through light scattering and recognizes the biological characteristics of microbial particles through ultraviolet-excited biofluorescence. Laser-induced fluorescence detection technology achieves characteristic counting of microbial concentration at a two-dimensional level, considering both morphological and biological parameters. Using high-power lasers or xenon lamps as ultraviolet excitation sources, countries such as the US and UK have developed mobile biological warfare agent detection and alarm systems for field conditions (TSl3312, BAWS, and VeroTect, etc.). Using UV-LEDs as ultraviolet excitation sources, TSI has miniaturized and ported biological particle counting devices, enabling monitoring and alarming of biosafety in enclosed building environments. However, because non-microbial particles such as cigarette smoke and pollen in the actual environment can also produce fluorescence under ultraviolet laser induction, this technology is easily interfered with by these particles. Furthermore, even miniaturized detection devices still cost over 100,000 RMB, hindering market adoption.
[0006] 4. Other techniques: Mass spectrometry uses variable wavelengths to measure the spectra of compounds or organic compounds. It can be used to detect a class of pathogens by identifying several major compounds or a group of compounds in microorganisms, but it is mainly used for liquid samples. Near-infrared short-wave spectroscopy can rapidly determine the spectra of hundreds of microorganisms and the number of microorganisms per unit area or volume. Based on this principle, ultraviolet resonance Raman spectroscopy can also detect bacteria and spores. However, these methods are relatively complex and still in the research stage, and are very expensive. They are suitable for protection fields with different needs, such as identifying microbial species.
[0007] Traditional airborne microbial detection technologies, such as culture methods, have limitations such as long detection cycles and inability to perform real-time detection, while real-time airborne microbial particle detectors are expensive and cannot be widely adopted. The distribution of microbial content in indoor air is related to environmental parameters such as particulate matter, temperature, humidity, and CO2. Based on extensive field measurement data, this study investigates this correlation, using statistical methods to establish a mathematical model of the real-time relationship between environmental parameters and airborne microbial content and particle size distribution. The model can then be used to calculate the real-time microbial content and particle size distribution from environmental parameters, providing a low-cost, easy-to-implement, and widely applicable technical means for real-time online monitoring. To establish this mathematical model, existing technical approaches include... Figure 1As shown in the diagram. The scheme first involves on-site measurements and data collection, followed by data classification, organization, statistical analysis, and correlation analysis. During on-site measurements, PM1, PM2.5, and PM10 data were collected for monitoring air particulate matter concentration; an Anderson six-stage impactor sampler was used for air microbial sampling; and temperature, humidity, and CO2 concentration data were collected for air quality. In the actual monitoring process, due to differences in human activity, ventilation and air conditioning systems, and indoor microbial pollution sources under different environmental types, researchers conducted on-site measurements for different environmental types and then performed correlation analysis to establish functional relationships between air microbial concentration and air particulate matter concentration, temperature, humidity, and CO2 concentration parameters.
[0008] In summary, the main technical defects in the existing technology can be summarized as follows:
[0009] 1. Traditional airborne microbial detection technologies, such as culture methods, have limitations such as long detection cycles and inability to detect in real time; while real-time airborne microbial particle detection instruments are expensive and cannot be widely used.
[0010] 2. Existing statistical methods often employ a single, highly correlated environmental data point to establish a functional relationship between airborne microbial concentration and this data. However, in reality, changes in airborne microbial concentration are the result of the combined influence of multiple factors, including outdoor and indoor air pollution levels, air quality parameters, and human activities—meaning they are correlated with multiple parameters. In specific environments, using multiple environmental data points that significantly impact airborne microbial pollution to establish a functional relationship between airborne microbial concentration and these data points can significantly reduce errors. Conversely, a functional relationship based on a single environmental data point may discard certain relevant factors, leading to substantial errors when airborne microbial concentration changes due to variations in these factors.
[0011] 3. Existing statistical methods-based solutions utilize data correlation analysis results to establish functional relationships for variables with strong correlations using single linear regression. For example, a literature study on regression analysis results of subway bacteria and particulate matter... Figure 2 As shown. From Figure 2 As can be seen from the graph, the linear regression results basically conform to the distribution characteristics of each point, and the bacterial content calculated using this functional relationship will not have a large error and is basically usable. However, after collecting a large amount of data, we found that in the real air environment, the concentration of airborne microorganisms is affected by various linear and nonlinear factors, and cannot be described by a simple linear regression equation. Linear regression will lose a lot of useful information, especially the causal relationship between particulate matter concentration and microbial concentration. Figure 3For data from a specific office in a certain location, the correlation between bacteria and PM2.5 is weak. Directly using linear regression cannot accurately reflect their intrinsic functional relationship, resulting in significant errors in the calculated bacterial content. Therefore, for data that cannot be easily analyzed using linear regression, it is necessary to explore other numerical analysis methods that can better uncover the underlying patterns of change and establish a functional relationship. This is crucial to establishing a usable functional relationship between airborne microbial concentration and environmental data even when the correlation is weak, thereby improving the accuracy of real-time monitoring of airborne microbial content. Summary of the Invention
[0012] Therefore, the technical problem to be solved by the present invention is to provide a method for constructing a real-time monitoring model for airborne microorganisms, so as to solve the problems of long detection cycle, poor timeliness, large error and low accuracy of existing airborne microorganism monitoring methods.
[0013] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0014] A method for constructing a real-time airborne microbial monitoring model includes the following steps:
[0015] Step (1): Collect airborne bacteria in the air at the location to be monitored, and calculate the concentration of airborne bacteria in the air at the location to be monitored using the culture counting method;
[0016] Step (2): During the same time period, monitor and collect air quality data in the air at the location to be monitored;
[0017] Step (3): Repeat steps (1) and (2) to collect data on the concentration of airborne bacteria and air quality at the monitoring locations during different time periods, and form a multidimensional data table.
[0018] Step (4): Based on the multidimensional data table obtained in step (3), analyze the correlation between the concentration of airborne bacteria in the air at the location to be monitored and the air quality data (generally the correlation is very poor, usually less than 0.5, so a Gaussian model is required).
[0019] Step (5): Based on the correlation analysis results from step (4), select the air quality data with a high correlation to the concentration of airborne bacteria [generally, the data with the highest correlation] as vector X in the Gaussian model (e.g., vector (x1, x2, x3)). T T points to the transpose of the vector, where x1 is the concentration of PM2.5, x2 is the temperature, and x3 is the humidity (the example below uses a one-dimensional vector). Gaussian mixture model is used to model [to obtain multiple ellipses for classification, that is, to wrap the data in different ellipses (hyperellipsoids), and high dimension is the hyperellipsoid], resulting in a Gaussian mixture model with a K value greater than or equal to 2. The K value represents the number of ellipses in the Gaussian mixture model, and the K value is a positive integer.
[0020] Step (6): Optimize the K value of the Gaussian mixture model to obtain the final Gaussian mixture model;
[0021] Step (7): Select the major axis of the corresponding ellipse as the constituent line segment of the piecewise function. The slope of the major axis is less than or equal to 1 (since the concentration of microorganisms in the actual environment will not suddenly increase, the major axis with an absolute value of slope not exceeding 1 is taken as the main line segment) to obtain a continuous piecewise linear function, which is the real-time air microbial monitoring model.
[0022] In real-world environments, the correlation between airborne microorganisms and single environmental parameters is generally weak. To address this, this invention designs a piecewise linear model construction method based on Gaussian mixture model preprocessing. By utilizing real-time monitored ambient air quality parameters and combining them with the piecewise linear model designed in this invention, the real-time microbial content can be calculated, providing a highly accurate, low-cost, and easily implemented and promoted technical means for real-time online monitoring of airborne microorganisms.
[0023] In the above method for constructing a real-time airborne microbial monitoring model, step (1) involves using a six-stage Anderson impactor sampler to collect airborne bacteria at the location to be monitored; then, the six plates of the six-stage Anderson sampler are cultured according to the national standard method, and the number of colonies on each of the six plates is recorded to calculate the concentration of airborne bacteria at the location to be monitored. The six-stage Anderson impactor sampler is a detection device used in the national standard method to collect and detect the total number of airborne bacteria in the air.
[0024] In the above method for constructing a real-time air microbial monitoring model, step (2) includes air environmental quality data including one or a combination of two or more of the following: particulate matter concentration, carbon dioxide concentration, temperature data, and humidity data; particulate matter concentration includes at least the concentration of PM2.5.
[0025] In the above method for constructing a real-time airborne microbial monitoring model, step (2) involves using an air quality monitor to monitor and collect air quality data at the location to be monitored. Impact sampling and air quality monitoring should, as far as possible, be conducted at the same time, at the same specific location (including altitude), and under the same environmental conditions.
[0026] The method for constructing the aforementioned real-time airborne microbial monitoring model requires an air quality monitor that is one or a combination of two or more of the following: a laser particle counting sensor, a carbon dioxide sensor, a temperature sensor, and a humidity sensor. The indoor air quality real-time continuous monitoring device must include at least PM2.5 monitoring functionality.
[0027] In the above method for constructing the real-time monitoring model of airborne microorganisms, in step (3), the total number of data groups of multidimensional data is n, and n is a positive integer greater than or equal to 45 [under the premise of having experimental conditions, in order to ensure that the obtained data is statistically significant, the number of samplings should be as many as possible in order to obtain more data]; the K value is greater than or equal to 2 and less than or equal to n / 4.
[0028] In the above method for constructing a real-time airborne microbial monitoring model, step (4) uses the correlation coefficient r, which is calculated using the following formula:
[0029]
[0030] In Formula 1, n is the total number of data sets; x i y represents the total number of colonies in the i-th sample; i For air quality data; This represents the average total number of colonies in n groups; This represents the average value of n sets of air quality data.
[0031] In the above method for constructing a real-time airborne microbial monitoring model, step (5) involves modeling a Gaussian mixture model, and the calculation formula for its probability function is as follows:
[0032]
[0033] In Formula 2, x represents air quality data; p(x) is the weighted sum of k p(x|μ,∑); k is the number of components (the number of components corresponds to the number of ellipses, i.e., the number of components is equal to the number of ellipses); α is the weight coefficient of each component; p(x|μ,∑) is the probability density function of an n-dimensional random vector x that follows a Gaussian distribution.
[0034] therefore,
[0035] Formula 3 is the multidimensional Gaussian distribution function, where n is the vector of x, μ is the expected value of x, μ = E(x), T denotes transpose, and ∑ is the covariance matrix; where ∑ = Cov(x) = E{(x-μ)(x-μ)} T}
[0036] The specific method for step (6) of the above-mentioned method for constructing the real-time airborne microbial monitoring model is as follows:
[0037] Step (6-1): Set the K value of the Gaussian mixture model to k; K is the number of components in the Gaussian mixture model [i.e. the number of ellipses]. The maximum value of K is one-quarter of the number of samples n, that is, k is greater than or equal to 2 and less than or equal to n / 4, and k is a positive integer.
[0038] Step (6-2): Based on the value of k, perform Gaussian mixture modeling, calculate the slope of the major axis using the two foci of the obtained ellipse, and obtain k linear equations;
[0039] Step (6-3): Substitute the raw air quality data into the linear equation obtained in step (6-2) to calculate the calculated concentration of airborne bacteria; compare the calculated concentration of airborne bacteria with its corresponding actual measured concentration, and obtain the sum of the absolute values of their errors as T. k ;
[0040] Step (6-4): Set the K value of the Gaussian mixture model to k+1, repeat steps (6-2) and (6-3), and obtain the sum of the absolute values of the errors between the calculated concentration of planktonic bacteria and the corresponding actual measured concentration of planktonic bacteria when K is k+1, which is T. k+1 ;
[0041] Step (6-5): Set the K value of the Gaussian mixture model to k+2, repeat steps (6-2) and (6-3) to obtain T. k+2 ...; Set the K value of the Gaussian mixture model to n / 4, round n / 4, and repeat steps (6-2) and (6-3) to obtain T. n / 4 Compare T k T k+1 T k+2 ...T n / 4 The optimal K value for the Gaussian mixture model is determined by the magnitude of the sum of the absolute values of the errors.
[0042] The above method for constructing a real-time airborne microbial monitoring model, when setting the K value of the Gaussian mixture model to k, denotes the data samples of the multidimensional data as D = {x1, x2, ..., x3}. m}; Therefore, the Gaussian mixture clustering algorithm is used to process data sample D as follows:
[0043] Input: Sample set D = {x1, x2, ... x} m};
[0044] process:
[0045]
[0046]
[0047] Output: Component partitioning C = {C1, C2, ..., C} k}
[0048] The technical solution of the present invention achieves the following beneficial technical effects:
[0049] 1. This invention uses one or more data such as microbial concentration and particulate matter concentration, CO2 concentration and temperature and humidity related to microbial concentration to establish a statistical function model. Since changes in particulate matter concentration, CO2 concentration and temperature and humidity in the real environment will affect the changes in microbial concentration over time, the model established in this invention uses a variety of data with high accuracy.
[0050] 2. When the correlation between microbial concentration and particulate matter concentration, CO2 concentration, temperature and humidity is weak, directly using a function model based on linear regression will result in a large error in the calculated microbial concentration. In contrast, the piecewise linear function model obtained by superimposing Gaussian mixture models in this invention has higher accuracy.
[0051] 3. Compared with traditional national standard culture methods and real-time airborne microbial particle detection instruments, the statistical function model established in this invention can calculate the current airborne microbial concentration using real-time monitored air quality parameters. This method is more accurate, less costly, and easier to implement and promote. Attached Figure Description
[0052] Figure 1 The technical approach for establishing a mathematical model for real-time online monitoring of airborne microorganisms in existing technologies;
[0053] Figure 2 Scatter plot of airborne bacteria and PM10 concentration in subways in existing technology;
[0054] Figure 3 A scatter plot of airborne bacteria and PM2.5 concentration in an office using existing technology;
[0055] Figure 4 Distribution diagram of total bacterial count and measured PM2.5 points in embodiments of the present invention;
[0056] Figure 5 The linear function graph of the linear regression between total bacterial count and PM2.5 in the embodiments of the present invention;
[0057] Figure 6 The Gaussian mixture model data processing component partitioning diagram obtained when k is 4 in this embodiment of the invention;
[0058] Figure 7 The sum of squared errors of the Gaussian mixture model after processing under different k values in the embodiments of the present invention;
[0059] Figure 8 The construction graph of the piecewise linear function when k is 4 in this embodiment of the invention; Detailed Implementation
[0060] In this embodiment, the method for constructing the real-time airborne microbial monitoring model specifically includes the following steps:
[0061] Step (1): Use a six-stage Anderson impact sampler to monitor and collect airborne bacteria at a specific location; after culturing the six plates of the six-stage Anderson sampler according to the national standard method, record the colony data of the six plates respectively, sum them to obtain a colony summary, and calculate the equivalent airborne bacteria concentration.
[0062] Step (2): Using a real-time air quality monitor, monitor and collect particulate matter concentrations (PM2.5), CO2 concentrations, temperature data, and humidity data in the air at the same time and at the same specific location, record the data, and take the average value.
[0063] Step (3): Next, the data collection work of steps (1) and (2) is carried out simultaneously at different time periods in the same location, and the large amount of data collected is made into a multidimensional data table. The multidimensional data table obtained in this embodiment is shown in Table 1. A total of 45 samples were collected in this embodiment. Due to space limitations, only a part of the data is listed.
[0064] Table 1
[0065]
[0066] Step (4) Analyze the correlation between the multidimensional data collected by the air quality monitor and the airborne bacteria concentration data, and draw a correlation conclusion.
[0067] According to national standards, the flow rate of a level VI Anderson impact sampler is 28.3 L / min, and the duration of a single impact sampling is set to 15 min. Therefore, the air volume collected in a single impact sampling is 28.3 × 15 = 424.5 L = 0.4245 m³. 3 Divide the colony count in the table by its corresponding air volume, which is 0.4245m³. 3 The corresponding airborne bacteria concentration can then be obtained, in units of CFU / m³. 3 ;
[0068] For the multidimensional data in the table, we first analyzed the correlation between the total bacterial count and PM2.5, CO2 concentration, temperature, and humidity, respectively, that is, we calculated the correlation coefficient (denoted by r), which, strictly speaking, should be called the "linear correlation coefficient." This is because the correlation coefficient only describes the degree of "linear" relationship between the two sets of variables X and Y; the formula for calculating r is as follows:
[0069]
[0070] In Formula 1, n represents the total number of data sets; in this example, n = 45; x i Let y be the total number of colonies in the i-th sample. i This refers to a specific air quality data point (i.e., one of PM2.5, CO2 concentration, temperature, and humidity). This represents the average total number of colonies in the 45 groups. It is the average value of 45 sets of certain air quality data.
[0071] The calculated correlation coefficient r ranges from -1 to 1. A positive r indicates a positive correlation between the two variables, a negative r indicates a negative correlation, and 0 indicates no correlation. The closer the absolute value of r is to 1, the stronger the correlation; the closer it is to 0, the weaker the correlation. The degree of correlation between variables can be roughly estimated using the following criteria: r < 0.2 indicates a weak correlation, 0.2 < r < 0.4 indicates a low correlation, 0.4 < r < 0.7 indicates a moderate correlation, and r > 0.7 indicates a high correlation. Generally, if the correlation is weak (< 0.5), linear regression analysis cannot be directly used to establish a functional relationship between airborne microbial concentration and environmental data.
[0072] For the data in this embodiment, the correlation coefficients between the total number of colonies and PM2.5, CO2 concentration, temperature, and humidity were calculated, and the results are shown in Table 2:
[0073] Table 2
[0074] Correlation PM2.5 CO2 temperature humidity Total bacterial count 0.209404088 0.078763433 -0.127728336 0.044934085
[0075] Table 2 shows that the total bacterial count is weakly correlated with PM2.5 and weakly correlated with CO2 concentration, temperature, and humidity; that is, there are other functional relationships between the total bacterial count and these variables, but not a simple direct linear relationship. To facilitate comparison between the method of this invention and existing methods, this embodiment still selects one variable to establish a linear regression model. Since the correlation coefficient of PM2.5 is relatively large, PM2.5 monitoring data is selected to obtain a distribution map of the total bacterial count and PM2.5 data points, as shown below. Figure 4 The linear function established directly using the linear regression method is: Figure 5 The line in the diagram is shown; the linear function is: Y = 0.09X² + 4.70; where Y is the total number of colonies and X² is PM2.5.
[0076] Step (5): Next, Gaussian mixture modeling analysis is performed on the total bacterial count and PM2.5. The Gaussian mixture model uses the Gaussian probability density function (normal distribution curve) to accurately quantify things. It decomposes things into several models based on the Gaussian probability density function (normal distribution curve). Using the Gaussian mixture model to model multidimensional data yields multiple ellipses for classification, that is, wrapping the data in different ellipses (hyperellipsoids). The higher dimension is the hyperellipsoid; its probability function and probability density function are as follows:
[0077]
[0078] Where x represents PM2.5 data, p(x) is the weighted sum of k p(x|μ,∑); k represents the number of components (corresponding to the number of ellipses); α is the weight coefficient of each component; p(x|μ,∑) is the probability density function of an n-dimensional random vector x that follows a Gaussian distribution;
[0079]
[0080] Formula 3 is the multidimensional Gaussian distribution function, where n is the vector of x, and μ is the expected value of x, μ = E.
[0081] (x), where T denotes the transpose, and ∑ is the covariance matrix; in Formula 3, ∑=Cov(x)=E{(x-μ)(x-μ)} T};
[0082] Assume that all data samples D are generated by a mixture distribution composed of k multivariate Gaussian distributions. The Gaussian mixture clustering algorithm is used to process data samples D, with the following steps [other processing methods in the existing technology can also be used]:
[0083] Input: Sample set D = {x1, x2, ... x} m};
[0084] process:
[0085]
[0086]
[0087] Output: Component partitioning C = {C1, C2, ..., C} k}
[0088] In this embodiment, the data sample D = {x1, x2}, where x1 is the total colony count and x2 is PM2.5. A K value needs to be specified during data processing; the size of K determines the number of component divisions. When representing the data processing results graphically, in the distribution map of total colony count and PM2.5, one component corresponds to one ellipse. When K is 4, the component divisions obtained after data processing using a mixture multivariate Gaussian distribution are as follows: Figure 6 As shown.
[0089] Based on the elliptical shape formed by component division, linear regression is then performed on the data within each ellipse to obtain the linear function of each component, i.e., the major axis of symmetry of each ellipse, such as... Figure 8 As shown.
[0090] The piecewise linear function is shown below:
[0091] x1 = -0.70x2 + 16.15, x2 takes values in the range [0, 20];
[0092] x1 = 0.48x2 - 6.31, x2 takes values in the range [20, 33];
[0093] x1 = 0.02x2 + 4.96, x2 takes values in the range [33, 57];
[0094] x1 = 1.28x2 - 66, x2 takes values in the range [57, 100];
[0095] Using the piecewise linear function described above, the total bacterial count x1 for each measured PM2.5 data set can be calculated. This count is then compared with the total bacterial count X from the national standard method to determine the error t of the piecewise linear function model. The sum of squared errors T for the 45 data sets can then be calculated using the following formula:
[0096] t = x1 - X;
[0097] T = t1*t1 + t2*t2 + ... + t 45 *t 45 ;
[0098] Step (6) Optimize the number of ellipses (hyperellipsoids). The specific operation method is as follows:
[0099] Step (6-1): Set the K value of the Gaussian mixture model to k;
[0100] Step (6-2): Based on the value of k, perform Gaussian mixture modeling, calculate the slope of the major axis using the two foci of the obtained ellipse, and obtain k linear equations;
[0101] Step (6-3): Substitute the raw air quality data into the linear equation obtained in step (6-2) to calculate the calculated concentration of airborne bacteria; compare the calculated concentration of airborne bacteria with its corresponding actual measured concentration, and obtain the sum of the absolute values of their errors as T. k ;
[0102] Step (6-4): Set the K value of the Gaussian mixture model to k+1, repeat steps (6-2) and (6-3), and obtain the sum of the absolute values of the errors between the calculated concentration of planktonic bacteria and the corresponding actual measured concentration of planktonic bacteria when K is k+1, which is T. k+1 ;
[0103] Step (6-5): Set the K value of the Gaussian mixture model to k+2, repeat steps (6-2) and (6-3) to obtain T. k+2 ...; Set the K value of the Gaussian mixture model to n / 4, round n / 4, and repeat steps (6-2) and (6-3) to obtain T. n / 4 Compare T k Tk+1 T k+2 ...T n / 4 The optimal K value for the Gaussian mixture model is determined by the magnitude of the sum of the absolute values of the errors.
[0104] When K is 1, it is equivalent to simple linear regression. Generally, the value of K can increase from 2. The maximum value of K should take into account the number of data sets in the data sample D. In this example, there are 45 data sets. To determine an ellipse, 4 to 5 data sets are needed, that is, 4 to 5 points determine an ellipse. Therefore, the maximum value of K is 9. Similar to the steps when K is 4, for other values of K from 1 to 9, data processing is performed using a mixture multivariate Gaussian distribution to obtain component partitions. Based on the elliptical graphs formed by the component partitions, linear regression is then performed on the data within each ellipse to obtain the linear function of each component. Finally, each piecewise linear function is generated, and the sum of squares of its errors is calculated. The sum of squares of the errors of the piecewise linear models established under different K values can be obtained as follows: Figure 7 As shown. By Figure 7 It can be seen that the sum of squared errors is minimized when K=4, so this embodiment ultimately adopts the piecewise linear model generated by K=4.
[0105] Step (7): Select the major axis of the corresponding ellipse as the constituent line segment of the piecewise linear function: Since the concentration of microorganisms in the actual environment will not suddenly increase, the major axis with an absolute value of slope not exceeding 1 is selected as the main line segment to form a continuous piecewise linear function.
[0106] In this embodiment, the linear function Y(X2) obtained directly from the linear regression model is: Y = 0.09X2 + 4.70; where Y is the total colony count and X2 is PM2.5. The piecewise linear function x1(x2) obtained after Gaussian mixture clustering analysis is taken as K = 4. Based on the measured PM2.5 data for each group, the corresponding total colony count Y is calculated using function Y(X2), and the corresponding total colony count x1 is calculated using function x1(x2). These are then compared with the total colony count X detected by the national standard method to obtain the error rates m0 and n0 of each function. By calculating and comparing the mean (i.e., average error rate) of the error rates of 45 groups of data, the accuracy of the linear function Y(X2) and the piecewise linear function x1(x2) can be determined. Specifically:
[0107] m0 = |(YX)| / X*100%;
[0108] n0 = |(x1-X)| / X*100%;
[0109] In this embodiment, the average value of m0 is 46.56%, and the average value of n0 is 26.48%; the average error rate of the piecewise linear function x1(x2) is much smaller than that of the linear function Y(X2).
[0110] This embodiment designs a method for monitoring the concentration of airborne microorganisms in enclosed indoor environments based on unsupervised learning and Gaussian mixture clustering analysis. The collected environmental data (particulate matter, temperature, humidity, CO2) are corrected and processed, and the data format is standardized through a preprocessing procedure. A Gaussian mixture model is used to obtain the multidimensional normal distribution of regionally related microbial concentrations. Multiple sub-hyperellipsoids of microbial concentrations are unsupervisedly partitioned. The objective function for microbial concentration is derived based on the principle of region minimization. Cluster analysis is then used to transform the microbial concentration prediction into a linear solution process. Finally, by combining data from different time periods (temperature, humidity, particulate matter, CO2), the final short-term regional microbial concentration prediction results are obtained. Experimental results demonstrate that, compared with existing prediction methods, the prediction results based on mixture clustering analysis are closer to the actual values and have higher accuracy.
[0111] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of the claims of this patent application.
Claims
1. A method for constructing an air microorganism real-time monitoring model, characterized in that, Includes the following steps: Step (1): Collect airborne bacteria in the air at the location to be monitored, and calculate the concentration of airborne bacteria in the air at the location to be monitored using the culture counting method; Step (2): During the same time period, monitor and collect air quality data in the air at the location to be monitored; the air quality data is one or a combination of two or more of the following: particulate matter concentration, carbon dioxide concentration, temperature data, and humidity data. Step (3): Repeat steps (1) and (2) to collect data on the concentration of airborne bacteria and air quality at the monitoring locations during different time periods, and form a multidimensional data table; Step (4): Based on the multidimensional data table obtained in step (3), analyze the correlation between the concentration of airborne bacteria in the air at the monitoring location and the air quality data; Step (5): Based on the correlation analysis results in step (4), select the air quality data that is highly correlated with the concentration of airborne bacteria as the vector in the Gaussian model, and use the Gaussian mixture model to model the model to obtain a Gaussian mixture model with a K value greater than or equal to 2. The K value represents the number of ellipses in the Gaussian mixture model, and the K value is a positive integer. Step (6): Optimize the K value of the Gaussian mixture model to obtain the final Gaussian mixture model; Step (7): Select the major axis of the corresponding ellipse as the line segment of the piecewise function. The slope of the major axis is less than or equal to 1 to obtain a continuous piecewise linear function, which is the real-time air microbial monitoring model. In step (5), when modeling the Gaussian mixture model, the probability function is calculated using the following formula: (Formula 2); In formula 2, x For air quality data; p ( x )for k indivual p ( x| The weighted sum of μ,∑); k The number of components corresponds to the number of ellipses; α is the weight coefficient for each component. p ( x| μ,∑) are Gaussian distributions. n 3D random vector x The probability density function; Thus, (Formula 3); Equation 3 is a multi-dimensional Gaussian distribution function, n is x a vector, μ is x an expected value μ = E(x), T denotes a transpose, and ∑ is a covariance matrix; in Equation 3, ∑ = Cov( x )= E{( x - μ )( x - μ ) T}; The specific method for step (6) is as follows: Step (6-1): Set the K value of the Gaussian mixture model to k ; k greater than or equal to 2 and less than or equal to n / 4, k is a positive integer; Step (6-2): Gaussian mixture modeling is performed according to the values of k The slope of the major axis is calculated using the two foci of the k obtained ellipse, resulting in a linear equation k Step (6-3): The linear equation is solved to find the intersection point of the major axis with the line Step (6-3): Substitute the raw air quality data into the linear equation obtained in step (6-2) to calculate the calculated concentration of airborne bacteria; compare the calculated concentration of airborne bacteria with its corresponding actual measured concentration, and obtain the sum of the absolute values of their errors as T. k ; Step (6-4): Set the K value of the Gaussian mixture model to be... k +1, repeat steps (6-2) and (6-3) to obtain the value of K. k When +1, the sum of the absolute values of the errors between the calculated concentration of planktonic bacteria and the corresponding actual measured concentration of planktonic bacteria is T. k+1 ; Step (6-5): Set the K value of the Gaussian mixture model to k+2, repeat steps (6-2) and (6-3) to obtain T. k+2 ...; Set the K value of the Gaussian mixture model to n / 4, round n / 4, and repeat steps (6-2) and (6-3) to obtain T. n / 4 Compare T k T k+1 T k+2 ...T n / 4 The optimal K value for the Gaussian mixture model is determined by the magnitude of the sum of the absolute values of the errors.
2. The method of claim 1, wherein the method further comprises: In step (1), an airborne bacteria in the air is collected at the location to be monitored using a six-stage Anderson impact sampler; then, the six plates of the six-stage Anderson sampler are cultured according to the national standard method, and the number of colonies in the six plates is recorded respectively, and the concentration of airborne bacteria in the air at the location to be monitored is calculated.
3. The method of claim 1, wherein the method further comprises: determining a concentration of the microorganism in the air. In step (2), an air quality monitor is used to monitor and collect air quality data at the location to be monitored.
4. The method of claim 3, wherein the method further comprises: Air quality monitors are one or more of the following: laser particle counting sensors, carbon dioxide sensors, temperature sensors, and humidity sensors.
5. The method of claim 1, wherein, In step (3), the total number of data groups of the multi-dimensional data is n , and n is a positive integer and is greater than or equal to 45.
6. The method of claim 5, wherein the method further comprises: In step (4), the correlation is represented by a correlation coefficient r , r The calculation formula is: (Formula 1); in Equation 1, x i for the first i total number of colonies for the first sampling; y i for the air environmental quality data; for n the average of the total number of colonies for the group; for n the average of the air environmental quality data for the group.
Citation Information
Patent Citations
On-line detection method for acquiring number of bacterial colonies in air, medium and electronic equipment thereof
CN113360846A