High-efficiency collection system for crowd aerosol data based on big data
By determining the target peak point and baseline sequence in gas chromatography analysis and using an improved Gaussian kernel weighted filter, the baseline fluctuation problem caused by noise interference was solved, thus improving the accuracy of aerosol component analysis.
Patent Information
- Application Number
- CN202510505132.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In existing technologies, when analyzing human aerosols using gas chromatography, noise interference causes baseline fluctuations, affecting the accuracy of peak shape and area information. Gaussian filtering is difficult to effectively filter these peaks, thus impacting the accuracy of gas chromatography analysis.
The data acquisition module acquires gas chromatographic data, the analysis module determines the target peak point and the cutoff point, the feature processing module constructs the baseline sequence and the benchmark value, and the data filtering module uses an improved Gaussian kernel weight to perform filtering, thereby improving the accuracy of baseline filtering.
It improves the accuracy of gas chromatography data filtering, enhances the accuracy of aerosol component analysis, and reduces filtering errors.
Smart Images

Figure CN120408084A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gas chromatography analysis, and particularly relates to an efficient acquisition system for population aerosol data based on big data. Background Art
[0002] Population aerosol mainly refers to a complex gaseous dispersion system generated by the human respiratory tract; some biomarker molecules produced by human metabolism, such as disease biomarker molecules, will be excreted from the body through the human respiratory system. The diameter of these particles is usually in the micron range and can float in the air for a long time. Measuring the biomarker molecules in population aerosol is a fast, accurate and non-contact disease diagnosis method.
[0003] In the prior art, gas chromatography analysis is usually used to analyze and measure the components of the collected population aerosol; however, in the process of obtaining the gas chromatography data of population aerosol, due to noise interference and other situations, there may be fluctuations such as noise in the baseline, which affects the accurate acquisition of information such as peak area; and different aerosols may have different influencing factors and components, and the degree of noise interference on their baseline data may also be different, making it difficult for the existing Gaussian filtering to accurately filter the gas chromatography data, ultimately affecting the accuracy of gas chromatography analysis of aerosol components. Summary of the Invention
[0004] In order to solve the technical problem that the accuracy of Gaussian filtering for filtering gas chromatography data is relatively low, which in turn affects the accuracy of analyzing aerosol components, the purpose of the present invention is to provide an efficient acquisition system for population aerosol data based on big data, and the specific technical solutions adopted are as follows: A data acquisition module, used to acquire gas chromatography data of population aerosol; A data analysis module, used to obtain target peak points according to the data characteristics of peak points and the difference characteristics of data distribution within a preset neighborhood range of peak points in the gas chromatography data; obtain segmentation points of the gas chromatography data according to the data change characteristics within the preset neighborhood range of the target peak points; obtain a baseline sequence of the gas chromatography data according to the segmentation points; obtain a reference value according to the data distribution characteristics of the baseline sequence; A feature processing module, used to construct a local sequence according to the baseline data points and preset connected data points in the baseline sequence; obtain baseline credibility according to the data change characteristics of the local sequence and the data difference characteristics between the local sequence and the reference value; obtain a reference sequence according to the baseline credibility; obtain the degree of noise interference according to the data difference characteristics and the difference characteristics of the degree of dispersion between the local sequence of any baseline data point and the reference sequence; A data filtering module, which is used to obtain the noise interference similarity degree of any baseline data point according to the noise interference degrees of the any baseline data point and other baseline data points; obtain the improved Gaussian kernel weight between the any baseline data point and other baseline data points according to the Gaussian filtering algorithm and the noise interference similarity degree; filter the any baseline data point according to the improved Gaussian kernel weight to obtain improved gas chromatography data.
[0005] Further, the step of obtaining the target peak point according to the data characteristics of the peak point in the gas chromatography data and the difference characteristics of the data distribution within the preset neighborhood range of the peak point includes: Taking the peak point as the center within the preset neighborhood range, calculating the average value of the absolute value of the difference in the amplitudes of the data points with the same rank before and after and performing a negative correlation mapping to obtain a symmetry eigenvalue; calculating the product of the symmetry eigenvalue and the amplitude of the peak point and normalizing it to obtain the peak confidence degree of the peak point; taking the peak point with the peak confidence degree exceeding the preset confidence threshold as the target peak point.
[0006] Further, the step of obtaining the segmentation point of the gas chromatography data according to the data change characteristics within the preset neighborhood range of the target peak point includes: Within the preset neighborhood range of the target peak point, linearly fitting the data points on both sides of the target peak point respectively, and taking the data points where the two obtained fitting lines intersect with the gas chromatography data as the segmentation point.
[0007] Further, the step of obtaining the baseline sequence of the gas chromatography data according to the segmentation point includes: Segmenting the gas chromatography data at the segmentation point to obtain different data segments; sorting the values in the data segments that do not contain the target peak point in chronological order to obtain the baseline sequence.
[0008] Further, the step of obtaining the reference value according to the data distribution characteristics of the baseline sequence includes: Taking the median of the baseline sequence as the reference value.
[0009] Further, the step of obtaining the baseline credibility according to the data change characteristics of the local sequence and the data difference characteristics between the local sequence and the reference value includes: Calculating the absolute value of the difference between the average value of the local sequence and the reference value to obtain a deviation eigenvalue; calculating the sum of the coefficient of variation of the local sequence and the deviation eigenvalue and performing a negative correlation mapping to obtain the baseline credibility of the local sequence.
[0010] Further, the step of obtaining the reference sequence according to the baseline credibility includes: Take the local sequence corresponding to the maximum value of the baseline credibility as the reference sequence.
[0011] Further, the step of obtaining the noise interference degree according to the data difference characteristics and the difference characteristics of the dispersion degree between the local sequence of any baseline data point and the reference sequence includes: Calculate the absolute value of the difference between the average value of the local sequence of any baseline data point and the average value of the reference sequence to obtain the baseline difference characteristic value; calculate the difference between the coefficient of variation of the local sequence of any baseline data point and the reference sequence and map it positively to obtain the dispersion degree difference value; calculate the product of the dispersion degree difference value and the baseline difference characteristic value to obtain the noise interference degree of any baseline data point.
[0012] Further, the step of obtaining the noise interference similarity degree of any baseline data point according to the noise interference degrees of any baseline data point and other baseline data points includes: Calculate the ratio of the noise interference degrees of any baseline data point and the other baseline data points to obtain the noise interference similarity degree of any baseline data point.
[0013] Further, the step of obtaining the improved Gaussian kernel weight between any baseline data point and other baseline data points according to the Gaussian filtering algorithm and the noise interference similarity degree includes: Obtain the initial Gaussian kernel weights of any baseline data point and the other baseline data points according to the Gaussian filtering algorithm; calculate the product of the noise interference similarity degree and the initial Gaussian kernel weights to obtain the improved Gaussian kernel weight between any baseline data point and the other baseline data points.
[0014] The present invention has the following beneficial effects: In the present invention, obtaining the target peak point can determine the region representing the detected substance in the gas chromatography data, and then determine the baseline region to be filtered according to the target peak point; obtaining the segmentation point can more accurately determine the start and end positions of the baseline. Obtaining the reference value can determine the baseline value representing the normal level according to the distribution characteristics of the baseline data, which is beneficial for subsequent steps to determine the noise interference degrees suffered by different baseline data. Calculating the baseline credibility can obtain the reference sequence according to the severity of the noise interference suffered by the baseline, and the noise interference degrees of other local sequences can be evaluated through the reference sequence; obtaining the noise interference degree can determine the severity of the noise interference suffered by the baseline data points at different positions, thereby improving the filtering accuracy. Obtaining the noise interference similarity degree can determine the magnitude of the noise interference degrees suffered between different baseline data points, so as to correct the Gaussian kernel weight between two baseline data points and reduce the filtering error. The finally obtained improved gas chromatography data improves the accuracy of aerosol component analysis compared with the gas chromatography data. Brief Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0016] Figure 1 It is a block diagram of an efficient acquisition system for population aerosol data based on big data provided by an embodiment of the present invention. Detailed Embodiments
[0017] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, detail the specific embodiments, structures, features and effects of an efficient acquisition system for population aerosol data based on big data proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0019] The following will specifically describe the specific solution of an efficient acquisition system for population aerosol data based on big data provided by the present invention with reference to the accompanying drawings.
[0020] Please refer to Figure 1 , which shows a block diagram of an efficient acquisition system for population aerosol data based on big data provided by an embodiment of the present invention. The system includes the following modules: A data acquisition module S1, configured to acquire gas chromatography data of population aerosol.
[0021] In the embodiment of the present invention, the implementation scenario is to filter the gas chromatography data of the analyzed population aerosol to improve the accuracy of analyzing the components of aerosol by gas chromatography. First, the gas chromatography data of the population aerosol is acquired. It should be noted that using a gas chromatograph to acquire the gas chromatography data in the aerosol sample belongs to the prior art, and the specific steps will not be elaborated here.
[0022] The data analysis module S2 is used to obtain target peak points based on the data characteristics of the peak points in the gas chromatography data and the difference characteristics of the data distribution within the preset neighborhood range of the peak points; obtain the segmentation points of the gas chromatography data based on the data change characteristics within the preset neighborhood range of the target peak points; obtain the baseline sequence of the gas chromatography data according to the segmentation points; and obtain the reference value according to the data distribution characteristics of the baseline sequence.
[0023] When analyzing the gas chromatography data of population aerosol, due to the influence of impurities in the aerosol sample and the background noise of the instrument, etc., it is easy to cause fluctuations in the baseline in the gas chromatography data, which affects the extraction of the peak characteristics in the gas chromatography data, resulting in inaccurate component analysis. Since there are certain differences in the components and contents in the aerosol sample, the heights and positions of the peaks in the generated gas chromatography data are different. In order not to affect the height and position of the peaks when correcting the baseline, it is necessary to improve the accuracy of baseline filtering. First, it is necessary to determine the peak region in the gas chromatography data that may be composed of the detected substance, and then the data segments that do not belong to the peak region may be the baseline data of noise interference. Since the peaks representing the detected substance have the characteristic of a certain symmetry and the amplitude characteristics of such peaks are relatively obvious, the target peak points can be obtained according to the data characteristics of the peak points in the gas chromatography data and the difference characteristics of the data distribution within the preset neighborhood range of the peak points.
[0024] Preferably, in the embodiment of the present invention, the steps of obtaining the target peak points include: taking the peak point as the center within the preset neighborhood range, calculating the average value of the absolute values of the differences in the amplitudes of the data points with the same rank before and after and performing a negative correlation mapping to obtain the symmetry eigenvalue; in the embodiment of the present invention, the preset neighborhood range is the range composed of 15 data points before and after the peak point respectively centered on the peak point, and the implementer can determine it according to the implementation scenario; for example, for the first adjacent data points before and after the peak point, the ranks are both 1. If the peak is relatively symmetric, the amplitudes of the data points with the same rank before and after the peak point are relatively close, and the smaller the absolute value of the difference, the larger the symmetry eigenvalue, which means that the peak is more likely to represent the detected substance. Calculate the product of the symmetry eigenvalue and the amplitude of the peak point and normalize it to obtain the peak confidence of the peak point; when the amplitude of the peak point is larger, it means that the peak is more likely to represent the detected substance; therefore, the larger the peak confidence, the more the peak can reflect the component characteristics of the aerosol. The peak points with the peak confidence exceeding the preset confidence threshold are used as the target peak points. The target peak points represent the positions in the gas chromatography data that can reflect the detected substance. In the embodiment of the present invention, the preset confidence threshold is 0.6, and the implementer can determine it according to the implementation scenario. The formula for obtaining the peak confidence includes: In the formula, P represents the peak confidence of the peak point, denotes the normalization function, and F denotes the amplitude of the peak point. denotes the exponential function with the natural constant as the base, and N denotes the number of data points on one side within the preset neighborhood range of the peak point. denotes the amplitude of the data point at the nth position before this peak point. denotes the amplitude of the data point at the nth position of the sum of this peak point. denotes the symmetric eigenvalue.
[0025] Further, after obtaining all the target peak points in the gas chromatography data, the relatively stable data between two target peak points may be baseline data. Therefore, the segmentation points of the gas chromatography data are obtained according to the data change characteristics within the preset neighborhood range of the target peak points; preferably, in the embodiments of the present invention, the steps of obtaining the segmentation points include: linearly fitting the data points on both sides of the target peak point within the preset neighborhood range of the target peak point, and taking the data points where the two obtained fitting lines intersect with the gas chromatography data as the segmentation points. It should be noted that linear fitting belongs to the prior art and the specific steps will not be elaborated. The fitting lines characterize the change trends on both sides of the wave peak, and further, the segmentation points characterize the approximate positions where the baseline and the wave peak are distinguished. After obtaining all the segmentation points, all the baseline data in the gas chromatography can be obtained. Therefore, the baseline sequence of the gas chromatography data is obtained according to the segmentation points; preferably, in the embodiments of the present invention, the steps of obtaining the baseline sequence include: segmenting the gas chromatography data at the segmentation points to obtain different data segments; sorting the values in the data segments that do not contain the target peak points in chronological order to obtain the baseline sequence; the baseline sequence characterizes the data characteristics of the baseline in the gas chromatography data.
[0026] In the gas chromatography data, since the baseline fluctuates due to interference from sample impurities and instrument background noise, etc., the fluctuation will affect the position of the average value of the baseline data and it is difficult to determine the normal level of the baseline; while the median can, to a certain extent, characterize the baseline level that avoids noise interference. Therefore, the reference value is obtained according to the data distribution characteristics of the baseline sequence; preferably, in the embodiments of the present invention, the steps of obtaining the reference value include: taking the median of the baseline sequence as the reference value; the reference value characterizes the normal baseline characteristics in the gas chromatography data.
[0027] The feature processing module S3 is used to construct a local sequence according to the baseline data points and the preset connected data points in the baseline sequence; obtain the baseline credibility according to the data change characteristics of the local sequence and the data difference characteristics between the local sequence and the reference value; obtain the reference sequence according to the baseline credibility; obtain the degree of noise interference according to the data difference characteristics and the difference characteristics of the dispersion degree between the local sequence of any baseline data point and the reference sequence.
[0028] In order to further improve the accuracy of baseline filtering, it is necessary to determine the degree of noise interference at different positions in the baseline. First, a local sequence is constructed based on the baseline data points in the baseline sequence and the preset connected data points. In the embodiments of the present invention, the preset connected data points are five other baseline data points before and after each of the baseline data points, and the implementer can determine them according to the implementation scenario. If the degree of noise interference on the baseline is greater, the baseline fluctuation is more obvious, and the difference from the reference value is greater. Therefore, the baseline credibility is obtained based on the data change characteristics of the local sequence and the data difference characteristics between the local sequence and the reference value.
[0029] Preferably, in the embodiments of the present invention, the steps of obtaining the baseline credibility include: calculating the absolute value of the difference between the average value of the local sequence and the reference value to obtain the deviation characteristic value; when the difference between the average value of the local sequence and the reference value is greater, the deviation characteristic value is greater, which means that the baseline at this local sequence deviates more from the median and is more affected by noise interference. Calculate the sum of the coefficient of variation of the local sequence and the deviation characteristic value and perform a negative correlation mapping to obtain the baseline credibility of this local sequence; it should be noted that the calculation of the coefficient of variation belongs to the prior art, and the specific calculation steps will not be elaborated. When the fluctuation of the local sequence is more obvious, the coefficient of variation is greater, and it is more likely to be affected by noise interference; therefore, the smaller the baseline credibility, the greater the degree of noise interference on this local sequence, and the less it conforms to the normal baseline characteristics.
[0030] Furthermore, when the baseline credibility is greater, it means that the degree of noise interference on this local sequence is greater, and it can better represent the normal baseline characteristics. Then, a reference sequence is obtained based on the baseline credibility; preferably, in the embodiments of the present invention, the steps of obtaining the reference sequence include: taking the local sequence corresponding to the maximum value of the baseline credibility as the reference sequence, and the reference sequence represents the normal baseline characteristics in the gas chromatography data. After obtaining the reference sequence, the degree of noise interference on the local sequence can be determined according to the difference characteristics between the local sequence and the reference sequence, thereby improving the accuracy of filtering; therefore, the degree of noise interference is obtained based on the data difference characteristics and the difference characteristics of the dispersion degree between the local sequence of any baseline data point and the reference sequence.
[0031] Preferably, in the embodiments of the present invention, the step of obtaining the noise interference degree includes: calculating the absolute value of the difference between the average value of the local sequence of any baseline data point and the average value of the reference sequence to obtain a baseline difference eigenvalue; when the baseline difference eigenvalue is larger, it means that the feature difference between the local sequence of the any baseline data point and the baseline sequence is larger, and the degree of noise interference is greater. Calculating the difference between the coefficient of variation of the local sequence of any baseline data point and the reference sequence and performing a positive correlation mapping to obtain a dispersion degree difference value; when the dispersion degree difference value is larger, it means that the volatility of the local sequence of the any baseline data point is greater than that of the baseline sequence, and the degree of noise interference is greater. Calculating the product of the dispersion degree difference value and the baseline difference eigenvalue to obtain the noise interference degree of any baseline data point; when the noise interference degree is larger, it means that the noise interference at the any baseline data point is more serious.
[0032] The data filtering module S4 is configured to obtain the noise interference similarity degree of any baseline data point according to the noise interference degrees of the any baseline data point and other baseline data points; obtain the improved Gaussian kernel weight between the any baseline data point and other baseline data points according to the Gaussian filtering algorithm and the noise interference similarity degree; filter the any baseline data point according to the improved Gaussian kernel weight to obtain improved gas chromatography data.
[0033] After obtaining the noise interference degrees of different baseline data points, adaptive Gaussian filtering can be performed on the baseline data to improve the filtering accuracy of the baseline data; in the Gaussian filtering algorithm, when filtering a certain baseline data point, it is necessary to calculate the Gaussian kernel weight between other baseline data points and the baseline data point. If the noise interference degree of other baseline data points is large, it means that the Gaussian kernel weight between the other data point and the baseline data point needs to be reduced to reduce the error of the filtering result of the baseline data point. Therefore, the noise interference similarity degree of any baseline data point is obtained according to the noise interference degrees of the any baseline data point and other baseline data points; preferably, in the embodiments of the present invention, the step of obtaining the noise interference similarity degree includes: calculating the ratio of the noise interference degrees of the any baseline data point and other baseline data points to obtain the noise interference similarity degree of the any baseline data point; when the noise interference similarity degree is larger, it means that the noise interference degree of the other baseline data point is less than that of the any baseline data point, and the data of the other baseline data point is more reliable, so the Gaussian kernel weight between the two can be larger; when the noise interference similarity degree is smaller, it means that the noise interference degree of the other baseline data point is greater than that of the any baseline data point, and the data of the other baseline data point is less reliable, and the Gaussian kernel weight between the two should be smaller.
[0034] Further, after obtaining the similarity degree of noise interference, the improved Gaussian kernel weight between any baseline data point and other baseline data points can be obtained according to the Gaussian filtering algorithm and the similarity degree of noise interference; preferably, in the embodiment of the present invention, the steps of obtaining the improved Gaussian kernel weight include: obtaining the initial Gaussian kernel weight between the any baseline data point and other baseline data points according to the Gaussian filtering algorithm; it should be noted that the calculation of the initial Gaussian kernel weight belongs to the prior art, and the specific calculation steps will not be elaborated. Calculate the product of the similarity degree of noise interference and the initial Gaussian kernel weight to obtain the improved Gaussian kernel weight between the any baseline data point and other baseline data points; when the similarity degree of noise interference is larger, the corresponding improved Gaussian kernel weight is larger; when the similarity degree of noise interference is smaller, the corresponding improved Gaussian kernel weight is smaller.
[0035] Compared with the initial Gaussian kernel weight, the improved Gaussian kernel weight combines the noise interference degree of the baseline data point, thereby improving the accuracy of Gaussian filtering; therefore, filter any baseline data point according to the improved Gaussian kernel weight, and recombine the filtered data with the unfiltered data at the target peak point to obtain the improved gas chromatography data; it should be noted that Gaussian filtering belongs to the prior art, and the specific steps will not be elaborated. The improved gas chromatography data is more accurate than the gas chromatography data, thereby improving the accuracy of the component analysis of the population aerosol.
[0036] In summary, the embodiment of the present invention provides an efficient acquisition system for population aerosol data based on big data; obtain the target peak point according to the data distribution characteristics of the gas chromatography data; obtain the segmentation point according to the data change characteristics within the preset neighborhood range of the target peak point; obtain the baseline sequence according to the segmentation point; obtain the reference value according to the baseline sequence; obtain the reference sequence according to the data difference characteristics between the local sequence, the local sequence and the reference value; obtain the noise interference degree according to the data difference characteristics and the difference characteristics of the discrete degree between the local sequence and the reference sequence; obtain the similarity degree of noise interference according to the noise interference degree of different baseline data points; obtain the improved Gaussian kernel weight according to the Gaussian filtering and the similarity degree of noise interference. The present invention filters any baseline data point according to the improved Gaussian kernel weight to obtain the improved gas chromatography data, thereby improving the accuracy of the aerosol component analysis.
[0037] It should be noted that: the above sequence of the embodiments of the present invention is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0038] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. An efficient acquisition system for crowd aerosol data based on big data, characterized in that, The system includes the following modules: A data acquisition module, configured to acquire gas chromatography data of crowd aerosol; A data analysis module, configured to obtain a target peak point according to the data characteristics of peak points in the gas chromatography data and the difference characteristics of data distribution within a preset neighborhood range of the peak points; obtain a segmentation point of the gas chromatography data according to the data change characteristics within the preset neighborhood range of the target peak point; obtain a baseline sequence of the gas chromatography data according to the segmentation point; and obtain a reference value according to the data distribution characteristics of the baseline sequence; A feature processing module, configured to construct a local sequence according to baseline data points and preset connected data points in the baseline sequence; obtain a baseline credibility according to the data change characteristics of the local sequence and the data difference characteristics between the local sequence and the reference value; obtain a reference sequence according to the baseline credibility; and obtain a noise interference degree according to the data difference characteristics and the difference characteristics of the dispersion degree between the local sequence of any baseline data point and the reference sequence; A data filtering module, configured to obtain a noise interference similarity degree of any baseline data point according to the noise interference degrees of any baseline data point and other baseline data points; Obtain an improved Gaussian kernel weight between any baseline data point and other baseline data points according to the Gaussian filtering algorithm and the noise interference similarity degree; Filter any baseline data point according to the improved Gaussian kernel weight to obtain improved gas chromatography data.
2. The efficient acquisition system for crowd aerosol data based on big data according to claim 1, characterized in that The step of obtaining a target peak point according to the data characteristics of peak points in the gas chromatography data and the difference characteristics of data distribution within a preset neighborhood range of the peak points includes: Taking the peak point as the center within the preset neighborhood range, calculating the average value of the absolute values of the differences in the amplitudes of data points with the same rank before and after and performing a negative correlation mapping to obtain a symmetry eigenvalue; calculating the product of the symmetry eigenvalue and the amplitude of the peak point and normalizing it to obtain the peak confidence of the peak point; and taking the peak point with the peak confidence exceeding a preset confidence threshold as the target peak point.
3. The efficient acquisition system for crowd aerosol data based on big data according to claim 1, characterized in that, The step of obtaining a segmentation point of the gas chromatography data according to the data change characteristics within the preset neighborhood range of the target peak point includes: Within the preset neighborhood range of the target peak point, linearly fit the data points on both sides of the target peak point respectively, and take the data points where the two obtained fitting lines intersect with the gas chromatography data as the segmentation point.
4. An efficient acquisition system for crowd aerosol data based on big data according to claim 1, characterized in that, The step of obtaining a baseline sequence of the gas chromatography data according to the segmentation point includes: Segment the gas chromatography data at the segmentation point to obtain different data segments; sort the values in the data segments that do not include the target peak point in chronological order to obtain the baseline sequence.
5. An efficient acquisition system for crowd aerosol data based on big data according to claim 1, characterized in that, The step of obtaining a reference value according to the data distribution characteristics of the baseline sequence includes: Taking the median of the baseline sequence as the reference value.
6. The high-efficiency acquisition system for crowd aerosol data based on big data according to claim 1, characterized in that The step of obtaining a baseline credibility according to the data change characteristics of the local sequence and the data difference characteristics between the local sequence and the reference value includes: Calculate the absolute value of the difference between the average value of the local sequence and the reference value to obtain the deviation eigenvalue; calculate the sum of the coefficient of variation of the local sequence and the deviation eigenvalue and perform a negative correlation mapping to obtain the baseline credibility of the local sequence.
7. An efficient acquisition system for crowd aerosol data based on big data according to claim 1, characterized in that, The step of obtaining the reference sequence according to the baseline credibility includes: Taking the local sequence corresponding to the maximum value of the baseline credibility as the reference sequence.
8. An efficient acquisition system for crowd aerosol data based on big data according to claim 1, characterized in that, The step of obtaining the noise interference degree according to the data difference feature and the difference feature of the dispersion degree between the local sequence of any baseline data point and the reference sequence includes: Calculate the absolute value of the difference between the average value of the local sequence of any baseline data point and the average value of the reference sequence to obtain the baseline difference eigenvalue; calculate the difference between the coefficient of variation of the local sequence of any baseline data point and the reference sequence and perform a positive correlation mapping to obtain the dispersion degree difference value; calculate the product of the dispersion degree difference value and the baseline difference eigenvalue to obtain the noise interference degree of any baseline data point.
9. An efficient acquisition system for crowd aerosol data based on big data according to claim 1, characterized in that, The step of obtaining the noise interference similarity degree of any baseline data point according to the noise interference degrees of any baseline data point and other baseline data points includes: Calculate the ratio of the noise interference degree of any baseline data point to the noise interference degree of the other baseline data points to obtain the noise interference similarity degree of any baseline data point.
10. An efficient acquisition system for crowd aerosol data based on big data according to claim 1, characterized in that, The step of obtaining the improved Gaussian kernel weight between any baseline data point and other baseline data points according to the Gaussian filtering algorithm and the noise interference similarity degree includes: Obtain the initial Gaussian kernel weights of any baseline data point and the other baseline data points according to the Gaussian filtering algorithm; calculate the product of the noise interference similarity degree and the initial Gaussian kernel weights to obtain the improved Gaussian kernel weight between any baseline data point and the other baseline data points.
Citation Information
Patent Citations
Automatic analysis method and system for expiration molecular analysis gas chromatography data
CN116242954A
Automatic monitoring method and system for biological aerosol
CN116698680A
Intelligent pollutant toxicity detection system
CN117007577A
Gas chromatograph data optimization storage method and system
CN117785818A
Method, device and system for monitoring breathing signal of patient
CN119112155A