Efficient collection system of crowd aerosol data based on big data

By identifying target peak points and baseline sequences in gas chromatography data and utilizing improved Gaussian kernel weighted filtering, the problem of noise interference in gas chromatography data was solved, thereby improving the accuracy of aerosol component analysis.

CN120408084BActive Publication Date: 2025-11-07BEIJING DITAN HOSPITAL CAPITAL MEDICAL UNIVERSTY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510505132.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-11-07
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

In existing technologies, Gaussian filtering has low accuracy in filtering gas chromatography data, affecting the accuracy of gas chromatography analysis of aerosol components. In particular, the degree of noise interference is inconsistent in human aerosols, making it difficult to effectively filter out noise interference.

Method used

The data acquisition module acquires gas chromatographic data, the analysis module determines the target peak points and cutoff points, the feature processing module constructs a baseline sequence, and the data filtering module uses an improved Gaussian kernel weight to perform filtering, thereby improving the accuracy of baseline filtering.

Benefits of technology

It improves the accuracy of gas chromatography data filtering, enhances the accuracy of aerosol component analysis, and reduces filtering errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408084B_ABST
    Figure CN120408084B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of gas chromatography analysis, and particularly relates to a high-efficiency collection system for crowd aerosol data based on big data; a target peak point is obtained according to the data distribution characteristics of the gas chromatography data; a segmentation point is obtained according to the data within the preset neighborhood range of the target peak point; a baseline sequence is obtained according to the segmentation point; a reference value is obtained according to the baseline sequence; a reference sequence is obtained according to the data difference characteristics of the local sequence and the reference value; a noise interference degree is obtained according to the data difference characteristics and the difference characteristics of the discrete degree of the local sequence and the reference sequence; a noise interference similarity degree is obtained according to the noise interference degrees of different baseline data points; and an improved Gaussian kernel weight is obtained according to the Gaussian filtering and the noise interference similarity degree. According to the improved Gaussian kernel weight, any baseline data point is filtered to obtain improved gas chromatography data, and the accuracy of aerosol component analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of gas chromatography analysis, and in particular to an efficient collection system for crowd aerosol data based on big data. BACKGROUND

[0002] Crowd aerosol mainly refers to a complex gaseous dispersion system generated by the respiratory tract of the human body; some biological characteristic molecules generated by human metabolism, such as disease marker molecules, are discharged out of the body with the human respiratory system, and the diameter of these microparticles is usually in the micron level and can float in the air for a long time. Measuring the marker molecules in the crowd aerosol is a fast, accurate and non-contact disease diagnosis method.

[0003] In the prior art, gas chromatography analysis is usually used to analyze and measure the composition of collected crowd aerosol; however, in the process of obtaining the gas chromatography data of the crowd aerosol, due to noise interference and the like, there may be fluctuations such as noise in the baseline, thereby affecting the accurate acquisition of information such as peak shape area; and different aerosols may have different influencing factors and components, and the degree of noise interference on the baseline data may also be different, so that the existing Gaussian filtering is difficult to accurately filter the gas chromatography data, thereby ultimately affecting the accuracy of the gas chromatography analysis of the aerosol composition. SUMMARY

[0004] In order to solve the technical problem that the accuracy of Gaussian filtering of gas chromatography data is low and thus affects the accuracy of analyzing the aerosol composition, the purpose of the present application is to provide an efficient collection system for crowd aerosol data based on big data, and the technical solution adopted is as follows:

[0005] The data acquisition module is configured to acquire gas chromatography data of the crowd aerosol.

[0006] The data analysis module is configured to obtain a target peak point according to the data characteristics of the peak point in the gas chromatography data and the difference characteristics of the data distribution in the preset neighborhood range of the peak point; obtain a segmentation point of the gas chromatography data according to the data change characteristics in the preset neighborhood range of the target peak point; obtain a baseline sequence of the gas chromatography data according to the segmentation point; and obtain a reference value according to the data distribution characteristics of the baseline sequence.

[0007] The feature processing module is configured to construct a local sequence according to the baseline data points and the preset connected data points in the baseline sequence; obtain a baseline credibility according to the data change characteristics of the local sequence and the data difference characteristics between the local sequence and the reference value; obtain a reference sequence according to the baseline credibility; and obtain a noise interference degree according to the data difference characteristics and the difference characteristics of the degree of dispersion between the local sequence of any baseline data point and the reference sequence.

[0008] The data filtering module is configured to obtain a noise interference similarity degree of an arbitrary baseline data point according to a noise interference degree of the arbitrary baseline data point and other baseline data points; obtain an improved Gaussian kernel weight between the arbitrary baseline data point and other baseline data points according to a Gaussian filtering algorithm and the noise interference similarity degree; and filter the arbitrary baseline data point according to the improved Gaussian kernel weight to obtain improved gas chromatography data.

[0009] Further, the step of obtaining a target peak point according to a data feature of a peak point in the gas chromatography data and a difference feature of data distribution in a preset neighborhood range of the peak point comprises:

[0010] The average value of the absolute value of the difference between the amplitudes of the data points with the same bit order before and after the peak point is calculated in the preset neighborhood range centered on the peak point, and a negative correlation mapping is performed to obtain a symmetry feature value; the product of the symmetry feature value and the amplitude of the peak point is calculated and normalized to obtain a peak confidence of the peak point; and the peak point with a peak confidence exceeding a preset confidence threshold is taken as the target peak point.

[0011] Further, the step of obtaining a segmentation point of the gas chromatography data according to a data change feature in a preset neighborhood range of the target peak point comprises:

[0012] The data points on both sides of the target peak point in the preset neighborhood range of the target peak point are linearly fitted respectively, and the data points at which the obtained two fitted straight lines intersect with the gas chromatography data are taken as the segmentation point.

[0013] Further, the step of obtaining a baseline sequence of the gas chromatography data according to the segmentation point comprises:

[0014] The gas chromatography data is segmented at the segmentation point to obtain different data segments; and the values in the data segment that does not contain the target peak point are sorted in time sequence to obtain the baseline sequence.

[0015] Further, the step of obtaining a reference value according to a data distribution feature of the baseline sequence comprises:

[0016] The median of the baseline sequence is taken as the reference value.

[0017] Further, the step of obtaining a baseline confidence according to a data change feature of the local sequence and a data difference feature between the local sequence and the reference value comprises:

[0018] Calculate the absolute value of the difference between the average value of the local sequence and the reference value to obtain a deviation characteristic value; calculate the sum of the variation coefficient of the local sequence and the deviation characteristic value and perform negative correlation mapping to obtain the baseline reliability of the local sequence.

[0019] Further, the step of obtaining a reference sequence according to the baseline reliability comprises:

[0020] The local sequence corresponding to the maximum value of the baseline reliability is taken as the reference sequence.

[0021] Further, the step of obtaining the noise interference degree according to the data difference characteristic of the local sequence of any baseline data point and the reference sequence and the difference characteristic of the discrete degree comprises:

[0022] Calculate the absolute value of the difference between the average value of the local sequence of the any baseline data point and the average value of the reference sequence to obtain a baseline difference characteristic value; calculate the difference between the variation coefficient of the local sequence of the any baseline data point and the reference sequence and perform positive correlation mapping to obtain a discrete degree difference value; calculate the product of the discrete degree difference value and the baseline difference characteristic value to obtain the noise interference degree of the any baseline data point.

[0023] Further, the step of obtaining the noise interference similarity degree of the any baseline data point according to the noise interference degrees of the any baseline data point and other baseline data points comprises:

[0024] Calculate the ratio of the noise interference degrees of the any baseline data point and the other baseline data points to obtain the noise interference similarity degree of the any baseline data point.

[0025] Further, the step of obtaining the improved Gaussian kernel weight between the any baseline data point and other baseline data points according to the Gaussian filtering algorithm and the noise interference similarity degree comprises:

[0026] Obtain the initial Gaussian kernel weight of the any baseline data point and the other baseline data points according to the Gaussian filtering algorithm; calculate the product of the noise interference similarity degree and the initial Gaussian kernel weight to obtain the improved Gaussian kernel weight between the any baseline data point and the other baseline data points.

[0027] The present application has the following beneficial effects:

[0028] In the present application, obtaining the target peak point can determine the region representing the detected substance in the gas chromatography data, and then determine the baseline region that needs to be filtered according to the target peak point; obtaining the segmentation point can more accurately determine the beginning and end positions of the baseline. Obtaining the reference value can determine the baseline value representing the normal level according to the distribution characteristics of the baseline data, which is beneficial to determining the noise interference degree of different baseline data in the subsequent step. Calculating the baseline credibility can obtain the reference sequence according to the severity of the noise interference of the baseline, and through the reference sequence, the noise interference degree of other local sequences can be evaluated; obtaining the noise interference degree can determine the severity of the noise interference of the baseline data points at different positions, thereby improving the filtering accuracy. Obtaining the noise interference similarity degree can determine the degree of noise interference between different baseline data points, so as to correct the Gaussian kernel weight between two baseline data points and reduce the filtering error. The improved gas chromatography data finally obtained improves the accuracy of aerosol component analysis compared with the gas chromatography data. BRIEF DESCRIPTION OF DRAWINGS

[0029] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0030] Figure 1 A block diagram of an efficient collection system of crowd aerosol data based on big data provided by an embodiment of the present application. DETAILED DESCRIPTION

[0031] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined object, the specific embodiments, structure, features and effects of the efficient collection system of crowd aerosol data based on big data according to the present application are described in detail as follows. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.

[0033] The specific scheme of the efficient collection system of crowd aerosol data based on big data provided by the present application is described in detail below with reference to the drawings.

[0034] Please refer to Figure 1It shows a high-efficiency collection system block diagram of crowd aerosol data based on big data provided by one embodiment of the present application, and the system comprises the following modules:

[0035] The data acquisition module S1 is configured to acquire the gas chromatography data of the crowd aerosol.

[0036] In the embodiment of the present application, the implementation scenario is to filter the gas chromatography data of the crowd aerosol, and improve the accuracy of the gas chromatography analysis of the aerosol composition. First, the gas chromatography data of the crowd aerosol is acquired. It should be noted that the acquisition of the gas chromatography data of the aerosol sample by the gas chromatograph is a prior art, and the specific steps are not described here.

[0037] The data analysis module S2 is configured to obtain a target peak point according to the data characteristics of the peak point in the gas chromatography data and the difference characteristics of the data distribution in the preset neighborhood range of the peak point; obtain a segmentation point of the gas chromatography data according to the data change characteristics in the preset neighborhood range of the target peak point; obtain a baseline sequence of the gas chromatography data according to the segmentation point; and obtain a reference value according to the data distribution characteristics of the baseline sequence.

[0038] When analyzing the gas chromatography data of the crowd aerosol, due to the influence of impurities in the aerosol sample and background noise of the instrument, etc., the baseline in the gas chromatography data is prone to fluctuation, which affects the extraction of the wave peak characteristics in the gas chromatography data, thereby causing inaccurate composition analysis. Since the composition and content of the aerosol sample have certain differences, the height and position of the wave peak in the generated gas chromatography data are not the same. In order to correct the baseline without affecting the height and position of the wave peak, the accuracy of the baseline filtering needs to be improved. First, the wave peak region in the gas chromatography data that may be composed of the detected substance needs to be determined, and then the data segment that does not belong to the wave peak region may be the baseline data disturbed by noise. Since the wave peak representing the detected substance has certain symmetry characteristics, and the amplitude characteristics of such wave peak are relatively obvious, the target peak point can be obtained according to the data characteristics of the peak point in the gas chromatography data and the difference characteristics of the data distribution in the preset neighborhood range of the peak point.

[0039] Preferably, in the embodiments of the present application, the step of obtaining the target peak point comprises: taking the peak point as the center, calculating the average value of the absolute value of the difference between the amplitudes of the data points with the same bit order before and after the peak point and negatively correlating mapping to obtain a symmetry characteristic value; in the embodiments of the present application, the preset neighborhood range is a range composed of 15 data points before and after the peak point as the center, and the implementer can determine it by himself according to the implementation scene; for example, the first adjacent data points before and after the peak point have the same bit order of 1, if the peak is relatively symmetrical, the amplitudes of the data points with the same bit order before and after the peak point are relatively close, the smaller the absolute value of the difference is, the larger the symmetry characteristic value is, which means that the peak is more likely to represent the detected substance. The product of the symmetry characteristic value and the amplitude of the peak point is calculated and normalized to obtain the peak confidence of the peak point; the larger the amplitude of the peak point is, the more likely the peak is to represent the detected substance; therefore, the larger the peak confidence is, the more the peak can reflect the component characteristics of the aerosol. The peak point with a peak confidence exceeding a preset confidence threshold is taken as the target peak point, and the target peak point represents the position in the gas chromatogram data that can reflect the detected substance; in the embodiments of the present application, the preset confidence threshold is 0.6, and the implementer can determine it by himself according to the implementation scene. The formula for obtaining the peak confidence comprises:

[0040]

[0041] wherein P represents the peak confidence of the peak point, represents a normalization function, F represents the amplitude of the peak point, represents an exponential function with a natural constant as the base, N represents the number of data points on one side in the preset neighborhood range of the peak point, represents the amplitude of the nth bit order data point before the peak point, represents the amplitude of the nth bit order data point after the peak point, represents the symmetry characteristic value.

[0042] Further, after all target peak points in the gas chromatogram data are obtained, the data between two target peak points that is relatively smooth can be baseline data, so the dividing points of the gas chromatogram data are obtained according to the data change characteristics in the preset neighborhood range of the target peak points; preferably, in the embodiment of the present application, the step of obtaining the dividing points comprises: linear fitting the data points on both sides of the target peak points in the preset neighborhood range of the target peak points respectively, and taking the data points obtained by intersecting the two fitted straight lines with the gas chromatogram data as the dividing points; it should be noted that linear fitting belongs to the prior art, and the specific steps will not be described here, the fitted straight line represents the change trend on both sides of the peak, and the dividing point represents the approximate position where the baseline is distinguished from the peak. After all the dividing points are obtained, all the baseline data in the gas chromatogram can be obtained, so the baseline sequence of the gas chromatogram data is obtained according to the dividing points; preferably, in the embodiment of the present application, the step of obtaining the baseline sequence comprises: dividing the gas chromatogram data at the dividing points to obtain different data segments; sorting the values in the data segments that do not contain target peak points in time sequence to obtain the baseline sequence; the baseline sequence represents the data characteristics of the baseline in the gas chromatogram data.

[0043] In the gas chromatogram data, the baseline fluctuates due to the interference of sample impurities and instrument background noise, etc., and the fluctuation affects the position of the average value of the baseline data, making it difficult to determine the normal level of the baseline; the median can represent the baseline level that avoids noise interference to a certain extent, so the reference value is obtained according to the data distribution characteristics of the baseline sequence; preferably, in the embodiment of the present application, the step of obtaining the reference value comprises: taking the median of the baseline sequence as the reference value; the reference value represents the normal baseline characteristics in the gas chromatogram data.

[0044] The feature processing module S3 is configured to construct a local sequence according to the baseline data points in the baseline sequence and preset connected data points; obtain baseline reliability according to the data change characteristics of the local sequence and the data difference characteristics of the local sequence and the reference value; obtain a reference sequence according to the baseline reliability; and obtain the noise interference degree according to the data difference characteristics and the difference in dispersion degree of the local sequence and the reference sequence of any baseline data point.

[0045] In order to further improve the accuracy of baseline filtering, it is necessary to determine the noise interference degree of different positions in the baseline. First, a local sequence is constructed according to the baseline data points in the baseline sequence and the preset connected data points, in the embodiment of the present application, the preset connected data points are five other baseline data points before and after the baseline data point, and the implementer can determine it according to the implementation scene. The greater the degree of noise interference of the baseline, the more obvious the fluctuation of the baseline, and the greater the difference from the reference value, so the baseline reliability is obtained according to the data change characteristics of the local sequence and the data difference characteristics of the local sequence and the reference value.

[0046] Preferably, in the embodiments of the present application, the step of obtaining baseline credibility comprises: calculating the absolute value of the difference between the average value of the local sequence and the reference value to obtain a deviation characteristic value; the greater the difference between the average value of the local sequence and the reference value, the greater the deviation characteristic value, which means that the baseline at the local sequence deviates from the median more and is more seriously interfered by noise. The sum of the coefficient of variation of the local sequence and the deviation characteristic value is calculated and negatively correlated to obtain the baseline credibility of the local sequence; it should be noted that the calculation of the coefficient of variation is a prior art and the specific calculation steps will not be repeated; the more obvious the fluctuation of the local sequence, the greater the coefficient of variation, the more likely it is to be interfered by noise; therefore, the smaller the baseline credibility, the greater the degree of interference of the local sequence by noise, and the more it deviates from the normal baseline characteristics.

[0047] Further, the greater the baseline credibility, the greater the degree of interference of the local sequence by noise, the more it can represent the normal baseline characteristics, and then the reference sequence is obtained according to the baseline credibility; preferably, in the embodiments of the present application, the step of obtaining the reference sequence comprises: taking the local sequence corresponding to the maximum value of the baseline credibility as the reference sequence, which represents the normal baseline characteristics in the gas chromatography data. After obtaining the reference sequence, the degree of noise interference of the local sequence can be determined according to the difference characteristics of the local sequence and the reference sequence, thereby improving the accuracy of filtering; therefore, the degree of noise interference is obtained according to the data difference characteristics and the difference characteristics of the discrete degree of the local sequence and the reference sequence of any baseline data point.

[0048] Preferably, in the embodiments of the present application, the step of obtaining the degree of noise interference comprises: calculating the absolute value of the difference between the average value of the local sequence of any baseline data point and the average value of the reference sequence to obtain a baseline difference characteristic value; the greater the baseline difference characteristic value, the greater the characteristic difference between the local sequence of any baseline data point and the baseline sequence, and the greater the degree of interference by noise. The difference between the coefficient of variation of the local sequence of any baseline data point and the reference sequence is calculated and positively correlated to obtain a discrete degree difference value; the greater the discrete degree difference value, the greater the volatility of the local sequence of any baseline data point than the baseline sequence, and the greater the degree of interference by noise. The product of the discrete degree difference value and the baseline difference characteristic value is calculated to obtain the degree of noise interference of any baseline data point; the greater the degree of noise interference, the more serious the noise interference at the any baseline data point.

[0049] The data filtering module S4 is configured to obtain the similarity degree of noise interference of any baseline data point according to the degree of noise interference of any baseline data point and other baseline data points; obtain the improved Gaussian kernel weight between any baseline data point and other baseline data points according to the Gaussian filtering algorithm and the similarity degree of noise interference; and filter any baseline data point according to the improved Gaussian kernel weight to obtain improved gas chromatography data.

[0050] After obtaining the noise interference degree of different baseline data points, the baseline data can be adaptively Gaussian filtered to improve the filtering accuracy of the baseline data. In the Gaussian filtering algorithm, when filtering a certain baseline data point, the Gaussian kernel weight of other baseline data points and the baseline data point needs to be calculated. If the noise interference degree of the other baseline data points is large, it means that the Gaussian kernel weight of the other data points and the baseline data point needs to be reduced to reduce the error of the filtering result of the baseline data point. Therefore, the noise interference similarity degree of any baseline data point is obtained according to the noise interference degree of the baseline data point and other baseline data points. Preferably, in the embodiment of the present application, the step of obtaining the noise interference similarity degree includes: calculating the ratio of the noise interference degree of any baseline data point and other baseline data points to obtain the noise interference similarity degree of the baseline data point. When the noise interference similarity degree is larger, it means that the noise interference degree of the other baseline data point is smaller than that of the baseline data point, and the data of the other baseline data point is more reliable, so the Gaussian kernel weight of the two can be larger. When the noise interference similarity degree is smaller, it means that the noise interference degree of the other baseline data point is larger than that of the baseline data point, and the data of the other baseline data point is less reliable, so the Gaussian kernel weight of the two should be smaller.

[0051] Further, after obtaining the noise interference similarity degree, the improved Gaussian kernel weight between any baseline data point and other baseline data points can be obtained according to the Gaussian filtering algorithm and the noise interference similarity degree. Preferably, in the embodiment of the present application, the step of obtaining the improved Gaussian kernel weight includes: obtaining the initial Gaussian kernel weight of the baseline data point and other baseline data points according to the Gaussian filtering algorithm. It should be noted that the calculation of the initial Gaussian kernel weight belongs to the prior art, and the specific calculation steps are not described again. The product of the noise interference similarity degree and the initial Gaussian kernel weight is calculated to obtain the improved Gaussian kernel weight between the baseline data point and other baseline data points. When the noise interference similarity degree is larger, the corresponding improved Gaussian kernel weight is larger. When the noise interference similarity degree is smaller, the corresponding improved Gaussian kernel weight is smaller.

[0052] The improved Gaussian kernel weight combines the noise interference degree of the baseline data point compared with the initial Gaussian kernel weight, thereby improving the accuracy of Gaussian filtering. Therefore, any baseline data point is filtered according to the improved Gaussian kernel weight, the filtered data and the unfiltered data at the target peak point are recombined to obtain improved gas chromatography data. It should be noted that Gaussian filtering belongs to the prior art, and the specific steps are not described again. The accuracy of the improved gas chromatography data is higher than that of the gas chromatography data, thereby improving the accuracy of the composition analysis of the population aerosol.

[0053] In summary, the embodiment of the present application provides a kind of high-efficiency collection system of crowd aerosol data based on big data;According to the data distribution characteristics of gas chromatography data, obtain target peak point;According to the data variation characteristics in the preset neighborhood range of target peak point, obtain segmentation point;According to the baseline sequence obtained by segmentation point;According to the baseline sequence, obtain reference value;According to the data difference characteristics of local sequence, local sequence and reference value, obtain reference sequence;According to the data difference characteristics of local sequence and reference sequence, the difference characteristics of discrete degree, obtain noise interference degree;According to the noise interference degree of different baseline data points, obtain noise interference similarity degree;According to the improved Gaussian kernel weight obtained by Gaussian filtering and noise interference similarity degree.This application carries out filtering to any baseline data point according to improved Gaussian kernel weight, obtains improved gas chromatography data, improves the accuracy of aerosol component analysis.

[0054] It should be noted that the above-mentioned embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0055] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other, and each embodiment mainly describes the difference from other embodiments.

Claims

1. An efficient collection system of crowd aerosol data based on big data, characterized in that, The system comprises the following modules: a data acquisition module, configured to acquire gas chromatogram data of aerosol of a crowd; a data analysis module, configured to acquire a target peak point according to a data feature of a peak point in the gas chromatogram data and a difference feature of data distribution in a preset neighborhood range of the peak point; acquire a segmentation point of the gas chromatogram data according to a data change feature in a preset neighborhood range of the target peak point; acquire a baseline sequence of the gas chromatogram data according to the segmentation point; and acquire a reference value according to a data distribution feature of the baseline sequence; a feature processing module, configured to construct a local sequence according to a baseline data point in the baseline sequence and a preset connected data point; acquire a baseline credibility according to a data change feature of the local sequence and a data difference feature of the local sequence and the reference value; acquire a reference sequence according to the baseline credibility; and acquire a noise interference degree according to a data difference feature of a local sequence of an arbitrary baseline data point and the reference sequence and a difference feature of a discrete degree; a data filtering module, configured to acquire a noise interference similarity degree of an arbitrary baseline data point according to a noise interference degree of the arbitrary baseline data point and other baseline data points; acquire an improved Gaussian kernel weight between the arbitrary baseline data point and other baseline data points according to a Gaussian filtering algorithm and the noise interference similarity degree; and filter the arbitrary baseline data point according to the improved Gaussian kernel weight to obtain improved gas chromatogram data; the step of acquiring a target peak point according to a data feature of a peak point in the gas chromatogram data and a difference feature of data distribution in a preset neighborhood range of the peak point comprises: calculating an average value of absolute values of differences of amplitude values of data points with same bit positions before and after the peak point in the preset neighborhood range and performing negative correlation mapping to obtain a symmetric characteristic value; calculating a product of the symmetric characteristic value and the amplitude value of the peak point and normalizing to obtain a peak confidence of the peak point; and taking a peak point with a peak confidence exceeding a preset confidence threshold as the target peak point; the step of acquiring a segmentation point of the gas chromatogram data according to a data change feature in a preset neighborhood range of the target peak point comprises: linearly fitting data points on both sides of the target peak point in the preset neighborhood range of the target peak point, and taking data points intersected by two obtained fitting straight lines and the gas chromatogram data as the segmentation point; the step of acquiring a baseline sequence of the gas chromatogram data according to the segmentation point comprises: segmenting the gas chromatogram data at the segmentation point to obtain different data segments; and sorting values in a data segment not containing the target peak point in time sequence to obtain the baseline sequence; the step of acquiring a reference value according to a data distribution feature of the baseline sequence comprises: taking a median of the baseline sequence as the reference value; the step of acquiring a noise interference degree according to a data difference feature of a local sequence of an arbitrary baseline data point and the reference sequence and a difference feature of a discrete degree comprises: ​ ​ The absolute value of the difference between the average value of the local sequence of the arbitrary baseline data point and the average value of the reference sequence is calculated to obtain a baseline difference characteristic value; the difference between the coefficient of variation of the local sequence of the arbitrary baseline data point and the reference sequence is calculated and positively correlated to obtain a dispersion difference value; and the product of the dispersion difference value and the baseline difference characteristic value is calculated to obtain the noise interference degree of the arbitrary baseline data point; The step of obtaining the improved Gaussian kernel weight between the arbitrary baseline data point and other baseline data points according to the Gaussian filtering algorithm and the noise interference similarity degree comprises: The initial Gaussian kernel weight of the arbitrary baseline data point and the other baseline data points is obtained according to the Gaussian filtering algorithm; and the product of the noise interference similarity degree and the initial Gaussian kernel weight is calculated to obtain the improved Gaussian kernel weight between the arbitrary baseline data point and the other baseline data points. 2.The efficient collection system of crowd aerosol data based on big data according to claim 1, wherein, The step of obtaining the baseline reliability according to the data change characteristic of the local sequence and the data difference characteristic of the local sequence and the reference value comprises: The absolute value of the difference between the average value of the local sequence and the reference value is calculated to obtain a deviation characteristic value; and the sum of the coefficient of variation of the local sequence and the deviation characteristic value is calculated and negatively correlated to obtain the baseline reliability of the local sequence. 3.The efficient collection system of crowd aerosol data based on big data according to claim 1, wherein, The step of obtaining the reference sequence according to the baseline reliability comprises: The local sequence corresponding to the maximum value of the baseline reliability is taken as the reference sequence. 4.The efficient collection system of crowd aerosol data based on big data according to claim 1, wherein, The step of obtaining the noise interference similarity degree of the arbitrary baseline data point according to the noise interference degrees of the arbitrary baseline data point and other baseline data points comprises: The ratio of the noise interference degrees of the arbitrary baseline data point and the other baseline data points is calculated to obtain the noise interference similarity degree of the arbitrary baseline data point.

Citation Information

Patent Citations

  • Automatic analysis method and system for expiration molecular analysis gas chromatography data

    CN116242954A

  • Intelligent pollutant toxicity detection system

    CN117007577A