Pollution source analysis chemical indicator fingerprint spectrum extraction method based on non-targeted screening
Through the fingerprint spectrum extraction method of pollution source analytical chemical indicators based on non-targeted screening, and using mass spectrometry analysis and target analysis software, random forest and linear discriminant analysis models are trained, the problem of low efficiency in point source pollution identification and traceability of complex water areas is solved, and efficient traceability and classification of source pollution compounds is achieved.
Patent Information
- Application Number
- CN202510007285.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology has low efficiency in identifying and traceability of point source pollution in complex waters, making it difficult to fully reflect the real situation of pollutants in industrial wastewater.
The fingerprint spectrum extraction method of pollution source analytical chemical indicators based on non-targeted screening is adopted. Through mass spectrometry analysis and target analysis software, the fingerprint spectrum of the target pollution category is obtained, the target characteristic compounds are determined, and the random forest classification model and linear discriminant analysis classification model are trained to achieve traceability of source pollution compounds in complex pollution fields.
It improves the traceability efficiency of source polluted compounds in complex pollution fields, can adapt to various environments and pollution situations, and achieves the classification and acquisition of point-source polluted compounds of different types of pollution sources from multiple different pollution categories, which is highly adaptable.
Smart Images

Figure CN119985818A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of water environment industrial point source pollution analysis, and specifically relates to a method for extracting chemical indicator fingerprint spectra for pollution source analysis based on non-targeted screening. Background Art
[0002] With the continuous acceleration of industrialization, the discharge of industrial wastewater is increasing day by day, and the complex pollutants contained in it pose a serious threat to the aquatic environment ecosystem. Traditional environmental monitoring methods usually rely on the pre-identification of target pollutants, which is difficult to fully reflect the true situation of pollutants in industrial wastewater, and limits the ability to identify new pollutants and unknown risks.
[0003] In recent years, non-targeted screening technology has been increasingly widely used in the field of environmental analysis due to its advantages of high throughput, high sensitivity, and no need to predetermine target pollutants. This technology, combined with high-resolution mass spectrometry, liquid chromatography and other analytical methods, can simultaneously detect and identify a large number of compounds in environmental samples, and then identify potential characteristic pollutants. Compared with traditional targeted analysis methods, non-targeted screening technology can fully understand the composition of pollutants in industrial wastewater, help discover new environmental risk substances, and provide a more scientific basis for the prevention and control of water environmental pollution. At present, non-targeted screening technology has been used to identify characteristic pollutants in industrial wastewater in the pharmaceutical, printing and dyeing, chemical and other industries, and combined with chemometric methods such as principal component analysis and cluster analysis, the source of water environmental pollution has been preliminarily analyzed, and the traceability and identification of industrial point source pollution has been initially realized. However, the identification and traceability of point source pollution in complex waters is still an issue that needs to be considered. Summary of the invention
[0004] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening, so as to solve the problem of low efficiency in identifying and tracing point source pollution in complex waters in the prior art.
[0005] According to one aspect of the present application, a method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening is provided, the method comprising: Acquire a plurality of target water samples corresponding to a plurality of target pollution categories, wherein the plurality of target pollution categories include at least two of a biotechnology plant, a petrochemical plant, a sewage treatment plant, or a textile printing and dyeing plant; Based on the mass spectrometer, non-targeted mass spectrometry analysis is performed on the target water sample corresponding to each target pollution category to obtain multiple fingerprint spectra of each target pollution category for characterizing the chemical indicators in the corresponding target water sample, and each target pollution category corresponds to one fingerprint spectrum; Analyzing and processing each of the fingerprint spectra based on target analysis software to obtain a sample compound set contained in each target pollution category, wherein the sample compound set includes a plurality of sample characteristic compounds; Determining a plurality of target characteristic compounds corresponding to a plurality of target pollution categories according to a plurality of sample compound sets of a plurality of target pollution categories; Based on the plurality of target characteristic compounds, a random forest classification model is trained; Combining the random forest classification model and the feature screening algorithm to perform feature selection on the plurality of target characteristic compounds to obtain a plurality of target characteristic pollutants; Based on the plurality of target characteristic pollutants and the plurality of target pollution categories, a linear discriminant analysis classification model is trained, wherein each target characteristic pollutant corresponds to at least one target pollution category label; Based on the linear discriminant analysis classification model, compound analysis is performed on the samples to be analyzed to obtain the pollution category and source pollution compound of each sample to be analyzed.
[0006] In some embodiments, obtaining a plurality of target water samples corresponding to each of the plurality of target pollution categories comprises: Obtain multiple original water samples corresponding to each of the multiple target pollution categories; Performing a filtration pretreatment on each of the original samples based on a 0.45 μm glass fiber filter membrane; After adjusting the pH value of each of the raw water samples to the target extraction value, multiple different solid phase extraction columns are connected in series to perform compound extraction treatment on each of the raw water samples to obtain multiple extracted water samples, each of which corresponds to the raw water sample. The target extraction value ranges from 6.3 to 6.8. During the compound extraction treatment, the raw water sample is loaded at a speed of 10 ml / min; The residual solution on the solid phase extraction column is eluted into the extracted water sample and mixed to obtain a mixed water sample; The mixed water sample is concentrated to obtain the target water sample.
[0007] In some embodiments, the non-targeted mass spectrometry analysis of the target water sample corresponding to each target pollution category is performed based on the mass spectrometer to obtain a plurality of fingerprint spectra for characterizing the chemical indicators in the corresponding target water sample of each target pollution category, including: Based on the positive and negative ion modes of the mass spectrometer, non-targeted mass spectrometry analysis is performed on the target water sample corresponding to each target pollution category to obtain multiple fingerprint spectra for characterizing the chemical indicators in the target water sample corresponding to each target pollution category, wherein, during the non-targeted mass spectrometry analysis, the compounds in the target water sample are separated based on the chromatographic column, and a pre-guard column is set on the front side of the chromatographic column.
[0008] In some embodiments, the target-based analysis software processes each of the fingerprint spectra to obtain a set of sample compounds contained in each target pollution category, including: Based on the target analysis software, compound feature extraction is performed on each of the fingerprint spectra to obtain multiple sample feature compounds in each target pollution category, and the multiple sample feature compounds are combined to form the sample compound set; wherein, during the extraction process, the baseline of the target analysis software is set to 500 counts, the loaded ion types are positive ions (+H, +Na, +K), negative ions (-H), the retention time tolerance of the peak alignment parameters is ±0.00% + 0.10 min, and the mass tolerance is ±20 ppm+2.00 mDa.
[0009] In some embodiments, the determining, based on the plurality of sample compound sets of the plurality of target pollution categories, a plurality of target characteristic compounds corresponding to the plurality of target pollution categories comprises: Processing the characteristic compounds of the sample by combining a 75% drift algorithm and a baseline algorithm to obtain intermediate characteristic compounds; Matching the intermediate characteristic compound with the target compound in the target database to obtain a matching characteristic compound whose matching degree meets a preset matching degree; The matching characteristic compounds are deduplicated to obtain the target characteristic compounds.
[0010] In some embodiments, the training of a random forest classification model based on a plurality of target characteristic compounds comprises: Proportionally dividing the plurality of target specific compounds to obtain a first training set and a first test set; Performing training based on the first training set to obtain an initial random forest classification model; The initial random forest classification model is tested based on the first test set to obtain the random forest classification model, where the random forest classification model is the initial random forest classification model that meets the first generalization capability.
[0011] In some embodiments, the combination of the random forest classification model and the feature screening algorithm performs feature selection on the plurality of target characteristic compounds to obtain a plurality of target characteristic pollutants including: Based on the random forest classification model, performing a first importance assessment on the target characteristic compounds to obtain ranked characteristic compounds; Creating shadows of the sorted characteristic compounds by using the characteristic screening algorithm to obtain target shadow compounds; Based on the feature screening algorithm, a second importance evaluation is performed on the ranked feature compounds and the target shadow compounds to obtain a plurality of the target feature pollutants.
[0012] In some embodiments, the training of a linear discriminant analysis classification model based on the plurality of target characteristic pollutants and the plurality of target pollution categories comprises: Proportionally dividing the plurality of target characteristic pollutants to obtain a second training set and a second test set; Based on the second training set, an initial linear discriminant analysis model is obtained; Verifying the initial linear discriminant analysis classification model based on the second test set to obtain a target linear discriminant analysis model, wherein the target linear discriminant analysis model is the initial linear discriminant analysis classification model that meets the second generalization capability; Obtain multiple categories of characteristic pollutants corresponding to each target pollution category; The linear discriminant analysis classification model is determined based on a plurality of category characteristic pollutants corresponding to each target pollution category and the target linear discriminant analysis model.
[0013] In some embodiments, the determining of the linear discriminant analysis classification model based on the multiple category characteristic pollutants corresponding to each target pollution category and the target linear discriminant analysis model includes: Construct the category label vector and the category feature matrix and category projection matrix corresponding to each target pollution category; The linear discriminant analysis classification model is determined based on the category label vector, the category feature matrix, the category projection matrix and the target linear discriminant analysis model.
[0014] In some embodiments, based on the plurality of target characteristic pollutants and the plurality of target pollution categories, training a linear discriminant analysis classification model comprises: Based on the multiple target characteristic pollutants and the multiple target pollution categories, combined with the 5-fold cross-validation method, a linear discriminant analysis classification model was trained.
[0015] The present invention includes but is not limited to the following beneficial effects: (1) The present invention can realize the identification of important features of a large number of compounds in complex pollution fields, and is suitable for tracing source pollution compounds in complex pollution fields. Furthermore, combined with a mass spectrometer and target analysis software, a large number of samples can be automatically processed and analyzed, thereby improving the efficiency of tracing source pollution compounds in complex pollution fields; (2) The present invention does not rely on specific target compounds for tracing source pollution compounds, and can adapt to various environments and pollution conditions, thereby achieving classification of point source pollution compounds of different categories of pollution sources from multiple different pollution categories, and has strong adaptability; (3) By using a 0.45μm glass fiber filter membrane for pre-filtration treatment, large particle impurities in the sample can be effectively removed, thereby improving the accuracy of subsequent analysis. In addition, adjusting the pH value of the sample to the target extraction value and using multiple different solid phase extraction columns for compound extraction treatment can ensure that more compounds are extracted from the sample, thereby improving the comprehensiveness of the analysis; (4) By concentrating the mixed water sample, the sensitivity of the analysis can be improved, thereby The possibility of detecting low-concentration pollutants is increased; (5) The present invention can enhance the detection and analysis capabilities of chemical indicators in target water samples by performing non-targeted mass spectrometry analysis in the positive and negative ion modes of the mass spectrometer. In particular, for some easily ionized compounds, the mass spectrometry analysis in the positive and negative ion modes can provide more accurate and detailed information, and by using a chromatographic column to separate the compounds in the target water sample, the interference in the sample can be more effectively eliminated, thereby improving the accuracy of mass spectrometry analysis; further, by arranging a pre-protection column at the front side of the chromatographic column, the chromatographic column can be effectively protected and its service life can be extended, thereby improving the stability and reliability of the analysis; (6) The present invention can more accurately identify and extract characteristic compounds by setting the baseline, loading ion type, retention time tolerance, mass tolerance and compound detection index score of the target analysis software, thereby improving the accuracy of the analysis, and by combining the 75% drift algorithm and the baseline algorithm to process the initial characteristic compounds, the true characteristic compounds can be more effectively screened out, thereby improving the efficiency of the analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art description: Figure 1 It is a flow chart of the method for extracting chemical indicator fingerprint spectrum for pollution source analysis based on non-targeted screening according to the present invention; Figure 2 It is another flow chart of the method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to the present invention; Figure 3It is another flow chart of the method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to the present invention; Figure 4 It is another flow chart of the method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to the present invention; Figure 5 It is another flow chart of the method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to the present invention; Figure 6 It is another flow chart of the method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to the present invention; Figure 7 It is a schematic structural diagram of the chemical indicator fingerprint spectrum extraction device for pollution source analysis based on non-targeted screening according to the present invention; Figure 8 is a schematic structural diagram of an electronic device of the present invention; Fig. 9 This is an example diagram of the importance ranking of target characteristic compounds in the present invention; Fig.10 It is the mass spectrum intensity of the target characteristic compound in the target pollution category in the present invention. DETAILED DESCRIPTION
[0017] The embodiment of the present invention provides a method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening, the method comprising: obtaining multiple target water samples corresponding to multiple target pollution categories, the multiple target pollution categories including at least two of a biotechnology plant, a petrochemical plant, a sewage treatment plant, or a textile printing and dyeing plant; performing non-targeted mass spectrometry analysis on the target water samples corresponding to each target pollution category based on a mass spectrometer to obtain multiple fingerprint spectra for characterizing chemical indicators in the corresponding target water samples of each target pollution category, each target pollution category corresponding to a fingerprint spectrum; analyzing and processing each fingerprint spectrum based on target analysis software to obtain a set of sample compounds contained in each target pollution category, the sample chemical indicators The compound set includes multiple sample characteristic compounds; according to multiple sample compound sets of multiple target pollution categories, multiple target characteristic compounds corresponding to multiple target pollution categories are determined; based on multiple target characteristic compounds, a random forest classification model is trained; multiple target characteristic compounds are feature selected in combination with the random forest classification model and a feature screening algorithm to obtain multiple target characteristic pollutants; based on multiple target characteristic pollutants and multiple target pollution categories, a linear discriminant analysis classification model is trained, and each target characteristic pollutant corresponds to at least one target pollution category label; based on the linear discriminant analysis classification model, compound analysis is performed on the sample to be analyzed to obtain the pollution category and source pollution compound of each sample to be analyzed. The present invention can solve the problem of tracing the source of complex point source pollution and improve the efficiency of tracing the source.
[0018] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0019] For ease of understanding, the specific process of the embodiment of the present invention is described below. Specifically, Figure 1 This is a flow chart of the method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening of the present application, the steps comprising: S100, obtaining a plurality of target water samples corresponding to a plurality of target pollution categories.
[0020] Specifically, the target pollution sources may include but are not limited to at least two of a biotechnology plant, a petrochemical plant, a sewage treatment plant, or a textile printing and dyeing plant.
[0021] Exemplarily, in this embodiment, the target pollution sources may include four factories, namely, a biotechnology plant, a petrochemical plant, a sewage treatment plant, and a textile printing and dyeing plant, wherein each target water sample is obtained by further processing 500 ml of raw water samples discharged from various factories collected as independent samples, and the raw water samples may include 30 biotechnology plant samples, 25 petrochemical plant samples, 20 sewage treatment plant samples, and 22 textile printing and dyeing plant samples, which are stored at 4°C and transported back to the laboratory, and the samples are pre-treated within 24 hours. Exemplarily, the sample pre-treatment may include filtering, extracting, and concentrating each raw water sample to obtain the desired target water sample.
[0022] For example, Figure 2 This is another flow chart of a method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening. This method is an exemplary description of step S100. Figure 2 As shown, the following steps are included: S200, obtaining multiple original water samples corresponding to multiple target pollution categories.
[0023] Specifically, the raw water sample is the original water sample in the target pollution category.
[0024] S202. Pre-treat each raw water sample by filtration based on a 0.45 μm glass fiber filter membrane.
[0025] S204, after adjusting the pH value of each original water sample to the target extraction value, multiple different solid phase extraction columns are connected in series to perform compound extraction processing on each original water sample to obtain multiple extracted water samples.
[0026] In this example, the extracted water sample corresponds to the original water sample, the target extraction value ranges from 6.3 to 6.8, and during the compound extraction process, the original water sample is loaded at a speed of 10 ml / min, and the water sample is passed through the solid phase extraction column at a specific speed so that the adsorption material can adsorb the compounds to be extracted. Among them, the solid phase extraction column can include 4, such as Oasis HLB solid phase extraction column, ISOLUTE ENV+ solid phase extraction column, Strata XCW solid phase extraction column, Strata XAW solid phase extraction column, wherein the mass of the filler of the Oasis HLB solid phase extraction column is 200 mg, the mass of the filler of the ISOLUTE ENV+ solid phase extraction column is 150 mg, the mass of the filler of the Strata XCW solid phase extraction column is 100 mg, and the mass of the filler of the Strata XAW solid phase extraction column is 100 mg.
[0027] S206, eluting the residual solution on the solid phase extraction column into the extracted water sample and mixing it to obtain a mixed water sample.
[0028] Further, after the extraction is completed to obtain multiple extracted water samples, the residual solution remaining on the solid phase extraction column during each extraction process is eluted into the extracted water sample and mixed to obtain a mixed water sample. Preferably, the eluent can be 6 ml of ethyl acetate: methanol (1:1) solution containing 0.5% ammonia and 3 ml of ethyl acetate: 3 ml of methanol (1:1) solution containing 1.7% formic acid.
[0029] S208, concentrating each mixed water sample to obtain a target water sample.
[0030] For example, the solution can be concentrated to 100 μL and reconstituted to 1 mL under nitrogen.
[0031] It can be understood that in this example, by using a 0.45μm glass fiber filter membrane for filtration pretreatment, large particle impurities in the sample can be effectively removed, thereby improving the accuracy of subsequent analysis. In addition, adjusting the pH value of the sample to the target extraction value and using multiple different solid phase extraction columns for compound extraction treatment can ensure that more compounds are extracted from the sample, thereby improving the comprehensiveness of the analysis; further, by concentrating the mixed water sample, the sensitivity of the analysis can be improved, thereby increasing the possibility of detecting low-concentration pollutants.
[0032] S102. Perform non-targeted mass spectrometry analysis on the target water samples corresponding to each pollution category based on a mass spectrometer to obtain multiple fingerprint spectra of each target pollution category for characterizing chemical indicators in the corresponding target water samples.
[0033] Specifically, each target pollution category corresponds to a fingerprint spectrum. Exemplarily, the mass spectrometer can be an Agilent 6546 LC / Q-TOF high-resolution mass spectrometer, and each target water sample is subjected to non-target mass spectrometry analysis by the Agilent 6546 LC / Q-TOF high-resolution mass spectrometer. Pre-selected, in the non-targeted mass spectrometry analysis process, the compounds in the target water sample are separated based on the chromatographic column, and a protective column is set on the front side of the chromatographic column. Preferably, the chromatographic column can select an ACQUITY UPLC BEH C18 (1.7 μm, 2.1×100 mm) chromatographic column, and the protective column can select an ACQUITY UPLCBEH C18 (1.7 μm) protective column, that is, "the chromatographic column used is ACQUITY UPLC BEH C18 (1.7 μm, 2.1×100 mm)": ACQUITY UPLC BEH C18 chromatographic column is a commonly used reverse phase chromatographic column suitable for separating most organic compounds. In this example, the column dimensions are 2.1 mm ID, 100 mm length, and the particle size of the filler is 1.7 μm. These parameters determine the separation efficiency and pressure of the column. The ACQUITY UPLC BEH C18 guard column is (1.7 μm). The filler of this guard column is the same as the main column, which can prevent impurities or large molecular compounds from entering the main column.
[0034] It can be understood that this technical solution can enhance the detection and analysis capabilities of chemical indicators in target water samples by performing non-targeted mass spectrometry analysis in the positive and negative ion modes of the mass spectrometer. In particular, for some easily ionized compounds, mass spectrometry analysis in positive and negative ion modes can provide more accurate and detailed information, and by using a chromatographic column to separate the compounds in the target water sample, interferences in the sample can be more effectively eliminated, thereby improving the accuracy of mass spectrometry analysis. Furthermore, by setting a pre-guard column on the front side of the chromatographic column, the chromatographic column can be effectively protected and its service life can be extended, thereby improving the stability and reliability of the analysis.
[0035] Furthermore, when performing non-targeted mass spectrometry analysis, the non-targeted analysis process is specifically completed in combination with the relevant data in Table 1 and Table 2 below, wherein Table 1 is the setting parameters of the Agilent 6546 LC / Q-TOF high-resolution mass spectrometer non-targeted analysis instrument. Table 2 is the mobile phase gradient setting data table: Table 1
[0036] Table 2
[0037] S104, analyzing and processing the multiple fingerprint spectra based on the target analysis software to obtain a set of sample compounds included in each target pollution category.
[0038] Specifically, each sample compound set includes a plurality of sample characteristic compounds.
[0039] It is understood that a fingerprint spectrum is a graphical representation used to display the results of mass spectrometry analysis. Fig.10As shown, the X-axis usually represents the mass-to-charge ratio (m / z), while the Y-axis represents the relative intensity or ion count. Each peak represents a specific ion, the position of the peak (m / z value) represents the mass-to-charge ratio of the ion, and the height of the peak represents the relative abundance of the ion. The mass spectrum obtained after sample processing and mass spectrometry analysis may contain tens to hundreds of peaks, each peak representing a compound or a fragment of a compound in the sample. These peaks can be used to identify compounds in the sample. For example, if you know the m / z value of a peak and its possible fragmentation pattern, you can infer which compound this peak may represent. The analysis of fingerprints requires specialized knowledge and tools. Exemplarily, in this example, the analysis is performed by target analysis software, such as AGILENT MassHunter software, or Mass ProfilerProfessional software. Operations such as compound feature extraction, alignment, normalization and matching are implemented. In this example, when multiple fingerprint spectra are analyzed and processed separately based on the target analysis software, the baseline of the target analysis software can be set to 500 counts, the type of charged ions can be positive ions (+H, +Na, +K), negative ions (-H), the retention time tolerance of the peak alignment parameter can be set to ±0.00% + 0.10 min, the mass tolerance can be set to ±20 ppm+2.00 mDa, the compound detection index score can be set to 75, and the EIC peak integration function can be set to Agile 2. Among them, 500 counts is the threshold for peak detection, and only peaks with signal intensity exceeding this value can be detected. The baseline is set to 500 counts: This is the threshold for peak detection, and only peaks with signal intensity exceeding this value can be detected. The type of charged ions can be positive ions (+H, +Na, +K), negative ions (-H): It is used to specify the ionization mode, that is, what type of ions are formed by the compounds in the sample in the ion source. The tolerance of the peak alignment parameter is used to determine whether the peaks of the same compound from different samples correspond. Including the tolerance of the retention time and the tolerance of the mass (mass-to-charge ratio). The compound detection index score is used to evaluate the quality of feature extraction, and only compounds with a score greater than this value will be retained. Select Agile 2 for the EIC peak integration function: This specifies the method for peak area calculation.
[0040] Furthermore, after the fingerprint spectrum is analyzed by the target software, the sample characteristic compound can be obtained, and then step S106 can be performed to obtain the target characteristic compound.
[0041] S106. Determine a plurality of target characteristic compounds corresponding to the plurality of target pollution categories according to the plurality of sample compound sets of the plurality of target pollution categories.
[0042] Specifically, Figure 3Another flow chart of the method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening is shown in FIG. 1 . The flow chart is an exemplary description of step S106. Figure 3 As shown, the following steps are included: S300, combined with the 75% drift algorithm and the baseline algorithm, processes the sample characteristic compounds to obtain intermediate characteristic compounds.
[0043] Exemplarily, a Z-Transform baseline algorithm may be selected to process the initial characteristic compounds to obtain intermediate characteristic compounds.
[0044] S302, matching the intermediate characteristic compound with the target compound in the target database to obtain a matching characteristic compound whose matching degree meets a preset matching degree.
[0045] Among them, the target database can be the Agilent EPA PCDL database, and the compounds whose matching degree meets the preset matching degree as the target characteristic compounds can refer to the intermediate characteristic compounds whose matching degree with the target compounds in the target database is greater than 90 points as the matching characteristic compounds. It can be understood that the Agilent EPAPCDL (Personal Compound Database and Library) database is an advanced compound database and compound library provided by Agilent Technologies. This database contains a large amount of compound information, including their fingerprint spectra, retention time, molecular structure, etc., and fingerprint spectra, retention time, molecular structure, etc. can be matched during matching.
[0046] S304, performing deduplication processing on the matching characteristic compounds to obtain target characteristic compounds.
[0047] Illustratively, in this example, after analyzing and processing the fingerprint spectra formed by 97 compounds using the target analysis software, 47 target characteristic compounds were obtained.
[0048] It can be understood that in this example, by setting the baseline of the target analysis software, the type of loaded ions, the retention time tolerance of the peak alignment parameters, the mass tolerance and the compound detection index score, the characteristic compounds can be more accurately identified and extracted, thereby improving the accuracy of the analysis, and by combining the 75% drift algorithm and the baseline algorithm to process the initial characteristic compounds, the real characteristic compounds can be more effectively screened out, thereby improving the efficiency of the analysis, and further by setting the pre-guard column to be located on the front side of the chromatographic column, the chromatographic column can be effectively protected and its service life can be extended, thereby improving the stability and reliability of the analysis.
[0049] S108. Based on multiple target characteristic compounds, a random forest classification model is trained.
[0050] Specifically, Figure 4 This is another flow chart of the method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening. This flow chart is an exemplary description of step S108. For details, see Figure 4 The steps include: S400, dividing the plurality of target characteristic compounds into proportions to obtain a first training set and a first test set.
[0051] Exemplarily, after obtaining 47 target characteristic compounds in step S104, a random forest classification model is trained based on the 47 characteristic compounds. Preferably, in this example, the 47 characteristic compounds can be divided into a ratio to obtain a first training set and a first test set, preferably divided into a ratio of 7:3 to obtain a first training set and a first test set, respectively.
[0052] S402: Perform training based on the first training set to obtain an initial random forest classification model.
[0053] S404: Verify the initial random forest classification model based on the first test set to obtain a random forest classification model.
[0054] Specifically, the random forest classification model is an initial random forest classification model that meets the first generalization capability. It can be understood that the first test set is used to evaluate the performance of the initial random forest classification model, performance indicators such as accuracy, recall rate, F1 score, etc. If the performance indicators meet the preset requirements, it is considered that the initial random forest classification model meets the first generalization capability and can be used as a random forest classification model.
[0055] S110. Combining the random forest classification model and the feature screening algorithm, feature selection is performed on multiple target characteristic compounds to obtain multiple target characteristic pollutants.
[0056] Specifically, each target characteristic pollutant corresponds to at least one target pollution category label.
[0057] In one example, if Figure 5 As shown, another flow chart of the method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening is disclosed. This method is an exemplary description of step S110. Figure 5 As shown, the following steps are included: S500: Based on the random forest classification model, a first importance evaluation is performed on the target characteristic compounds to obtain ranked characteristic compounds.
[0058] S502, creating shadows of sorted characteristic compounds through a characteristic screening algorithm to obtain target shadow compounds.
[0059] S504: Perform a second importance evaluation on the ranked characteristic compounds and target shadow compounds based on a characteristic screening algorithm to obtain a plurality of target characteristic pollutants.
[0060] Specifically, the first importance assessment may refer to an initial assessment of the first contribution of the target characteristic compound based on the random forest classification model. During the assessment, the random forest model may provide an importance score for each target characteristic pollutant, which is based on the contribution of the target characteristic pollutant in the model prediction. Further, after the initial assessment, a ranking will be obtained according to the degree of importance. Further, a second importance assessment is performed based on the feature screening algorithm. When performing the second importance assessment, step S502 is first performed, and the shadow of the sorted characteristic compound is created based on the feature algorithm to obtain the target shadow compound. Step S504 is further performed to compare the target shadow compound with the sorted characteristic compound. Preferably, if the importance of a characteristic compound in the sorted characteristic compound exceeds the importance of all target shadow compounds, then this feature can be considered to be important and recorded as the target characteristic pollutant. It can be understood that since there are multiple sorted characteristic compounds, the comparison process will be repeated many times until all features are confirmed to be important or unimportant.
[0061] It can be understood that in this example, the use of the random forest classification model for initial importance assessment can quickly and effectively sort a large number of characteristic compounds, thereby enhancing the reliability of feature selection. Furthermore, by creating shadow features and performing another importance assessment, truly meaningful characteristic compounds can be further screened out, which enhances the accuracy of feature screening. Among a large number of characteristic compounds, the combined use of random forest and feature screening algorithms can efficiently identify target characteristic pollutants, greatly improving the analysis efficiency.
[0062] In this example, after step S110, 27 target characteristic pollutants are screened out from the 47 target characteristic compounds. The 27 target characteristic pollutants are shown in Table 3. The importance ranking of the 27 target characteristic pollutants is as follows: Fig. 9 As shown, the mass spectrum intensity in the target contamination category is Fig.10 shown.
[0063] Table 3
[0064] S112. Based on multiple target characteristic pollutants and multiple target pollution categories, a linear discriminant analysis classification model is trained.
[0065] In one example, Figure 6 This is another flow chart of the method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening. This flow chart is an exemplary description of step S112. Figure 6 , including the following steps: S600, dividing the plurality of target characteristic pollutants into proportions to obtain a second training set and a second test set.
[0066] Specifically, the 27 target characteristic pollutants may be first divided proportionally to obtain a second training set and a second test set.
[0067] S602: Obtain an initial linear discriminant analysis model based on the second training set training.
[0068] It is understandable that during the model training process, data expansion can be performed on the second training set to obtain sufficient training samples.
[0069] S604: Verify the initial linear discriminant analysis classification model based on the second test set to obtain a target linear discriminant analysis model.
[0070] Specifically, the target linear discriminant analysis model is an initial linear discriminant analysis classification model that meets the second generalization capability. It is understandable that the second test set is used to evaluate the performance of the initial linear discriminant analysis classification model, performance indicators such as accuracy, recall rate, F1 score, etc. If the performance indicators meet the preset requirements, it is considered that the initial linear discriminant analysis classification model meets the second generalization capability and can be used as a random forest classification model.
[0071] Exemplarily, in this example, the evaluation is based on the ROC classification accuracy, area under the curve (AUC), recall, precision and F1 score. It can be understood that accuracy is the most intuitive evaluation indicator, which indicates the proportion of samples correctly predicted by the model to the total number of samples. After the test, the accuracy is 0.9262, that is, the linear discriminant analysis classification model made correct predictions on approximately 97% of the samples. Further, AUC (Area Under Curve) is the area under the ROC curve (receiver operating characteristic curve), which is used to evaluate the performance of the classifier. The value of AUC is between 0.5 and 1, and the closer the value is to 1, the better the performance of the model. Further recall is also called the true positive rate, which is the proportion of the number of positive samples correctly predicted by the model to all samples that are actually positive. After the test, the recall is 0.9262, that is, the linear discriminant analysis classification model can correctly identify 92.62% of the pollution source samples. Precision is also called positive prediction value. It is the ratio of the number of positive samples correctly predicted by the model to the number of samples predicted as positive. After the test, the precision is 0.9437, which means that 94.37% of the samples predicted as positive by the linear discriminant analysis classification model are correct. The F1 score is the harmonic mean of precision and recall, which is used to consider both precision and recall. After the test, the F1 score is 0.9194, indicating that the linear discriminant analysis classification model performs well in both precision and recall. Kappa is a measure of the performance of a classifier relative to random classification. After the test, the Kappa coefficient is 0.8991, indicating that the performance of the linear discriminant analysis classification model is significantly higher than random classification. Matthews Correlation Coefficient (MCC) is an indicator to measure the performance of a binary classification model. Its value range is -1 to 1, with 1 indicating perfect prediction, 0 indicating random prediction, and -1 indicating completely reverse prediction. After the test, the MCC is 0.9115, indicating that the performance of the linear discriminant analysis classification model is very good. It is understandable that the evaluation test can also be performed based on other indicators, which are not specifically limited here.
[0072] S606: Obtain multiple categories of characteristic pollutants corresponding to each target pollution category.
[0073] It is understandable that, among the multiple sample characteristic compounds corresponding to the sample compound set of each target pollution category obtained in step S104, each sample characteristic compound will correspond to a target pollution category label. After the sub-steps S300-S304 of step S106 are subsequently executed, the deduplication process performed only removes the repeated characteristic compounds. After removing the repeated characteristic compounds, the remaining characteristic compounds will simultaneously merge the target pollution category labels corresponding to the deduplicated characteristic compounds, so that the number of category characteristic pollutants corresponding to each target pollution category can be determined from the 27 target characteristic pollutants.
[0074] S608: Determine a linear discriminant analysis classification model based on multiple category characteristic pollutants corresponding to each target pollution category and the target linear discriminant analysis model.
[0075] Specifically, we can first construct a category label vector and a category feature matrix and a category projection matrix corresponding to each target pollution category; then determine the linear discriminant analysis classification model based on the category label vector, the category feature matrix, the category projection matrix and the target linear discriminant analysis model.
[0076] Specifically, constructing the category label vector and the category feature matrix and category projection matrix corresponding to each target pollution category can be achieved based on the following steps: The multiple category characteristic pollutants corresponding to each target pollution category are used as category samples of the corresponding target pollution category, and the multiple category characteristic pollutants corresponding to each target pollution category are sorted into a feature matrix X, assuming that the number of category samples is N. Create a category label vector Y, whose length is N, and each element corresponds to the target pollution category label of the category sample in the matrix feature.
[0077] The category mean vector and the intra-category covariance matrix are further calculated, where the category mean vector is determined as follows: For each target pollution category C k , calculate the category mean vector of the category characteristic pollutants corresponding to each target pollution category: ; in The target pollution category is C k The number of class samples, It belongs to the target pollution category C k The category feature vector of Within-class covariance matrix Determined based on the following steps: Compute the covariance matrix for each target pollution class and sum them: ; in, The target pollution category is C k The covariance matrix of is calculated as: ; Furthermore, the inter-category covariance matrix of the target pollution category is calculated: Between-class covariance matrix : Compute the difference between the category mean vector and the overall mean vector: ; The covariance matrix between target pollution categories is: ; in, The target pollution category The number of class samples; Further, solve the eigenvectors and eigenvalues: Specifically, by solving the generalized eigenvalue problem: ; Calculate eigenvalues and the corresponding eigenvector .
[0078] Furthermore, a linear discriminant function is constructed: Specifically, the projection matrix W is used to project the feature matrix X into a low-dimensional space: Z = XW; The linear discriminant function is: ; in, is the weight extracted from the projection matrix, and b is the bias term, which can be set in advance.
[0079] During the training process, Z and label vector Y obtained based on feature matrix X are input into the target linear discriminant analysis model for training to obtain a linear discriminant analysis classification model.
[0080] In another example, a linear discriminant analysis classification model can be trained based on multiple target characteristic pollutants in combination with a 5-fold cross-validation method.
[0081] In yet another example, a linear discriminant analysis classification model may be trained based on multiple target characteristic pollutants and multiple target pollution categories in combination with a 5-fold cross-validation method.
[0082] S114. Analyze the compounds in the samples to be analyzed based on the linear discriminant analysis classification model to obtain the pollution category and source pollution compounds of each sample to be analyzed.
[0083] Specifically, after the linear discriminant analysis classification model is determined, the sample to be analyzed that needs to be tested can be input into the linear discriminant analysis classification model to obtain the output pollution category of the sample to be analyzed and the source pollution compounds in the sample.
[0084] Exemplarily, the source contaminant compounds can be displayed based on a source fingerprint spectrum.
[0085] It can be understood that the present invention can realize the identification of important characteristics of a large number of compounds in complex pollution fields, and is suitable for tracing source pollution compounds in complex pollution fields. Furthermore, combined with a mass spectrometer and target analysis software, a large number of samples can be automatically processed and analyzed, thereby improving the efficiency of tracing source pollution compounds in complex pollution fields; further, the present invention does not rely on specific target compounds for tracing source pollution compounds, and can adapt to various environments and pollution conditions, thereby realizing the classification of point source pollution compounds of different categories of pollution sources from multiple different pollution categories, and has strong adaptability.
[0086] Further, according to one aspect of the present application, a device for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening is also disclosed, such as Figure 7 As shown, the device comprises: A target water sample acquisition module is used to acquire a plurality of target water samples corresponding to a plurality of target pollution categories, wherein the plurality of target pollution categories include at least two of a biotechnology plant, a petrochemical plant, a sewage treatment plant, or a textile printing and dyeing plant; A non-targeted mass spectrometry analysis module is used to perform non-targeted mass spectrometry analysis on the target water sample corresponding to each target pollution category based on a mass spectrometer to obtain a plurality of fingerprint spectra for each target pollution category for characterizing the chemical indicators in the corresponding target water sample, and each target pollution category corresponds to one fingerprint spectrum; A fingerprint spectrum analysis and processing module, used for analyzing and processing each of the fingerprint spectrum based on the target analysis software to obtain a sample compound set contained in each target pollution category, wherein the sample compound set includes a plurality of sample characteristic compounds; A target characteristic compound determination module, used to determine a plurality of target characteristic compounds corresponding to a plurality of target pollution categories according to a plurality of sample compound sets of a plurality of target pollution categories; A random forest classification model training module, used for training a random forest classification model based on a plurality of target characteristic compounds; A target characteristic pollutant module is used to perform feature selection on a plurality of target characteristic compounds in combination with the random forest classification model and the characteristic screening algorithm to obtain a plurality of target characteristic pollutants; A linear discriminant analysis classification model training module, used to train a linear discriminant analysis classification model based on a plurality of target characteristic pollutants and a plurality of target pollution categories, wherein each target characteristic pollutant corresponds to at least one target pollution category label; The compound analysis module is used to perform compound analysis on the samples to be analyzed based on the linear discriminant analysis classification model to obtain the pollution category and source pollution compound of each sample to be analyzed.
[0087] The application introduction of the relevant modules of the device in this example can refer to the relevant introduction of the principle of the above method, which will not be repeated here.
[0088] According to another aspect of the present application, the present application also discloses an electronic device, which includes a memory and at least one processor, wherein instructions are stored in the memory; at least one processor calls the instructions in the memory so that the electronic device executes the various steps of the above-mentioned method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening.
[0089] above Figure 7 The chemical indicator fingerprint extraction device for pollution source analysis based on non-targeted screening in the embodiment of the present invention is described in detail from the perspective of modular functional entities, and the electronic device in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0090] Figure 8 8 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device 800 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 810 (for example, one or more processors) and a memory 820, and one or more storage media 830 (for example, one or more mass storage devices) storing application programs 833 or data 832. Among them, the memory 820 and the storage medium 830 can be short-term storage or permanent storage. The program stored in the storage medium 830 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the electronic device 800. Furthermore, the processor 810 may be configured to communicate with the storage medium 830 to execute a series of instruction operations in the storage medium 830 on the electronic device 800.
[0091] The electronic device 800 may also include one or more power supplies 840, one or more wired or wireless network interfaces 850, one or more input and output interfaces 860, and / or one or more operating systems 831, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be appreciated by those skilled in the art that Figure 8 The structure of the electronic device shown does not constitute a limitation on the electronic device, and may include more or less components than shown in the figure, or combine some components, or arrange the components differently.
[0092] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of a method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening.
[0093] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device, or unit can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0094] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0095] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening, characterized in that: The method comprises: Acquire a plurality of target water samples corresponding to a plurality of target pollution categories, wherein the plurality of target pollution categories include at least two of a biotechnology plant, a petrochemical plant, a sewage treatment plant, or a textile printing and dyeing plant; Based on the mass spectrometer, non-targeted mass spectrometry analysis is performed on the target water sample corresponding to each target pollution category to obtain multiple fingerprint spectra of each target pollution category for characterizing the chemical indicators in the corresponding target water sample, and each target pollution category corresponds to one fingerprint spectrum; Analyzing and processing each of the fingerprint spectra based on target analysis software to obtain a sample compound set contained in each target pollution category, wherein the sample compound set includes a plurality of sample characteristic compounds; Determining a plurality of target characteristic compounds corresponding to a plurality of target pollution categories according to a plurality of sample compound sets of a plurality of target pollution categories; Based on the plurality of target characteristic compounds, a random forest classification model is trained; Combining the random forest classification model and the feature screening algorithm to perform feature selection on the plurality of target characteristic compounds to obtain a plurality of target characteristic pollutants; Based on the plurality of target characteristic pollutants and the plurality of target pollution categories, a linear discriminant analysis classification model is trained, wherein each target characteristic pollutant corresponds to at least one target pollution category label; Based on the linear discriminant analysis classification model, compound analysis is performed on the samples to be analyzed to obtain the pollution category and source pollution compound of each sample to be analyzed.
2. The method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to claim 1, characterized in that: The step of obtaining a plurality of target water samples corresponding to a plurality of target pollution categories comprises: Obtain multiple original water samples corresponding to each of the multiple target pollution categories; Performing a pre-treatment of each of the raw water samples by filtration based on a 0.45 μm glass fiber filter membrane; After adjusting the pH value of each of the raw water samples to the target extraction value, multiple different solid phase extraction columns are connected in series to perform compound extraction treatment on each of the raw water samples to obtain multiple extracted water samples, each of which corresponds to the raw water sample. The target extraction value ranges from 6.3 to 6.
8. During the compound extraction treatment, the raw water sample is loaded at a speed of 10 ml / min; The residual solution on the solid phase extraction column is eluted into the extracted water sample and mixed to obtain a mixed water sample; The mixed water sample is concentrated to obtain the target water sample.
3. The method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to claim 1, characterized in that: The non-targeted mass spectrometry analysis is performed on the target water sample corresponding to each target pollution category based on the mass spectrometer to obtain multiple fingerprint spectra for characterizing the chemical indicators in the corresponding target water sample of each target pollution category, including: Based on the positive and negative ion modes of the mass spectrometer, non-targeted mass spectrometry analysis is performed on the target water sample corresponding to each target pollution category to obtain multiple fingerprint spectra for characterizing the chemical indicators in the target water sample corresponding to each target pollution category, wherein, during the non-targeted mass spectrometry analysis, the compounds in the target water sample are separated based on the chromatographic column, and a pre-guard column is set on the front side of the chromatographic column.
4. The method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to claim 1, characterized in that: The target-based analysis software processes each of the fingerprint spectra to obtain a set of sample compounds contained in each target pollution category, including: Based on the target analysis software, compound feature extraction is performed on each of the fingerprint spectra to obtain multiple sample feature compounds in each target pollution category, and the multiple sample feature compounds are combined to form the sample compound set; wherein, during the extraction process, the baseline of the target analysis software is set to 500 counts, the loaded ion types are positive ions (+H, +Na, +K), negative ions (-H), the retention time tolerance of the peak alignment parameters is ±0.00% + 0.10 min, and the mass tolerance is ±20 ppm+2.00 mDa.
5. The method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to claim 1, characterized in that: Determining a plurality of target characteristic compounds corresponding to a plurality of target pollution categories according to a plurality of sample compound sets of a plurality of target pollution categories comprises: Processing the characteristic compounds of the sample by combining a 75% drift algorithm and a baseline algorithm to obtain intermediate characteristic compounds; Matching the intermediate characteristic compound with the target compound in the target database to obtain a matching characteristic compound whose matching degree meets a preset matching degree; The matching characteristic compounds are deduplicated to obtain the target characteristic compounds.
6. The method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to claim 1, characterized in that: The training of a random forest classification model based on the plurality of target characteristic compounds comprises: Proportionally dividing the plurality of target characteristic compounds to obtain a first training set and a first test set; Performing training based on the first training set to obtain an initial random forest classification model; The initial random forest classification model is tested based on the first test set to obtain the random forest classification model, where the random forest classification model is the initial random forest classification model that meets the first generalization capability.
7. The method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to claim 1, characterized in that: The feature selection of the plurality of target characteristic compounds is performed by combining the random forest classification model and the feature screening algorithm to obtain a plurality of target characteristic pollutants including: Based on the random forest classification model, performing a first importance assessment on the target characteristic compounds to obtain ranked characteristic compounds; Creating shadows of the sorted characteristic compounds by using the characteristic screening algorithm to obtain target shadow compounds; Based on the feature screening algorithm, a second importance evaluation is performed on the ranked feature compounds and the target shadow compounds to obtain a plurality of the target feature pollutants.
8. The method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to claim 1, characterized in that: The training of a linear discriminant analysis classification model based on the plurality of target characteristic pollutants and the plurality of target pollution categories comprises: Proportionally dividing the plurality of target characteristic pollutants to obtain a second training set and a second test set; Based on the second training set, an initial linear discriminant analysis model is obtained; Verifying the initial linear discriminant analysis classification model based on the second test set to obtain a target linear discriminant analysis model, wherein the target linear discriminant analysis model is the initial linear discriminant analysis classification model that meets the second generalization capability; Obtain multiple categories of characteristic pollutants corresponding to each target pollution category; The linear discriminant analysis classification model is determined based on a plurality of category characteristic pollutants corresponding to each target pollution category and the target linear discriminant analysis model.
9. The method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to claim 8, characterized in that: The step of determining the linear discriminant analysis classification model based on the multiple category characteristic pollutants corresponding to each target pollution category and the target linear discriminant analysis model includes: Construct the category label vector and the category feature matrix and category projection matrix corresponding to each target pollution category; The linear discriminant analysis classification model is determined based on the category label vector, the category feature matrix, the category projection matrix and the target linear discriminant analysis model.
10. The method for extracting chemical indicator fingerprints for pollution source analysis based on non-targeted screening according to claim 1, characterized in that: Based on the plurality of target characteristic pollutants and the plurality of target pollution categories, the linear discriminant analysis classification model obtained by training includes: Based on the multiple target characteristic pollutants and the multiple target pollution categories, combined with the 5-fold cross-validation method, a linear discriminant analysis classification model was trained.
Citation Information
Patent Citations
Method, system and terminal for screening anti-breast cancer candidate drug molecule descriptors
CN114334033A
Variable screening method and system based on breast cancer data, and readable storage medium
CN115346682A
Pollutant tracing method, system, device and medium
CN116297936A
Data feature selection method and device
CN116738194A
HPV infection outcome individualized prediction model construction method
CN116978567A