Receptor Model Evaluation Method and Device Based on Synthetic Dataset of Road Sediments
By constructing a road sediment synthesis data set, generating an unknown pollution source component spectrum and inputting the receptor model, the problem of difficulty in quantifying the accuracy of the receptor model in road sediment tracing is solved, and the accuracy evaluation of the traceability results is achieved.
Patent Information
- Application Number
- CN202411290182.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-09-14
AI Technical Summary
The accuracy of existing receptor models is difficult to quantify and evaluate in the traceability of road sediments. The traditional method is mainly used for atmospheric particulate matter, and there is a lack of accuracy verification when applied to road sediments.
By obtaining the component spectrum of known pollution sources, generating the component spectrum of unknown pollution sources, randomly log-normally generate the pollution source contribution mass concentration of road sediments, constructing a synthetic data set and inputting the receptor model, and evaluating the traceability results and pollution source contribution rate to improve accuracy.
The accuracy of the traceability results of the receptor model on road sediment is quantified and evaluated, and the accuracy of the traceability results of the model is improved.
Smart Images

Figure CN119355234B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sediment source tracing, and in particular to a receptor model evaluation method and device based on a synthetic dataset of road sediments. Background Art
[0002] Urban road pollution has attracted increasing attention. In the control and management of road sediment pollution, the identification and analysis of pollution sources are the key to taking effective control measures. This process includes tracing and analyzing pollution sources based on the quantitative contribution of pollution sources to the urban road environmental medium. Among them, the receptor model is a widely used source tracing method and has been widely applied in the fields of atmospheric particulate matter, soil, etc. However, traditional receptor models have all been developed in the study of air pollution, and existing methods are also used to verify the tracing accuracy of the receptor model for atmospheric particulate matter. Moreover, when comparing the receptor models, the number of receptor models is often only two. There is no accuracy of the mass contribution rate of each pollution source obtained by source analysis when the receptor model is used for road sediment source tracing. In addition, the application of the receptor model in road sediments is relatively limited, resulting in difficulty in quantifying and evaluating the accuracy of the results of applying the receptor model to road sediment source tracing. Summary of the Invention
[0003] In view of this, the main purpose of the embodiments of the present invention is to provide a receptor model evaluation method and device based on a synthetic dataset of road sediments to improve the accuracy of evaluating the source tracing results of the receptor model.
[0004] To achieve the above object, on the one hand, an embodiment of the present invention provides a receptor model evaluation method based on a synthetic dataset of road sediments, and the method includes the following steps:
[0005] Obtain a known pollution source composition spectrum;
[0006] Obtain an unknown pollution source composition spectrum according to the known pollution source composition spectrum;
[0007] Randomly generate the mass concentrations of pollution source contributions of several road sediments by lognormal distribution;
[0008] Generate a synthetic dataset according to the known pollution source composition spectrum, the unknown pollution source composition spectrum, and the mass concentrations of pollution source contributions;
[0009] Input the synthetic dataset into the receptor model to obtain the source tracing results of the receptor model and the pollution source contribution rates in the synthetic dataset;
[0010] Obtain the accuracy of the receptor model according to the source tracing results and the pollution source contribution rates in the synthetic dataset.
[0011] In some embodiments, the obtaining of the known pollution source component spectrum includes the following steps:
[0012] Obtain the component spectrum data of the actual pollution sources of road sediments;
[0013] Perform normalization processing on the component spectrum data to obtain the known pollution source component spectrum.
[0014] In some embodiments, the formula for performing normalization processing on the component spectrum data to obtain the known pollution source component spectrum is:
[0015]
[0016] In the formula, D ik represents the known pollution source component spectrum; x 1k , x 2k , …, x ik respectively represent the kth known pollution source data of the first element, the kth known pollution source data of the second element, ……, the kth known pollution source data of the ith element in the known pollution source component spectrum.
[0017] In some embodiments, the obtaining of the unknown pollution source component spectrum according to the known pollution source component spectrum includes the following steps:
[0018] Obtain the concentration average value of all known pollution sources according to the known pollution source component spectrum;
[0019] Multiply the concentration average value by a random number to obtain the concentration data of the unknown pollution source;
[0020] Perform normalization processing on the concentration data of the unknown pollution source to obtain the unknown pollution source component spectrum.
[0021] In some embodiments, the formula for performing normalization processing on the concentration data of the unknown pollution source to obtain the unknown pollution source component spectrum is:
[0022]
[0023] In the formula, E iuns represents the unknown pollution source component spectrum; x 1uns , x 2uns , …, x iuns respectively represent the concentration data of the unknown pollution source of the first element, the concentration data of the unknown pollution source of the second element, ……, the concentration data of the unknown pollution source of the ith element in the unknown pollution source component spectrum.
[0024] In some embodiments, generating a synthetic dataset according to the known pollution source composition spectrum, the unknown pollution source composition spectrum, and the mass concentration of pollution source contribution includes the following steps:
[0025] Generating the metal element pollution source composition spectrum in the road sediment according to the known pollution source composition spectrum and the unknown pollution source composition spectrum;
[0026] Obtaining the mass fractions of several metals according to the metal element pollution source composition spectrum;
[0027] Generating the metal element concentration according to the mass fractions of the metals and the mass concentration of pollution source contribution;
[0028] Obtaining a set of the metal element concentrations of several road sediments to obtain a synthetic dataset.
[0029] To achieve the above object, on the other hand, an embodiment of the present invention provides a receptor model evaluation device based on a synthetic dataset of road sediments, and the device includes:
[0030] A first module for obtaining a known pollution source composition spectrum;
[0031] A second module for obtaining an unknown pollution source composition spectrum according to the known pollution source composition spectrum;
[0032] A third module for randomly generating the mass concentrations of pollution source contributions of several road sediments by lognormal distribution;
[0033] A fourth module for generating a synthetic dataset according to the known pollution source composition spectrum, the unknown pollution source composition spectrum, and the mass concentration of pollution source contribution;
[0034] A fifth module for inputting the synthetic dataset into a receptor model to obtain the source tracing result of the receptor model and the pollution source contribution rate in the synthetic dataset;
[0035] A sixth module for obtaining the accuracy of the receptor model according to the source tracing result and the pollution source contribution rate in the synthetic dataset.
[0036] To achieve the above object, on the other hand, an embodiment of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned receptor model evaluation method based on a synthetic dataset of road sediments.
[0037] To achieve the above object, another aspect of the embodiments of the present invention provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-described receptor model evaluation method based on a synthetic dataset of road sediments.
[0038] To achieve the above object, another aspect of the embodiments of the present invention provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device executes the above-described receptor model evaluation method based on a synthetic dataset of road sediments.
[0039] The embodiments of the present invention at least include the following beneficial effects: The present invention provides a receptor model evaluation method and device based on a synthetic dataset of road sediments. The solution obtains the known pollution source component spectrum; obtains the unknown pollution source component spectrum according to the known pollution source component spectrum; randomly generates the mass concentrations of the pollution source contributions of several road sediments by lognormal distribution; generates a synthetic dataset according to the known pollution source component spectrum, the unknown pollution source component spectrum, and the mass concentrations of the pollution source contributions; inputs the synthetic dataset into the receptor model to obtain the source tracing result of the receptor model and the pollution source contribution rate in the synthetic dataset; and obtains the accuracy of the receptor model according to the source tracing result and the pollution source contribution rate in the synthetic dataset, so as to quantify and evaluate the accuracy of the source tracing result of the receptor model in road sediments. The present invention improves the accuracy of the evaluation of the source tracing result of the receptor model. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0041] Figure 1 is a flowchart of a receptor model evaluation method based on a synthetic dataset of road sediments provided by an embodiment of the present invention;
[0042] Figure 2 is a schematic diagram of the normalized known and unknown pollution source component spectra provided by an embodiment of the present invention;
[0043] Figure 3It is a schematic diagram of the source contribution results obtained by inputting the generated 6 synthetic data sets into the PMF and Unmix receptor models respectively and running them according to the embodiments of the present invention;
[0044] Figure 4 It is a schematic diagram of the source contribution results obtained by inputting the generated 6 synthetic data sets into the CMB and SCMD receptor models respectively and running them according to the embodiments of the present invention;
[0045] Figure 5 It is a schematic diagram of the overall steps of a receptor model evaluation method based on a synthetic data set of road sediments according to the embodiments of the present invention;
[0046] Figure 6 It is a schematic diagram of the hardware structure of an electronic device according to the embodiments of the present invention. Detailed implementation manners
[0047] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention described in detail in the appended claims.
[0048] It should be noted that although the functional modules are divided in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the system or the order in the flowchart. The terms "first / S100" and "second / S200" in the specification, claims and the above-mentioned drawings may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "when" as used herein may be interpreted as "when" or "when" or "in response to determining".
[0049] The terms "at least one", "multiple", "each", "any one", etc. used in the present invention, at least one includes one, two or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any one refers to any one of the multiple.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used herein are for the purpose of describing embodiments of the invention only and are not intended to limit the invention.
[0051] Receptor models are a widely used method for tracing the sources of road sediments and have been applied in many fields such as atmospheric particulate matter and soil. However, it is difficult to quantify and evaluate the accuracy of the results of current receptor models applied to road sediment source tracing.
[0052] In view of this, as Figure 1 shown, embodiments of the present invention provide a method for evaluating a receptor model based on a synthetic dataset of road sediments, and this method includes but is not limited to steps S100 to S600:
[0053] Step S100, obtain the composition spectrum of known pollution sources;
[0054] Step S200, obtain the composition spectrum of unknown pollution sources according to the composition spectrum of known pollution sources;
[0055] Step S300, randomly generate the mass concentration of pollution source contributions of several road sediments by lognormal distribution;
[0056] Step S400, generate a synthetic dataset according to the composition spectrum of known pollution sources, the composition spectrum of unknown pollution sources, and the mass concentration of pollution source contributions;
[0057] Step S500, input the synthetic dataset into the receptor model to obtain the source tracing result of the receptor model and the contribution rate of pollution sources in the synthetic dataset;
[0058] Step S600, obtain the accuracy of the receptor model according to the source tracing result and the contribution rate of pollution sources in the synthetic dataset.
[0059] In some embodiments, step S100 may include but is not limited to steps S110 to S120:
[0060] Step S110, obtain the composition spectrum data of the actual pollution sources of road sediments;
[0061] Step S120, perform normalization processing on the composition spectrum data to obtain the composition spectrum of known pollution sources.
[0062] In steps S110 to S120 of some embodiments, the composition spectrum data x of the actual pollution sources of road sediments may be obtained according to previous research data ik . Since the concentration differences of different metal pollutants emitted by road sediment pollution sources are relatively large, the order of magnitude of low-concentration elements is 10-2 , the order of magnitude of high-concentration elements can reach 10 5 , so for the component spectrum data x ik , perform mass normalization to obtain D ik . In the component spectrum data of the actual pollution source, the ratio of the concentration of each metal element to the total concentration of metal elements is defined as the normalized concentration (dimensionless unit) of the element, which reduces the differences caused by the differences in pollutant concentrations emitted by different pollution sources, thereby highlighting the contour characteristics of metal element pollution, reducing the standard deviation of the input data, and helping to improve the fitting effect of the subsequent synthetic data set in the receptor model. The calculation formula for normalization is as follows:
[0063]
[0064] In the formula, D ik represents the component spectrum of the known pollution source, dimensionless unit; x 1k , x 2k , …, x ik respectively represent the kth known pollution source data of the first element, the kth known pollution source data of the second element, ……, the kth known pollution source data of the ith element in the component spectrum of the known pollution source, and the units are all mg / kg. Calculate the normalized data of i elements and k pollution sources respectively, that is, D ik , which together form the total known pollution source component spectrum D after normalization.
[0065] In some embodiments, step S200 may include but is not limited to steps S210 to S230:
[0066] Step S210, obtain the average concentration of all known pollution sources according to the component spectrum of the known pollution source;
[0067] Step S220, multiply the average concentration by a random number to obtain the concentration data of the unknown pollution source;
[0068] Step 8230, perform normalization on the concentration data of the unknown pollution source to obtain the component spectrum of the unknown pollution source.
[0069] In steps S210 to S230 of some embodiments, by generating the component spectra of unknown pollution sources, unknown pollution sources in the real world are simulated. The generation of the unknown sources is based on the following criteria: 1) The component spectra of the unknown sources are not similar to those of the real sources; 2) The characteristic elements of the known sources are not obvious in the unknown sources; 3) The average contribution of the unknown sources should account for about 15% of the total mass percentage. Exemplarily, in order to meet the above conditions and conform to the situation of unknown sources in reality, the average metal element concentration of all known pollution sources is obtained, multiplied by a random number of 0.2 - 0.8 generated randomly, and then normalized. Then the expression for the average concentration of all known pollution sources is:
[0070]
[0071] In the formula, x iave is the average concentration of all known pollution sources of the i-th element; p is the number of known pollution sources; x i1 , x i2 ,..., x ik are the data of the first known pollution source of the i-th element, the data of the second known pollution source of the i-th element,..., the data of the k-th known pollution source of the i-th element in the known pollution source component spectra, and the unit is mg / kg for all.
[0072] The expression for the concentration data of the unknown pollution source is:
[0073] x iuns = x iave × ran
[0074] In the formula, x iuns is the concentration data of the unknown pollution source of the i-th element; ran is a random number, optionally ran is a random number between 0.2 and 0.8.
[0075] The expression for the unknown pollution source component spectrum is:
[0076]
[0077] In the formula, E iuns represents the unknown pollution source component spectrum, dimensionless unit; x 1uns , x 2uns ,..., x iuns represent the concentration data of the unknown pollution source of the first element, the concentration data of the unknown pollution source of the second element,..., the concentration data of the unknown pollution source of the i-th element in the unknown pollution source component spectrum, and the unit is mg / kg for all. Among them, each unknown pollution source component spectrum E iuns together constitute the normalized total unknown source component spectrum E.
[0078] In step S300 of some embodiments, by sampling real road sediment samples, the mass and standard deviation of metal elements contained in the road sediment in the sample can be obtained, and the total mass concentration of metal elements contained in j synthetic road sediments is randomly generated in a lognormal manner. Set the contribution rate of the real pollution source and the contribution rate of the unknown pollution source, and multiply the total mass concentration of metal elements contained in the j synthetic road sediments randomly generated in a lognormal manner by the set contribution rate of the real pollution source and the contribution rate of the unknown pollution source to obtain the contribution masses of several pollution sources with the same contribution rate. Then, obtain the average value and standard deviation of the logarithms of the contribution masses of the several pollution sources with the same contribution rate. According to the average value and standard deviation, randomly generate in Matlab the pollution source contribution mass concentrations (i.e., the total mass concentration of metal elements) B of j synthetic road sediments with different pollution source contribution rates kj Then, the set of pollution source contribution mass concentrations of j synthetic road sediments with different pollution source contribution rates is denoted as B.
[0079] In some embodiments, step S400 may include but is not limited to steps S410 to S440:
[0080] Step S410, generate the metal element pollution source composition spectrum in the road sediment according to the known pollution source composition spectrum and the unknown pollution source composition spectrum;
[0081] Step S420, obtain the mass fractions of several metals according to the metal element pollution source composition spectrum;
[0082] Step S430, generate the metal element concentration according to the mass fraction of the metal and the pollution source contribution mass concentration;
[0083] Step S440, obtain the set of the metal element concentrations of several road sediments to obtain a synthetic data set.
[0084] In step S410 of some embodiments, the normalized data of the i elements and k pollution sources calculated respectively in the above steps, that is, D ik , jointly form the normalized total known pollution source composition spectrum D; the respective unknown pollution source composition spectra E iuns jointly form the normalized total unknown source composition spectrum E; the total known pollution source composition spectrum D and the total unknown source composition spectrum E jointly form the metal element pollution source composition spectrum A in the environmental medium road sediment.
[0085] In steps S420 to S430 of some embodiments, according to the metal element pollution source composition spectrum A, obtain the mass fraction A of the i-th metal in the k-th source of the j-th sample ik . According to the mass fraction A of the i-th metal in the k-th source of the j-th sampleik and the mass concentration contributed by the pollution source, the concentration C of the i-th metal element in the j-th sample can be generated ij Then there is the following formula:
[0086]
[0087] In the formula, C ij is the concentration of the i-th metal element in the j-th sample; p is the number of pollution sources; A ik is the mass fraction (dimensionless unit) of the i-th metal in the k-th source of the j-th sample; B kj is the total mass concentration of metal elements in the k-th source of the j-th sample (mg / m 2 ).
[0088] In step S440 of some embodiments, by obtaining a set of metal element concentrations of j synthetic road sediments and using this set as the synthetic data set C
[0089] In steps S500 to S600 of some embodiments, the synthetic data set is incorporated into the source apportionment of each receptor model. According to the source apportionment results obtained from the operation of the receptor model and the actual contribution rates of the pollution sources in the synthetic data set, the accuracy of each receptor model can be compared, and the accuracy of the source apportionment results of the receptor model for road sediments can be quantified and evaluated
[0090] Exemplarily, as shown in Table 1, according to existing research, for the judgment and selection of urban road sediment pollution sources, brake wear (BW), vehicle exhaust (VE), tire wear (TW), and road surrounding soil (RS) are selected as the main pollution sources. The metal element concentration data of the source component spectra of these four pollution sources in the existing research are averaged and multiplied by a random number randomly generated between 0.2 and 0.8 to obtain the source component spectrum of the unknown source (UnS) before normalization. The source component spectra of the four real known pollution sources and one unknown pollution source are subjected to mass normalization processing, and the obtained source component spectra are shown in Table 2 below, including 19 metal elements: Pb, Zn, Cr, Cu, Ni, Fe, Mn, Ba, Na, K, Mg, Ca, Al, Cd, Sb, Co, Sn, Mo, and Sr, obtaining 5×19 data, denoted as A. Then the source component spectra of the known pollution sources and the unknown pollution source before normalization are shown in Table 1 below
[0091] Table 1
[0092] Before normalization BW VE TW RS UnS Pb 152.2309 499.2755 17.6959 37.7865 121.7495 Zn 6289.8393 9387.8817 2050.1772 65.6244 3307.2605 Cr 856.1989 22029.4453 1.7103 14.1412 1581.3030 Cu 28941.4582 3009.5111 7.4758 8.7061 5978.0612 Ni 139.8922 10242.5283 2.4438 11.9021 1506.0120 Fe 254618.8051 139842.9427 169.2343 22453.0923 26956.5861 Mn 2605.6781 13282.8676 4.0369 430.4262 1498.0398 Ba 6512.8363 1248.0949 5.0180 84.9609 1036.5732 Na 1337.3546 4480.7808 471.7219 330.9283 1281.9564 K 4253.7440 6642.3245 189.9953 4121.6820 2961.4542 Mg 5132.0784 4230.1195 88.0469 1620.5969 815.2785 Ca 18847.1646 36749.9050 477.3804 4320.7732 11812.6363 Al 9897.5888 43604.3538 343.2844 68752.2676 23731.8152 Cd 6.4998 0.3804 1.8735 0.0565 1.0819 Sb 195.0762 513.3319 0.2655 5.7377 121.4800 Co 37.2553 268.3863 1.0424 7.2677 22.3794 Sn 873.2989 74.7717 1.0187 2.5071 107.7818 Mo 1472.9887 474.1755 0.7834 1.1617 365.1858 Sr 610.8831 175.4944 1.3736 21.2053 136.5770
[0093] Reference Figure 2 and the source component spectra of the known pollution sources and the unknown pollution source after normalization are as Figure 2As shown, correspondingly, the source component spectra of the normalized known pollution sources and unknown pollution sources are as shown in Table 2 below:
[0094] Table 2
[0095]
[0096]
[0097] In some embodiments, it is intended to generate 100 synthetic road sediments. According to the sampling situation of actual road sediments, it is assumed that every 1 m 2The mass of the collected road sediment sample is 20 g. The 20 g of road sediment contains 935.90 mg of metal elements, and the standard deviation is 270.55 mg. Using the two data of 935.90 and 270.55, 100 total mass concentration data of metal elements in synthetic road sediments are randomly generated in a lognormal distribution in Matlab. To simulate real-world situations, two different urban road traffic scenarios are established in this embodiment, namely, a high traffic flow scenario and a low traffic flow scenario. Assuming that the impacts of brake wear (BW), vehicle exhaust (VE), and tire wear (TW) are more obvious on high traffic flow roads, while the contribution of road surrounding soil (RS) is relatively low. Therefore, the contributions of BW, VE, TW, and RS are set to approximately 20%, 25%, 35%, and 20% respectively. On the contrary, on low traffic flow roads, the contributions of brake wear, vehicle exhaust emissions, and tire wear decrease, and the contribution of road surrounding soil becomes more prominent. Therefore, the contribution rates of BW, VE, TW, and RS are set to 10%, 20%, 30%, and 40% respectively. The contribution rates of unknown pollution sources are set to 10%, 15%, and 20%. As the unknown pollution source increases, the contribution rates of the four real pollution sources decrease proportionally. Multiply the 100 total mass concentration data of metal elements in the synthetic road sediments randomly generated in a lognormal distribution by the contribution rates of the five pollution sources to obtain 5×100 metal element data (i.e., contribution masses) of pollution sources with the same contribution rates. Calculate the average value and standard deviation for these data respectively. Among them, the contribution rates of the five pollution sources refer to the contribution rates of BW, VE, TW, and RS and the contribution rate of one set unknown pollution source. According to the average value and standard deviation, 100 pollution source contribution mass concentrations of synthetic road sediments with different pollution source contribution rates are randomly generated in a lognormal distribution in Matlab. Finally, 100×5 total mass concentration data of metal elements of pollution sources with random contribution rates are obtained, denoted as B. The metal element concentration C of the synthetic road sediment = A×B. Because there are two urban road traffic scenarios and three contribution rates of unknown pollution sources, a total of 6 synthetic data sets of 100×19 are finally generated, numbered SD1 - SD6. SD1 is the high traffic flow scenario with a 10% unknown source contribution, SD2 is the low traffic flow scenario with a 10% unknown source contribution, SD3 is the high traffic flow scenario with a 15% unknown source contribution, SD4 is the low traffic flow scenario with a 15% unknown source contribution, SD5 is the high traffic flow scenario with a 20% unknown source contribution, and SD6 is the low traffic flow scenario with a 20% unknown source contribution. The following Table 3 shows the contribution rates of each pollution source of the 6 generated synthetic data sets.
[0098] Table 3
[0099]
[0100]
[0101] The six generated synthetic datasets are respectively input into four receptor models, namely PMF, Unmix, CMB, and SCMD, and the obtained source contribution results are as Figure 3 , Figure 4 shown.
[0102] To quantify the errors of the four models, in this embodiment, relative error (RE), relative prediction error (RPE), and symmetric mean absolute percentage error (SMAPE) are used for evaluation. The calculation formulas of the three error metrics are as follows:
[0103]
[0104]
[0105]
[0106] where p is the number of pollution sources; Y calculated is the model calculation result; Y SD is the data of the synthetic dataset. In the case of perfect fitting between the measured value and the calculated value, RE = 0, RPE = 0, and SMAPE = 0. The error table of the operation results of the four models is shown in Table 4 below.
[0107] Table 4
[0108]
[0109] The errors of the source-unknown models PMF and Unmix are significantly higher than those of the source-known models CMB and SCMD. Except for the RE and SMAPE values of SD1, the error values of the SCMD source tracing results are better than those of CMB. As the contribution of unknown sources increases, the RE, RPE, and SMAPE of SCMD and CMB all increase, but the stability and accuracy of the SCMD results are significantly higher than those of CMB. Therefore, it is considered that SCMD is a more suitable receptor model for tracing urban road sediments among the four models, and its source tracing results are also more accurate.
[0110] This embodiment of the present invention also provides a receptor model evaluation device based on a synthetic dataset of road sediments, which can implement the above-mentioned receptor model evaluation method based on a synthetic dataset of road sediments. The device includes:
[0111] The first module is used to obtain the known pollution source component spectra;
[0112] The second module is used to obtain the unknown pollution source component spectra according to the known pollution source component spectra;
[0113] The third module is used to randomly generate the mass concentrations of the pollution source contributions of several road sediments in a lognormal manner;
[0114] The fourth module is used to generate a synthetic dataset according to the known pollution source component spectrum, the unknown pollution source component spectrum, and the mass concentration of the pollution source contribution;
[0115] The fifth module is used to input the synthetic dataset into the receptor model to obtain the source tracing result of the receptor model and the pollution source contribution rate in the synthetic dataset;
[0116] The sixth module is used to obtain the accuracy of the receptor model according to the source tracing result and the pollution source contribution rate in the synthetic dataset.
[0117] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0118] An embodiment of the present invention also provides an electronic device, which includes a processor and a memory. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned method for evaluating a receptor model based on a synthetic dataset of road sediments. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0119] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0120] Refer to Figure 6 , Figure 6 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0121] A processor 701, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention;
[0122] The memory 702 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 702 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 702 and are called by the processor 701 to execute a receptor model evaluation method based on a synthetic dataset of road sediments according to an embodiment of the present invention;
[0123] The input / output interface 703 is used to implement information input and output;
[0124] The communication interface 704 is used to implement communication interaction between this device and other devices. Communication can be achieved through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);
[0125] The bus 705 transmits information between various components of the device (such as the processor 701, the memory 702, the input / output interface 703, and the communication interface 704);
[0126] Among them, the processor 701, the memory 702, the input / output interface 703, and the communication interface 704 achieve communication connections with each other inside the device through the bus 705.
[0127] An embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned receptor model evaluation method based on a synthetic dataset of road sediments.
[0128] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0129] An embodiment of the present invention also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned receptor model evaluation method based on a synthetic dataset of road sediments.
[0130] In summary, a receptor model evaluation method and device according to an embodiment of the present invention have the following advantages:
[0131] Receptor models are rarely used in the field of urban road sediment source tracing, and the accuracy of the tracing results of these receptor models is not clear. For example Figure 5 As shown, in the embodiments of the present invention, first, the composition spectra of real pollution sources are obtained, the composition spectra of unknown pollution sources are randomly generated, and all are subjected to mass normalization processing. Then, the contribution values of the pollution sources are randomly generated logarithmically normally. Finally, a synthetic data set is calculated and input into the receptor model, so as to compare the accuracy of each model and quantify and evaluate the model. The synthetic data set generated by using the method of this embodiment can more intuitively obtain the accuracy of the receptor model tracing results, and provide guidance for the application of the receptor model in the field of urban road sediment source tracing.
[0132] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.
[0133] In addition, although the present invention has been described in the context of functional modules, it should be understood that unless otherwise stated to the contrary, one or more of the functions and / or features described may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0134] If the above-described functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., all of which can store program codes.
[0135] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0136] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), optical fiber devices, and portable compact disc read-only memories (CDROMs). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0137] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0138] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0139] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
[0140] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. A receptor model evaluation method based on a synthetic dataset of road sediments, characterized in that, Including the following steps: Obtain the known pollution source component spectra; According to the known pollution source component spectra, obtain the unknown pollution source component spectra; Randomly generate several source contribution mass concentrations of road sediments by lognormal distribution; According to the known pollution source component spectra, the unknown pollution source component spectra, and the source contribution mass concentrations, generate a synthetic dataset; Input the synthetic dataset into the receptor model to obtain the source tracing results of the receptor model and the source contribution rates in the synthetic dataset; According to the source tracing results and the source contribution rates in the synthetic dataset, obtain the accuracy of the receptor model; Among them, the obtaining of the known pollution source component spectra includes the following steps: Obtain the component spectrum data of the actual pollution sources of road sediments; Perform normalization processing on the component spectrum data to obtain the known pollution source component spectra, and the calculation formula is: where D ik represents the known source profile; x 1k , x 2k , …, x ik respectively represent the k-th known source data of the first element, the k-th known source data of the second element, ……, the k-th known source data of the i-th element in the known source profile; The obtaining of the unknown pollution source component spectra according to the known pollution source component spectra includes the following steps: According to the known pollution source component spectra, obtain the concentration average values of all known pollution sources; Multiply the concentration average values by random numbers to obtain the concentration data of unknown pollution sources; Perform normalization processing on the concentration data of unknown pollution sources to obtain the unknown pollution source component spectra, and the calculation formula is: where, E iuns represents the composition spectrum of the unknown pollution source; x 1uns , x 2uns , …, x iuns represent the concentration data of the unknown pollution source of the first element, the concentration data of the unknown pollution source of the second element, ……, the concentration data of the unknown pollution source of the i-th element in the composition spectrum of the unknown pollution source, respectively.
2. The receptor model evaluation method based on the synthetic data set of road sediments according to claim 1, wherein, The generating of the synthetic dataset according to the known pollution source component spectra, the unknown pollution source component spectra, and the source contribution mass concentrations includes the following steps: According to the known pollution source component spectra and the unknown pollution source component spectra, generate the metal element pollution source component spectra in the road sediments; According to the metal element pollution source component spectra, obtain the mass fractions of several metals; According to the mass fractions of the metals and the source contribution mass concentrations, generate metal element concentrations; Obtain a set of the metal element concentrations of several road sediments to obtain a synthetic dataset.
3. An evaluation device for receptor models based on a synthetic dataset of road sediments, characterized in that, Including: The first module is used to obtain the known pollution source component spectra; The second module is used to obtain the unknown pollution source component spectra according to the known pollution source component spectra; The third module is used to randomly generate several source contribution mass concentrations of road sediments by lognormal distribution; The fourth module is used to generate a synthetic dataset according to the known pollution source component spectra, the unknown pollution source component spectra, and the source contribution mass concentrations; The fifth module is used to input the synthetic dataset into the receptor model to obtain the source tracing results of the receptor model and the source contribution rates in the synthetic dataset; The sixth module is used to obtain the accuracy of the receptor model according to the source tracing results and the source contribution rates in the synthetic dataset; Among them, the first module is specifically used for: Obtain the component spectrum data of the actual pollution sources of road sediments; Perform normalization processing on the component spectrum data to obtain the known pollution source component spectra, and the calculation formula is: where D ik represents the known source composition spectrum; x 1k , x 2k , …, x ik respectively represent the k-th known source data of the first element, the k-th known source data of the second element, ……, the k-th known source data of the i-th element in the known source composition spectrum; The second module is specifically used for: According to the known pollution source component spectra, obtain the concentration average values of all known pollution sources; Multiply the concentration average values by random numbers to obtain the concentration data of unknown pollution sources; Normalize the concentration data of the unknown pollution source to obtain the component spectrum of the unknown pollution source. The calculation formula is as follows: where, E iuns represents the composition spectrum of the unknown pollution source; x 1uns , x 2uns , …, x iuns respectively represent the concentration data of the unknown pollution source of the first element, the concentration data of the unknown pollution source of the second element, ……, the concentration data of the unknown pollution source of the i-th element in the composition spectrum of the unknown pollution source.
4. An electronic device, characterized in that, including a processor and a memory; the memory is used to store programs; the processor executes the program to implement the method according to any one of claims 1 to 2.
5. A computer-readable storage medium, characterized in that, The storage medium stores a program, and the program is executed by the processor to implement the method according to any one of claims 1 to 2.
6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Atmospheric pollution monitoring and management method as well as system based on high-density deployment of sensors
CN105181898A
Soil heavy metal pollutant traceability and concentration prediction method, equipment and medium
CN114548598A