Health effect-oriented atmospheric particulate pollution source management and control scheme design method
By integrating data such as isotope ratios and meteorological parameters, and combining machine learning and Bayesian algorithms, the health impact of pollution sources is quantified, solving the problem of identifying the contribution of atmospheric particulate matter (PM) emission sources to health effects, and enabling the formulation of precise pollution source control and prevention strategies.
Patent Information
- Application Number
- CN202510349830.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies lack a clear identification of the health effects of PM emissions, making it impossible to achieve precise pollution source control guided by health effects, thus limiting the effectiveness and targeted nature of air pollution prevention and control efforts.
By integrating isotope ratios, meteorological parameters, ecological and environmental factors, socioeconomic parameters, and health data, an isotope ratio prediction model is established using machine learning algorithms. Combined with Bayesian and Monte Carlo algorithms, the contribution of different pollution sources to health impacts is quantified, and a health-effect-oriented precision control plan is formulated.
It enables precise tracking and analysis of PM health effect emission sources, provides health effect-oriented air pollution source control solutions, supports the scientific formulation of prevention and control strategies, and improves the pertinence and effectiveness of pollution prevention and control.
Smart Images

Figure CN120878249A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of atmospheric pollution source tracing technology, specifically to a design method for atmospheric particulate matter pollution source control schemes based on isotope ratios to achieve health effect orientation. Background Technology
[0002] According to the latest Global Burden of Disease (GBD) study, air pollution has become the leading cause of global health burden, especially fine particulate matter (PM2.5). 2.5 The impact of air pollution on human health is becoming increasingly serious. The incidence and mortality rates of respiratory and cardiovascular diseases caused by air pollution continue to rise, posing a major challenge that public health systems worldwide urgently need to address. my country's air pollution control is at a critical stage of transformation, shifting from a traditional, simple quality-standard-meeting model to a health-effect-oriented control strategy. Currently, control measures for particulate matter (PM) are mainly based on the overall pollution level of PM, lacking analysis of the public health effects of specific components. That is, the contribution of emission sources to the health effects of PM is unclear, making it impossible to accurately carry out health-effect-oriented PM source control. This traditional control strategy cannot accurately identify and precisely intervene in the specific threats to public health posed by pollutants from different sources, limiting the effectiveness and targeting of air pollution control efforts. This study constructs a direct link between PM source emissions and health outcomes based on isotope ratio data of a certain component (or element) in PM and its emission sources. Utilizing the high precision and sensitivity of isotope tracing technology, the sources and composition of air pollutants can be accurately tracked and analyzed, further revealing the specific contributions of different pollution sources to health effects. This research will provide a theoretical basis for health-effect-oriented management of air pollution sources and promote the scientific transformation of air pollution prevention and control strategies. Summary of the Invention
[0003] This application provides a method for designing a health-effect-oriented atmospheric particulate matter pollution source control scheme based on isotope ratio data. The aim is to clarify the contribution of emission sources to the health effects of PM, with the core being the establishment of a direct link between PM source emissions and health outcomes, thereby achieving health-effect-oriented PM source control.
[0004] To solve the above-mentioned technical problems, the technical solution proposed in this application is as follows:
[0005] In a first aspect, the present invention provides a method for designing a health-effect-oriented atmospheric particulate matter pollution source control scheme, comprising the following steps:
[0006] Step 1: Integrate isotope ratios, meteorological parameters, ecological and environmental factors, socioeconomic parameters, and health data, and perform one-to-one spatiotemporal matching. Then, establish an isotope ratio prediction model using machine learning algorithms.
[0007] ISO=f(x,θ) (1)
[0008] In the formula, ISO represents the predicted isotope ratio; f represents the model's prediction function, the specific form of which depends on the machine learning algorithm used; x represents the vector of independent variables, including time, geographical location, meteorological parameters, ecological and environmental factors, and socio-economic parameters; θ represents the model parameters, the specific form of which depends on the type of model.
[0009] Step 2: Establish a quasi-Poisson distribution generalized additive model of isotope ratios with population disease burden and mortality, and analyze the changes in health status corresponding to unit isotope changes.
[0010] logE(Y i )=β0+f1(ISO)+f2(X1)+f3(X2)+…(2)
[0011] In the formula E(Y) i ) represents the expected value of the number of deaths or incidence rate Y for disease i; β0 represents the constant term (intercept) of the model; f1(ISO) represents the regression coefficient used to capture Y. i The nonlinear relationship between the isotope ratio (ISO) and the isotope ratio (ISO); f1(X1) and f2(X2) represent the natural spline functions used to process meteorological parameters X1 and X2.
[0012] Step 3: A Bayesian algorithm coupled with a Monte Carlo algorithm is used to analyze the sources of health risks, quantify the contribution of different pollution sources to health impacts, and then formulate a health-effect-oriented precision control plan.
[0013]
[0014] f i Indicates the contribution of a typical emission source; i refers to the serial number of the emission source;
[0015] In the formula, coal, dust, vehicle, bio., and ind. correspond to five common sources of atmospheric particulate matter: coal combustion, dust, vehicle exhaust, biomass combustion, and industrial emissions.
[0016] Preferably, in step 1, the atmospheric particulate matter components or elements that can be used to achieve this method include: organic carbon (OC), inorganic carbon (EC), and nitrate (NO3). - ), ammonium ions (NH4) + ), sulfate (SO4 2- ), silicon (Si), iron (Fe), nickel (Ni), copper (Cu), zinc (Zn), strontium (Sr), neodymium (Nd), hafnium (Hf), lead (Pb), and mercury (Hg).
[0017] Preferably, in step 1, the isotope types corresponding to different components or elements include: δ 13 C and f M -14C (used for organic and inorganic carbon), δ 15 N (used for nitrate and ammonium ions), δ 34 S (for sulfate), δ 30 Si (used for silicon), δ 56 Fe (used in iron), δ 60 Ni (used for nickel), δ 65 Cu (used in copper), δ 66 Zn (used in zinc), δ 87 Sr (used for strontium), δ 144 Nd (used for neodymium), δ 177 Hf (used for hafnium) 207 Pb / 206 Pb (used for lead) and δ 202 Hg (used for mercury).
[0018] Preferably, in step 1, the meteorological parameters used to establish the isotope ratio prediction model include temperature, relative humidity, precipitation, wind speed, and air pressure; the ecological environment factors include enhanced vegetation index, normalized difference vegetation index, gross primary productivity, and net primary productivity; and the socio-economic parameters include gross domestic product and population.
[0019] Preferably, in step 1, the spatial resolution of isotope ratios, meteorological parameters, ecological environment factors, socioeconomic parameters, and health data includes 0.1°×0.1°, 0.01°×0.01°, and 1km×1km, and the temporal resolution includes hourly, daily, monthly, and yearly data.
[0020] Preferably, in step 1, the machine learning algorithm model used for isotope ratio prediction includes extreme gradient boosting, lightweight gradient boosting machine, random forest, and support vector machine.
[0021] Preferably, in step 2, the method for obtaining population disease burden and mortality data is as follows: establish a population cohort, and collect individual health, lifestyle, and environmental exposure data through follow-up studies. The data includes basic demographic information, physiological and biochemical indicators, medical records, genetic information, and personal behavioral habits.
[0022] Preferably, in step 2, the population disease burden and mortality data include the incidence rate, number of deaths, and mortality rate of different disease types.
[0023] Preferably, in step 2, the disease types can be classified according to organ systems into respiratory system diseases, circulatory system diseases, genitourinary system diseases, nervous system diseases, and digestive system diseases.
[0024] Preferably, in step 2, the gam() function is used to construct a generalized additive model for the datasets of each region, which is used to model the nonlinear relationship between the number of deaths from the disease and the isotope ratio.
[0025] Preferably, gam() adopts a Poisson distribution and uses log as the link function to adapt to the characteristics of continuous data and response variables. In the model, the number of deaths from disease is processed by the smoothing function s(), and meteorological parameters are processed by the natural spline function ns() to capture their nonlinear relationship with isotope ratios.
[0026] Preferably, in step 3, if there is a lack of prior values for the proportion of emission sources in the Bayesian model, an information-free prior can be used; atmospheric particulate matter and the metal fingerprints of emission sources are used as "data after mixing of multi-source emissions" and "initial data of source emissions", respectively; the running length parameter of the Monte Carlo Markov chain is set to "very long", the number of iterations is 1 million, and the number of aging times is 500,000.
[0027] According to a second aspect of the present invention, an electronic device is claimed, comprising: one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the processor to implement the method described in any of the preceding claims.
[0028] According to a third aspect of the present invention, the present invention claims protection for a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps of the method described in any of the preceding claims.
[0029] Compared with existing technologies, the present invention's method for designing a health-effect-oriented atmospheric particulate matter pollution source control scheme based on isotope ratio data achieves the following beneficial technical effects:
[0030] This invention combines the source tracing advantages of isotope fingerprinting with population health effects, quantitatively analyzing the nonlinear relationship between isotope ratios and population health effects, and enabling the design of health-effect-oriented atmospheric particulate matter pollution source control schemes. This method can be flexibly configured according to needs, providing technical support for the scientific formulation of health-effect-oriented atmospheric pollution prevention and control strategies. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 A flowchart illustrating the design method for a health-effect-oriented atmospheric particulate matter pollution source control scheme provided in this embodiment of the invention.
[0033] Figure 2 The map showing the distribution of emissions contribution from different pollution sources in the Beijing-Tianjin-Hebei region of China from 2016 to 2020 is provided for the purposes of this invention.
[0034] Figure 3 This is a structural module diagram of the computing device for a health-effect-oriented atmospheric particulate matter pollution source control scheme provided in an embodiment of the present invention. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] This invention provides a health-effect-oriented design method for controlling atmospheric particulate matter pollution sources. Based on isotope ratio data of a certain component (or element) in atmospheric particulate matter and its emission sources, a direct link between PM emission and health outcomes is established, thereby achieving health-effect-oriented PM source control. This is described in detail below.
[0037] This invention comprises three parts: integrating isotope ratios, meteorological parameters, ecological and environmental factors, socioeconomic parameters, and health data, and performing one-to-one spatiotemporal matching; establishing an isotope ratio prediction model using machine learning algorithms; establishing a quasi-Poisson distribution generalized additive model of isotope ratios and population disease burden and mortality rates to analyze changes in health status corresponding to unit isotope variations; and employing a Bayesian algorithm coupled with a Monte Carlo algorithm to perform health risk tracing analysis, quantifying the contribution of different pollution sources to health impacts, and thereby formulating a health effect-oriented precision control plan.
[0038] The isotope ratio of atmospheric particulate matter components (or elements) is usually expressed as the deviation (‰) of the isotope ratio of the sample relative to the standard substance, as shown in equation (4).
[0039]
[0040] Where E represents the selected element, and x and y represent the mass numbers of isotopes of element E.
[0041] In particular, Pb isotope ratio data are typically not expressed as a thousandth of a degree of deviation relative to standard isotope ratios, but rather as the ratios of different types of isotopes directly, for example... 207 Pb / 206 Pb, 208 Pb / 207 Pb.
[0042] for 14 This special class of radioactive isotopes uses the contemporary carbon fraction (f... M ) is used to represent, as in equation (5).
[0043]
[0044] in,( 14 C / 12 C) sample and( 14 C / 12 C) standard These represent the isotope ratios of the sample and the standard (NIST oxalate II 4990C), respectively.
[0045] Step 1: Integrate isotope ratios, meteorological parameters, ecological and environmental factors, socioeconomic parameters, and health data, and perform one-to-one spatiotemporal matching. Then, establish an isotope ratio prediction model using machine learning algorithms.
[0046] ISO=f(x,θ) (1)
[0047] In the formula, ISO represents the predicted isotope ratio; f represents the model's prediction function, the specific form of which depends on the machine learning algorithm used; x represents the vector of independent variables, including time, geographical location, meteorological parameters, ecological and environmental factors, and socio-economic parameters; θ represents the model parameters, the specific form of which depends on the type of model.
[0048] The method for obtaining atmospheric particulate matter isotopic fingerprint data is as follows: relevant research articles published in the database are retrieved, and the isotopic fingerprint data and other relevant information (such as particulate matter type, sampling time, sampling location, etc.) are extracted from them.
[0049] The method for spatiotemporal matching of isotope ratios, meteorological parameters, ecological and environmental factors, socioeconomic parameters, and health data is as follows: Taking a spatial resolution of 0.1°×0.1° as an example, ArcGIS is used to populate the data into 6,480,000 rasters globally, and each raster is assigned a unique Object ID. Data with the same time coordinates and Object IDs are concatenated to form a new dataset, thereby achieving one-to-one matching of spatiotemporal locations.
[0050] The time coordinate is the decimal form of the date of atmospheric particulate matter sampling, as shown in equation (6).
[0051]
[0052] Where Y represents the year in which atmospheric particulate matter was sampled, DoY represents the day of the year in which the sampling occurred, and NoY represents the total number of days in the sampling year.
[0053] To improve the spatiotemporal resolution of atmospheric particulate matter isotope ratio data, four machine learning algorithms—limited gradient boosting, lightweight gradient boosting machine, random forest, and support vector machine—were employed for isotope ratio prediction. Limited gradient boosting and lightweight gradient boosting machine, based on gradient boosting decision trees, efficiently handle large-scale data and capture complex nonlinear relationships. Random forest uses multiple decision trees to construct an ensemble model, exhibiting strong anti-overfitting capabilities and is suitable for high-dimensional feature data. Support vector machine maps to a high-dimensional space through kernel functions, making it suitable for small sample sizes and nonlinear regression tasks. Model training utilizes an ensemble model of the four machine learning algorithms. This involves exhaustively searching to find the optimal combination of the four algorithms within the ensemble model, with the proportions of limited gradient boosting, lightweight gradient boosting machine, random forest, and support vector machine being 0.25, 0.25, 0.5, and 0.5 respectively, achieving R² of the inherited model. 2 Both the RMSE and RMSE are optimal. Taking the random forest algorithm as an example, the collected isotope ratio data are divided into training and testing sets in a 7:3 ratio, and the cross-multiplication method is used for validation. After the model is trained, meteorological parameters, ecological environment factors, socio-economic parameters, and other data are input to complete the missing isotope ratio data at the corresponding spatiotemporal locations, thereby improving the spatiotemporal resolution of the isotope ratio data.
[0054] Step 2: Establish a quasi-Poisson distribution generalized additive model of isotope ratios with population disease burden and mortality, and analyze the changes in health status corresponding to unit isotope changes.
[0055] logE(Y i )=β0+f1(ISO)+f1(X1)+f2(X2)+…(2)
[0056] In the formula E(Y) i ) represents the expected value of the number of deaths or incidence rate Y for disease i; β0 represents the constant term (intercept) of the model; f1(ISO) represents the regression coefficient used to capture Y. i The nonlinear relationship between the isotope ratio (ISO) and the isotope ratio (ISO); f1(X1) and f2(X2) represent natural spline functions for processing meteorological parameters X1 and X2.
[0057] The method for obtaining population disease burden and mortality data is as follows: establish a population cohort, and collect individual health, lifestyle, and environmental exposure data through follow-up studies. The data includes basic demographic information, physiological and biochemical indicators, medical records, genetic information, and personal behavioral habits.
[0058] Population disease burden and mortality data include incidence, number of deaths, and mortality rates for different disease types.
[0059] Diseases can be classified according to organ system (ICD-10 criteria) into respiratory system diseases, circulatory system diseases, genitourinary system diseases, nervous system diseases, and digestive system diseases.
[0060] The `gam()` function was used to construct a generalized additive model for datasets from various regions, modeling the nonlinear relationship between disease-related deaths and isotope ratios. Specifically, `gam()` employed a Poisson distribution and used `log` as the link function to accommodate the characteristics of continuous data and the response variable. In the model, disease-related deaths were processed using the smoothing function `s()`, and meteorological and environmental factors were processed using the natural spline function `ns()` to capture their nonlinear relationship with isotope ratios.
[0061] Step 3: A Bayesian algorithm coupled with a Monte Carlo algorithm is used to analyze the sources of health risks, quantify the contribution of different pollution sources to health impacts, and then formulate a health-effect-oriented precision control plan.
[0062]
[0063] f i This indicates the contribution of a typical emission source; i refers to the serial number of the emission source.
[0064] In the formula, coal, dust, vehicle, bio., and ind. correspond to five common sources of atmospheric particulate matter: coal combustion, dust, vehicle exhaust, biomass combustion, and industrial emissions. The types and quantities of emission sources can be adjusted according to the actual emission sources.
[0065] In the Bayesian model, if there is a lack of prior values for the proportion of emission sources, an information-free prior can be used; atmospheric particulate matter and the metal fingerprints of emission sources are used as "data after mixing of multi-source emissions" and "initial data of source emissions", respectively; the running length parameter of the Monte Carlo Markov chain is set to "very long", the number of iterations is 1 million, and the number of aging times is 500,000.
[0066] In another embodiment of this application, referring to the figures, this application also provides an electronic device, specifically:
[0067] The electronic device may include components such as a processor with one or more processing cores, a memory with one or more computer-readable storage media, a power supply, and an input unit.
[0068] in:
[0069] The processor is the control center of the electronic device. It connects all parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in memory, and by calling data stored in memory, thereby providing overall monitoring of the electronic device. Optionally, the processor may include one or more processing cores; the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. Preferably, the processor may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may also not be integrated into the processor.
[0070] Memory can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. Memory can primarily include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area can store data created based on the use of the electronic device. Furthermore, memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory can also include a memory controller to provide the processor with access to the memory.
[0071] The electronic device also includes a power supply for powering the various components. Preferably, the power supply can be connected to the processor logic through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply may also include one or more DC or AC power sources, a recharging system, a power fault detection circuit, a power converter or inverter, a power status indicator, or any other components.
[0072] The electronic device may also include an input unit, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0073] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory according to the following instructions, and the processor runs the applications stored in the memory to realize various functions.
[0074] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0075] In some embodiments of this application, a computer-readable storage medium is also provided, which may include: a read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc. A computer program is stored thereon, which is loaded by a processor to execute the steps in the order processing method provided in the embodiments of this application.
[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for designing a health-effect-oriented atmospheric particulate matter pollution source control scheme, characterized in that, Includes the following steps: Step 1: Integrate isotope ratios, meteorological parameters, ecological and environmental factors, socioeconomic parameters, and health data, and perform one-to-one spatiotemporal matching. Then, establish an isotope ratio prediction model using machine learning algorithms. ISO=f(x,θ) (1) In the formula, ISO represents the predicted isotope ratio; f represents the model's prediction function, the specific form of which depends on the machine learning algorithm used; x represents the vector of independent variables, including time, geographical location, meteorological parameters, ecological and environmental factors, and socio-economic parameters; θ represents the model parameters, the specific form of which depends on the type of model. Step 2: Establish a quasi-Poisson distribution generalized additive model of isotope ratios with population disease burden and mortality, and analyze the changes in health status corresponding to unit isotope changes. logE(Y i )=β0+f1(ISO)+f2(X1)+f3(X2)+…(2) In the formula E(Y) i ) represents the expected value of the number of deaths or incidence rate Y for disease i; β0 represents the constant term (intercept) of the model; f1(ISO) represents the regression coefficient used to capture Y. i The nonlinear relationship between the isotope ratio (ISO) and the isotope ratio (ISO); f1(X1) and f2(X2) represent the natural spline functions used to process meteorological parameters X1 and X2. Step 3: A Bayesian algorithm coupled with a Monte Carlo algorithm is used to analyze the sources of health risks, quantify the contribution of different pollution sources to health impacts, and then formulate a health-effect-oriented precision control plan. f i Indicates the contribution of a typical emission source; i refers to the serial number of the emission source; In the formula, coal, dust, vehicle, bio., and ind. correspond to five common sources of atmospheric particulate matter: coal combustion, dust, vehicle exhaust, biomass combustion, and industrial emissions.
2. The method according to claim 1, characterized in that, In step 1, the atmospheric particulate matter components or elements that can be used to achieve this method include: organic carbon (OC), inorganic carbon (EC), and nitrate (NO3). - ), ammonium ions (NH4) + ), sulfate (SO4 2- ), silicon (Si), iron (Fe), nickel (Ni), copper (Cu), zinc (Zn), strontium (Sr), neodymium (Nd), hafnium (Hf), lead (Pb), and mercury (Hg).
3. The method according to claim 1, characterized in that, In step 1, the isotope types corresponding to different components or elements include: δ 13 C and f M -14C (used for organic and inorganic carbon), δ 15 N (used for nitrate and ammonium ions), δ 34 S (for sulfate), δ 30 Si (used for silicon), δ 56 Fe (used in iron), δ 60 Ni (used for nickel), δ 65 Cu (used in copper), δ 66 Zn (used in zinc), δ 87 Sr (used for strontium), δ 144 Nd (used for neodymium), δ 177 Hf (used in hafnium) 207 Pb / 206 Pb (used for lead) and δ 202 Hg (used for mercury).
4. The method according to claim 1, characterized in that, In step 1, the meteorological parameters used to establish the isotope ratio prediction model include temperature, relative humidity, precipitation, wind speed, and air pressure; the ecological environment factors include enhanced vegetation index, normalized difference vegetation index, gross primary productivity, and net primary productivity; and the socio-economic parameters include gross domestic product and population.
5. The method according to claim 1, characterized in that, In step 1, the spatial resolution of isotope ratios, meteorological parameters, ecological and environmental factors, socioeconomic parameters, and health data includes 0.1°×0.1°, 0.01°×0.01°, and 1km×1km, while the temporal resolution includes hourly, daily, monthly, and yearly data.
6. The method according to claim 1, characterized in that, In step 1, the machine learning algorithm models used for isotope ratio prediction include extreme gradient boosting, lightweight gradient boosting machine, random forest, and support vector machine.
7. The method according to claim 1, characterized in that, In step 2, the method for obtaining population disease burden and mortality data is as follows: establish a population cohort, and collect individual health, lifestyle, and environmental exposure data through follow-up studies. The data includes basic demographic information, physiological and biochemical indicators, medical records, genetic information, and personal behavioral habits.
8. The method according to claim 1, characterized in that, In step 2, population disease burden and mortality data include incidence, number of deaths, and mortality rates for different disease types.
9. The method according to claim 1, characterized in that, In step 2, disease types can be classified according to organ system into respiratory system diseases, circulatory system diseases, genitourinary system diseases, nervous system diseases, and digestive system diseases.
10. The method according to claim 1, characterized in that, In step 2, the gam() function is used to construct a generalized additive model for the datasets of each region, which is used to model the nonlinear relationship between the number of disease deaths and the isotope ratio.
11. The method according to claim 10, characterized in that, gam() employs a Poisson distribution and uses log as the link function to accommodate the characteristics of continuous data and response variables. In the model, the number of deaths from disease is processed using the smoothing function s(), and meteorological parameters are processed using the natural spline function ns() to capture their nonlinear relationship with isotope ratios.
12. The method according to claim 1, characterized in that, In step 3, in the Bayesian model, if there is a lack of prior values for the proportion of emission sources, an information-free prior can be used; atmospheric particulate matter and the metal fingerprints of emission sources are used as "data after mixing of multi-source emissions" and "initial data of source emissions", respectively; the running length parameter of the Monte Carlo Markov chain is set to "very long", the number of iterations is 1 million, and the number of aging times is 500,000.
13. An electronic device, characterized in that, The electronic device includes: one or more processors; a memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1 to 12.
14. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to perform the steps of the method according to any one of claims 1 to 12.