Hybrid data regression model-based pm 2.5 influence factor analysis method and system

By constructing a hybrid data regression model, the problem of fusing functional and component data was solved, revealing the dynamic influencing factors of PM2.5 concentration and analyzing the impact of urban industrial structure, economic level and temperature on air quality.

CN122047706AActive Publication Date: 2026-05-15CAPITAL UNIV OF ECONOMICS & BUSINESS
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-13
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively integrate functional and component data to analyze the dynamic impact of PM2.5 concentrations on industrial structure, economic factors, and temperature.

Method used

A mixed data regression model was established. Through equidistant logarithmic ratio transformation, function basis expansion, and iterative reweighted least squares method, a mixed data regression model was constructed with component data and numerical data as covariates and functional data as dependent variables. The dynamic impact of the proportion of the three industries, per capita GDP, and temperature on PM2.5 in various cities was analyzed.

Benefits of technology

It enables precise analysis of the dynamic influencing factors of PM2.5 concentration, revealing how industrial structure, economic level, and temperature dynamically affect air quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122047706A_ABST
    Figure CN122047706A_ABST
Patent Text Reader

Abstract

The invention provides a mixed data regression model-based pm2.5 influence factor analysis method and system, and the method comprises the steps: building a mixed data regression model of which covariables are component data and numerical data and dependent variables are functional data: the mixed data regression model is a pm2.5 concentration curve of a city, is the functional data, is the proportion of first yield, second yield and third yield of the city, and is called the proportion of third yield for short; the data is component data, logarithm of per capita GDP, average temperature and numerical data, is a to-be-estimated component type coefficient changing along with time, is distributed to a component type covariable at any moment, is a to-be-estimated function type coefficient and is a function type residual error; obtaining robust M-estimation based on equidistant logarithmic ratio transformation, functional basis expansion and an iterative reweighted least square method; and according to the estimated values, analyzing the influence of the three-yield ratio, the per capita GDP and the average temperature of each city on the pm 2.5. The method can be used for analyzing the pm 2.5 influence factors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of PM2.5 influencing factor analysis technology, and specifically refers to a method and system for PM2.5 influencing factor analysis based on a mixed data regression model. Background Technology

[0002] Many factors influence PM2.5 concentrations in various cities. Among them, economic factors such as industrial structure (the proportion of primary, secondary, and tertiary industries, referred to as the tertiary industry ratio) and per capita GDP, as well as natural factors such as temperature, are particularly important. However, PM2.5 concentration curves are function-based data, industrial structure is component-based data, and per capita GDP and temperature are numerical data. How to integrate these different types of data for effective analysis is an urgent problem to be solved. Summary of the Invention

[0003] To address the technical problems existing in the prior art, this invention provides a method and system for analyzing PM2.5 influencing factors based on a mixed data regression model. The technical solution is as follows: On the one hand, a method for analyzing PM2.5 influencing factors based on a mixed data regression model is provided, which includes: S1. Collect data on the output value of the tertiary industry, per capita GDP, average temperature, and daily average PM2.5 concentration of various cities in 2023; S2. Convert the output value of the three industries into component data, take the average of the daily average PM2.5 concentration every seven days to obtain the weekly average PM2.5 concentration, and use a third-order B-spline function to fit the weekly average PM2.5 concentration to obtain the PM2.5 concentration curve. S3. Establish a mixed data regression model with covariates consisting of component data and numerical data, and dependent variable consisting of functional data: ,in, For the city The PM2.5 concentration curve is function-based data. For the city The proportions of primary, secondary, and tertiary industries, referred to as the tertiary industry ratio, are component data. For the city The logarithm of GDP per capita is a numerical data point. For the city The average temperature is numerical data. It is a time-varying component coefficient to be estimated, which is at any given time... Assigned to constituent covariates A coefficient , , The coefficients of the function to be estimated are: It is a functional residual. This refers to the inner product operation in simplex space. The calculation results are functional data; S4. Based on the equidistant logarithmic ratio transformation, function basis expansion, and iterative reweighted least squares method, we obtain... , , Robust M-estimation; S5, according to , , The estimated values ​​were used to analyze the dynamic impact of the proportion of the three industries, per capita GDP, and average temperature on PM2.5 in various cities.

[0004] On the other hand, a PM2.5 influencing factor analysis system based on a mixed data regression model is provided, the system comprising: The data collection module is used to collect data on the output value of the tertiary industry, per capita GDP, average temperature, and daily average PM2.5 concentration of various cities in 2023. The preprocessing module is used to convert the output value of the three industries into component data, take the average of the daily average PM2.5 concentration every seven days to obtain the weekly average PM2.5 concentration, and use a third-order B-spline function to fit the weekly average PM2.5 concentration to obtain the PM2.5 concentration curve. The module is used to build a mixed data regression model where the covariates are component data and numerical data, and the dependent variable is functional data. ,in, For the city The PM2.5 concentration curve is function-based data. For the city The proportions of primary, secondary, and tertiary industries, referred to as the tertiary industry ratio, are component data. For the city The logarithm of GDP per capita is a numerical data point. For the city The average temperature is numerical data. It is a time-varying component coefficient to be estimated, which is at any given time... Assigned to constituent covariates A coefficient , , The coefficients of the function to be estimated are: It is a functional residual. This refers to the inner product operation in simplex space. The calculation results are functional data; The estimation module is used to obtain the results based on the equidistant logarithmic ratio transformation, function basis expansion, and iterative reweighted least squares method. , , Robust M-estimation; Analysis module, used to analyze based on , , The estimated values ​​were used to analyze the dynamic impact of the proportion of the three industries, per capita GDP, and average temperature on PM2.5 in various cities.

[0005] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the above-described method for analyzing PM2.5 influencing factors based on a mixed data regression model.

[0006] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, the at least one instruction being loaded and executed by a processor to implement the above-described method for analyzing PM2.5 influencing factors based on a mixed data regression model.

[0007] The beneficial effects of the technical solution provided by this invention include at least the following: This invention establishes a mixed data regression model with component data and numerical data as covariates and functional data as dependent variables. Based on the logarithmic ratio transformation, function basis expansion, and iterative reweighted least squares method, the model parameters are estimated. Based on the estimated values ​​of the model parameters, the dynamic impact of the proportion of the three industries, per capita GDP, and average temperature on PM2.5 in various cities can be analyzed. This invention can effectively analyze how the industrial structure, economic level, and temperature of a city dynamically affect air quality. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a flowchart of a PM2.5 influencing factor analysis method based on a mixed data regression model provided by an embodiment of the present invention; Figure 2 This is the PM2.5 concentration curve provided in the embodiments of the present invention; Figure 3 This is provided by the embodiments of the present invention. The curve of the estimated value; Figure 4 This is provided by the embodiments of the present invention. The curve of the estimated value; Figure 5 This is provided by the embodiments of the present invention. The curve of the estimated value; Figure 6 This is a block diagram of a PM2.5 influencing factor analysis system based on a mixed data regression model provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0010] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0011] Component data focuses on the structural (or proportional) information of the research objects, such as industrial structure (the proportion of primary, secondary, and tertiary industries, referred to as the tertiary industry ratio), consumption structure, and industry income gap structure. The PM2.5 concentration curve is functional data. Establishing a regression model of the tertiary industry structure on the PM2.5 concentration curve involves component covariates and functional dependent variables. In addition, per capita GDP and temperature data that affect PM2.5 concentration are numerical data. How to integrate multiple different types of data into a linear regression model to meet the needs of practical applications is an urgent problem to be solved.

[0012] In fact, componential covariates are constrained to sum to 1 and belong to an Atchison space; functional dependent variables are infinite-dimensional and belong to a square-integrable function space. The algebraic operations that can be performed on these two types of data, which reside in two different spaces, are different. Defining a new operation to convert componential data to functional data is challenging. This invention addresses this by defining a novel coefficient similar to that of componential time series data—the functional component coefficient. Furthermore, by utilizing the Atchison inner product operation, a mixed data regression model is proposed, in which the dependent variable is functional data and the covariates include component data and numerical data. This model can be used to analyze the changes in the influence of the proportion of each component of the component covariate on the functional dependent variable over time. The mixed data regression model proposed in this embodiment of the invention, because it fully considers the data characteristics of each variable during modeling, can better characterize practical problems and obtain richer data analysis results.

[0013] Furthermore, considering the sensitivity of least squares estimation to outliers, this embodiment of the invention uses equidistant logarithmic ratio transformation, function basis expansion, and iterative reweighted least squares method to obtain robust M-estimations of model parameters. The proposed model is applied to the analysis of the impact of industrial structure on PM2.5 concentration, illustrating the working process and application value of the proposed method. The following detailed description, in conjunction with the accompanying drawings, illustrates a method for analyzing PM2.5 influencing factors based on a mixed data regression model provided by this embodiment of the invention.

[0014] This invention provides a method for analyzing PM2.5 influencing factors based on a mixed data regression model. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The diagram shown is a flowchart of the method. The processing flow may include the following steps: S1. Collect data on the output value of the tertiary industry, per capita GDP, average temperature, and daily average PM2.5 concentration of various cities in 2023; This invention collects daily average PM2.5 concentration data for various cities in 2023 from Weather Post Network, and obtains data on the output value of the tertiary industry, per capita GDP and average temperature in 2023 from the statistical yearbooks of various provinces and municipalities (this invention can also obtain these data from other sources, which is not limited and is within the protection scope of this invention).

[0015] S2. Convert the output value of the three industries into component data, take the average of the daily average PM2.5 concentration every seven days to obtain the weekly average PM2.5 concentration (to reduce irregular changes in the function curve), and use a third-order B-spline function to fit the weekly average PM2.5 concentration to obtain the PM2.5 concentration curve. Figure 2 The curve is shown (12 B-spline bases can represent the PM2.5 concentration curve well). S3. Establish a mixed data regression model with covariates consisting of component data and numerical data, and dependent variable consisting of functional data: ,in, For the city The PM2.5 concentration curve is function-based data. For the city The proportions of primary, secondary, and tertiary industries, referred to as the tertiary industry ratio, are component data. For the city The logarithm of GDP per capita is a numerical data point. For the city The average temperature is numerical data. It is a time-varying component coefficient to be estimated, which is at any given time... Assigned to constituent covariates A coefficient , , The coefficients of the function to be estimated are: It is a functional residual. This refers to the inner product operation in simplex space. The calculation results are functional data; Optionally, the component data of the transformation of the output value of the three industries form a simplex space. ,in yes Dimensional component data, each component Non-negative, with a cumulative sum of 1, for the proportion data of the three industries, Define addition operation in simplex space Sum of inner products This yields the Atchison space, for any component data. and Their addition operation is ,in For closure operations , and The inner product is ,in It is the geometric mean subscript This indicates that the operation is performed in Atchison space, and the composition data is derived from the inner product operation. norm ,in It is the natural logarithm.

[0016] The following explains function-type data: Assuming that the embodiments of the present invention are derived from samples High-frequency data observed If the high-frequency data is potentially continuously changing, it can be... Consider it as a functional data sample At different times The discrete observations are used. Considering the noise at each observation point, the following model is established:

[0017] in, Representing the A functional data sample in time The noise. And the sample It is a second-order stochastic process In time period The previous implementation, and has .

[0018] Discrete observation data Transform into a function curve Commonly used methods include basis function expansion and kernel function method. This embodiment of the invention uses basis function expansion: Select a set of basis functions , So This can be represented as a linear combination of these basis functions, and the model can be written as:

[0019] in, where are the expansion coefficients of the basis functions. In applications, common basis functions include Fourier basis, B-spline basis, wavelet basis, polynomial basis, and Bernstein basis. The basis function representation can be flexibly selected based on the characteristics of the data itself. The embodiments of this invention employ a B-spline basis.

[0020] S4. Based on the equidistant logarithmic ratio transformation, function basis expansion, and iterative reweighted least squares method, we obtain... , , Robust M-estimation; Because of the limitations of component data, directly treating component data as... Analyzing dimensional vectors using a standard linear regression model often yields misleading results; therefore, this embodiment of the invention uses an isometric log-ratio transformation.

[0021] Optionally, the equidistant logarithmic ratio transformation is:

[0022] in The coordinates are after logarithmic ratio transformation, for the first coordinate. It was discovered to be the first ingredient. Geometric mean of all remaining components The logarithm of the ratio multiplied by a factor Therefore This indicates the magnitude of the first component; in estimating... coefficient Afterwards, according to Determine the positive or negative value. Add one unit, that is The increase affects the dependent variable; at this time for and The logarithm of the ratio multiplied by a factor, excluding It cannot represent the magnitude of the second component, corresponding to It is not easy to explain, especially for each component of interest. Rearrange Each component was obtained Marked as Then use transformation

[0023] according to coefficient ,judge Add one unit, that is The increase affects the dependent variable, thus, a total of The second isochronous logarithmic ratio transformation yields all The influence of the proportion of each component on the dependent variable, making To the raw data The inverse logarithmic ratio transform of the interval is: .

[0024] Optionally, the hybrid data regression model , represented as: ,in , , , .

[0025] The following describes a more general model: .

[0026] It has The observation object, from the first The data observed on each object are denoted as . ,in, It is functional data exist The observed value at time; It contains Component data for each ingredient; It is numerical data.

[0027] mark , , dependent variable and covariates The following relationship must be satisfied: ,in, It is a function of component data, when time is fixed. hour, It is The component data of dimension, where the sum of each component is 1; when time When changing, A pie chart that represents changes over time is called functional component data. In this model, It is a time-varying component coefficient to be estimated, which is at any given time... Assigned to constituent covariates A coefficient , The coefficients of the function to be estimated are: It is a functional residual. This refers to the inner product operation in simplex space. The calculation result is functional data.

[0028] Next, we will explain the parameters in the model: First, the functional coefficients It means at any time ,when When adding a unit, Will increase One unit. Therefore It is more flexible than the commonly used fixed coefficient.

[0029] For functional component coefficients Its meaning can be explained from two perspectives: (1) Component covariates Overall changes The impact; At any time ,like Add to simplex space ,So This will increase by 1 unit, that is: .

[0030] (2) Component covariates Each component Changes in the dependent variable The impact; With the first component For example, according to the logarithmic ratio transformation, Therefore, a fixed time ,when Add one unit (equivalent to) Compared to When the proportion increases, Increase: , here The first component curve after the logarithmic ratio transformation method exist The value of . Because It is a variable coefficient that changes over time, and it demonstrates How does the change in the proportion of the first component affect the effect at different times? When you want to view Other components, such as How changes affect rearrange of The first component makes the second component... The first component is listed, thus the estimated value is... That is, it shows How to influence in the time dimension Similarly, to know all The impact of changes in the proportion of each component on the dependent variable requires a total of [number] [tests / analyses]. Secondary logarithmic ratio transformation and parameter estimation.

[0031] S4 specifically includes: S41. The components are not independent. and Transform into linearly independent vectors; Using logarithmic ratio transformation and After removing the constraints, they become respectively and ,because It is unknown. At this point, the parameter is to be estimated. According to the inner product-preserving property of the logarithmic ratio transform, we obtain: At this point, the model is represented as ; S42, to and Perform a basic expansion; Let a set of basis functions in the space containing the functional data be denoted as... ,So and Represent as Using a finite number of basis functions, the true function can be approximated quite well, therefore for By truncating the first K basis functions, we obtain , , At this point, the model is represented as ; S43. Perform robust M-estimation on the new model; The robust M-estimation is obtained by optimizing the following objective function:

[0032] in , , Huber loss function , It is a positive constant; S44. Use the iterative reweighted least squares method to obtain the numerical solution of the objective function; make , , , , ,in for The derivative of the objective function is obtained by taking the derivative of the objective function:

[0033] In the above formula Treat them as weights and solve using the iterative reweighted least squares method. ,remember:

[0034] Then the estimated value It is obtained by the following algorithm: (1) Initialize parameters , recorded as ; (2) Using the r-th step The estimated value ,calculate ,in Thus, updates are made:

[0035] (3) Repeat step (2) until Less than a pre-defined threshold; get The estimated value Afterwards, further acquisition and They are given by the following formula:

[0036] Then, through the inverse logarithmic ratio transformation, we obtain... The estimated values ​​are obtained, and thus all the parameters in the model are solved.

[0037] S5, according to , , The estimated values ​​were used to analyze the dynamic impact of the proportion of the three industries, per capita GDP, and average temperature on PM2.5 in various cities.

[0038] Optionally, S5 specifically includes: Observe the primary industry The curve of the coefficient, such as Figure 3 As shown, it was found that it fluctuated around 0, indicating that the change in the proportion of primary production has little effect on PM2.5 concentration; Observing the secondary industry The curve of the coefficient, such as Figure 3 As shown, except for early June and late September to early October, the ratio of secondary industry has a positive impact on PM2.5 concentration (the curve values ​​are non-zero positive numbers, indicating that during this period, cities with a higher ratio of secondary industry have higher PM2.5 concentrations). Observing the tertiary industry The curve of the coefficient, such as Figure 3 As shown, it was found that the proportion of the three industries had a positive impact on PM2.5 concentration in early June and from late September to early October (the higher the proportion of the three industries in a city, the higher the PM2.5 concentration, which may be related to the fact that this period is the Dragon Boat Festival and the National Day Golden Week). observe The curve, such as Figure 4 As shown, except for April to early May and July, per capita GDP has a negative impact on PM2.5 concentration. A negative value indicates that during this period, cities with higher per capita GDP had lower PM2.5 concentrations. This may be related to pandemic lockdowns, the May Day holiday, and effective pollution control in economically developed regions. observe The curve, such as Figure 5 As shown, it was found that temperature has a positive impact on PM2.5 concentration throughout the year (its value is always positive, indicating that the higher the temperature in a city, the higher the PM2.5 concentration).

[0039] like Figure 6 As shown in the figure, this embodiment of the invention also provides a PM2.5 influencing factor analysis system based on a mixed data regression model, the system comprising: The data collection module 610 is used to collect data on the output value of the tertiary industry, per capita GDP, average temperature, and daily average PM2.5 concentration of various cities in 2023. The preprocessing module 620 is used to convert the output value of the three industries into component data, take the average of the daily average PM2.5 concentration every seven days to obtain the weekly average PM2.5 concentration, and use a third-order B-spline function to fit the weekly average PM2.5 concentration to obtain the PM2.5 concentration curve. Module 630 is used to build a mixed data regression model where the covariates are component data and numerical data, and the dependent variable is functional data. ,in, For the city The PM2.5 concentration curve is function-based data. For the city The proportions of primary, secondary, and tertiary industries, referred to as the tertiary industry ratio, are component data. For the city The logarithm of GDP per capita is a numerical data point. For the city The average temperature is numerical data. It is a time-varying component coefficient to be estimated, which is at any given time... Assigned to constituent covariates A coefficient , , The coefficients of the function to be estimated are: It is a functional residual. This refers to the inner product operation in simplex space. The calculation results are functional data; Estimation module 640 is used to obtain, based on equidistant logarithmic ratio transformation, function basis expansion, and iterative reweighted least squares method, the following results: , , Robust M-estimation; Analysis module 650, used to analyze based on , , The estimated values ​​were used to analyze the dynamic impact of the proportion of the three industries, per capita GDP, and average temperature on PM2.5 in various cities.

[0040] Optionally, the component data of the transformation of the output value of the three industries form a simplex space. ,in yes Dimensional component data, each component Non-negative, with a cumulative sum of 1, for the proportion data of the three industries, Define addition operation in simplex space Sum of inner products This yields the Atchison space, for any component data. and Their addition operation is ,in For closure operations , and The inner product is ,in It is the geometric mean subscript This indicates that the operation is performed in Atchison space, and the composition data is derived from the inner product operation. norm ,in It is the natural logarithm.

[0041] Optionally, the equidistant logarithmic ratio transformation is:

[0042] in The coordinates are after logarithmic ratio transformation, for the first coordinate. It was discovered to be the first ingredient. Geometric mean of all remaining components The logarithm of the ratio multiplied by a factor Therefore This indicates the magnitude of the first component; in estimating... coefficient Afterwards, according to Determine the positive or negative value. Add one unit, that is The increase affects the dependent variable; at this time for and The logarithm of the ratio multiplied by a factor, excluding It cannot represent the magnitude of the second component, corresponding to It is not easy to explain, especially for each component of interest. Rearrange Each component was obtained Marked as Then use transformation

[0043] according to coefficient ,judge Add one unit, that is The increase affects the dependent variable, thus, a total of The second isochronous logarithmic ratio transformation yields all The influence of the proportion of each component on the dependent variable, making To the raw data The inverse logarithmic ratio transform of the interval is: .

[0044] Optionally, the hybrid data regression model , represented as: ,in , , , The estimation module is specifically used for: S41. The components are not independent. and Transform into linearly independent vectors; Using logarithmic ratio transformation and After removing the constraints, they become respectively and ,because It is unknown. At this point, the parameter is to be estimated. According to the inner product-preserving property of the logarithmic ratio transform, we obtain: At this point, the model is represented as ; S42, to and Perform a basic expansion; Let a set of basis functions in the space containing the functional data be denoted as... ,So and Represent as Using a finite number of basis functions, the true function can be approximated quite well, therefore for By truncating the first K basis functions, we obtain , , At this point, the model is represented as ; S43. Perform robust M-estimation on the new model; The robust M-estimation is obtained by optimizing the following objective function:

[0045] in , , Huber loss function , It is a positive constant; S44. Use the iterative reweighted least squares method to obtain the numerical solution of the objective function; make , , , , ,in for The derivative of the objective function is obtained by taking the derivative of the objective function:

[0046] In the above formula Treat them as weights and solve using the iterative reweighted least squares method. ,remember:

[0047] Then the estimated value It is obtained by the following algorithm: (1) Initialize parameters , recorded as ; (2) Using the r-th step The estimated value ,calculate ,in Thus, updates are made:

[0048] (3) Repeat step (2) until Less than a pre-defined threshold; get The estimated value Afterwards, further acquisition and They are given by the following formula:

[0049] Then, through the inverse logarithmic ratio transformation, we obtain... The estimated values ​​are obtained, and thus all the parameters in the model are solved.

[0050] Optionally, the analysis module is specifically used for: Observe the primary industry The curve of the coefficient shows that the proportion of primary industry has little effect on PM2.5 concentration; Observing the secondary industry The curve of the coefficient shows that, except for early June and late September to early October, the ratio of secondary industry has a positive impact on PM2.5 concentration. Observing the tertiary industry The coefficient curve shows that the ratio of the three industries has a positive impact on PM2.5 concentration in early June and from late September to early October. observe The curve shows that, except for April to early May and July, per capita GDP has a negative impact on PM2.5 concentration. observe The curve shows that temperature has a positive impact on PM2.5 concentration throughout the year.

[0051] The PM2.5 influencing factor analysis system based on a mixed data regression model provided in this embodiment of the invention has a functional structure that corresponds to the PM2.5 influencing factor analysis method based on a mixed data regression model provided in this embodiment of the invention, and will not be described again here.

[0052] Figure 7 This is a schematic diagram of the structure of an electronic device 700 provided in an embodiment of the present invention. The electronic device 700 may vary considerably due to different configurations or performance. It may include one or more processors (CPUs) 701 and one or more memories 702. The memory 702 stores at least one instruction, which is loaded and executed by the processor 701 to implement the steps of the above-mentioned PM2.5 influencing factor analysis method based on a mixed data regression model.

[0053] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to complete the above-described method for analyzing PM2.5 influencing factors based on a hybrid data regression model. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0054] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0055] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for analyzing PM2.5 influencing factors based on a mixed data regression model, characterized in that, The method includes: S1. Collect data on the output value of the tertiary industry, per capita GDP, average temperature, and daily average PM2.5 concentration of various cities in 2023; S2. Convert the output value of the three industries into component data, take the average of the daily average PM2.5 concentration every seven days to obtain the weekly average PM2.5 concentration, and use a third-order B-spline function to fit the weekly average PM2.5 concentration to obtain the PM2.5 concentration curve. S3. Establish a mixed data regression model with covariates consisting of component data and numerical data, and dependent variable consisting of functional data: ,in, For the city The PM2.5 concentration curve is function-based data. For the city The proportions of primary, secondary, and tertiary industries, referred to as the tertiary industry ratio, are component data. For the city The logarithm of GDP per capita is a numerical data point. For the city The average temperature is numerical data. It is a time-varying component coefficient to be estimated, which is at any given time... Assigned to constituent covariates A coefficient , , The coefficients of the function to be estimated are: It is a functional residual. This refers to the inner product operation in simplex space. The calculation results are functional data; S4. Based on the equidistant logarithmic ratio transformation, function basis expansion, and iterative reweighted least squares method, we obtain... , , Robust M-estimation; S5, according to , , The estimated values ​​were used to analyze the dynamic impact of the proportion of the three industries, per capita GDP, and average temperature on PM2.5 in various cities.

2. The method according to claim 1, characterized in that, The component data of the transformation of the output value of the three industries constitute a simplex space. ,in yes Dimensional component data, each component Non-negative, with a cumulative sum of 1, for the proportion data of the three industries, Define addition operation in simplex space Sum of inner products This yields the Atchison space, for any component data. and Their addition operation is ,in For closure operations , and The inner product is ,in It is the geometric mean subscript This indicates that the operation is performed in Atchison space, and the composition data is derived from the inner product operation. norm ,in It is the natural logarithm.

3. The method according to claim 2, characterized in that, The equidistant logarithmic ratio transformation is as follows: in The coordinates are after logarithmic ratio transformation, for the first coordinate. It was discovered to be the first ingredient. Geometric mean of all remaining components The logarithm of the ratio multiplied by a factor Therefore This indicates the magnitude of the first component; in estimating... coefficient Afterwards, according to Determine the positive or negative value. Add one unit, that is The increase affects the dependent variable; at this time for and The logarithm of the ratio multiplied by a factor, excluding It cannot represent the magnitude of the second component, corresponding to It is not easy to explain, especially for each component of interest. Rearrange Each component was obtained Marked as Then use transformation according to coefficient ,judge Add one unit, that is The increase affects the dependent variable, thus, a total of The second isochronous logarithmic ratio transformation yields all The influence of the proportion of each component on the dependent variable, making To the raw data The inverse logarithmic ratio transform of the interval is: 。 4. The method according to claim 3, characterized in that, The mixed data regression model , represented as: ,in , , , S4 specifically includes: S41. The components are not independent. and Transform into linearly independent vectors; Using logarithmic ratio transformation and After removing the constraints, they become respectively and ,because It is unknown. At this point, the parameter is to be estimated. According to the inner product-preserving property of the logarithmic ratio transform, we obtain: At this point, the model is represented as ; S42, to and Perform a basic expansion; Let a set of basis functions in the space containing the functional data be denoted as... ,So and Represent as Using a finite number of basis functions, the true function can be approximated quite well, therefore for By truncating the first K basis functions, we obtain , , At this point, the model is represented as ; S43. Perform robust M-estimation on the new model; The robust M-estimation is obtained by optimizing the following objective function: in , , Huber loss function , It is a positive constant; S44. Use the iterative reweighted least squares method to obtain the numerical solution of the objective function; make , , , , ,in for The derivative of the objective function is obtained by taking the derivative of the objective function: In the above formula Treat them as weights and solve using the iterative reweighted least squares method. ,remember: Then the estimated value It is obtained by the following algorithm: (1) Initialize parameters , recorded as ; (2) Using the r-th step The estimated value ,calculate ,in Thus, updates are made: (3) Repeat step (2) until Less than a pre-defined threshold; get The estimated value Afterwards, further acquisition and They are given by the following formula: Then, through the inverse logarithmic ratio transformation, we obtain... The estimated values ​​are obtained, and thus all the parameters in the model are solved.

5. The method according to claim 4, characterized in that, S5 specifically includes: Observe the primary industry The curve of the coefficient shows that the proportion of primary industry has little effect on PM2.5 concentration; Observing the secondary industry The curve of the coefficient shows that, except for early June and late September to early October, the ratio of secondary industry has a positive impact on PM2.5 concentration. Observing the tertiary industry The coefficient curve shows that the ratio of the three industries has a positive impact on PM2.5 concentration in early June and from late September to early October. observe The curve shows that, except for April to early May and July, per capita GDP has a negative impact on PM2.5 concentration. observe The curve shows that temperature has a positive impact on PM2.5 concentration throughout the year.

6. A PM2.5 influencing factor analysis system based on a mixed data regression model, characterized in that, The system includes: The data collection module is used to collect data on the output value of the tertiary industry, per capita GDP, average temperature, and daily average PM2.5 concentration of various cities in 2023. The preprocessing module is used to convert the output value of the three industries into component data, take the average of the daily average PM2.5 concentration every seven days to obtain the weekly average PM2.5 concentration, and use a third-order B-spline function to fit the weekly average PM2.5 concentration to obtain the PM2.5 concentration curve. The module is used to build a mixed data regression model where the covariates are component data and numerical data, and the dependent variable is functional data. ,in, For the city The PM2.5 concentration curve is function-based data. For the city The proportions of primary, secondary, and tertiary industries, referred to as the tertiary industry ratio, are component data. For the city The logarithm of GDP per capita is a numerical data point. For the city The average temperature is numerical data. It is a time-varying component coefficient to be estimated, which is at any given time... Assigned to constituent covariates A coefficient , , The coefficients of the function to be estimated are: It is a functional residual. This refers to the inner product operation in simplex space. The calculation results are functional data; The estimation module is used to obtain the results based on the equidistant logarithmic ratio transformation, function basis expansion, and iterative reweighted least squares method. , , Robust M-estimation; Analysis module, used to analyze based on , , The estimated values ​​were used to analyze the dynamic impact of the proportion of the three industries, per capita GDP, and average temperature on PM2.5 in various cities.

7. The system according to claim 6, characterized in that, The component data of the transformation of the output value of the three industries constitute a simplex space. ,in yes Dimensional component data, each component Non-negative, with a cumulative sum of 1, for the proportion data of the three industries, Define addition operation in simplex space Sum of inner products This yields the Atchison space, for any component data. and Their addition operation is ,in For closure operations , and The inner product is ,in It is the geometric mean subscript This indicates that the operation is performed in Atchison space, and the composition data is derived from the inner product operation. norm ,in It is the natural logarithm.

8. The system according to claim 7, characterized in that, The equidistant logarithmic ratio transformation is as follows: in The coordinates are after logarithmic ratio transformation, for the first coordinate. It was discovered to be the first ingredient. Geometric mean of all remaining components The logarithm of the ratio multiplied by a factor Therefore This indicates the magnitude of the first component; in estimating... coefficient Afterwards, according to Determine the positive or negative value. Add one unit, that is The increase affects the dependent variable; at this time for and The logarithm of the ratio multiplied by a factor, excluding It cannot represent the magnitude of the second component, corresponding to It is not easy to explain, especially for each component of interest. Rearrange Each component was obtained Marked as Then use transformation according to coefficient ,judge Add one unit, that is The increase affects the dependent variable, thus, a total of The second isochronous logarithmic ratio transformation yields all The influence of the proportion of each component on the dependent variable, making To the raw data The inverse logarithmic ratio transform of the interval is: 。 9. The system according to claim 8, characterized in that, The mixed data regression model , represented as: ,in , , , The estimation module is specifically used for: S41. The components are not independent. and Transform into linearly independent vectors; Using logarithmic ratio transformation and After removing the constraints, they become respectively and ,because It is unknown. At this point, the parameter is to be estimated. According to the inner product-preserving property of the logarithmic ratio transform, we obtain: At this point, the model is represented as ; S42, to and Perform a basic expansion; Let a set of basis functions in the space containing the functional data be denoted as... ,So and Represent as Using a finite number of basis functions, the true function can be approximated quite well, therefore for By truncating the first K basis functions, we obtain , , At this point, the model is represented as ; S43. Perform robust M-estimation on the new model; The robust M-estimation is obtained by optimizing the following objective function: in , , Huber loss function , It is a positive constant; S44. Use the iterative reweighted least squares method to obtain the numerical solution of the objective function; make , , , , ,in for The derivative of the objective function is obtained by taking the derivative of the objective function: In the above formula Treat them as weights and solve using the iterative reweighted least squares method. ,remember: Then the estimated value It is obtained by the following algorithm: (1) Initialize parameters , recorded as ; (2) Using the r-th step The estimated value ,calculate ,in Thus, updates are made: (3) Repeat step (2) until Less than a pre-defined threshold; get The estimated value Afterwards, further acquisition and They are given by the following formula: Then, through the inverse logarithmic ratio transformation, we obtain... The estimated values ​​are obtained, and thus all the parameters in the model are solved.

10. The system according to claim 9, characterized in that, The analysis module is specifically used for: Observe the primary industry The curve of the coefficient shows that the proportion of primary industry has little effect on PM2.5 concentration; Observing the secondary industry The curve of the coefficient shows that, except for early June and late September to early October, the ratio of secondary industry has a positive impact on PM2.5 concentration. Observing the tertiary industry The coefficient curve shows that the ratio of the three industries has a positive impact on PM2.5 concentration in early June and from late September to early October. observe The curve shows that, except for April to early May and July, per capita GDP has a negative impact on PM2.5 concentration. observe The curve shows that temperature has a positive impact on PM2.5 concentration throughout the year.