A method for attributing regional water environment antibiotic spatial heterogeneity based on a GWR-GeoDetector coupling model
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NAT RESERACH CENT OF GEOANALYSIS
- Filing Date
- 2026-04-27
- Publication Date
- 2026-08-04
AI Technical Summary
[0003]即使并列使用地理加权回归与地理探测器形成互补,也难以应对城市水环境中抗生素分布的复杂空间异质性,导致q统计量偏低
本发明通过耦合地理加权回归与地理探测器两种模型,实现了对抗生素空间分布的精准解析。首先,突破了传统统计模型对空间非平稳性的限制,能够识别因子在不同区域的局部作用差异。其次,显著提升归因解释力,通过“先分区后探测”策略,有效消除全局空间干扰,使解释力较全局分析提升。第三,实现污染源精准溯源,可明确不同地理位置如城区与农耕区的主导污染因子,为精准治理提供科学依据。最后,揭示因子间的协同效应,表明多类污染源的叠加作用常呈非线性增强,为环境风险预警提供更全面的参考。
Smart Images

Figure CN122508532A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water environment monitoring and spatial statistical analysis technology, specifically to a method for attribution analysis of spatial heterogeneity of antibiotics in regional water environment based on the GWR-GeoDetector coupled model. Background Technology
[0002] In recent years, antibiotics, as a typical emerging pollutant, have been frequently detected in urban water environments, posing a potential threat to aquatic ecosystems and public health. Existing analytical methods still have limitations in regional-scale studies. Principal component analysis and multiple linear regression are limited by linear assumptions and global characteristics, making it difficult to quantify the contribution of local pollution point sources and easily neglecting nonlinear interactions. While geographically weighted regression can identify spatial nonstationarity, its reliability decreases when the sample size is limited or local fluctuations are severe. Geographic detectors have advantages in nonlinear analysis, but they are highly dependent on boundary accuracy, and local point source effects are easily averaged out globally.
[0003] Even when geographically weighted regression and geographic detectors are used in parallel to complement each other, it is still difficult to cope with the complex spatial heterogeneity of antibiotic distribution in urban water environments, resulting in low q statistics. For example, in the context of complex multi-source inputs, traditional linear models such as ordinary least squares cannot handle spatial nonstationarity, while single geographic models are insufficient to explain point source pollution such as medical, aquaculture, and pharmaceutical industries, and are difficult to quantify nonlinear interactions.
[0004] Therefore, there is an urgent need to develop a quantitative method that can accurately analyze the driving mechanisms of water pollution sources under complex spatial heterogeneity conditions, in order to solve the problem that antibiotics in regional water environments are difficult to accurately identify their driving mechanisms due to complex spatial heterogeneity. Summary of the Invention
[0005] This invention aims to provide a method for attributing the spatial heterogeneity of antibiotics in regional water environments based on a GWR-GeoDetector coupled model. This method constructs an analytical framework that couples geographic weighted regression with a geographic detector, enabling the zoning of spatial distribution patterns of antibiotic pollution sources in regional water environments and the identification and quantitative analysis of key driving factors within different zones. This provides a scientific basis for precise zoning management and source control of regional water environments.
[0006] To solve the above-mentioned technical problems, the technical solution proposed in this application is as follows: This invention provides a method for attribution analysis of spatial heterogeneity of antibiotics in regional aquatic environments based on a GWR-GeoDetector coupled model, comprising the following steps: (1) Data acquisition and indicator quantification: Obtain antibiotic concentration data from surface water sampling points in the target area, and construct a database containing multiple environmental driving factors based on the geographic information system; (2) Geographically weighted regression spatial modeling and feature partitioning: Using total antibiotic concentration as the dependent variable and each environmental driving factor as the independent variable, a geographically weighted regression model is constructed to obtain the local regression coefficients of each driving factor; based on the spatial distribution characteristics of the local regression coefficients, the study area is divided into several feature partitions. (3) Quantitative analysis and attribution of single factors in geographic detector: Within each of the aforementioned feature regions, the factor detector module in the geographic detector is used to calculate the q value of each environmental driving factor. The q value is used to characterize the explanatory power of the driving factor on the spatial distribution differences of antibiotics, thereby identifying the dominant driving factor in each region. (4) Analysis of multi-factor interaction effects: Using the interaction detector module in the geographic detector, the different environmental driving factors are combined in pairs for analysis, the q value after interaction is calculated, and the interaction type between the driving factors is determined by comparing the q value after interaction with the corresponding single factor q value.
[0007] Furthermore, the antibiotic concentration data in step (1) is quantitatively detected by online solid-phase extraction-ultra-high performance liquid chromatography-high resolution mass spectrometry, and the antibiotics include at least one class of compounds among sulfonamides, fluoroquinolones and macrolides.
[0008] Furthermore, the environmental driving factor database mentioned in step (1) includes spatial variables characterizing the intensity of human activities, wherein the spatial variables are selected from at least one of population density, land use type, distribution of medical institutions, distribution of pharmaceutical companies, distribution of livestock and poultry farming, and distribution of aquaculture.
[0009] Furthermore, in step (2), a Gaussian kernel function is used to weight the spatial interaction between sampling points, and the optimal bandwidth of each local regression point is iteratively searched using the AICc information criterion. The geographic weighted regression model is then constructed using the optimal bandwidth.
[0010] Furthermore, the method for dividing the feature partitions in step (2) is as follows: the upper quartile of the local regression coefficient of each driving factor is selected as the threshold of the influence range, and the influence range of different driving factors is comprehensively divided by the spatial superposition analysis method.
[0011] Further, in step (3), the q value is 1 minus the ratio of the sum of variances within each stratified unit of environmental variables to the total variance of the whole region; the q values are sorted in descending order, and the factor with the significance level p<0.05 and the first q value is taken as the dominant driving factor of the partition.
[0012] Furthermore, the interaction types described in step (4) include nonlinear enhancement, two-factor enhancement, independence, nonlinear reduction, or single-factor nonlinear reduction.
[0013] Furthermore, step (4) further includes: statistically analyzing the frequency of occurrence of each driving factor in the combination of strong interaction factors with an interaction q value greater than or equal to 0.9 in each partition, and determining the most frequently occurring factor as the core synergistic factor driving the spatial distribution of antibiotics in the region.
[0014] Furthermore, the study area described in step (2) is divided into four feature partitions.
[0015] Further, step (1) further includes: the environmental driving factor database includes distance data and density data; the distance data is used as input to the geographic weighted regression model, and the density data is used as input to the geographic detector.
[0016] Furthermore, the distance data is the weighted Euclidean distance from each sampling point to each pollution source type, and the pollution source types include at least one of hospitals, residential areas, farmland, livestock stations, aquaculture farms, and pharmaceutical factories.
[0017] Furthermore, the density data includes population density, land use ratio, and kernel density values for each type of pollution source.
[0018] Furthermore, step (3) further includes: before performing geographic detector analysis on each feature partition, discretizing the original data and converting it into classification data; calculating the q statistics under each combination of levels by iteratively testing different grading methods and the number of levels, and selecting the discretization scheme that maximizes the q value as the optimal grading strategy for that partition.
[0019] Compared with the prior art, the present invention achieves the following beneficial technical effects: This invention achieves precise analysis of the spatial distribution of antibiotics by coupling two models: geographic weighted regression and geographic detectors. First, it overcomes the limitations of traditional statistical models regarding spatial non-stationarity, enabling the identification of differences in the local effects of factors across different regions. Second, it significantly enhances attribution explanatory power by effectively eliminating global spatial interference through a "regional-first, detector-later" strategy, thus improving explanatory power compared to global analysis. Third, it enables precise source tracing of pollution, clearly identifying the dominant pollutants in different geographical locations, such as urban and agricultural areas, providing a scientific basis for precise governance. Finally, it reveals synergistic effects among factors, showing that the cumulative effects of multiple pollution sources often exhibit non-linear enhancement, providing a more comprehensive reference for environmental risk early warning. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A schematic diagram of the spatial distribution of pollution sources used in a GWR model in one embodiment of the present invention includes pharmaceutical companies (a), hospitals (b), residential areas (c), livestock and veterinary stations (d), aquaculture sites (e), and farmland (f). Figure 2 A schematic diagram of the spatial distribution of pollution sources and population density for the GeoDetector model in one embodiment of the present invention includes pharmaceutical companies (a), hospitals (b), population (c), livestock and veterinary stations (d), aquaculture sites (e), and land use types (f). Figure 3 A schematic diagram showing the concentration distribution and detection frequency of antibiotics detected in surface water in Suzhou in one embodiment of the present invention; Figure 4 A schematic diagram of the spatial distribution of antibiotic concentrations in surface water in Suzhou, according to one embodiment of the present invention; Figure 5 In one embodiment of the present invention, a schematic diagram of the spatial zoning of pollution sources based on the GWR model is shown, including the spatial distribution and regression coefficient reclassification results of pharmaceutical companies (a), hospitals (b), residential areas (c), livestock and veterinary stations (d), aquaculture sites (e), and farmland (f); Figure 6 A schematic diagram of Geodetector factor detection results in one embodiment of the present invention shows the driving factor q value and proportional contribution of different antibiotic types and total concentrations; Figure 7 A schematic diagram of Geodetector interaction detection results in one embodiment of the present invention shows the strong interaction effect and frequency distribution between different antibiotic types of driving factors.
[0022] Figure 8 A flowchart illustrating the overall process in one embodiment of the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] This embodiment takes Suzhou City as an example to explain in detail the specific implementation steps of the regional water environment antibiotic spatial heterogeneity attribution analysis method based on the GWR-GeoDetector coupling model.
[0025] In one embodiment of this application, a method for attributing spatial heterogeneity of antibiotics in regional aquatic environments based on a GWR-GeoDetector coupled model is provided. The overall flowchart of the method in this embodiment is shown below. Figure 8 First, data acquisition and indicator quantification steps were performed. Specifically, surface water samples were collected from scientifically selected sampling points within Suzhou City. An online solid-phase extraction-ultra-high performance liquid chromatography-high resolution mass spectrometry (SPE-UPLC-Q-ExactiveOrbitrap HRMS) system was used to analyze the collected water samples and quantitatively determine the concentrations of 27 antibiotics. The results showed that the main antibiotics detected included sulfonamides, fluoroquinolones, and macrolides. Second, a driving factor database was constructed based on a geographic information system (GIS). This database contains two types of data: distance data and density data. Distance data was used as input for the subsequent geographic weighted regression model, specifically calculating the weighted Euclidean distance from each sampling point to hospitals, residential areas, farmland, livestock stations, aquaculture farms, and pharmaceutical factories. Figure 1 The diagram shows the spatial distribution of pollution sources used in the GWR model, including the geographical distribution of pharmaceutical companies (a), hospitals (b), residential areas (c), livestock and veterinary stations (d), aquaculture sites (e), and farmland (f). Density data is used as input to the geographic detector, specifically including the extraction of population density, land use percentage, and kernel density values for each pollution source such as pharmaceutical companies, hospitals, livestock and veterinary stations, and aquaculture sites. Figure 2 The diagram shows the spatial distribution of pollution sources and population density used in the GeoDetector model, including the spatial density of pharmaceutical companies (a), hospitals (b), population (c), livestock and veterinary stations (d), aquaculture sites (e), and land use types (f). Through analysis of water samples, the concentration distribution and detection frequency of antibiotics in Suzhou surface water were obtained, as shown below. Figure 3 As shown. Figure 3In the diagram, the boxes represent the quartile range (25th to 75th percentile), the midline represents the median, the red squares represent the mean, the whisker lines extend to the 5th to 95th percentiles, the black diamonds represent outliers (beyond the 5th to 95th percentiles), and the blue diamonds represent the data distribution characteristics, showing the distribution of total antibiotics, various antibiotics, and the top two compounds detected in each category. The spatial distribution of antibiotic concentrations in Suzhou surface water is shown below. Figure 5 As shown, where Figure 4 (a) shows the total antibiotic concentration detected at each sampling point. Figure 4 (b) Shows the spatial distribution of total concentrations of major antibiotic classes (sulfonamides, fluoroquinolones, macrolides), Figure 4 (c) shows the antibiotic concentrations at each sampling point, providing a basis for subsequent analysis of spatial heterogeneity.
[0026] Next, the geographically weighted regression spatial modeling and feature partitioning steps are performed. This step aims to transform complex global spatial heterogeneity into relatively homogeneous local regions. Using total antibiotic concentration as the dependent variable and distance factors from each pollution source (distance to hospitals, residential areas, farmland, livestock stations, aquaculture farms, and pharmaceutical factories) as explanatory variables, a geographically weighted regression equation is constructed. A fixed bandwidth model is adopted, and the optimal bandwidth is selected using AICc (Modified Akaike Information Criterion) to maximize the model's ability to capture spatial non-stationarity. The local regression coefficients of each driving factor are calculated and spatially interpolated. The sign and magnitude of the local regression coefficients are used to assess the spatial intensity and extent of the influence of different pollution sources. Specifically, the upper quartile (i.e., the top 25%) of the local regression coefficients of each driving factor is selected as the threshold for the scope of influence. Figure 5 (a) to Figure 5 As shown in (f), the dot symbols represent the spatial distribution of each explanatory variable (pharmaceutical companies, hospitals, residential areas, livestock and veterinary stations, aquaculture sites, and farmland), and the red lines represent the boundaries of the top 25% of the regression coefficients for each factor. Then, using spatial overlay analysis, the influence range of different factors is comprehensively divided, and four typical influence areas are identified based on the spatial overlay of multi-factor boundaries, such as... Figure 5 As shown in (g). These four partitions will serve as independent units for subsequent geographic detector analysis, thereby reducing statistical interference caused by excessive spatial scale.
[0027] Then, the single-factor quantitative analysis and attribution steps of the geospatial detector are executed. Within each feature partition defined by the GWR, the geospatial detector is run for attribution analysis. To enable the geospatial detector to process the data, the raw data is first discretized, transforming continuous data into categorical data. Within each partition, different grading methods (such as equal interval method, quantile method, etc.) and the number of gradations are iteratively tested to calculate the q-statistic for each gradation combination. The discretization scheme that maximizes the q-value is selected as the optimal grading strategy for that partition. Subsequently, the factor detector module is used to quantitatively evaluate the contribution rate of each environmental factor to the spatial heterogeneity of antibiotics. The q-value is calculated as: q = 1 - (Sum of variances within strata / Total variance of the whole region), that is, the q-value equals 1 minus the ratio of the sum of variances within each stratified unit of environmental variables to the total variance of the whole region. The larger the q-value, the higher the explanatory power of the factor for the variation in antibiotic distribution. The q-values are sorted in descending order, and the factor with the highest q-value at a significance level of p < 0.05 is selected as the "dominant driving factor" for that partition. Factor detector results as follows Figure 6 As shown, the driving factors q-values and their proportional contributions for different antibiotic types and total concentrations in four typical affected areas are illustrated. Fluoroquinolones and macrolides in the aquaculture affected area were not included in the analysis due to their extremely low detection rates. This step clarifies the primary reasons for the spatial distribution differences of antibiotics within each zone.
[0028] Finally, the multi-factor interaction effect analysis step was performed. To understand the combined action mechanism among different pollution sources, the interaction detector module in the geographic detector was used to calculate the interaction explanatory power q(X1∩X2) when any two driving factors (X1 and X2) act together. The calculated interaction q value was compared with the corresponding two single-factor q values to determine the interaction type. Possible interaction types include: nonlinear enhancement, two-factor enhancement, independence, nonlinear attenuation, or single-factor nonlinear attenuation. The interaction detection results are as follows: Figure 8 As shown, the strong interaction effect (interaction q value ≥ 0.9) of different antibiotic types driving factors in four typical affected areas is demonstrated. Figure 7The outer ring displays the various driving factors and their sum of interaction q-values; the middle ring displays the frequency of occurrence of strongly interacting driving factors among the three classes of antibiotics; and the inner ring displays the frequency of occurrence of strongly interacting driving factors in the total antibiotic concentration. By statistically analyzing the frequency of occurrence of each driving factor in high explanatory power factor combinations (interaction q-value ≥ 0.9) in each region, the most frequently occurring factors were identified as the core synergistic factors driving the spatial distribution of antibiotics in the region. Finally, the identified core synergistic factors were combined with regional environmental and geographical characteristics to provide a final environmental interpretation of the analysis results from the perspectives of pollutant input sources, migration pathways, and spatial enrichment mechanisms.
[0029] In one embodiment of this application, the antibiotic detection technology used in the above data acquisition step is online solid-phase extraction-ultra-high performance liquid chromatography-high resolution mass spectrometry (SPE-UPLC-HRMS). Specifically, the target antibiotic is quantitatively detected using an online SPE-UPLC-HRMS system. The detected antibiotics include at least one class of compounds from sulfonamides, fluoroquinolones, and macrolides. For example, in this embodiment, a total of 27 antibiotics from three major classes—sulfonamides, fluoroquinolones, and macrolides—were detected simultaneously, and their concentration distribution and detection frequency are detailed in [details omitted]. Figure 3 .
[0030] In one embodiment of this application, the constructed environmental driving factor database includes spatial variables characterizing the intensity of human activities. These spatial variables are selected from at least one of population density, land use type, distribution of medical institutions, distribution of pharmaceutical companies, distribution of livestock and poultry farming, and distribution of aquaculture. In this embodiment, the driving factor database comprehensively covers all the above-mentioned variables, wherein distance data (hospitals, residential areas, farmland, livestock stations, aquaculture farms, pharmaceutical factories) is used as input to the geographically weighted regression model, such as... Figure 1 As shown; density data (population density, land use ratio, and kernel density values of pharmaceutical companies, hospitals, livestock and veterinary stations, and aquaculture sites) are used as input for the geographic detector, such as... Figure 2 As shown.
[0031] In one embodiment of this application, a Gaussian kernel function is used to weight the spatial interactions between sampling points when constructing the geographically weighted regression model. Simultaneously, the optimal bandwidth for each local regression point is iteratively searched across the entire domain using the AICc information criterion, and this optimal bandwidth is used to construct the final geographically weighted regression model to obtain the regression coefficients between antibiotic concentration and environmental variables as a function of spatial location, thereby achieving a quantitative characterization of spatial nonstationarity.
[0032] In one embodiment of this application, the above-mentioned feature partitioning is achieved through the following specific method: First, the upper quartile (i.e., the top 25% of the values) of the local regression coefficients of each driving factor is selected as the spatial threshold of the influence range of that factor; then, spatial overlay analysis is performed using geographic information system software to comprehensively overlay the boundaries of the influence ranges of different driving factors, thereby dividing the entire study area into several feature partitions with relatively consistent internal characteristics. For example... Figure 5 As shown in (g), based on the superposition of multi-factor boundary space, this embodiment successfully identified four typical impact zones with distinct environmental significance.
[0033] In one embodiment of this application, the q-value used by the geographic detector is calculated as follows: the q-value is equal to one minus a ratio, where the numerator of the ratio is the sum of the products of the variance within each stratified unit and its corresponding sample size, and the denominator of the ratio is the product of the total variance of the entire study area and the total sample size of the entire area. The q-value ranges from 0 to 1, with a larger value indicating a stronger explanatory power of the driving factor for the spatial distribution differences in antibiotic concentrations within the region. When identifying the dominant driving factor, the calculated q-values of each driving factor are arranged in descending order of numerical value, and the factor that is statistically significant (i.e., a significance level p-value less than 0.05) and has the highest q-value is selected as the final dominant driving factor for that partition.
[0034] In one embodiment of this application, the specific criteria for determining the interaction type between driving factors are as follows: if q(X1∩X2) > q(X1) + q(X2), it is nonlinear enhancement; if q(X1∩X2) > Max(q(X1), q(X2)), it is two-factor enhancement; if q(X1∩X2) = q(X1) + q(X2), it is independent; if q(X1∩X2) < Min(q(X1), q(X2)), it is nonlinear attenuation; if q(X1∩X2) is between the maximum value of a single factor and the sum, it is single-factor nonlinear attenuation. In this embodiment, as... Figure 7 The results of the related interaction detection show that a variety of strong interaction effects were identified.
[0035] In one embodiment of this application, to identify the core synergistic factors across the entire region, based on the analysis of multi-factor interaction effects, the frequency of occurrence of each driving factor within factor combinations with high explanatory power (e.g., strong interaction combinations with interaction q values greater than or equal to 0.9) in each feature partition was further statistically analyzed. The driving factor that frequently appears across different partitions and different antibiotic types, and has the highest overall frequency, was ultimately determined as the core synergistic factor driving the spatial distribution of antibiotics in the entire regional aquatic environment. Figure 7The statistical results of the inner and outer loops show that a few key factors that play a core synergistic role in the total antibiotic concentration and the distribution of various antibiotics can be quantitatively identified.
[0036] In one embodiment of this application, the study area is accurately divided into four feature partitions using the above-described partitioning method, such as... Figure 5 As shown in (g).
[0037] In one embodiment of this application, the constructed driver factor database explicitly distinguishes the data types input to different models. Specifically, distance data, i.e., the weighted Euclidean distance from each sampling point to each pollution source type, is used as input to the geographically weighted regression model. These pollution source types include at least one of hospitals, residential areas, farmland, livestock stations, aquaculture farms, and pharmaceutical factories.
[0038] In one embodiment of this application, the aforementioned density data specifically includes population density, land use proportion, and kernel density values for each pollution source type. In this embodiment, these data are explicitly used as input to a geographic detector to characterize the activity intensity or coverage ratio at different spatial locations.
[0039] In one embodiment of this application, rigorous preprocessing is performed on the data before geographic detector analysis. Specifically, the original continuous data is discretized to transform it into categorical data. During discretization, different grading methods (such as natural breakpoint method, equal interval method, quantile method, etc.) and different numbers of grading levels (such as 4-10 levels) are iteratively tested through programming or software operation. The q-statistic is calculated for each combination of grading schemes, and finally, the discretization scheme that maximizes the q-value of the target factor is selected as the optimal grading strategy for that specific partition. This step ensures the accuracy and reliability of the geographic detector analysis results.
[0040] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for attributing spatial heterogeneity of antibiotics in regional aquatic environments based on a GWR-GeoDetector coupled model, characterized in that, Includes the following steps: (1) Data acquisition and indicator quantification: Obtain antibiotic concentration data from surface water sampling points in the target area, and construct a database containing multiple environmental driving factors based on the geographic information system; (2) Geographically weighted regression spatial modeling and feature partitioning: Using total antibiotic concentration as the dependent variable and each environmental driving factor as the independent variable, a geographically weighted regression model is constructed to obtain the local regression coefficients of each driving factor; based on the spatial distribution characteristics of the local regression coefficients, the study area is divided into several feature partitions. (3) Quantitative analysis and attribution of single factors in geographic detector: Within each of the aforementioned feature regions, the factor detector module in the geographic detector is used to calculate the q value of each environmental driving factor. The q value is used to characterize the explanatory power of the driving factor on the spatial distribution differences of antibiotics, thereby identifying the dominant driving factor in each region. (4) Analysis of multi-factor interaction effects: Using the interaction detector module in the geographic detector, the different environmental driving factors are combined in pairs for analysis, the q value after interaction is calculated, and the interaction type between the driving factors is determined by comparing the q value after interaction with the corresponding single factor q value.
2. The method according to claim 1, characterized in that, The antibiotic concentration data in step (1) is quantitatively detected by online solid phase extraction-ultra-high performance liquid chromatography-high resolution mass spectrometry. The antibiotics include at least one class of compounds among sulfonamides, fluoroquinolones and macrolides.
3. The method according to claim 1, characterized in that, The environmental driving factor database mentioned in step (1) includes spatial variables characterizing the intensity of human activities. The spatial variables are selected from at least one of population density, land use type, distribution of medical institutions, distribution of pharmaceutical companies, distribution of livestock and poultry farming, and distribution of aquaculture.
4. The method according to claim 1, characterized in that, In step (2), a Gaussian kernel function is used to weight the spatial interaction between sampling points, and the optimal bandwidth of each local regression point is iteratively searched using the AICc information criterion. The geographic weighted regression model is then constructed using the optimal bandwidth.
5. The method according to claim 1, characterized in that, The method for dividing the feature partitions in step (2) is as follows: the upper quartile of the local regression coefficient of each driving factor is selected as the threshold of the influence range, and the influence range of different driving factors is comprehensively divided by the spatial superposition analysis method.
6. The method according to claim 1, characterized in that, In step (3), the q value is 1 minus the ratio of the sum of variances within each stratified unit of environmental variables to the total variance of the whole region; the q values are sorted in descending order, and the factor with the first q value and a significance level of p<0.05 is taken as the dominant driving factor of the partition.
7. The method according to claim 1, characterized in that, The interaction types mentioned in step (4) include nonlinear enhancement, two-factor enhancement, independence, nonlinear reduction, or single-factor nonlinear reduction.
8. The method according to claim 1, characterized in that, Step (4) further includes: statistically analyzing the frequency of occurrence of each driving factor in the combination of strong interaction factors with interaction q values greater than or equal to 0.9 in each partition, and determining the most frequently occurring factor as the core synergistic factor driving the spatial distribution of antibiotics in the region.
9. The method according to claim 1, characterized in that, The study area described in step (2) is divided into four feature partitions.
10. The method according to claim 1, characterized in that, Step (1) further includes: the environmental driving factor database includes distance data and density data; the distance data is used as input to the geographic weighted regression model, and the density data is used as input to the geographic detector.
11. The method according to claim 10, characterized in that, The distance data is the weighted Euclidean distance from each sampling point to each pollution source type, and the pollution source types include at least one of hospitals, residential areas, farmland, livestock stations, aquaculture farms and pharmaceutical factories.
12. The method according to claim 10, characterized in that, The density data includes population density, land use ratio, and kernel density values for each type of pollution source.
13. The method according to claim 1, characterized in that, Step (3) further includes: before performing geographic detector analysis on each feature partition, discretizing the original data and converting it into classification data; calculating the q statistics under each combination of levels by iteratively testing different grading methods and the number of levels, and selecting the discretization scheme that maximizes the q value as the optimal grading strategy for the partition.