A pollution source contribution analysis method based on nonlinear dimension reduction of multi-dimensional environmental data

CN122839097APending Publication Date: 2026-09-29CHINESE RES ACAD OF ENVIRONMENTAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610557992.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-24
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]有鉴于此,本申请的目的在于提供一种基于多维环境数据非线性降维的污染源贡献分析方法、计算机设备和可读存储介质,能够通过引入多稳定同位素、环境理化性质、环境因子等多种环境因素整合构成多源样本数据体系(其对应多维环境数据),来提升污染源贡献分析过程中的污染源区分度,并利用UMAP(Uniform Manifold Approximation andProjection,统一流形近似和投影)降维技术(其属于非线性降维技术)提取多源样本数据体系的低维核心特征进行污染源贡献预测,来改善传统多源数据分析时存在的高维特征数据冗余、高维特征噪声大、高维特征分析效率低、高维特征分析结果精准度差等问题,从而实现污染源区分度良好且分析精准度高的污染源贡献率快速分析效果

Benefits of technology

本申请通过引入多稳定同位素、环境理化性质、环境因子等多种环境因素整合构成多源样本数据体系,获取目标流域在任意监测时段内的多个环境监测样本数据(其中,每个环境监测样本数据包括该目标流域中单个样本采集点在该监测时段内监测到的多稳定同位素数据、环境理化性质数据和环境因子数据),以利用多源样本数据体系在污染源贡献分析过程中凸显描述不同污染源之间的特征差异性,来提升污染源贡献分析过程中的污染源区分度,接着在提取多个环境监测样本数据各自的高维样本特征向量的基础上,对多个高维样本特征向量进行UMAP降维处理,得到多个低维样本核心特征向量,并同步构建与目标流域适配的目标贝叶斯同位素混合模型以对多个低维样本核心特征向量进行污染源贡献预测,得到目标流域涉及到的所有污染源各自在该监测时段内的实际污染贡献率,从而通过UMAP降维技术有效改善传统多源数据分析时存在的高维特征数据冗余、高维特征噪声大、高维特征分析效率低、高维特征分析结果精准度差等问题,使整个污染源贡献分析操作具备分析效率高、污染源区分度良好且分析精准度高等特点。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839097A_ABST
    Figure CN122839097A_ABST
Patent Text Reader

Abstract

The application provides a pollution source contribution analysis method based on nonlinear dimension reduction of multi-dimensional environmental data, and relates to the technical field of environmental governance. The application highlights the characteristic difference between different pollution sources in the pollution source contribution analysis process by introducing multiple environmental factors such as multi-stable isotopes, environmental physicochemical properties and environmental factors to form a multi-source sample data system, thereby improving the pollution source differentiation degree. Then, on the basis of extracting the high-dimensional sample feature vectors of multiple environmental monitoring sample data (which are used to represent the aforementioned multi-source sample data system), the UMAP dimension reduction technology is used to extract low-dimensional sample core features for pollution source contribution prediction, so as to improve the problems of high-dimensional feature data redundancy, high-dimensional feature noise, low high-dimensional feature analysis efficiency, low high-dimensional feature analysis result accuracy and the like in the traditional multi-source data analysis, thereby realizing the effect of fast analysis of pollution source contribution rate with good pollution source differentiation degree and high analysis accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of environmental governance technology, and more specifically, to a method for analyzing pollution source contributions based on nonlinear dimensionality reduction of multidimensional environmental data. Background Technology

[0002] With the continuous advancement of science and technology, the water environment of river basins is under multiple pressures from industrial emissions, agricultural non-point source pollution, domestic sewage, and sudden pollution accidents, resulting in an increasingly severe overall development trend of water environment problems in river basins worldwide, which seriously threatens human health and water security.

[0003] Against this backdrop, strengthening watershed ecological environmental protection has become a general consensus in the international community. Quantitative analysis of pollution sources in watersheds is a key technology for preventing water environment accidents, ensuring water quality safety, and maintaining ecological security. It not only provides a scientific basis for maintaining ecosystem balance but also lays an important foundation for formulating scientific and reasonable watershed ecological environmental protection policies. However, it is worth noting that current traditional pollution source quantitative analysis technologies generally suffer from poor source discrimination, low analysis efficiency, and poor analysis accuracy, failing to meet the needs of real-world watershed ecological environment assessment. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a pollution source contribution analysis method, computer equipment, and readable storage medium based on nonlinear dimensionality reduction of multidimensional environmental data. This method can improve the pollution source discrimination in the pollution source contribution analysis process by integrating multiple environmental factors such as multi-stable isotopes, environmental physicochemical properties, and environmental factors to form a multi-source sample data system (corresponding to multidimensional environmental data). It also utilizes UMAP (Uniform Manifold Approximation and Projection) dimensionality reduction technology (which belongs to nonlinear dimensionality reduction technology) to extract low-dimensional core features of the multi-source sample data system for pollution source contribution prediction. This improves the problems of high-dimensional feature data redundancy, high high-dimensional feature noise, low efficiency of high-dimensional feature analysis, and poor accuracy of high-dimensional feature analysis results in traditional multi-source data analysis, thereby achieving a rapid analysis effect of pollution source contribution rate with good pollution source discrimination and high analysis accuracy.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: Firstly, this application provides a pollution source contribution analysis method based on nonlinear dimensionality reduction of multidimensional environmental data, the method comprising: Acquire multiple environmental monitoring sample data of the target watershed during any monitoring period, wherein each environmental monitoring sample data includes multi-stable isotope data, environmental physicochemical property data and environmental factor data monitored by a single sample collection point in the target watershed during the monitoring period; High-dimensional feature quantization extraction is performed on the multiple environmental monitoring sample data to obtain the high-dimensional sample feature vectors of each of the multiple environmental monitoring sample data; UMAP dimensionality reduction is performed on multiple high-dimensional sample feature vectors to obtain multiple low-dimensional sample core feature vectors. Based on the core feature vectors of the multiple low-dimensional samples, an adaptive Bayesian isotope mixing model adapted to the target watershed is constructed. The target Bayesian isotope mixture model is invoked to predict the pollution source contribution based on the core feature vectors of the multiple low-dimensional samples, thereby obtaining the actual pollution contribution rate of each pollution source involved in the target watershed during the monitoring period.

[0006] In an optional implementation, the step of performing high-dimensional feature quantization extraction on the plurality of environmental monitoring sample data to obtain the high-dimensional sample feature vectors of each of the plurality of environmental monitoring sample data includes: Data cleaning was performed on the multiple environmental monitoring sample data to obtain multiple multidimensional sample index data, where each multidimensional sample index data corresponds to a single environmental monitoring sample data. Detect whether there are any missing data in the multiple multidimensional sample indicator data; If no missing data is detected in the multiple multidimensional sample index data, the multiple multidimensional sample index data is dimensionless, and the dimensionless multidimensional sample index data is transformed into feature vectors to obtain multiple high-dimensional sample feature vectors.

[0007] In an optional implementation, the step of performing high-dimensional feature quantization extraction on the plurality of environmental monitoring sample data to obtain the high-dimensional sample feature vectors of each of the plurality of environmental monitoring sample data further includes: If data loss is detected in the multiple multidimensional sample indicator data, it is determined whether the overall data loss rate of the multiple multidimensional sample indicator data exceeds a preset loss rate threshold. When the overall data missing rate is detected to exceed the preset missing rate threshold, all candidate indicators of samples that simultaneously have valid data at the multiple multidimensional sample indicator data are determined. For each multidimensional sample indicator data, extract the target dataset that corresponds to all the candidate indicators of the samples from that multidimensional sample indicator data; Multiple target datasets are dimensionless, and feature vector transformation is performed on the dimensionless target datasets to obtain multiple high-dimensional sample feature vectors.

[0008] In an optional implementation, the step of performing high-dimensional feature quantization extraction on the plurality of environmental monitoring sample data to obtain the high-dimensional sample feature vectors of each of the plurality of environmental monitoring sample data further includes: When the overall data missing rate is detected to be no more than the preset missing rate threshold, all first sample indicator data with missing data and all second sample indicator data without missing data are determined from the multiple environmental monitoring sample data. For each first sample indicator data, missing data is filled in based on all the second sample indicator data to obtain the corresponding third sample indicator data. All second-sample indicator data and all third-sample indicator data are dimensionless, and feature vector transformation is performed on the dimensionless second-sample indicator data and all third-sample indicator data respectively to obtain multiple high-dimensional sample feature vectors.

[0009] In an optional implementation, for each first sample indicator data, the step of performing missing data completion processing on the first sample indicator data based on all the second sample indicator data to obtain the corresponding third sample indicator data includes: Based on the relative positional relationship between all sample collection points within the target watershed, select all target sample indicator data whose corresponding sample collection point is close to the reference collection point from all the second sample indicator data, wherein the reference collection point is the sample collection point to which the first sample indicator data belongs; Based on all the target sample indicator data, a transfer learning completion model is trained, and the trained transfer learning completion model is used to complete the missing data indicators involved in the first sample indicator data to obtain the third sample indicator data corresponding to the first sample indicator data.

[0010] In an optional implementation, the step of adaptively constructing a target Bayesian isotope mixture model adapted to the target watershed based on the multiple low-dimensional sample core feature vectors includes: An initial Bayesian isotope mixing model is constructed, and the model parameter constraints of the initial Bayesian isotope mixing model at the target watershed are determined. The ant colony optimization algorithm is used to optimize the model parameters based on the core feature vectors of the multiple low-dimensional samples to obtain the optimal model parameters that satisfy the model parameter constraints. The initial Bayesian isotope mixing model is updated according to the optimal model parameters to obtain the target Bayesian isotope mixing model.

[0011] In an optional implementation, the method further includes: Obtain environmental sample datasets corresponding to the target watershed in multiple historical monitoring periods, wherein each environmental sample dataset includes multiple low-dimensional sample core feature vectors of the target watershed in the corresponding historical monitoring period, and the historical pollution contribution rate of each pollution source involved in the target watershed in the historical monitoring period. A support vector regression model was trained based on multiple environmental sample datasets to obtain a prediction model for the target pollution source contribution of the target watershed.

[0012] In an optional implementation, the method further includes: Acquire multiple sample data to be analyzed in the target watershed during the target monitoring period, wherein each sample data to be analyzed includes multi-stable isotope data, environmental physicochemical property data and environmental factor data monitored by a single sample collection point in the target watershed during the target monitoring period; High-dimensional feature quantization extraction is performed on the multiple sample data to be analyzed to obtain the high-dimensional sample feature vectors of each of the multiple sample data to be analyzed. UMAP dimensionality reduction is performed on multiple high-dimensional sample feature vectors to obtain multiple low-dimensional sample core feature vectors. The target pollution source contribution prediction model is invoked to predict the pollution source contribution based on the core feature vectors of the multiple low-dimensional samples, so as to obtain the theoretical pollution contribution rate of each pollution source involved in the target watershed during the target monitoring period.

[0013] Secondly, this application provides a computer device, including a processor and a memory, wherein the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the pollution source contribution analysis method based on nonlinear dimensionality reduction of multidimensional environmental data as described in any of the foregoing embodiments.

[0014] Thirdly, this application provides a readable storage medium storing a computer program thereon, which, when executed by a computer device, implements the pollution source contribution analysis method based on nonlinear dimensionality reduction of multidimensional environmental data as described in any of the foregoing embodiments.

[0015] In this case, the beneficial effects of the embodiments of this application may include the following: This application integrates multiple environmental factors, such as multi-stable isotopes, environmental physicochemical properties, and environmental factors, to form a multi-source sample data system. This system acquires multiple environmental monitoring sample data from a target watershed during any monitoring period (each environmental monitoring sample data includes multi-stable isotope data, environmental physicochemical property data, and environmental factor data monitored at a single sampling point within the target watershed during that monitoring period). The aim is to leverage this multi-source sample data system to highlight the characteristic differences between different pollution sources during pollution source contribution analysis, thereby improving the source discrimination in the analysis process. Then, based on the extraction of high-dimensional sample feature vectors from each of the multiple environmental monitoring sample data... UMAP dimensionality reduction is performed on multiple high-dimensional sample feature vectors to obtain multiple low-dimensional sample core feature vectors. Simultaneously, a target Bayesian isotope mixture model adapted to the target watershed is constructed to predict the pollution source contribution of multiple low-dimensional sample core feature vectors. This yields the actual pollution contribution rate of each pollution source involved in the target watershed during the monitoring period. Thus, UMAP dimensionality reduction technology effectively improves the problems of high-dimensional feature data redundancy, high noise in high-dimensional features, low efficiency of high-dimensional feature analysis, and poor accuracy of high-dimensional feature analysis results in traditional multi-source data analysis. This makes the entire pollution source contribution analysis operation have the characteristics of high analysis efficiency, good pollution source discrimination, and high analysis accuracy.

[0016] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A schematic diagram of the composition of a computer device provided in the embodiments of this application; Figure 2 One of the flowcharts for the pollution source contribution analysis method based on nonlinear dimensionality reduction of multidimensional environmental data provided in the embodiments of this application; Figure 3 for Figure 2 One of the flowcharts illustrating the sub-steps included in step S220; Figure 4 for Figure 2 The second flowchart of the sub-steps included in step S220; Figure 5 for Figure 2The third flowchart of the sub-steps included in step S220; Figure 6 for Figure 2 A flowchart illustrating the sub-steps included in step S240; Figure 7 The second flowchart illustrates the pollution source contribution analysis method based on nonlinear dimensionality reduction of multidimensional environmental data provided in this application embodiment. Figure 8 The third flowchart illustrates the pollution source contribution analysis method based on nonlinear dimensionality reduction of multidimensional environmental data provided in this application embodiment.

[0019] Icons: 10-Computer equipment; 11-Memory; 12-Processor; 13-Communication unit. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0021] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0022] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0023] In the description of this application, it should be understood that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are used only for the convenience of describing this application and simplifying the description, and are not intended to indicate or imply that the equipment or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0024] In the description of this application, it should also be noted that, unless otherwise expressly specified and limited, the terms "set up," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0025] Furthermore, it is understood in the description of this application that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Those skilled in the art will understand the specific meaning of the above terms in this application based on the specific circumstances.

[0026] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0027] Please refer to Figure 1 , Figure 1 This is a schematic diagram of the computer device 10 provided in this embodiment of the application. In this embodiment, the computer device 10 can be used to quickly analyze the pollution source contribution rate of any watershed with good differentiation and high analysis accuracy, so that environmental governance personnel can formulate reasonable watershed ecological and environmental protection policies based on the pollution source contribution analysis results of the corresponding watershed. The computer device 10 may be, but is not limited to, a server, a personal computer, a laptop computer, etc.

[0028] In this embodiment of the application, the computer device 10 may include a memory 11, a processor 12, and a communication unit 13. The memory 11, the processor 12, and the communication unit 13 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines.

[0029] In this embodiment, the memory 11 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. The memory 11 is used to store computer programs, and the processor 12 can execute the computer programs accordingly after receiving execution instructions.

[0030] In this embodiment, the processor 12 can be an integrated circuit chip with signal processing capabilities. The processor 12 can be a general-purpose processor, including at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Network Processor (NP), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in this embodiment.

[0031] In this embodiment, the communication unit 13 is used to establish a communication connection between the computer device 10 and other electronic devices via a network, and to send and receive data via the network, wherein the network includes wired communication networks and wireless communication networks. For example, the computer device 10 can obtain environmental monitoring sample data from different sampling points within the same watershed at different monitoring periods through the communication unit 13, wherein each environmental monitoring sample data belongs to a multi-stable isotope (which may involve, but is not limited to, multiple light element stable isotopes (e.g., nitrogen stable isotope¹)). 5 N, oxygen stable isotopes¹ 8 O, carbon stable isotope ¹³C, sulfur stable isotope ³ 4 S, hydrogen stable isotope²H, boron stable isotope¹¹B, etc.) and various heavy metal stable isotopes (e.g., lead stable isotope²). 06 Pb,²07 Pb and² 08 Pb, stable isotopes of cadmium¹¹ 4 CD&¹¹ 0 Cd, zinc stable isotopes 66 Zn and 64 Zn, stable isotope of mercury² 0 ²Hg and ¹ 98 The data system integrates various environmental factors, including soil physicochemical properties (such as soil pH, soil organic matter content, soil total nitrogen content, soil total phosphorus content, soil total potassium content, soil moisture content, soil bulk density, and soil texture type) or water physicochemical properties (such as water pH, dissolved oxygen concentration, water conductivity, water turbidity, chemical oxygen demand, biochemical oxygen demand, ammonia nitrogen concentration, nitrate nitrogen concentration, total phosphorus concentration, and water temperature), environmental factors (which may include, but are not limited to, natural ecological factors (such as precipitation, temperature, runoff, and evaporation), and human activity factors (such as land use type, fertilizer application intensity, pesticide application intensity, wastewater discharge, and industrial activity intensity indicators), to form a multi-source sample data system (i.e., corresponding multi-dimensional environmental data).

[0032] In this embodiment, the computer device 10 may pre-store a specific computer program related to the pollution source contribution analysis function in the memory 11. By driving the processor 12 to execute the specific computer program, a multi-source sample data system is formed by integrating multiple environmental factors such as multiple stable isotopes, environmental physicochemical properties, and environmental factors to improve the pollution source discrimination in the pollution source contribution analysis process. On the basis of improving the pollution source discrimination in the pollution source contribution analysis process, nonlinear dimensionality reduction technology (e.g., UMAP dimensionality reduction technology) is used to extract the low-dimensional core features of the multi-source sample data system for pollution source contribution prediction. This improves the problems of high-dimensional feature data redundancy, high-dimensional feature noise, low efficiency of high-dimensional feature analysis, and poor accuracy of high-dimensional feature analysis results in traditional multi-source data analysis. Thus, a rapid analysis effect of pollution source contribution rate with good pollution source discrimination and high analysis accuracy is achieved.

[0033] Understandable, Figure 1 The block diagram shown is only a schematic diagram of one configuration of the computer device 10. The computer device 10 may also include components such as... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.

[0034] In this application, to ensure that the computer device 10 can achieve rapid analysis of pollution source contribution rates with good pollution source discrimination and high analysis accuracy, this application embodiment provides a pollution source contribution analysis method based on nonlinear dimensionality reduction of multidimensional environmental data to achieve the aforementioned objective. The pollution source contribution analysis method provided in this application will be described in detail below.

[0035] Please refer to Figure 2 , Figure 2 This is one of the flowcharts illustrating the pollution source contribution analysis method based on nonlinear dimensionality reduction of multidimensional environmental data provided in this application embodiment. In this application embodiment, Figure 2 The pollution source contribution analysis method shown may include steps S210 to S250.

[0036] Step S210: Obtain multiple environmental monitoring sample data of the target watershed during any monitoring period, wherein each environmental monitoring sample data includes multi-stable isotope data, environmental physicochemical property data and environmental factor data monitored by a single sample collection point in the target watershed during the monitoring period.

[0037] In this embodiment, the target watershed is the watershed where pollution source contribution analysis needs to be performed; multiple environmental monitoring sample data each correspond to a sample collection point in the target watershed, and the multi-stable isotope data in a single environmental monitoring sample data may include nitrogen stable isotopes¹ 5 N, oxygen stable isotopes¹ 8 O, carbon stable isotope ¹³C, sulfur stable isotope ³ 4 The isotopic δ values ​​of S, hydrogen stable isotope²H, and boron stable isotope¹¹B, as well as the lead stable isotope² 06 Pb relative to lead stable isotopes² 07 Pb isotope ratios, lead stable isotopes² 08 Pb relative to lead stable isotopes² 06 Pb isotope ratios, cadmium stable isotopes¹¹ 4 Cd relative to stable cadmium isotopes¹¹ 0 Cd isotope ratios, zinc stable isotopes 66 Zn stable isotopes relative to zinc 64 Zn isotope ratios, stable mercury isotopes² 0 ²Hg relative to stable mercury isotopes¹ 98 Isotope ratios of Hg, etc.

[0038] Therefore, this application can integrate multiple environmental factors such as stable isotopes, environmental physicochemical properties, and environmental factors to form a multi-source sample data system, so as to highlight the characteristic differences between different pollution sources in the process of pollution source contribution analysis, thereby improving the differentiation of pollution sources in the target watershed and avoiding the phenomenon that "traditional single isotope index system is difficult to distinguish pollution sources with similar isotope characteristics".

[0039] Step S220: High-dimensional feature quantization extraction is performed on multiple environmental monitoring sample data to obtain the high-dimensional sample feature vectors of each environmental monitoring sample data.

[0040] In this embodiment, the computer device 10 can obtain the high-dimensional sample feature vector of each environmental monitoring sample data in the high-dimensional feature space by performing unified data modeling processing on the multiple environmental monitoring sample data in the same high-dimensional feature space.

[0041] Alternatively, please refer to Figure 3 , Figure 3 yes Figure 2 One of the flowcharts for the sub-steps included in step S220. In this embodiment of the application, step S220 may include sub-steps S221 to S223 to achieve reliable high-dimensional sample feature extraction for a complete environmental monitoring sample dataset (which consists of multiple environmental monitoring sample data from the same monitoring period).

[0042] Sub-step S221 involves cleaning the data from multiple environmental monitoring samples to obtain multiple multidimensional sample indicator data.

[0043] In this embodiment, the data cleaning operations involved in sub-step S221 may include, but are not limited to, low-quality data removal and outlier correction. Each multidimensional sample indicator data corresponds to a separate environmental monitoring sample data.

[0044] Sub-step S222 checks whether there are missing data in the multidimensional sample index data.

[0045] Sub-step S223: If no missing data is detected in multiple multidimensional sample index data, the multiple multidimensional sample index data is dimensionless, and the dimensionless multidimensional sample index data is transformed into feature vectors to obtain multiple high-dimensional sample feature vectors.

[0046] In this embodiment, the dimensionless effect between the multiple multidimensional sample index data can be achieved by using standardization or normalization operations.

[0047] Therefore, this application can achieve reliable high-dimensional sample feature extraction function for complete environmental monitoring sample data by executing the above sub-steps S221 to S223.

[0048] Alternatively, please refer to Figure 4 , Figure 4 yes Figure 2 The second flowchart illustrates the sub-steps included in step S220. In this embodiment, [the flowchart is] related to... Figure 3 In comparison, step S220 is in Figure 4 The process may also include sub-steps S224 to S227 to adaptively reduce the high-dimensional feature dimension for environmental monitoring sample datasets with severely missing data, so that the extracted high-dimensional sample feature vectors can ensure the stability and reliability of the final pollution source contribution analysis results.

[0049] Sub-step S224: If multiple multidimensional sample indicator data are found to have missing data, check whether the overall missing data rate of the multiple multidimensional sample indicator data exceeds the preset missing rate threshold.

[0050] Sub-step S225: When the overall data missing rate is detected to exceed the preset missing rate threshold, determine all candidate indicators of samples that simultaneously have valid data at multiple multidimensional sample indicator data.

[0051] In this embodiment, when the overall data missing rate exceeds a preset missing rate threshold (e.g., 40%), it indicates that the corresponding environmental monitoring sample dataset actually has a serious data missing phenomenon. At this time, effective data statistics can be performed on the multiple multidimensional sample indicator data to determine all sample candidate indicators that simultaneously have valid data at multiple multidimensional sample indicator data.

[0052] Sub-step S226: For each multidimensional sample indicator data, extract the target dataset corresponding to all the candidate indicators of the samples in the multidimensional sample indicator data.

[0053] In this embodiment, a single target dataset consists of the valid data corresponding to each of the candidate indicators of the samples at the same multidimensional sample indicator data; each multidimensional sample indicator data separately constructs a target dataset.

[0054] Sub-step S227 involves performing dimensionless processing on multiple target datasets, and then performing feature vector transformation on the dimensionless target datasets to obtain multiple high-dimensional sample feature vectors.

[0055] Therefore, by executing the above sub-steps S224 to S227, this application can adaptively modify the high-dimensional feature dimension of the environmental monitoring sample dataset with severely missing data, so that the extracted high-dimensional sample feature vector can ensure the stability and reliability of the final pollution source contribution analysis results.

[0056] Alternatively, please refer to Figure 5 , Figure 5 yes Figure 2 The third flowchart illustrates the sub-steps included in step S220. In this embodiment, [the process is similar to...] Figure 4 In comparison, step S220 is in Figure 5 The process may also include sub-steps S228 to S2210 to perform adaptive data completion for environmental monitoring sample datasets with slight data gaps, so that the extracted high-dimensional sample feature vectors can effectively maintain the desired feature dimension.

[0057] Sub-step S228: When the overall data missing rate is detected to be no more than the preset missing rate threshold, determine all first sample indicator data with data missing phenomena and all second sample indicator data without data missing phenomena in multiple environmental monitoring sample data.

[0058] In this embodiment, when the overall data missing rate does not exceed a preset missing rate threshold (e.g., 40%), it indicates that the corresponding environmental monitoring sample dataset only has a slight data missing phenomenon.

[0059] Sub-step S229: For each first sample indicator data, perform missing data completion processing on the first sample indicator data based on all second sample indicator data to obtain the corresponding third sample indicator data.

[0060] In this embodiment, the feature correlation of source data can be used to effectively complete missing data, thereby ensuring the data reliability of the corresponding third sample indicator data. Each first sample indicator data corresponds to a separate third sample indicator data. Specifically, for each first sample indicator data, the step "perform missing data completion processing on the first sample indicator data based on all second sample indicator data to obtain the corresponding third sample indicator data" may include: Based on the relative positional relationship between all sample collection points within the target watershed, select all target sample indicator data whose corresponding sample collection point is close to the reference collection point from all the second sample indicator data, wherein the reference collection point is the sample collection point to which the first sample indicator data belongs; Based on all the target sample indicator data, a transfer learning completion model is trained, and the trained transfer learning completion model is used to complete the missing data indicators involved in the first sample indicator data to obtain the third sample indicator data corresponding to the first sample indicator data.

[0061] Therefore, this application can reliably complete missing data by executing the above sub-step S229 and utilizing the feature correlation of the same source data.

[0062] Sub-step S2210 involves performing dimensionless processing on all second-sample index data and all third-sample index data, and then performing feature vector transformation on all second-sample index data and all third-sample index data after dimensionless processing to obtain multiple high-dimensional sample feature vectors.

[0063] Therefore, by executing the above sub-steps S228 to S2210, this application can perform adaptive data completion for environmental monitoring sample datasets with slightly missing data, so that the extracted high-dimensional sample feature vectors can effectively maintain the desired feature dimension without reducing the high-dimensional feature dimension.

[0064] Step S230: Perform UMAP dimensionality reduction on multiple high-dimensional sample feature vectors to obtain multiple low-dimensional sample core feature vectors.

[0065] In this embodiment, since UMAP dimensionality reduction technology is a nonlinear dimensionality reduction technology, it can construct a neighborhood topology graph of high-dimensional feature data and ensure the global and local structural correlation of feature data in the low-dimensional feature space. It has the characteristics of high computational efficiency, adaptability to large-scale datasets, and the ability to realize feature dimensionality reduction and core information extraction of multi-source high-dimensional data. As a result, the computer device 10 can use UMAP dimensionality reduction technology to learn the global topological structure and local neighborhood relationship between multiple high-dimensional sample feature vectors, and map multiple high-dimensional sample feature vectors to the low-dimensional feature space. This allows the corresponding multiple low-dimensional sample core feature vectors to retain the core distinguishing information of the data while significantly reducing the feature dimension, thereby achieving accurate screening of core features and providing high-quality input features for subsequent pollution source contribution prediction operations.

[0066] Step S240: Adaptively construct a target Bayesian isotope hybrid model that is compatible with the target watershed based on the core feature vectors of multiple low-dimensional samples.

[0067] In this embodiment, the Bayesian isotope mixture model, as a mainstream technique for quantitative analysis of pollution sources, constructs the isotopic probability distributions of pollution sources and receptor samples in a watershed, and calculates the contribution rate of each pollution source in the watershed by combining fractionation effect parameters. This model employs the Markov Chain Monte Carlo (MCMC) algorithm for parameter iteration to ensure reliable pollution source contribution rate analysis results. Specifically, a suitable target Bayesian isotope mixture model needs to be constructed for each target watershed at different monitoring periods.

[0068] Alternatively, please refer to Figure 6 , Figure 6 yes Figure 2 A flowchart illustrating the sub-steps included in step S240. In this embodiment, step S240 may include sub-steps S241 to S243 to achieve automatic global optimization of the model parameters of the Bayesian isotope mixture model, ensuring that the target Bayesian isotope mixture model constructed for a single monitoring period can effectively adapt to the watershed environmental characteristics of the target watershed during that monitoring period.

[0069] Sub-step S241: Construct an initial Bayesian isotope mixing model and determine the model parameter constraints of the initial Bayesian isotope mixing model at the target watershed.

[0070] In this embodiment, the model parameter constraints are used to limit the distribution of parameter values ​​of all model parameters when the initial Bayesian isotope mixing model is applied to the target watershed. Since all model parameters involved in the traditional Bayesian isotope mixing model can include "pollution source prior parameters", "stable isotope fractionation parameters", "stable isotope weights" and "MCMC parameters", the model parameter constraints can include "the distribution range of Dirichlet prior parameters of all pollution sources that may be involved in the target watershed (including but not limited to: nitrate pollution sources, organic matter pollution sources, heavy metal pollution sources, etc.) (e.g., [0.1, 10])", "the distribution range of isotope fractionation coefficients of all stable isotopes (e.g., [-5‰, 5‰])", "the sum of isotope weights of all stable isotopes is 1, and the distribution range of isotope weights of each stable isotope is [0, 1]", and "the distribution range of the number of MCMC single chain iterations (e.g., [10000, 200000]), the distribution range of the number of MCMC chains (e.g., [2, 6]), and the parameter distribution range of the MCMC iteration step size (e.g., [0.01, 1])".

[0071] Sub-step S242 uses the ant colony optimization algorithm to optimize the model parameters based on the core feature vectors of multiple low-dimensional samples, and obtains the optimal model parameters that satisfy the model parameter constraints.

[0072] In this embodiment, when the ant colony optimization algorithm is applied to the model parameter optimization process of the Bayesian isotope mixture model, each "ant" represents a set of optimized model parameters (which may include the Dirichlet prior parameter values ​​of all pollution sources that may be involved in the target watershed, the isotope fractionation coefficient values ​​of all stable isotopes, the isotope weight values ​​of all stable isotopes, the number of iterations of a single MCMC chain, the number of MCMC chains, the MCMC iteration step size, etc.). It can search for the optimal combination of model parameters in the parameter space based on the pheromone state transition rules. In the model parameter optimization process, by reasonably setting parameters such as the number of ants, the number of iterations, the pheromone evaporation coefficient, and the heuristic function weight, it can be ensured that the output optimal model parameters can effectively adapt to the watershed environmental characteristics of the target watershed in the corresponding monitoring period (i.e., the monitoring period to which the core feature vectors of the multiple low-dimensional samples belong), thus avoiding the problems of strong configuration subjectivity, slow convergence speed, and difficulty in obtaining the global optimal solution in the traditional Bayesian isotope mixture model parameter configuration process.

[0073] Sub-step S243 updates the parameters of the initial Bayesian isotope mixture model according to the optimal model parameters to obtain the target Bayesian isotope mixture model.

[0074] Therefore, by executing the above sub-steps S241 to S243, this application can achieve automatic global optimization of the model parameters of the Bayesian isotope mixture model, ensuring that the target Bayesian isotope mixture model constructed for a single monitoring period can effectively adapt to the watershed environmental characteristics of the target watershed during that monitoring period.

[0075] Step S250: Call the target Bayesian isotope mixture model to predict the pollution source contribution based on multiple low-dimensional sample core feature vectors, and obtain the actual pollution contribution rate of each pollution source involved in the target watershed during the monitoring period.

[0076] Therefore, by executing the above steps S210 to S250, this application introduces a multi-source sample data system composed of multiple environmental factors such as stable isotopes, environmental physicochemical properties, and environmental factors to improve the pollution source discrimination in the pollution source contribution analysis process. It also uses UMAP dimensionality reduction technology to extract the low-dimensional core features of the multi-source sample data system and combines it with Bayesian isotope mixture model adaptive construction operation to predict pollution source contributions. This improves the problems of high-dimensional feature data redundancy, high-dimensional feature noise, low efficiency of high-dimensional feature analysis, and poor accuracy of high-dimensional feature analysis results in traditional multi-source data analysis, thereby achieving a rapid analysis effect of pollution source contribution rate with good pollution source discrimination and high analysis accuracy.

[0077] In this application, because the target Bayesian isotope mixture model is essentially only applicable to pollution source contribution rate analysis for a single monitoring period, it cannot be directly transferred to other monitoring periods (including historical or future monitoring periods) for pollution source contribution rate analysis. Therefore, it is practically difficult to adapt to pollution source contribution rate analysis operations in dynamic watershed environments. Thus, this application can... Figure 2 Based on the pollution source contribution analysis method shown, a further pollution source contribution analysis method is provided to train a rapid pollution source contribution analysis model with good dynamic prediction capabilities that can be reused in different monitoring periods.

[0078] Alternatively, please refer to Figure 7 , Figure 7 This is the second schematic flowchart of the pollution source contribution analysis method based on nonlinear dimensionality reduction of multidimensional environmental data provided in this application embodiment. In this application embodiment, with Figure 2 Compared to the pollution source contribution analysis methods shown, Figure 7 The pollution source contribution analysis method shown may also include steps S310 to S320 to train a rapid pollution source contribution analysis model with good dynamic prediction capabilities that can be reused in different monitoring periods.

[0079] Step S310: Obtain environmental sample datasets corresponding to the target watershed in multiple historical monitoring periods. Each environmental sample dataset includes multiple low-dimensional sample core feature vectors of the target watershed in the corresponding historical monitoring period, as well as the historical pollution contribution rate of all pollution sources involved in the target watershed in that historical monitoring period.

[0080] Step S320: Training a support vector regression model based on multiple environmental sample datasets to obtain a target pollution source contribution prediction model for the target watershed.

[0081] Therefore, by executing the above steps S310 to S320, this application can train a rapid analysis model of pollution source contribution with good dynamic prediction capabilities and reusable to different monitoring periods.

[0082] Alternatively, please refer to Figure 8 , Figure 8 This is the third flowchart illustrating the pollution source contribution analysis method based on nonlinear dimensionality reduction of multidimensional environmental data provided in this application embodiment. In this application embodiment, [the method is related to...]. Figure 7 Compared to the pollution source contribution analysis methods shown, Figure 8 The pollution source contribution analysis method shown may also include steps S410 to S440, so as to quickly achieve the pollution source contribution analysis effect of the target watershed under any monitoring period by using the trained target pollution source contribution prediction model (i.e. the above-mentioned rapid pollution source contribution analysis model).

[0083] Step S410: Obtain multiple sample data to be analyzed in the target watershed during the target monitoring period. Each sample data to be analyzed includes multi-stable isotope data, environmental physicochemical property data, and environmental factor data monitored by a single sample collection point in the target watershed during the target monitoring period.

[0084] In this embodiment, the target monitoring period is any period to be analyzed.

[0085] Step S420: High-dimensional feature quantization extraction is performed on multiple sample data to be analyzed to obtain high-dimensional sample feature vectors for each of the multiple sample data to be analyzed.

[0086] In this embodiment, the computer device 10 can refer to the specific execution details of sub-steps S221 to S2210 above to perform the high-dimensional sample feature vector extraction operation on the multiple sample data to be analyzed.

[0087] Step S430: Perform UMAP dimensionality reduction on multiple high-dimensional sample feature vectors to obtain multiple low-dimensional sample core feature vectors.

[0088] Step S440: Call the target pollution source contribution prediction model to predict the pollution source contribution based on multiple low-dimensional sample core feature vectors, and obtain the theoretical pollution contribution rate of each pollution source involved in the target watershed during the target monitoring period.

[0089] Therefore, by executing the above steps S410 to S440, this application can quickly achieve the pollution source contribution analysis effect of the target watershed under any monitoring period by utilizing the trained target pollution source contribution prediction model.

[0090] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of the apparatus, methods, and computer program products according to embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0091] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part. If the function is implemented as a software functional module and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.

[0092] The above descriptions are merely various embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for analyzing pollution source contributions based on nonlinear dimensionality reduction of multidimensional environmental data, characterized in that, The method includes: Acquire multiple environmental monitoring sample data of the target watershed during any monitoring period, wherein each environmental monitoring sample data includes multi-stable isotope data, environmental physicochemical property data and environmental factor data monitored by a single sample collection point in the target watershed during the monitoring period; High-dimensional feature quantization extraction is performed on the multiple environmental monitoring sample data to obtain the high-dimensional sample feature vectors of each of the multiple environmental monitoring sample data; UMAP dimensionality reduction is performed on multiple high-dimensional sample feature vectors to obtain multiple low-dimensional sample core feature vectors. Based on the core feature vectors of the multiple low-dimensional samples, an adaptive Bayesian isotope mixing model adapted to the target watershed is constructed. The target Bayesian isotope mixture model is invoked to predict the pollution source contribution based on the core feature vectors of the multiple low-dimensional samples, thereby obtaining the actual pollution contribution rate of each pollution source involved in the target watershed during the monitoring period.

2. The method according to claim 1, characterized in that, The step of extracting high-dimensional features from the plurality of environmental monitoring sample data to obtain the high-dimensional sample feature vectors of each of the plurality of environmental monitoring sample data includes: Data cleaning was performed on the multiple environmental monitoring sample data to obtain multiple multidimensional sample index data, where each multidimensional sample index data corresponds to a single environmental monitoring sample data. Detect whether there are any missing data in the multiple multidimensional sample indicator data; If no missing data is detected in the multiple multidimensional sample index data, the multiple multidimensional sample index data is dimensionless, and the dimensionless multidimensional sample index data is transformed into feature vectors to obtain multiple high-dimensional sample feature vectors.

3. The method according to claim 2, characterized in that, The step of extracting high-dimensional features from the plurality of environmental monitoring sample data to obtain the high-dimensional sample feature vectors of each of the plurality of environmental monitoring sample data further includes: If data loss is detected in the multiple multidimensional sample indicator data, it is determined whether the overall data loss rate of the multiple multidimensional sample indicator data exceeds a preset loss rate threshold. When the overall data missing rate is detected to exceed the preset missing rate threshold, all candidate indicators of samples that simultaneously have valid data at the multiple multidimensional sample indicator data are determined. For each multidimensional sample indicator data, extract the target dataset that corresponds to all the candidate indicators of the samples from that multidimensional sample indicator data; Multiple target datasets are dimensionless, and feature vector transformation is performed on the dimensionless target datasets to obtain multiple high-dimensional sample feature vectors.

4. The method according to claim 3, characterized in that, The step of extracting high-dimensional features from the plurality of environmental monitoring sample data to obtain the high-dimensional sample feature vectors of each of the plurality of environmental monitoring sample data further includes: When the overall data missing rate is detected to be no more than the preset missing rate threshold, all first sample indicator data with missing data and all second sample indicator data without missing data are determined from the multiple environmental monitoring sample data. For each first sample indicator data, missing data is filled in based on all the second sample indicator data to obtain the corresponding third sample indicator data. All second-sample indicator data and all third-sample indicator data are dimensionless, and feature vector transformation is performed on the dimensionless second-sample indicator data and all third-sample indicator data respectively to obtain multiple high-dimensional sample feature vectors.

5. The method according to claim 4, characterized in that, For each first sample indicator data, the step of performing missing data completion processing on the first sample indicator data based on all the second sample indicator data to obtain the corresponding third sample indicator data includes: Based on the relative positional relationship between all sample collection points within the target watershed, select all target sample indicator data whose corresponding sample collection point is close to the reference collection point from all the second sample indicator data, wherein the reference collection point is the sample collection point to which the first sample indicator data belongs; Based on all the target sample indicator data, a transfer learning completion model is trained, and the trained transfer learning completion model is used to complete the missing data indicators involved in the first sample indicator data to obtain the third sample indicator data corresponding to the first sample indicator data.

6. The method according to claim 1, characterized in that, The step of adaptively constructing a target Bayesian isotope mixing model adapted to the target watershed based on the core feature vectors of the multiple low-dimensional samples includes: An initial Bayesian isotope mixing model is constructed, and the model parameter constraints of the initial Bayesian isotope mixing model at the target watershed are determined. The ant colony optimization algorithm is used to optimize the model parameters based on the core feature vectors of the multiple low-dimensional samples to obtain the optimal model parameters that satisfy the model parameter constraints. The initial Bayesian isotope mixing model is updated according to the optimal model parameters to obtain the target Bayesian isotope mixing model.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Obtain environmental sample datasets corresponding to the target watershed in multiple historical monitoring periods, wherein each environmental sample dataset includes multiple low-dimensional sample core feature vectors of the target watershed in the corresponding historical monitoring period, and the historical pollution contribution rate of each pollution source involved in the target watershed in the historical monitoring period. A support vector regression model was trained based on multiple environmental sample datasets to obtain a prediction model for the target pollution source contribution of the target watershed.

8. The method according to claim 7, characterized in that, The method further includes: Acquire multiple sample data to be analyzed in the target watershed during the target monitoring period, wherein each sample data to be analyzed includes multi-stable isotope data, environmental physicochemical property data and environmental factor data monitored by a single sample collection point in the target watershed during the target monitoring period; High-dimensional feature quantization extraction is performed on the multiple sample data to be analyzed to obtain the high-dimensional sample feature vectors of each of the multiple sample data to be analyzed. UMAP dimensionality reduction is performed on multiple high-dimensional sample feature vectors to obtain multiple low-dimensional sample core feature vectors. The target pollution source contribution prediction model is invoked to predict the pollution source contribution based on the core feature vectors of the multiple low-dimensional samples, so as to obtain the theoretical pollution contribution rate of each pollution source involved in the target watershed during the target monitoring period.

9. A computer device, characterized in that, The system includes a processor and a memory, the memory storing a computer program executable by the processor, which can execute the computer program to implement the pollution source contribution analysis method based on nonlinear dimensionality reduction of multidimensional environmental data as described in any one of claims 1-8.

10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a computer device, it implements the pollution source contribution analysis method based on nonlinear dimensionality reduction of multidimensional environmental data as described in any one of claims 1-8.