Industrial park pollution source rapid identification method based on fuzzy clustering
By combining the possibility C-mean and fuzzy rough C-mean algorithm and improved differential evolution algorithm to optimize the fuzzy clustering parameters, the pollution diffusion model is constructed, and the adaptability and accuracy of the pollution source identification method in the existing technology in complex environments is solved, and efficient and accurate pollution source identification and traceability analysis are achieved.
Patent Information
- Application Number
- CN202510766423.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing pollution source identification methods have limitations in processing uncertain pollution data, optimizing fuzzy clustering parameters, and combining pollution diffusion models for traceability analysis. Especially in complex environments, it is poorly adaptable and difficult to achieve high-precision and high-rootability pollution source identification.
Combining the possibility C-mean and fuzzy rough C-mean algorithms, the uncertainty of the pollution data is processed, the fuzzy clustering parameters are optimized through the improved differential evolution algorithm, and the pollution diffusion model is constructed to optimize the spatial distribution identification of pollution sources, and the membership is dynamically adjusted to adapt to the fluctuations of pollution data.
It improves the accuracy and robustness of pollution source identification, enhances pollution traceability capabilities, can quickly and accurately identify pollution sources and optimize diffusion paths in complex environments, and improves the stability and real-time identification.
Smart Images

Figure CN120277448A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of environmental monitoring and pollution source identification, and particularly to a rapid pollution source identification method for industrial parks based on fuzzy clustering. Background Art
[0002] Industrial parks are important carriers of modern industrial production. However, with the acceleration of the industrialization process, the pollution problems in industrial parks have become increasingly serious and have become one of the important challenges in environmental protection and ecological security. The pollution sources in industrial parks are complex, including various types such as air pollution, water pollution, and solid waste pollution. These pollutants not only affect the surrounding ecological environment but also pose potential threats to human health. Therefore, how to efficiently and accurately identify pollution sources and conduct pollution source tracing analysis has become the research focus in the field of environmental monitoring. Currently, pollution source identification methods mainly include traditional statistical analysis methods, pollution source tracing methods based on physical and chemical models, and data-driven intelligent computing methods.
[0003] Traditional pollution source identification methods usually rely on pollutant concentration monitoring data and combine empirical models or statistical analysis methods for pollution source identification. For example, multivariate statistical analysis methods such as principal component analysis (PCA), factor analysis (FA), etc. can be used to identify the main sources of pollutants. However, such methods often ignore the non-linear diffusion characteristics of pollutants in the environment and are difficult to handle the uncertainty in pollution data. In addition, physical and chemical models such as the Lagrangian particle diffusion model, Gaussian plume model, etc. can simulate the diffusion process of pollutants. However, these methods usually rely on high-precision meteorological data and emission inventories and are less adaptable in complex environments due to the limitations of data integrity and computational complexity.
[0004] In recent years, machine learning and data mining technologies have been widely applied in pollution source identification. Among them, the fuzzy clustering algorithm has become an important tool for pollution source identification because it can handle the uncertainty and ambiguity of pollution data. Fuzzy C-means is a commonly used unsupervised learning method that can assign pollution data points to multiple pollution source categories. However, it is sensitive to abnormal data and is easily interfered by noise. To address this problem, the possibility C-means algorithm was proposed. By introducing a possibility distribution parameter, it can reduce the influence of abnormal data on the clustering center and improve the robustness of clustering. However, the possibility C-means algorithm still has a membership contraction problem when dealing with high-dimensional data and is difficult to effectively distinguish overlapping pollution sources. At the same time, the fuzzy rough C-means method combines rough set theory on the basis of fuzzy clustering and uses upper and lower approximation sets to identify data, thereby enhancing the ability to process uncertain data. However, the fuzzy rough C-means method has a strong dependence on parameters and is difficult to automatically adjust the optimal parameters in practical applications, which affects the accuracy of pollution source identification.
[0005] In terms of pollution source tracing, pollution diffusion models can be used to simulate the propagation path of pollutants in the atmosphere or water bodies to improve the spatial resolution of pollution source identification. The commonly used Gaussian plume model can better simulate the diffusion process of atmospheric pollutants. However, the accuracy of this model depends on external environmental factors such as the pollutant emission rate, wind speed, and wind direction. In an environment with complex terrain or large variations in meteorological conditions, the model accuracy decreases. In addition, pollution diffusion models are usually used separately from pollution source identification methods, lacking a joint optimization strategy, resulting in possible deviations between the pollutant diffusion path and the pollution source identification result, affecting the reliability of pollution source tracing.
[0006] Regarding the problem of optimizing the parameters of the fuzzy clustering algorithm, in recent years, intelligent optimization algorithms such as genetic algorithms, particle swarm optimization, and differential evolution have been widely used to optimize fuzzy clustering parameters. However, the traditional differential evolution algorithm usually adopts fixed mutation factors and crossover probabilities, which are prone to falling into local optima in the initial search stage, resulting in limited accuracy of pollution source identification. In addition, the standard DE algorithm adopts a fixed mutation strategy in different optimization stages, lacking a dynamic adjustment mechanism for global search and local convergence. In the case of complex pollution data distribution and overlapping pollution source categories, the convergence speed is slow. In addition, the population initialization of the traditional DE algorithm uses random sampling, which may lead to uneven distribution of the initial solutions, affecting the global search ability and thus the final optimization effect of pollution source identification.
[0007] In summary, the existing pollution source identification methods still have limitations in dealing with uncertain pollution data, optimizing fuzzy clustering parameters, and combining pollution diffusion models for tracing analysis. Although traditional statistical analysis and physical-chemical models can provide a certain ability for pollution source tracing, they have poor adaptability in complex pollution environments. Data-driven intelligent computing methods such as fuzzy clustering algorithms can better handle the fuzziness and uncertainty of pollution data, but the existing methods still have problems such as being sensitive to abnormal data, difficult parameter optimization, and low computational efficiency. In addition, there is no effective joint optimization mechanism between pollution diffusion models and pollution source identification methods, affecting the spatial resolution and identification accuracy of pollution source identification. Therefore, it is necessary to propose a pollution source identification method that combines an improved fuzzy clustering algorithm, a pollution diffusion model, and an adaptive optimization strategy to improve the accuracy, robustness, and real-time performance of pollution source identification. Summary of the Invention
[0008] An object of the present invention is to propose a rapid identification method for pollution sources in industrial parks based on fuzzy clustering. The present invention combines the Possibilistic C-Means and Fuzzy Rough C-Means algorithms, uses the fuzzy clustering method to handle the uncertainty of pollution data, and optimizes the spatial distribution recognition accuracy of pollution sources through a pollution diffusion model. At the same time, an improved differential evolution algorithm is used to optimize the clustering parameters of the Possibilistic C-Means and Fuzzy Rough C-Means, realizing dynamic adjustment of membership degrees and improving the stability and accuracy of pollution source identification.
[0009] A rapid identification method for pollution sources in industrial parks based on fuzzy clustering according to an embodiment of the present invention includes the following steps: S1. Collect pollution data in the industrial park and preprocess the pollution data to form standardized pollution data; S2. Use a fuzzy clustering algorithm combining Possibilistic C-Means and Fuzzy Rough C-Means to identify the standardized pollution data, calculate the possibilistic membership degree of data points using Possibilistic C-Means, and at the same time use the Fuzzy Rough C-Means method to identify pollution sources based on the upper approximation set and lower approximation set of pollution data to obtain a preliminary identification result of pollution sources; S3. Construct a pollution diffusion model, calculate the diffusion path of pollutants according to meteorological data, geographical information, and pollutant diffusion characteristics, and correct it in combination with the preliminary identification result of pollution sources to optimize the spatial distribution recognition accuracy of pollution sources; S4. Use an improved differential evolution algorithm to optimize the fuzzy clustering parameters, set a fitness function, use the accuracy of pollution source identification and the matching degree of pollution diffusion as optimization objectives, dynamically adjust the weight between Possibilistic C-Means and Fuzzy Rough C-Means, and optimize the distribution of clustering centers; S5. When abnormal fluctuations occur in the pollution data, dynamically adjust the membership degree, and recompute the pollutant diffusion path in combination with the pollution diffusion model to correct the pollution source identification result; S6. Output the optimized pollution source identification result and generate a pollution source tracing analysis report.
[0010] Optionally, the pollution data includes air pollution data, water pollution data, solid waste pollution data, meteorological data, and enterprise emission inventories.
[0011] Optionally, the preprocessing includes data cleaning, normalization processing, and feature extraction.
[0012] Optionally, the S2 specifically includes: S21. Initialize the fuzzy clustering parameters, set the standardized pollution data as , where represents the th pollution data point, set the dimension of the standardized pollution data as , and each pollution data point contains Based on the pollution characteristics, set the number of clustering categories as , and initialize the clustering center matrix , where represents the clustering center of the th category, and initialize the fuzzy membership degree , where represents the membership degree of the data point to the th clustering. Set the convergence threshold and the maximum number of iterations as the termination conditions; S22. Calculate the Euclidean distance from each data point in the standardized pollution data to each clustering center : ; Among them, represents the Euclidean distance from the data point to the clustering center , represents the data point, represents the clustering center, represents the dimension of the standardized pollution data, represents the pollution characteristic value of the data point in the th dimension, represents the value of the clustering center in the th dimension; S23. Use the possibility C-means method to calculate the possibility membership degree of the data points, set the possibility distribution parameter for constraint, and calculate the objective function of the possibility C-means method: ; Among them, represents the objective function of the possibility C-means method, represents the fuzzy coefficient, which controls the fuzziness of the membership degree assignment, represents the possibility distribution parameter, which is used to reduce the influence of abnormal data points on the clustering result, represents the possibility membership degree of the data point , represents the number of clustering categories, represents the data point to the clustering center the square of the Euclidean distance, represents the data point the th power of the possibility membership degree; Calculate the data point Possibility membership degree in class : ; Wherein, represents the possibility membership degree of the data point in class . The larger the value, the more the data point belongs to class . represents the fuzzy coefficient, which controls the fuzziness of the membership degree distribution. represents the possibility distribution parameter, which is used to reduce the influence of abnormal data points on the clustering result. represents the data point to the clustering center the square of the Euclidean distance; S24. The uncertainty of contaminated data is processed by the fuzzy rough C-means method, and the upper approximation set and the lower approximation set of the contaminated data are defined, and the upper approximation membership degree and the lower approximation membership degree of the data point are calculated: ; Wherein, represents the upper approximation membership degree of the data point, represents the lower approximation membership degree of the data point, represents the upper approximation set, which describes the set of sample points that clearly belong to class in the contaminated data set, represents the set of sample points with uncertain attribution, represents the set the number of data points in, represents the set the number of data points in, represents the data point, represents belonging to the contaminated data set or the data point in in class membership degree; S25. The final membership degree is calculated by combining the possibility C-means and the fuzzy rough C-means using a weighted fusion strategy: ; Wherein, represents the final membership degree, represents the weight coefficient, which controls the contribution degrees of the possibility C-means and the fuzzy rough C-means in the pollution source identification. represents the possibility membership degree of the possibility C-means method, Represents the upper approximate membership degree of the data point, Represents the lower approximate membership degree of the data point; S26. Use the final membership degree To calculate the updated cluster center: ; Wherein, Represents the cluster center after the th iteration, Represents the final membership degree, Represents the fuzzy coefficient, which controls the fuzziness of the membership degree assignment, Represents the data point; S27. Judge the convergence condition. If it is satisfied: ; Wherein, Represents the distance between the cluster center after the current iteration and the cluster center of the previous round. If it is less than the set convergence threshold , or the maximum number of iterations is reached, terminate the calculation and output the preliminary identification result of the pollution source. Otherwise, return to S22 to continue the iterative calculation until the convergence condition is satisfied or the set maximum number of iterations is reached.
[0013] Optionally, the S3 specifically includes: S31. Construct a pollution diffusion model, and set the pollution source location set as , wherein represents the coordinate position of the th pollution source, set the pollutant concentration distribution matrix as , wherein represents the pollutant concentration at the monitoring point , set the meteorological data set as , wherein represents the wind speed, represents the wind direction, represents the environmental temperature, represents the relative humidity, represents the atmospheric stability parameter, and set the geographical information set to include the terrain features, altitude data and land use information around the pollution source; S32. Based on the meteorological data and geographical information, calculate the diffusion path of the pollutant in space, define that the pollutant diffusion conforms to the Gaussian plume model, and calculate the concentration of the pollutant at the monitoring point : ; Wherein, represents the concentration of the pollutant at the monitoring point , Represents the pollution source emission rate, and represent the lateral and vertical diffusion coefficients respectively, represents the wind speed, represents the natural exponential function, represents the offset of the pollutant along the wind direction, represents the pollutant diffusion height, represents the pollution source emission height; S33. Based on the preliminary identification result of the pollution source, correct the pollution diffusion path and set the pollution source identification correction matrix: ; where, represents the pollution source identification correction matrix, represents the pollution source the correction contribution value of the pollutant concentration at the data point , represents the final membership degree, represents the concentration of the pollutant at the monitoring point ; S34. Calculate the pollutant diffusion correction factor , and adjust the pollution diffusion model based on the pollution source identification result: ; where, represents the pollutant diffusion correction factor, describing the correction contribution ratio of the pollution source at the data point , represents the pollution source the correction contribution value of the pollutant concentration at the data point , represents the total pollution contribution of all pollution sources to the data point ; S35. Update the pollutant diffusion path and set the corrected pollutant concentration: ; where, represents the corrected pollutant concentration value, represents the pollutant diffusion correction factor, represents the concentration of the pollutant at the monitoring point ; S36. Calculate the pollution source matching degree , to measure the rationality of the pollution source in spatial distribution: ; where, Indicates the pollution source matching degree and describes the pollution source The contribution ratio to the observed pollution concentration; S37. Based on the pollution source matching degree Modify the spatial distribution of the pollution source. If It is lower than the set threshold , then adjust the set of pollution source positions Recalculate the pollutant concentration distribution using the pollution diffusion path and update the identification result of the pollution source.
[0014] Optionally, the S4 specifically includes: S41. Initialize the optimization parameters and set the initial optimization parameter set of the improved differential evolution algorithm as , where Indicates the set of fuzzy clustering parameters to be optimized. The set of fuzzy clustering parameters includes the fuzzy coefficient , the clustering center matrix And the adjustment weight of the membership degree matrix , set the population size , the maximum number of iterations , the adaptive mutation factor And the crossover probability ; S42. Initialize the population using Latin hypercube sampling and define the population matrix , where Indicates the th candidate solution, Indicates the population size. Generate the initial solution using the Latin hypercube sampling method: ; Among them, Indicates the initial solution, And Respectively represent the upper and lower bounds of the optimization parameters, Indicates the permutation order of the population index, Indicates a uniformly distributed random number to ensure uniform coverage of the search space, Indicates the population size; S43. Define the fitness function. The fitness function comprehensively considers the accuracy of pollution source identification and the pollution diffusion matching degree: ; Among them, Indicates the fitness function, Indicates the objective function of fuzzy clustering, Indicates the pollution diffusion matching degree error, And Indicates the weighting coefficient, Indicates the data point The square of the Euclidean distance to the cluster center , represents the fuzziness coefficient, which controls the fuzziness of the membership degree assignment represents the number of clustering categories represents the final membership degree represents the pollution source matching degree, which describes the contribution ratio of the pollution source to the observed pollution concentration S44. An adaptive mutation factor and crossover probability are adopted, and an adaptive adjustment mechanism for the mutation factor and crossover probability is defined: ; ; wherein represents the adaptive mutation factor and represent the range of the mutation factor, which is gradually adjusted with the number of iterations represents the maximum number of iterations represents the crossover probability and represent the range of the crossover probability, which is dynamically adjusted according to the population standard deviation represents the population standard deviation represents the population mean S45. A hybrid mutation strategy is adopted for the mutation operation. According to the iteration process, the global search strategy and the local search strategy are alternately used; In the first half the global search strategy is used: ; wherein represents the mutated individual of the th candidate solution in the th iteration represents the candidate solution with the best fitness in the current population and represent two different randomly selected candidate solutions represents the adaptive mutation factor; In the second half the local search strategy is used: ; wherein represents the mutated individual of the th candidate solution in the th iteration represents the candidate solution with the best fitness in the current population and denotes two different candidate solutions randomly selected, represents the adaptive mutation factor, Indicates the number of Candidate solutions; S46. Perform crossover operation to generate test individuals : ; in, Indicates the test individual No. parameter values, that is, the new individuals obtained after the crossover operation, Represents a variant individual No. parameter values, Represents the original individual No. parameter values, Represents the randomly selected index, ensuring that each test individual has at least one parameter from the variant individual; S47, execute selection operation, calculate test individuals The fitness value of ,like Better than the original , then replace the original value: ; in, Indicates The final updated value of each individual, Indicates the test individual The fitness value of Indicates The original value of each individual; S48, repeat S44 to S47 until the termination condition is met: ; in, Indicates The fitness value after iterations is the fitness value of the optimized fuzzy clustering algorithm. Indicates The fitness value after iterations is Represents the convergence threshold, which is used to measure whether the optimization process has reached the convergence condition; Update the fuzzy clustering parameter set and apply it to the fuzzy clustering algorithm.
[0015] The beneficial effects of the present invention are: First, by combining the Possibilistic C-Means and Fuzzy Rough C-Means algorithms, the present invention improves the processing ability of pollution source identification for uncertain pollution data, making the pollution source identification more accurate and robust. The Possibilistic C-Means method is used to reduce the interference of abnormal data and enhance the adaptability to the ambiguity and noise of pollution data. At the same time, the Fuzzy Rough C-Means method is used to combine the upper and lower approximation sets to improve the identification accuracy of pollution sources and reduce the misidentification problem caused by the overlap of pollution sources.
[0016] Secondly, the present invention constructs a pollution diffusion model. Based on meteorological data, geographical information, and pollutant diffusion characteristics, it optimizes the diffusion path of pollutants and corrects it in combination with the preliminary identification results of fuzzy clustering, improving the identification accuracy of the spatial distribution of pollution sources and enhancing the pollution source tracing ability. Aiming at the problems of strong parameter dependence and difficult optimization of traditional fuzzy clustering methods, the present invention uses an improved differential evolution algorithm to optimize the fuzzy clustering parameters, sets the fitness function, takes the accuracy of pollution source identification and the matching degree of pollution diffusion as the optimization objectives, adaptively adjusts the weights of the Possibilistic C-Means and Fuzzy Rough C-Means, and optimizes the distribution of clustering centers to make the pollution source identification more stable and accurate.
[0017] In addition, the present invention adopts an improved strategy of adaptive mutation factor and crossover probability to improve the global search ability and local convergence speed of the optimization process, avoid the optimization falling into local optimum, and improve the generalization ability of the pollution source identification model. When abnormal fluctuations occur in pollution data, the present invention can dynamically adjust the fuzzy membership degree, optimize the pollution source identification strategy in real time, and recalculate the pollutant diffusion path in combination with the pollution diffusion model to ensure the timeliness and accuracy of the pollution source identification results.
[0018] Finally, the present invention can effectively improve the accuracy, anti-interference ability, and adaptive optimization ability of pollution source identification in a complex industrial park environment, providing scientific and reliable technical support for pollution monitoring, pollution control, and environmental management. Description of the Drawings
[0019] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a flowchart of a method for rapid identification of pollution sources in an industrial park based on fuzzy clustering proposed by the present invention; Figure 2 is a schematic diagram of optimizing fuzzy clustering parameters by using an improved differential evolution algorithm for a method for rapid identification of pollution sources in an industrial park based on fuzzy clustering proposed by the present invention. Detailed Embodiments
[0020] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, and therefore only showing the components related to the present invention.
[0021] Reference Figure 1 and Figure 2 , a rapid identification method for pollution sources in industrial parks based on fuzzy clustering, comprising the following steps: S1. Collect pollution data in the industrial park and preprocess the pollution data to form standardized pollution data; S2. Use a fuzzy clustering algorithm combining possibilistic C-means and fuzzy rough C-means to identify the standardized pollution data, calculate the possibilistic membership degree of data points using possibilistic C-means, and at the same time use the fuzzy rough C-means method to identify pollution sources based on the upper approximation set and lower approximation set of the pollution data, obtaining a preliminary identification result of pollution sources; S3. Construct a pollution diffusion model, calculate the diffusion path of pollutants according to meteorological data, geographical information and pollutant diffusion characteristics, and correct it in combination with the preliminary identification result of pollution sources to optimize the identification accuracy of the spatial distribution of pollution sources; S4. Use an improved differential evolution algorithm to optimize the fuzzy clustering parameters, set a fitness function, take the accuracy of pollution source identification and the matching degree of pollution diffusion as the optimization objectives, dynamically adjust the weight between possibilistic C-means and fuzzy rough C-means, and optimize the distribution of clustering centers; S5. When abnormal fluctuations occur in the pollution data, dynamically adjust the membership degree, and recalculate the pollutant diffusion path in combination with the pollution diffusion model to correct the pollution source identification result; S6. Output the optimized pollution source identification result and generate a pollution source tracing analysis report.
[0022] In this embodiment, the pollution data includes air pollution data, water pollution data, solid waste pollution data, meteorological data and enterprise emission inventories.
[0023] In this embodiment, the preprocessing includes data cleaning, normalization processing and feature extraction.
[0024] In this embodiment, the S2 specifically includes: S21. Initialize the fuzzy clustering parameters, set the standardized pollution data as , where represents the th pollution data point, set the dimension of the standardized pollution data as , each pollution data point contains dimensional pollution features, set the number of clustering categories as , initialize the clustering center matrix , where Represent the cluster center of the category, and initialize the fuzzy membership , where represents the data point For the membership degree of the cluster, set the convergence threshold and the maximum number of iterations as the termination condition; S22. Calculate the Euclidean distance from each data point in the standardized pollution data to each cluster center : ; Among them, represents the Euclidean distance from the data point to the cluster center , represents the data point, represents the cluster center, represents the dimension of the standardized pollution data, represents the data point in the pollution eigenvalue of the dimension, represents the value of the cluster center in the S23. Use the possibility C-means method to calculate the possibility membership degree of the data point, and set the possibility distribution parameter for constraint, and calculate the objective function of the possibility C-means method: ; Among them, represents the objective function of the possibility C-means method, represents the fuzzy coefficient, which controls the fuzziness of the membership degree assignment, represents the possibility distribution parameter, which is used to reduce the influence of abnormal data points on the clustering result, represents the possibility membership degree of the data point , represents the number of clustering categories, represents the data point to the cluster center the square of the Euclidean distance, represents the data point of the possibility membership degree of power; Calculate the possibility membership degree of the data point in the category: ; Among them, represents the possibility membership degree of the data point in the class. The larger the value, the more the data point belongs to the class. represents the fuzzy coefficient, which controls the fuzziness of the membership degree assignment. represents the possibility distribution parameter, which is used to reduce the influence of abnormal data points on the clustering result. represents the data point to the clustering center the square of the Euclidean distance; S24. The uncertainty of the contaminated data is processed by the fuzzy rough C - means method, and the upper approximation set and the lower approximation set of the contaminated data are defined, and the upper approximation membership degree and the lower approximation membership degree of the data point are calculated: ; Among them, represents the upper approximation membership degree of the data point, represents the lower approximation membership degree of the data point, represents the upper approximation set, which describes the set of sample points that clearly belong to the class in the contaminated data set, represents the set of sample points with uncertain attribution, represents the set the number of data points in it, represents the set the number of data points in it, represents the data point, represents belonging to the contaminated data set or the data points in in the class membership degree; S25. The final membership degree is calculated by combining the possibility C - means and the fuzzy rough C - means using a weighted fusion strategy: ; Among them, represents the final membership degree, represents the weight coefficient, which controls the contribution degrees of the possibility C - means and the fuzzy rough C - means in the pollution source identification, represents the possibility membership degree of the possibility C - means method, represents the upper approximation membership degree of the data point, represents the lower approximation membership degree of the data point; S26. Using the final membership degree Calculate the updated cluster center: ; Wherein, represents the cluster center after the -th iteration, represents the final membership degree, represents the fuzzy coefficient, which controls the fuzziness of membership degree assignment, represents the data point; S27. Judge the convergence condition. If it is satisfied: ; Wherein, represents the distance between the cluster center after the current iteration and the cluster center of the previous round. If it is less than the set convergence threshold , or the maximum number of iterations is reached, terminate the calculation and output the preliminary identification result of the pollution source. Otherwise, return to S22 to continue the iterative calculation until the convergence condition is satisfied or the set maximum number of iterations is reached.
[0025] In this embodiment, the specific steps of S3 are as follows: S31. Construct a pollution diffusion model, and set the set of pollution source locations as , wherein represents the coordinate position of the -th pollution source. Set the pollutant concentration distribution matrix as , wherein represents the pollutant concentration at the monitoring point . Set the meteorological data set as , wherein represents the wind speed, represents the wind direction, represents the environmental temperature, represents the relative humidity, represents the atmospheric stability parameter, and set the geographical information set to include the topographic features, altitude data and land use information around the pollution source; S32. Based on the meteorological data and geographical information, calculate the diffusion path of the pollutant in space. Define that the pollutant diffusion conforms to the Gaussian plume model, and calculate the concentration of the pollutant at the monitoring point : ; Wherein, represents the concentration of the pollutant at the monitoring point , represents the emission rate of the pollution source , and respectively represent the horizontal and vertical diffusion coefficients, represents the wind speed, represents the natural exponential function, represents the offset of the pollutant propagation along the wind direction, represents the pollutant diffusion height, represents the pollution source emission height; S33. Based on the preliminary identification result of the pollution source, correct the pollution diffusion path and set the pollution source identification correction matrix: ; Among them, represents the pollution source identification correction matrix, represents the pollution source for the data point the correction contribution value of the pollutant concentration at the location, represents the final membership degree, represents the pollutant at the monitoring point the concentration at the location; S34. Calculate the pollutant diffusion correction factor , and adjust the pollution diffusion model based on the pollution source identification result: ; Among them, represents the pollutant diffusion correction factor, describing the pollution source at the data point the correction contribution ratio at the location, represents the pollution source for the data point the correction contribution value of the pollutant concentration at the location, represents the total pollution contribution of all pollution sources to the data point at the location; S35. Update the pollutant diffusion path and set the corrected pollutant concentration: ; Among them, represents the corrected pollutant concentration value, represents the pollutant diffusion correction factor, represents the pollutant at the monitoring point the concentration at the location; S36. Calculate the pollution source matching degree , to measure the rationality of the pollution source in the spatial distribution: ; Among them, represents the pollution source matching degree, describing the pollution source the contribution ratio to the observed pollution concentration; S37. Based on the pollution source matching degree Modify the spatial distribution of the pollution sources. If it is lower than the set threshold , then adjust the set of pollution source positions , recalculate the pollutant concentration distribution using the pollution diffusion path, and update the identification result of the pollution sources.
[0026] In this embodiment, the specific steps of S4 include: S41. Initialize the optimization parameters. Set the initial optimization parameter set of the improved differential evolution algorithm as , where represents the set of fuzzy clustering parameters to be optimized. The set of fuzzy clustering parameters includes the fuzzy coefficient , the clustering center matrix and the adjustment weight of the membership degree matrix . Set the population size , the maximum number of iterations , the adaptive mutation factor and the crossover probability ; S42. Initialize the population using Latin hypercube sampling. Define the population matrix , where represents the th candidate solution, represents the population size. Generate the initial solution using the Latin hypercube sampling method: ; Among them, represents the initial solution, and respectively represent the upper and lower bounds of the optimization parameters, represents the permutation order of the population index, represents a uniformly distributed random number to ensure uniform coverage of the search space, represents the population size; S43. Define the fitness function. The fitness function comprehensively considers the accuracy of pollution source identification and the pollution diffusion matching degree: ; Among them, represents the fitness function, represents the objective function of fuzzy clustering, represents the pollution diffusion matching degree error, and represent the weighting coefficients, represents the data point to the clustering center the square of the Euclidean distance, denotes the fuzziness coefficient, which controls the fuzziness of membership assignment, denotes the number of clustering categories, denotes the final membership, denotes the pollution source matching degree, which describes the contribution ratio of the pollution source to the observed pollution concentration; S44. An adaptive mutation factor and crossover probability are adopted to define an adaptive adjustment mechanism for the mutation factor and crossover probability: ; ; wherein, denotes the adaptive mutation factor, and denotes the range of the mutation factor, which is gradually adjusted with the number of iterations denotes the maximum number of iterations, denotes the crossover probability, and denotes the range of the crossover probability, which is dynamically adjusted according to the population standard deviation denotes the population standard deviation, denotes the population mean; S45. A hybrid mutation strategy is adopted for mutation operation. According to the iteration process, the global search strategy and the local search strategy are alternately used; In the first half the global search strategy is used: ; wherein, denotes the mutated individual of the th candidate solution in the th iteration, denotes the candidate solution with the best fitness in the current population, and denote two different randomly selected candidate solutions, denotes the adaptive mutation factor; In the second half the local search strategy is used: ; wherein, denotes the mutated individual of the th candidate solution in the th iteration, denotes the candidate solution with the best fitness in the current population, and denote two different randomly selected candidate solutions, denotes the adaptive mutation factor, represents the th candidate solution in the current population; S46. Perform the crossover operation to generate a trial individual : ; wherein, represents the th parameter value of the trial individual , that is, the new individual obtained after the crossover operation, represents the th parameter value of the mutant individual , represents the th parameter value of the original individual , represents a randomly selected index to ensure that each trial individual has at least one parameter from the mutant individual; S47. Perform the selection operation, calculate the fitness value of the trial individual , if is better than the original individual , then replace the original value: ; wherein, represents the final updated value of the th individual, represents the fitness value of the trial individual , represents the th original value of the individual; S48. Repeat S44 to S47 until the termination condition is met: ; wherein, represents the fitness value after the th iteration, that is, the fitness value of the optimized fuzzy clustering algorithm, represents the fitness value after the th iteration, represents the convergence threshold, which is used to measure whether the optimization process reaches the convergence condition; Update the fuzzy clustering parameter set and apply it to the fuzzy clustering algorithm.
[0027] Example 1: To verify the feasibility of the present invention in implementation, the present invention is applied to a coastal industrial park, which contains multiple manufacturing enterprises, including the chemical industry, metal smelting, pharmaceuticals, and power industries. There are a wide variety of pollutant emissions in the area. Air pollution mainly comes from chemical waste gas, coal combustion soot, and VOCs. Water pollution is mainly caused by wastewater discharge. There is also a certain degree of pollution from solid waste. Due to the complex types of pollution sources and the diffusion and cross - influence of different pollutants, traditional pollution source identification methods are difficult to accurately distinguish the contribution of each pollution source to environmental pollution and are also difficult to adapt to the dynamic changes of pollutant emissions.
[0028] The present invention has established a pollution monitoring network in the industrial park, including 10 air quality monitoring stations, 15 water quality monitoring points, 3 meteorological data collection stations, and an enterprise emission inventory data collection system. The air quality monitoring stations are used to collect the concentrations of pollutants such as PM2.5, PM10, , , CO, , etc., sampling once per hour. The water quality monitoring points are used to detect water pollution data such as COD, ammonia nitrogen, and heavy metal ions, sampling four times a day. At the same time, the meteorological data collection stations provide data such as wind speed, wind direction, temperature, and humidity, recording once every 30 minutes, providing support for pollution diffusion modeling.
[0029] In this scenario, the present invention pre - processes the collected pollution data, including data cleaning, normalization, and feature extraction, removing missing data and abnormal data, and constructing a standardized pollution data set. Then, a fuzzy clustering method combining possibilistic C - means and fuzzy - rough C - means is used for pollution source identification. The possibilistic C - means method is used to reduce the interference of abnormal pollution data and improve the adaptability to the fuzziness and noise of pollutants. At the same time, the fuzzy - rough C - means method combines upper and lower approximation sets to enhance the accuracy of pollution source identification. Subsequently, based on the pollution diffusion model, the diffusion path of pollutants is calculated and corrected in combination with the preliminary pollution source identification results to improve the accuracy of the spatial distribution identification of pollution sources. During the optimization process, an improved differential evolution algorithm is used to optimize the fuzzy clustering parameters, and by adaptively adjusting the mutation factor and crossover probability, the convergence speed and optimization ability of the pollution source identification model are improved.
[0030] During the experiment, data from different time periods were selected for comparative analysis. The test results show that during the 3-month test period from July to September 2024, the main pollution sources in the industrial park could be accurately identified. For example, a chemical enterprise had a large VOC emission, and it was difficult for traditional pollution source identification methods to accurately trace the source. However, the method of the present invention can establish the correlation of pollutant diffusion between multiple monitoring points and accurately identify the contribution ratio of this enterprise as the main pollution source. In the water pollution identification experiment, the wastewater discharge of a metal smelting enterprise led to an abnormal increase in the COD concentration, and it was difficult for traditional methods to distinguish the pollution source. However, this method successfully identified the leading role of this enterprise in the pollution incident by combining the pollution diffusion model.
[0031] Table 1 Experimental comparison data table ; The experimental data show that the present invention is superior to traditional methods in terms of the time, accuracy, and stability of pollution source identification.
[0032] In terms of the pollution source identification time, the traditional method on average takes 26.8 minutes, while the average identification time of the method of the present invention is only 8.8 minutes, significantly improving the identification efficiency. Especially at the PM10 and monitoring points, the identification time is shortened by nearly 70%. This improvement is crucial for the rapid response to pollution emergencies, enabling the environmental protection department to take corresponding emergency measures more quickly.
[0033] In terms of the identification accuracy, the average accuracy of the method of the present invention reaches 93.2%, which is more than 10% higher than that of the traditional method. For individual pollutants, such as and VOC, the improvement in the identification accuracy is particularly obvious, up to 94.6% at most. This improvement shows that this method has more advantages in dealing with the uncertainty of pollution data, can more accurately identify pollutants from different sources, and improve the reliability of pollution source tracing.
[0034] In terms of stability, the stability index of the traditional method is between 85.2% - 87.5%, while the stability of the method of the present invention exceeds 95% and reaches 96.5% at most. This indicates that this method can maintain higher identification reliability in the face of complex pollution environments and fluctuations in pollution data and is not easily affected by abnormal data fluctuations. Especially in the VOC exceeding standard incident on August 15, 2024, this method completed the pollution source identification in only 10 minutes and corrected the pollution diffusion path, while the traditional method took 30 minutes and the stability was only 85.2%, making it difficult to quickly and accurately respond to sudden pollution incidents.
[0035] Overall, the present invention not only improves the efficiency of pollution source identification, but also shows superiority in accuracy and stability, accurately identifying pollution sources and providing more reliable data support for optimizing the pollutant diffusion path. These experimental results fully demonstrate the applicability and reliability of the method in a complex industrial park environment.
[0036] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A rapid identification method for pollution sources in industrial parks based on fuzzy clustering, characterized in that It includes the following steps: S1. Collect pollution data in the industrial park, and preprocess the pollution data to form standardized pollution data; S2. Use a fuzzy clustering algorithm that combines possibilistic C-means and fuzzy rough C-means to identify the standardized pollution data. Calculate the possibilistic membership degree of data points using possibilistic C-means, and at the same time use the fuzzy rough C-means method to identify pollution sources based on the upper approximation set and lower approximation set of pollution data, and obtain a preliminary pollution source identification result; S3. Construct a pollution diffusion model, calculate the diffusion path of pollutants according to meteorological data, geographical information and pollutant diffusion characteristics, and correct it in combination with the preliminary pollution source identification result to optimize the identification accuracy of the spatial distribution of pollution sources; S4. Use an improved differential evolution algorithm to optimize the fuzzy clustering parameters, set a fitness function, take the accuracy of pollution source identification and the matching degree of pollution diffusion as the optimization objectives, dynamically adjust the weight between possibilistic C-means and fuzzy rough C-means, and optimize the distribution of clustering centers; S5. When the pollution data shows abnormal fluctuations, dynamically adjust the membership degree, and recalculate the pollutant diffusion path in combination with the pollution diffusion model to correct the pollution source identification result; S6. Output the optimized pollution source identification result and generate a pollution source tracing analysis report.
2. The rapid identification method of industrial park pollution sources based on fuzzy clustering according to claim 1, characterized in that, The pollution data includes air pollution data, water pollution data, solid waste pollution data, meteorological data and enterprise emission inventories.
3. A rapid identification method for pollution sources in industrial parks based on fuzzy clustering according to claim 1, characterized in that, The preprocessing includes data cleaning, normalization processing and feature extraction.
4. A rapid identification method for pollution sources in industrial parks based on fuzzy clustering according to claim 1, characterized in that, The specific content of S2 includes: S21. Initialize the fuzzy clustering parameters, and set the standardized pollution data as , where represents the -th pollution data point. Set the dimension of the standardized pollution data as . Each pollution data point contains -dimensional pollution features. Set the number of clustering categories as . Initialize the clustering center matrix , where represents the clustering center of the -th class, and initialize the fuzzy membership degree , where represents the membership degree of the data point for the -th cluster. Set the convergence threshold and the maximum number of iterations as the termination conditions; S22. Calculate each data point in the standardized pollution data to each cluster center for the Euclidean distance: ; Among them, represents the Euclidean distance from the data point to the cluster center ; represents the data point, represents the cluster center, represents the dimension of the standardized contaminated data, represents the data point at the th dimension of the contamination eigenvalue, represents the cluster center at the th dimension value; S23. Calculate the possibility membership degree of data points using the possibility C-means method and set the possibility distribution parameters. Apply constraints and calculate the objective function of the possibility C-means method: ; Among them, represents the objective function of the Possibilistic C-Means method, represents the fuzzy coefficient, which controls the fuzziness of membership assignment, represents the possibility distribution parameter, which is used to reduce the influence of abnormal data points on the clustering result, represents the data point 's possibilistic membership, represents the number of clustering categories, represents the data point to the cluster center the square of the Euclidean distance, represents the data point 's possibilistic membership of the power; Calculated data point In the Class, the possibility membership degree is: ; Among them, represents the possibility membership degree of the data point in the class. The larger the value, the more the data point belongs to the class. represents the fuzzy coefficient, which controls the fuzziness of the membership degree distribution. represents the possibility distribution parameter, which is used to reduce the influence of abnormal data points on the clustering result. represents the data point to the clustering center the square of the Euclidean distance; S24. Deal with the uncertainty of contaminated data by using the fuzzy rough C-means method, and define the upper approximation set and the lower approximation set of the contaminated data and the lower approximation set , calculate the upper approximation membership degree of the data point and the lower approximation membership degree : ; Among them, represents the upper approximation membership degree of the data point, represents the lower approximation membership degree of the data point, represents the upper approximation set, describing the set of sample points that clearly belong to the th class in the contaminated data set, represents the set of sample points with uncertain attribution, represents the set the number of data points in, represents the set the number of data points in, represents the data point, represents belonging to the contaminated data set or the data points in, in the th class membership degree; S25. Use a weighted fusion strategy that combines possibilistic C-means and fuzzy rough C-means to calculate the final membership degree: ; Among them, represents the final membership degree, represents the weight coefficient, which controls the contribution degrees of the possibilistic C-means and the fuzzy-rough C-means in the pollution source identification, represents the possibilistic membership degree of the possibilistic C-means method, represents the upper approximation membership degree of the data point, represents the lower approximation membership degree of the data point; S26. Using the final membership degree Calculate the updated cluster center: ; Among them, represents the cluster center after the -th iteration, represents the final membership degree, represents the fuzzy coefficient, which controls the fuzziness of membership degree assignment, represents the data point; S27. Judge the convergence condition. If it is satisfied: ; Among them, represents the distance between the cluster center after the current iteration and the cluster center of the previous round. If it is less than the set convergence threshold , or reaches the maximum number of iterations , terminate the calculation and output the preliminary identification result of the pollution source. Otherwise, return to S22 to continue the iterative calculation until the convergence condition is met or the set maximum number of iterations is reached.
5. A rapid identification method for pollution sources in industrial parks based on fuzzy clustering according to claim 1, characterized in that, The specific content of S3 includes: S31. Construct a pollution diffusion model. Set the set of pollution source locations as , where represents the coordinate position of the -th pollution source. Set the pollutant concentration distribution matrix as , where represents the pollutant concentration at the monitoring point . Set the meteorological data set as , where represents the wind speed, represents the wind direction, represents the environmental temperature, represents the relative humidity, represents the atmospheric stability parameter, and set the geographic information set to include the terrain features, altitude data, and land use information around the pollution sources; S32. Based on meteorological data and geographical information, calculate the diffusion path of pollutants in space, define that the pollutant diffusion conforms to the Gaussian plume model, and calculate the concentration of pollutants at the monitoring point location : ; Among them, represents the concentration of pollutants at the monitoring point ; represents the emission rate of the pollution source ; and respectively represent the lateral and vertical diffusion coefficients, represents the wind speed, represents the natural exponential function, represents the offset of the pollutant propagating along the wind direction, represents the diffusion height of the pollutant, represents the emission height of the pollution source; S33. Correct the pollution diffusion path based on the preliminary pollution source identification result, and set a pollution source identification correction matrix: ; Among them, represents the pollution source identification correction matrix, represents the pollution source 's correction contribution value to the pollutant concentration at the data point ; represents the final membership degree, represents the concentration of the pollutant at the monitoring point ; S34. Calculate the pollutant diffusion correction factor , and adjust the pollution diffusion model based on the pollution source identification result: ; Among them, represents the pollutant diffusion correction factor, describing the correction contribution ratio of the pollution source at the data point ; represents the correction contribution value of the pollution source to the pollutant concentration at the data point ; represents the total pollution contribution of all pollution sources to the data point ; S35. Update the pollutant diffusion path and set the corrected pollutant concentration: ; Among them, represents the corrected pollutant concentration value, represents the pollutant diffusion correction factor, represents the pollutant at the monitoring point at the concentration; S36. Calculate the pollution source matching degree , and measure the pollution source for its rationality in spatial distribution: ; Among them, represents the pollution source matching degree, describing the pollution source contribution ratio to the observed pollution concentration; S37. Based on the pollution source matching degree Correct the spatial distribution of pollution sources. If it is lower than the set threshold , then adjust the pollution source location set , recalculate the pollutant concentration distribution using the pollution diffusion path, and update the identification results of pollution sources.
6. The rapid identification method of industrial park pollution sources based on fuzzy clustering according to claim 1, characterized in that The specific content of S4 includes: S41. Initialize the optimization parameters, and set the initial optimization parameter set of the improved differential evolution algorithm as , where represents the set of fuzzy clustering parameters to be optimized, and the set of fuzzy clustering parameters includes the fuzzy coefficient , the clustering center matrix , and the adjustment weight of the membership degree matrix . Set the population size , the maximum number of iterations , the adaptive mutation factor , and the crossover probability ; S42. Initialize the population using Latin hypercube sampling and define the population matrix , where represents the th candidate solution, represents the population size, and generate the initial solution using the Latin hypercube sampling method: ; Among them, represents the initial solution, and respectively represent the upper and lower bounds of the optimization parameters, represents the permutation order of the population index, represents a uniformly distributed random number to ensure uniform coverage of the search space, represents the population size; S43. Define a fitness function, and the fitness function comprehensively considers the accuracy of pollution source identification and the matching degree of pollution diffusion: ; Among them, represents the fitness function, represents the objective function of fuzzy clustering, represents the pollution diffusion matching degree error, and represents the weighting coefficient, represents the data point to the cluster center the square of the Euclidean distance, represents the fuzzy coefficient, controlling the fuzziness of the membership degree assignment, represents the number of clustering categories, represents the final membership degree, represents the pollution source matching degree, describing the contribution ratio of the pollution source to the observed pollution concentration; S44. Use an adaptive mutation factor and crossover probability, and define an adaptive adjustment mechanism for the mutation factor and crossover probability: ; ; Among them, represents the adaptive mutation factor, and represent the range of the mutation factor, which is adjusted step by step with the number of iterations gradually, represents the maximum number of iterations, represents the crossover probability, and represent the range of the crossover probability, which is dynamically adjusted according to the population standard deviation performed, represents the population standard deviation, represents the population mean; S45. Use a hybrid mutation strategy to perform mutation operations, and alternately use a global search strategy and a local search strategy according to the iteration process; In the first half Use the global search strategy: ; Among them, represents the mutant individual of the th candidate solution in the th iteration, represents the candidate solution with the best fitness in the current population, and represent two randomly selected different candidate solutions, represents the adaptive mutation factor; In the second half Use a local search strategy: ; Among them, represents the mutant individual of the th candidate solution in the th iteration, represents the candidate solution with the optimal fitness in the current population, and represent two randomly selected different candidate solutions, represents the adaptive mutation factor, represents the th candidate solution in the current population; S46. Perform crossover operation to generate trial individuals : ; Among them, represents the nth parameter value of the test individual, that is, the new individual obtained after the crossover operation, represents the nth parameter value of the mutated individual, represents the nth parameter value of the original individual, represents a randomly selected index to ensure that each test individual has at least one parameter from the mutated individual; S47. Perform a selection operation and calculate the fitness value of the test individual ; if it is better than the original individual , then replace the original value with it: ; Among them, represents the final updated value of the th individual, represents the fitness value of the test individual th individual's original value; S48. Repeat S44 to S47 until the termination condition is met: ; Among them, represents the fitness value after the th iteration, that is, the fitness value of the optimized fuzzy clustering algorithm, represents the fitness value after the th iteration, represents the convergence threshold, which is used to measure whether the optimization process reaches the convergence condition; Update the fuzzy clustering parameter set and apply it to the fuzzy clustering algorithm.
Citation Information
Cited By
Multi-dimensional visual display method and system for oil and gas field environmental protection data
CN121278644A
High-precision risk screening method for pollutants in industrial cluster area
CN121481273A