Method for mining the effect of industrial pollution on cancer based on improved spatial co-location pattern
By using an improved spatial isomorphic pattern mining method, the proximity relationship between pollution sources and cancer is calculated. Taking into account the concentration and carcinogenic category of the pollution source, a star-shaped influence instance table is generated. This solves the shortcomings of traditional methods in calculating proximity relationships and participation, and improves the accuracy and efficiency of mining the relationship between pollution sources and cancer.
Patent Information
- Application Number
- CN202310427552.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-20
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2043-04-20
AI Technical Summary
Traditional spatial isotope pattern mining methods have several drawbacks in discovering the isotope relationship between pollution sources and cancer. These include the unreasonable use of a single distance threshold to determine proximity, the failure of participation-based frequency measurement methods to satisfy the relationship between pollution sources and cancer risk, and the inability of features in the pattern to satisfy a specific order.
By calculating the offset coordinates, local and global concentrations, influence radius, star-shaped neighbor set, weighted influence rate, and kernel density estimation model of the pollution source instance set, a star-shaped influence instance table of candidate patterns is generated, and the influence degree of candidate patterns is calculated. The carcinogenicity category and diffusion effect of the pollution source are taken into account, and the influence of distance decay is adjusted.
It enables a more accurate measurement of the proximity relationship between pollution sources and cancer, takes into account the differences in pollution sources and carcinogenic risks, improves the rationality and accuracy of mining results, and can quickly calculate the influence of candidate patterns.
Smart Images

Figure CN116578946B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of spatial data mining technology, and in particular relates to a method for mining the impact of industrial pollution on cancer based on improved spatial isomorphic patterns. Background Technology
[0002] In February 2021, *CA-CANCER J CLIN* published *Global Cancer Statistics 2020*, which compiled incidence and mortality data for 36 types of cancer in 185 countries worldwide. As early as 1960, Percy Stocks demonstrated a significant relationship between airborne particulate matter and lung and stomach cancer. Lynge et al. demonstrated that residents living near polluting factories had a 20% higher risk of cancer than those living in unpolluted fishing communities. Current research on the relationship between environmental pollution and cancer tends to explore the effects of specific pollutants on the basic structure of human cells in clinical settings, which is challenging and time-consuming. Furthermore, the diverse types of cancer and environmental pollution sources limit the scope for comprehensive analysis of their correlations.
[0003] Tobler's First Geographical Law states that things that are geographically close are often more related than things that are geographically distant. This led to the development of spatial co-location pattern mining. A spatial co-location pattern, also known as a spatial feature pattern, is a non-empty subset of a spatial feature set where feature instances are frequently adjacent in geographical space. For example, schools and stationery stores are often frequently co-located. In spatial data, spatial features are used to represent different spatial things, such as species or urban facilities. A spatial instance is an object of a spatial feature at a specific geographical location. For example, if a school is a spatial feature, then Yunnan University is an instance of that feature. Spatial co-location pattern mining requires feature instances to meet a given distance threshold to form a proximity relationship. Participation measures the frequency of instances satisfying this proximity relationship appearing in close proximity in geographical space. A participation value greater than or equal to the user-given frequency threshold is considered a frequent co-location pattern. Spatial co-location pattern mining can automatically extract unknown but highly instructive patterns hidden in massive spatial data and has been widely applied in various fields such as species distribution, public health, and environmental management for many years. Chemical pollution accounts for nearly 50% of environmental pollution, and this air pollution can easily enter people's living areas due to airflow. This invention uses an improved spatial isotope model mining theory to explore the relationship between industrial air pollutants and various cancers, hoping to provide some reference for cancer prevention and treatment.
[0004] Spatial data is diverse, and traditional mining methods are not always applicable to different data types. The concept of a buffer was introduced in "XIONG H, SHEKHAR S, HUANG Y, et al. A framework for discovering co-location patterns in data sets with extended spatial objects[C]. Proc. SIAM International Conference on Data Mining, Florida, United States, 2004:78-89." This breaks the limitation of traditional methods that can only handle spatial point data, allowing the mining of spatial isolocation patterns from linear, polygonal, and other data. Because instrumental errors introduce many uncertainties into spatial data, "LIUZ, HUANG Y. Mining Co-locations under Uncertainty[C]. Proc. International Conference on Advances in Spatial and Temporal Databases, Munich, Germany, 2013:429-46." This proposes probabilistic participation, effectively mining spatial isolocation patterns from spatial instances with uncertain locations. In reality, spatial distribution is not uniform. Traditional methods cannot measure the differences caused by distance and direction between adjacent instances, which may lead to deviations between the final mining results and reality. "YAO X, CHEN L, PENG L, et al. A co-location pattern-mining algorithm with adensity-weighted distance thresholding consideration[J]. Information Sciences, 2017, 396: 144-61." Considering the influence of distance between instances, a new algorithm for density weighting based on distance is defined. With the popularity of spatial colocation pattern mining, many studies have applied this mining theory to practical applications. For example, spatial colocation pattern mining based on POI data has obtained the spatial association structure between urban service industries. "Hu Tian, Liu Tao, Du Ping, et al. Discovery and feature analysis of urban service industry associations supported by spatial colocation patterns[J]. Journal of Geoinformation Science, 2021, 23(6): 10."A probability-based co-location pattern mining method was used to explore the relationship between childhood cancer and pollutants. The method, "LI J, ADILMAGAMBETOV A, JABBAR MSM, et al. On discovering co-location patterns in datasets: a case study of pollutants and child cancers[J]. Geoinformatica, 2016, 20(4): 651-92," treats pollution sources as uncertain data and models them in the real world, discovering significant associations between combinations of chemical pollutants and certain cancers. However, the mining process is highly sensitive to the selection of grid granularity, resulting in significant differences in results with different granularities. The method, "Xie Wang, Wang Lizhen, Chen Hongmei, et al. Mining the Relationship between Pollution Sources and Cancer Cases Based on Spatial Ordered Couple Patterns[J]. Data Analysis and Knowledge Discovery, 2021, 5(02): 14-31," first proposed using spatial ordered couple patterns to mine the influence of pollution sources on cancer. However, this method first mines frequent spatial ordered couple patterns based on participation, then calculates the influence of the patterns to mine frequent strong spatial ordered couple patterns, making the calculation process cumbersome. The impact calculation simply sums up the effects of pollution sources on cancer cases, without considering the differences between the pollution sources themselves or external interference with the spread of pollution sources, making the results unreasonable. Furthermore, the data mining results are significantly affected by the number of cancer cases; when the number of cases is unevenly distributed, patterns with high impact are almost concentrated in cancer characteristics with fewer cases.
[0005] Despite extensive research over the years on methods for mining the relationship between pollution sources and cancer, existing mainstream spatial isotope pattern mining methods have the following limitations in uncovering this relationship: First, they typically use a distance threshold to determine instance proximity, and the frequency of a pattern is measured only by counting the occurrences of nearby pattern instances. However, for both pollution sources and cancer cases, the closer a cancer case is to a pollution source, the higher its risk of developing cancer. Therefore, the frequency measurement criteria must consider both the frequency of frequent isotope occurrences and the varying impacts caused by different distances between pollution source instances and cancer cases. Furthermore, different pollution source emission concentrations result in different impact ranges, making it unreasonable to determine proximity solely based on a distance threshold. Second, participation-based frequency measurement methods require all feature instances to contribute equally to the pattern. Different pollution source instances belong to different carcinogenic categories, leading to varying cancer risks in humans. Therefore, when calculating the "contribution" of pollution source instances, they should not be treated equally. Finally, the features in a pattern do not need to satisfy a specific order. For example, the patterns {matsutake, pine} and {pine, matsutake} can both show the isotopic relationship of the two species' habitats. However, for pollutants and cancer, studying patterns such as {pollutant A, pollutant B} or {cancer A, cancer B} is not practically meaningful. It is more appropriate to study coexistence relationships like {pollutant A, cancer A}, which requires a specific method for generating candidate patterns. Summary of the Invention
[0006] The purpose of this invention is to provide a method for mining the impact of industrial pollution on cancer based on improved spatial isomorphism patterns, in order to solve the problems of traditional spatial isomorphism pattern mining methods in the discovery of the isomorphism relationship between pollution sources and cancer, such as unreasonable use of only a single distance threshold to determine the proximity relationship, the requirement of the frequency measurement method based on participation not meeting the relationship between pollution sources and cancer risk, and the inability of the features in the pattern to meet a specific order.
[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is a method for mining the impact of industrial pollution on cancer based on an improved spatial isotope model, comprising the following steps:
[0008] Step S1: Calculate the offset coordinates of the pollution source instance set based on wind direction and wind speed, and calculate the cancer instance c based on the offset coordinates. s .t and examples of pollution sources p i The Euclidean distance between .j is dist(c s .t,p i .j);
[0009] Step S2: Calculate and process the concentration of the input pollution source instance set to obtain the local average concentration and global average concentration of the pollution source.
[0010] Step S3: Based on the relationship between the local average concentration and the global average concentration of the pollution source, determine the size of the pollution source's influence radius, and then determine the cancer case c. s .t and examples of pollution sources p i Does the j-value have a spatial proximity relationship?
[0011] Step S4: Generate a star-shaped neighbor set for the cancer instance set;
[0012] Step S5: Generate a star-shaped influence instance table of candidate patterns based on the star-shaped neighbor set;
[0013] Step S6: Calculate the improved impact rate of candidate modes based on the KDE kernel density estimation model.
[0014] Step S7: Calculate the weighted impact rate (WER) of candidate patterns based on the carcinogenic category of the pollution source. pi (c s );
[0015] Step S8: Calculate the influence of the candidate pattern and determine the relationship between the influence of the candidate pattern WEI(SOPP_c) and the influence threshold min_pii. If WEI(SOPP_c)≥min_pii, then output the frequent pattern.
[0016] Further, step S2 specifically includes:
[0017] Step S21, Example of a pollution source p i .j Concentration values at different time periods Local average concentration for The average value is calculated as follows:
[0018]
[0019] in, conc t Let t represent the concentration value during time period t, where t∈[1,q].
[0020] Step S22, Global Average Concentration For pollution source characteristics p i Instance set The local average concentration and the average value of are calculated as follows:
[0021]
[0022] in, Represents instance set Pollution source characteristics p i The number of instances.
[0023] Further, step S3 specifically includes:
[0024] Step S31: Based on the relationship between the local average concentration and the global average concentration of the pollution source, the influence radius of the pollution source is divided into three levels, denoted as r. min r mid r max , and r min <r mid <r max , and r min r mid r max These correspond to low, medium, and high concentration levels of pollution sources, respectively.
[0025] Step S32, Example of a pollution source p i Radius of influence of .j The judgment is as follows:
[0026]
[0027] Step S33: Determine the spatial proximity relationship R between instances. c_p ,if Then it is called cancer instance c s .t and examples of pollution sources p i .j has a neighbor relationship R c_p , remember in For example, cancer case c s The radius of activity of .t.
[0028] Furthermore, the star-shaped neighbor set in step S4 refers to the set of cancer instances c. s .t is a set of pollution source instances that satisfy spatial proximity.
[0029] Further, step S5 specifically includes:
[0030] Step S51, in cancer instance c s Among the star-shaped neighbors of .t, the characteristic of a single pollution source p i The set of instances is called p i For c s The star-shaped influence instance set of .t is denoted as
[0031] Step S52, the star-shaped influence instance table is defined as cancer feature c s A subset of instances and its corresponding pollution source feature set F p The set of instances affected by the star schema, denoted as
[0032] in, For pollution source characteristics pi The set of instances, SN(c s .t) represents c s The star-shaped neighbor set of .t Let k represent the set of instances of cancer features, and let F be the feature set of pollution sources. p The number of pollution source characteristics in the data.
[0033] Furthermore, the improved influence rate in step S6 The calculation is as follows:
[0034]
[0035] Where n represents cancer characteristic c s The total number of instances, n(c _ave ) represents the average of the sum of all cancer feature instances, δ is called the smoothing factor, and -1 < δ < 0; the impact rate before improvement. The calculation is as follows:
[0036]
[0037] Where InsT(SOPP_c) is the star-shaped influence instance table, and cancer instance c is the cancer instance. s .t instance set affected by star schema Impact The calculation is as follows:
[0038]
[0039] Wherein, n(p) i ) indicates pollution source p i The total number of instances, Represents cancer instance c s .t corresponds to the pollution source characteristic p i The number of instances is affected by the star schema. Example of a pollution source p i The radius of influence of .j For example, cancer case c s The radius of activity of .t, exp(·) is an exponential function with the natural constant e as its base.
[0040] Furthermore, the weighted influence rate in step S7 The calculation is as follows:
[0041]
[0042] in, For the improved impact rate, c s Indicates cancer characteristics, v(p) i ) represents the characteristics of the pollution source p iThe corresponding carcinogenic coefficients are ε1, ε2, and ε3, and 0 < ε3 < ε2 < ε1 ≤ 1.
[0043] Furthermore, the carcinogenicity coefficient v(p) i The method for determining ) is as follows:
[0044]
[0045] Among them, category(1) represents a Group 1 carcinogen, category(2) represents a Group 2 carcinogen, and category(3) represents a Group 3 carcinogen.
[0046] Furthermore, step S8 also includes: generating k-order candidate patterns cyclically based on the star-shaped neighbor set, continuously judging the relationship between the influence degree WEI (SOPP_c) of the candidate pattern and the influence degree threshold min_pii, until the k+1-order candidate patterns are empty.
[0047] Furthermore, the influence WEI(SOPP_c) is calculated as follows:
[0048]
[0049] in, This is the weighted influence rate obtained in step S7.
[0050] The beneficial effects of this invention are
[0051] 1) This invention proposes influence as a metric for patterns. Influence measures both the frequency of frequent co-occurrences of patterns and the attenuation of influence between instance pairs with distance. Furthermore, when calculating influence, instead of generating traditional table instances, a new star-shaped influence instance table is generated to facilitate influence calculation, enabling rapid calculation of the influence of candidate patterns.
[0052] 2) Fully consider the differences between pollution source instances. First, different levels of influence radii are set according to the pollutant emission concentration, and the activity radius of cancer cases is given. The two are organically combined to identify the proximity relationship between pollution source instances and cancer cases. Second, according to the different carcinogenic categories to which each type of pollution source belongs, each pollution source characteristic is weighted and distinguished when calculating the influence rate.
[0053] 3) Based on the relationship between atmospheric pollutant diffusion and meteorological conditions, the influence of wind direction and wind speed on pollutant diffusion is taken into account, and the concept of offset distance is proposed to restore the diffusion process of pollution sources in reality as much as possible. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0055] Figure 1 This is a flowchart of a method for mining the impact of industrial pollution sources on cancer based on an improved spatial isomorphic pattern, according to an embodiment of the present invention.
[0056] Figure 2 This is a schematic diagram of the spatial feature distribution according to an embodiment of the present invention.
[0057] Figure 3 This is a schematic diagram of the wind rose in Yunnan Province according to an embodiment of the present invention.
[0058] Figure 4 This is a schematic diagram illustrating the calculation process of the coordinate offset of a pollution source instance affected by wind direction, according to an embodiment of the present invention.
[0059] Figure 5 This is a schematic diagram of spatial proximity relationships in an embodiment of the present invention.
[0060] Figure 6 This is a diagram of the three-dimensional Gaussian kernel function model according to an embodiment of the present invention.
[0061] Figure 7 This is a schematic diagram illustrating the improvement of the measurement criteria in an embodiment of the present invention.
[0062] Figure 8 This is a distribution map of the actual dataset in an embodiment of the present invention.
[0063] Figure 9 These are the distance distribution frequency histograms of two patterns calculated in a real dataset according to an embodiment of the present invention, where (a) is the pattern [particulate matter, colorectal cancer] and (b) is the pattern [particulate matter, uterine cancer].
[0064] Figure 10 This refers to the impact of wind direction and speed changes on the embodiments of the present invention.
[0065] Figure 11 The influence of the difference in the radius of influence gradient corresponding to the pollution source concentration on the embodiments of the present invention is shown in (a), where (a) is the result on the real dataset and (b) is the result on the synthetic dataset 1.
[0066] Figure 12 The results show the impact of different proportions of high-concentration pollution sources on the embodiments of the present invention, where (a) is the result under windless conditions and (b) is the result under the influence of wind speed of 15 m / s.
[0067] Figure 13 The results show the effect of cancer-specific carcinogenicity coefficients on the embodiments of the present invention, where (a) is the result under windless conditions and (b) is the result under the influence of wind speed of 15 m / s. Detailed Implementation
[0068] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0069] like Figure 1 As shown, this invention provides a method for mining the impact of industrial pollution on cancer based on an improved spatial isomorphic pattern. The specific steps are as follows:
[0070] Step S1: Determine the spatial dataset, which includes the feature set F = {PF, CF}, where PF represents the pollution source feature set and CF represents the cancer feature set.
[0071] Step S11: Pollution source feature set PF = {p1, p2, ... p i ,...,p n}, the i-th pollution source feature p i The j-th pollution source instance is represented by p i .j represents i∈[1,n]. The set of all pollution source instances is called the pollution source instance set, denoted as .j. p represents the characteristic of the i-th pollution source. i (p i A set of instances of (∈PF). In addition to feature type, instance ID, and spatial location, each pollution source instance also carries concentration and carcinogenic category information, which can be uniquely identified by a 5-tuple, namely <feature type, instance ID, concentration, carcinogenic category, spatial location>.
[0072] Step S12, Cancer feature set CF = {c1, c2, ..., c s ,...,c m A cancer instance is represented by c. s The .t symbol represents the set of all cancer instances, denoted as .t. Indicates cancer characteristics c s (c s A set of instances of ∈CF, where a cancer instance can be represented by a triple, namely <feature type, instance ID, spatial location>.
[0073] like Figure 2The given example describes the pollution source feature set PF = {A, B} and the instance set O. p = {{A.1, A.2, A.3, A.4, A.5}, {B.1, B.2, B.3, B.5}}, cancer feature set CF = {a}, instance set O c ={{a.1,a.2,a.3}}.
[0074] Step S2: Calculate the offset coordinates of the pollution source instance set based on wind direction and wind speed.
[0075] Step S21: Analyze regional wind direction based on the wind direction characteristics of typical areas. Wind direction frequency is expressed as: Wind direction frequency = Number of occurrences of a certain wind direction / Total number of observations of a wind direction × 100%. The wind rose diagram is drawn based on the multi-year average statistical values of various wind direction frequencies and wind speeds for a certain region, according to a certain ratio. The wind direction on the rose diagram refers to the direction from the outside towards the center of the region. Taking Yunnan Province as an example, the wind rose diagram of Yunnan Province is as follows: Figure 3 As shown in the diagram, the wind rose shows that the prevailing wind direction in Yunnan Province is southwest, with an average wind speed of no more than 15 m / s. Under the influence of the southwest wind, the spatial location of pollution sources gradually shifts to the northeast, resulting in a shift distance.
[0076] Step S22: Given a pollution source instance Instance p i Under the influence of wind conditions, the offset distance from the northeast is denoted as d, where d = λ|v|. Here, λ is the stretching coefficient; the larger the stretching coefficient, the greater the offset distance. v is the southwest wind speed in Yunnan Province, varying from 0-25 m / s. Referring to the relationship between atmospheric pollutant diffusion and meteorological conditions, the stretching coefficient λ is fixed at 300. Assume pollution source example p... i The spatial coordinates of j are (x, y), and p i The offset coordinates (x', y') of .j are represented as follows:
[0077]
[0078] Step S23: Calculate the Euclidean distance between the cancer instance and the contamination source instance based on the offset coordinates. For example... Figure 4 As shown, the actual distance between cancer instance a.1 and pollution source instance A.1 is dist(a.1,A.1').
[0079] Step S3: Calculate and process the concentration of the input pollution source dataset. Based on the discharge concentration of each pollutant at different time periods recorded by the monitoring system of key polluting enterprises in the region, the higher the concentration, the greater the discharge volume, and the wider the impact range. This invention links concentration with the impact range of pollution sources, defining a new spatial proximity relationship division standard related to concentration.
[0080] Step S31: Given a pollution source instance Example of a pollution source p i .j Concentration values at different time periods Local average concentration is The average value, denoted as The calculation is as follows:
[0081]
[0082] Among them, conc t Let t represent the concentration value during time period t, where t∈[1,q].
[0083] For example, Conc A.1 ={0.1,0.2,0.3}, unit mg / Nm³ 3 Then LMConc A.1 = (0.1 + 0.2 + 0.3) / 3 = 0.2 mg / Nm 3 .
[0084] Step S32: Given a pollution source feature p i (p i ∈PF), the global average concentration is p i Instance set The local average concentration and the average value of are denoted as . The calculation is as follows:
[0085]
[0086] For example, LMConc A ={0.2,0.3,0.5,0.1,0.8}, unit: mg / Nm³ 3 Then GMConc A = (0.2 + ... + 0.8) / 5 = 0.4 mg / Nm 3 .
[0087] Step S4: Determine the size of the pollution source's influence radius based on the relationship between the local average concentration and the global average concentration.
[0088] Step S41: Based on the relationship between the local average concentration and the global average concentration of the pollution source, the influence radius of the pollution source is divided into three levels, denoted as r. min r mid r max .
[0089] Step S42: Given a pollution source instance Radius of influence The definition is as follows:
[0090]
[0091] r min r mid r max Define a user-defined threshold, and r min <r mid <r max , and r min r mid r max These correspond to the low, medium, and high concentration levels of the pollution source, respectively.
[0092] Step S43, given a cancer instance Activity radius The value R is used to determine the spatial proximity relationship between instances. c_p ,if Then it is called cancer instance c s .t and examples of pollution sources p i .j has a neighbor relationship R c_p , remember
[0093] like Figure 5 As shown, dist(a.2, A.5) = 2.83, r a.2 +r A.5 =2+2=4,dist(a.2,A.5)<r a.2 +r A.5 Instances a.2 and A.5 satisfy the proximity relationship. Connect all pairs of instances that satisfy the proximity relationship with dashed lines.
[0094] Step S5: Generate a star-shaped neighbor set for the cancer instance set. Given a cancer instance... c s The star-shaped neighbor set of .t is defined as a set of contaminant instances, all of which are associated with cancer instances c. s .t satisfies the spatial proximity relation. Denoted as SN(c s .t)={p i .j∈O p |R c_p (c s .t,p i .j)}, where O p It is a set of pollution source instances, and p i The .j index is sorted lexicographically by instance name. For example... Figure 5 The star-shaped neighborhood set of the cancer instance a.1 is SN(a.1)={A.1,A.2,B.1,B.2,B.3}.
[0095] Step S6: Generate a star-shaped influence instance table of candidate patterns based on the star-shaped neighbor set.
[0096] Step S61, in cancer case c s In the star-shaped neighbors of .t, the individual pollution source characteristics p i The set of instances is called p i For c s The star-shaped influence instance set of .t is denoted as exist Figure 5 In the example set, the star-shaped impact of pollution source characteristics A and B on cancer example a.1 is SEI. A (a.1)={A.1,A.2} and SEI B (a.1) = {B.1, B.2, B.3}.
[0097] Step S62: Given a spatial order pair pattern SOPP_c = [F p ,c s ], c s ∈CF, the star-shaped influence instance table of SOPP_c is defined as cancer feature c s A subset of instances and its corresponding pollution source feature set F p The set of instances affected by the star schema, denoted as Figure 5 Table 1 shows an example of the star-shaped influence of all candidate modes.
[0098] Table 1 shows examples of the star-shaped influence of candidate patterns.
[0099]
[0100] Step S7: Based on the KDE kernel density estimation model, calculate the influence rate of candidate modes starting from the second-order mode.
[0101] Kernel density estimation (KDE) is a nonparametric estimation method for estimating probability density functions. It is an effective mathematical tool for analyzing spatial autocorrelation and density distribution. For example, the KDE model can generate a smooth three-dimensional bell-shaped surface from a given set of two-dimensional spatial data. This surface displays the clustering information of the point spatial data, and the desired clusters can be extracted based on a specified density threshold. In point spatial data, the KDE model can be represented as:
[0102]
[0103] in Based on the dataset O'={O1,O2,...,O t} space instance O x The density estimation, O' is based on O x Let t be the number of instances within a circle with radius h, and O' be the number of instances. tFor the t-th instance, h is a smoothing parameter called bandwidth, and dist(·) is used to calculate the Euclidean distance between two spatial instances. K is used to simulate instance pairs (O x O i The kernel function contributes more to the distance; the larger the distance, the smaller the contribution. The kernel function K satisfies non-negativity and has an integral of 1. Assume O x O' is a set of instances of cancer, and O' is a set of instances of pollution sources. The KDE model can consider not only O' and O', but also other elements. x The KDE model considers the impact of distance decay on the model's performance and uses density estimation to reasonably measure the frequency of patterns. The KDE model has two important parameters: the kernel function K and the bandwidth h. K can be a Gaussian function, an exponential function, a power function, etc. In this invention, since the trend of the Gaussian kernel function better reflects the impact decay, a normal Gaussian kernel function is chosen.
[0104]
[0105] Here, the independent variable x refers to the value of the input sample, parameter a determines the peak height of the surface, b determines the position of the peak on the horizontal axis, and c determines the width of the curve. In the normal Gaussian kernel function, a = c = 1 and b = 0. exp(·) is an exponential function with the natural constant e as its base. The curve of the normal Gaussian kernel function in three-dimensional space is shown below. Figure 6 As shown. Assuming the center point is a cancer instance, the set of contaminant instances satisfying the distance threshold can generate a smooth 3D bell-shaped surface based on their Euclidean distance to the cancer instance. This surface clearly shows the degree to which the influence of the contaminant on the cancer case decays with distance. Therefore, the above formula can be transformed as follows:
[0106]
[0107] cancer instance c s .t sets of instances affected by a certain star shape The impact is denoted as The calculation is as follows:
[0108]
[0109] Where n(p) i ) indicates pollution source p i The total number of instances, Represents cancer instance c s .t corresponds to pollution source p i The number of instances is affected by the star schema. Kernel functions are used to simulate instance p i .j to c sThe impact of the .t parameter decreases with increasing distance. The impact is greatest when the distance between two instances is 0; otherwise, the impact of the source instance on the cancer instance decays with distance. The effect decays to 0 if and only if the pollution source characteristic p i All instances appear in c s When the star-shaped influence instance set of .t is in c s The kernel density estimate of .t is 1, i.e., p i For c s The effect of .t is 1, but in reality, p i It is almost impossible for all instances to appear in c. s Since it is around .t, the influence value is generally less than 1.
[0110] exist Figure 5 In the context of the pattern [{A,B},a], the star-shaped influence instance table InsT([{A,B},a]) can be used for calculation.
[0111]
[0112] Similarly, we can obtain SEIE A (a.3) = 0.352, SEIE B (a.1) = 0.413, SEIE B (a.3) = 0.144.
[0113] Pollution source characteristics p i (p i ∈F p ) characteristics of cancer s (c s ∈F c The impact rate is recorded as It is a characteristic of cancer. s The instance is in the star-shaped impact instance table of pattern SOPP_c, corresponding to pollution source p. i The sum of the effects is calculated as follows:
[0114]
[0115] This value is similar to the participation rate (PR) in traditional methods, except that ER places greater emphasis on the unidirectional influence relationship between two features. Because the influence of contaminant instances in the star-shaped influence instance table is counted repeatedly when calculating the ER value, the impact of the number of cancer instances on the ER value is uncertain; the more cancer instances, the larger the ER value. An exponential function can be used to improve this, mitigating the excessive bias in the impact rate caused by differences in the number of cancer instances. The improved impact rate is denoted as... The calculation is as follows.
[0116]
[0117] like Figure 5 In the pattern [{A,B},a], the ER A (a) = SEIE A (a.1)+SEIE A (a.3)=0.355+0.352=0.707,
[0118] Where n represents cancer characteristic c s The total number of instances, n(c _ave ) represents the average of the sum of all cancer feature instances, δ is called the smoothing factor, and -1 < δ < 0. Changing images such as Figure 7 As shown.
[0119] when c s There are a lot of them. Compare Slightly smaller, n(c) s The larger the ) The smaller the value, the better. Conversely, when... c s The quantity is small. Compare Slightly larger, n(c) s The smaller, The larger the value.
[0120] Regarding improvements in impact rate, for the same type of cancer, it does not affect the generation and combination of pattern antecedents, because for the same type of cancer... They are the same. Appropriately selecting values for δ can effectively avoid pattern mining defects caused by differences in the number of cancer instances.
[0121] Step S8: Calculate the weighted impact rate of candidate patterns based on the carcinogenic category to which the pollution source belongs.
[0122] The International Agency for Research on Cancer (IARC) of the World Health Organization classifies carcinogens into four different categories: Group 1, Group 2, Group 3, and Group 4, denoted as category(1),...,category(4). From Group 1 to Group 4, the sufficiency of evidence for carcinogenicity in humans decreases sequentially. The more substantial the evidence for a carcinogen, the more likely it is to cause cancer in humans, and the higher its carcinogenicity coefficient; conversely, the less substantial the evidence, the lower the carcinogenicity coefficient.
[0123] Since pollution sources generally do not contain Group 4 carcinogens, this invention only discusses the first three groups of carcinogens. The carcinogenicity coefficient of a pollution source characteristic is denoted as v(p). i ).
[0124]
[0125] ε1, ε2, and ε3 are carcinogenicity coefficients, and 0 < ε3 < ε2 < ε1 ≤ 1. The larger the carcinogenicity coefficient, the more sufficient the evidence of carcinogenicity in humans, and the greater the impact.
[0126] The weighted impact rate is denoted as
[0127]
[0128] Where v(p) i ) represents the characteristics of the pollution source p i The corresponding carcinogenicity coefficients, ε1 + ε2 + ε3, are the sum of the coefficients for the three carcinogenic groups. For example... Figure 5 In the pattern [{A,B},a] Similarly, WER B (a) = 0.08355.
[0129] Step S9: Calculate the influence of the candidate pattern and determine the relationship between the influence of the candidate pattern WEI(SOPP_c) and the human-given influence threshold min_pii. If WEI(SOPP_c)≥min_pii, then output the frequent pattern.
[0130] Based on the star-shaped neighbor set, k-order candidate patterns are generated cyclically. The influence measurement method given in this invention (i.e., the method of judging the relationship between the influence of the candidate pattern WEI(SOPP_c) and the influence threshold min_pii) is used until the k+1-order candidate patterns are empty.
[0131] The impact is denoted as WEI(SOPP_c):
[0132]
[0133] like Figure 5 In the middle, WEI AB (a)=1-(1-WER A (a))×(1-WER B (a))=1-(1-0.3535)×(1-0.08355)=0.4075. Therefore, the influence of the final pattern [{A,B},a] is 0.4075. If the influence threshold min_pii=0.3, then the pattern [{A,B},a] is a frequent ordered even pattern. The influence of ordered even patterns increases with the pattern order; higher-order patterns may have a greater influence than lower-order patterns.
[0134] For example Figure 5 In the middle, WEI B (a) = 0.08355, WEI AB(a) = 0.4075. Therefore, unlike the participation measure of traditional spatial isomorphic patterns, the influence measure of the frequency of ordered even patterns does not satisfy the downward closure property.
[0135] The algorithm of this invention is described as follows:
[0136]
[0137]
[0138] The algorithm process is explained as follows:
[0139] Step 1: Calculate the offset coordinates of the pollution source instance based on the input stretching coefficient λ and wind speed v.
[0140] Steps 2-3: Calculate the local average concentration and global average concentration of the pollution source instance.
[0141] Step 4: Determine the corresponding radius of influence of the pollution source instance based on the relationship between the local average concentration and the global average concentration, and then generate a star-shaped neighbor set of cancer instances.
[0142] Step 5: Generate second-order candidate patterns based on the star-shaped neighbor set. Since the antecedent of a pattern corresponding to a certain cancer will only appear in the features corresponding to the star-shaped neighbor set of the cancer, the candidate patterns can be generated based on the star-shaped neighbor set.
[0143] Steps 6-8: Starting from the second order, generate a star-shaped influence instance table for each pattern in a loop, then calculate the influence degree of the pattern based on the instance table, and retain the patterns that meet the influence degree threshold. Continue until the (k+1)th order candidate patterns are empty, then exit the loop and output the results.
[0144] Time complexity analysis of Algorithm 1: Steps 1, 2, and 3 all require traversing the set of contaminant instances, with a complexity of O(n). Step 4 requires traversing all contaminant instances to determine proximity relationships for each cancer instance, hence the complexity is O(m×n), where n is the number of contaminant instances and m is the number of cancer instances. Step 5 has the same complexity as step 4. In step 7, the most time-consuming step is the star-shaped influence instance table ins_T. k The generation process requires O(|SNS|) to generate an instance table of candidate patterns, therefore it requires O(|C|). k |×|SNS|). Generating frequent patterns P k The required time is O(|C k The time required to generate a k+1 order candidate pattern is O(|C). k |×(O(|C k |)-1)). A distance threshold that is too large or a large number of instances will negatively impact the algorithm's efficiency.
[0145] Example:
[0146] To verify the correctness of this invention, this embodiment is based on one real dataset and five artificially synthesized datasets. The real dataset includes two parts: cancer case data and pollution source data. The cancer case data was provided by a hospital in a certain province and contains 28 features and 28,797 instances. The pollution source data comes from the monitoring system of key polluting enterprises in a certain province. This embodiment only retains air pollutants that are carcinogenic to humans, containing 19 features and 9,294 instances. The distribution of the real data is as follows: Figure 8 As shown, lowercase letters a-beta represent cancer, and uppercase letters AS represent pollution sources. Synthetic dataset 1 modifies the geographic spatial distribution of pollution source data to a random distribution based on the real dataset. Other synthetic datasets are randomly generated, mainly used for algorithm efficiency analysis. Each dataset has 20 features, and the number of instances is 20,000, 40,000, 60,000, and 80,000 respectively. Concentration information is randomly generated, and the data is uniformly distributed in a space between longitude 97° and 107° and latitude 20° and 30°. All experiments in this embodiment are implemented in C++, with an Intel Core i7 processor, 16GB of RAM, and Visual Studio 2019 as the runtime environment.
[0147] This embodiment divides the experiment into three parts. First, it verifies the effectiveness of the embodiment of the present invention in mining the relationship between pollution sources and cancer on a real dataset. Second, it analyzes the changes in results caused by different environmental factors considered in the present invention and the influencing factors of the instance itself on real and synthetic datasets. Third, it analyzes the efficiency comparison results of the embodiment of the present invention with other algorithms on a synthetic dataset.
[0148] Validity analysis of influence measurement:
[0149] Since spatial ordered even pattern mining requires low-level modeling of pollution source data and cancer data, and the pattern generation differs from traditional methods, comparing the algorithm of this invention with traditional algorithms is difficult. To demonstrate the effectiveness of the algorithm of this invention, this embodiment improves upon the traditional participation algorithm, calling it the PI_SOPPMA algorithm. PI_SOPPMA generates candidate patterns in the manner of the Kde_SOPPMA algorithm. The main difference lies in that PI_SOPPMA generates table instances in the traditional way, calculates the participation rate and participation degree of candidate patterns based on the table instances, and then generates frequent patterns. This embodiment compares the two algorithms and demonstrates the effectiveness of Algorithm 1 in mining the relationship between pollution sources and cancer from both macroscopic and microscopic perspectives.
[0150] Experiments were conducted using both real and synthetic datasets, without considering the influence of wind direction. The selection of the distance threshold was guided by the relationship between atmospheric pollutant diffusion and meteorological conditions, with the maximum distance threshold not exceeding a small-scale range of 10 km. In Algorithm 1, the pollutant influence radius r... min r mid r max The activity radii r of cancer cases are 4.5km, 5km, and 5.5km respectively. c The distance is 4.5km. The carcinogenicity coefficients for Group I, Group II, and Group III are set to 1, 0.7, and 0.3, respectively, and the smoothing exponent is set to 0.6. The distance threshold for the PI_SOPPMA algorithm is the same as that for Algorithm 1.
[0151] Macroeconomic analysis:
[0152] Table 2 records the minimum, maximum, and average values of the second-order pattern frequency index obtained by the two methods. As can be seen from Table 2, under the same dataset and distance threshold, the influence index is larger than the participation index. This is because the influence index emphasizes the unilateral impact of contaminants on cancer, calculating the contribution of patterns using distance-weighted calculations between neighboring pairs, transforming interest into a density estimation problem. For participation, in addition to contaminants, the contribution of cancer must also be evaluated. In the materialization model of this embodiment, contaminants are always distributed around cancer instances, and there will be a certain gap between the contaminant participation rate and the cancer participation rate. Therefore, the calculated participation index value is often smaller than the influence index value. Nevertheless, the overall trend of influence and participation is the same; both the minimum, maximum, and average values of influence and participation on the real dataset are smaller than those on the synthetic dataset 1. From a macroscopic perspective, it can be seen that the influence index calculation based on the KDE model can reflect the frequent co-location relationships of patterns in space.
[0153] Table 2. Extreme values of the frequency index for second-order frequent patterns.
[0154]
[0155] Microscopic analysis:
[0156] Although the improved participation-based approach effectively leverages the advantages of this embodiment in modeling the underlying pollution source data, the frequency index results of the two methods differ under the same dataset and distance threshold, making comparison impossible with the same frequency threshold. Therefore, this embodiment uses the Top_k frequent patterns to compare the results of the two algorithms. Table 3 records the top ten second-order frequent spatially ordered pairs of patterns with the highest influence obtained by Algorithm 1 on the real dataset, as well as the rank and participation magnitude of the relevant patterns in the PI_SOPPMA algorithm.
[0157] Table 3 Comparison of Top 10 Patterns
[0158] Top 10 Kde_SOPPMA Sort PI-SOPPMA Sort {Particulate matter, TBL (tracheal, bronchial, and lung cancer)} 0.628 <![CDATA[ 1 ]]> 0.443 <![CDATA[ 1 ]]> {Acid mist, liver cancer} 0.610 2 0.074 292 {Cobalt and its compounds, TBL} 0.601 3 0.036 441 {Acid mist, TBL} 0.596 4 0.061 343 {Smoke and Dust, TBL} 0.575 5 0.166 70 {Benzo[a]pyrene, TBL} 0.569 6 0.023 497 {Particulate matter, colorectal cancer} 0.563 <![CDATA[ 7 ]]> 0.341 <![CDATA[ 9 ]]> Benzo[a]pyrene, bone cancer 0.559 8 0.032 463 {Particulate matter, uterine cancer} 0.550 <![CDATA[ 9 ]]> 0.367 <![CDATA[ 4 ]]> {Particulate matter, bone cancer} 0.541 10 0.302 22
[0159] Table 3 presents several pieces of information. First, among the top ten frequent spatial order pairs, the pollutants that cause lung cancer and bronchial malignancies are particulate matter, cobalt and its compounds, acid mist, soot, and benzo[a]pyrene, respectively. Except for cobalt and its compounds, all of the others have evidence to show a causal relationship with the onset of lung cancer, which indicates that the results obtained by the method of this invention are consistent with reality.
[0160] Secondly, comparing the results from the two algorithms, three patterns—{particulate matter, TBL}, {particulate matter, colorectal cancer}, and {particulate matter, uterine cancer}—were all in the top 10. {particulate matter, TBL} ranked first in both methods, indicating that particulate matter has a high probability of causing lung and tracheal cancer. It is noteworthy that the {particulate matter, colorectal cancer} and {particulate matter, uterine cancer} patterns received different rankings in the two methods. Figure 9 The distance distribution frequency histograms between two pattern instances show that the instance distance distribution of {particulate matter, colorectal cancer} is similar to that of {particulate matter, uterine cancer}, but the former's average distance is lower than the latter's, indicating that instances in the former are closer together than those in the latter. Algorithm 1 effectively captures the impact of this distance difference and assigns this pattern a lower frequency ranking than the PI_SOPPMA algorithm. Third, patterns such as {benzo[a]pyrene, TBL} have a large impact but a small participation rate. This is mainly because benzo[a]pyrene instances are relatively rare. Almost all lung cancers are located around benzo[a]pyrene, but a large number of lung cancers are located around very few benzo[a]pyrene instances. The participation rate of cancer is relatively low, and the participation rate is also very low. This means that although benzo[a]pyrene has a significant impact on lung cancer, the PI_SOPPMA algorithm cannot detect such patterns, while Algorithm 1 effectively handles this situation, enabling the identification of potentially affected cancers regardless of the number of pollutant instances.
[0161] In summary, from both macroscopic and microscopic perspectives, the method proposed in this invention is more effective than traditional methods in detecting the impact of pollution sources on cancer.
[0162] Analysis of the influence of wind direction and wind speed:
[0163] The actual data distribution shows that the cancer population is more concentrated in the northeastern part of Yunnan Province. Although this region is economically developed and densely populated, and the proportion of cancer cases is expected to be higher, the density of cancer cases in the northeastern part far exceeds normal levels. There are not many polluting enterprises around the city, making it difficult to analyze the causes of cancer. This section analyzes the impact of wind direction and speed on frequent pattern generation, considering the relationship between atmospheric pollutant diffusion and wind activity. Given that it is extremely rare in real life for more than five chemical pollutants to interact and cause cancer, this embodiment will only discuss the results of mining up to the 5th order frequent patterns, except for efficiency comparisons, and will not elaborate further. Figure 10 This shows the generation of frequent patterns in the real dataset and synthetic dataset 1 at different wind speeds, and the pollutant influence radius r. min r mid r max The activity radii r of cancer cases are 5km, 5.5km, and 6km respectively. c The range is 2km. The carcinogenicity coefficients for Group I, Group II, and Group III are set to 1, 0.7, and 0.3, respectively. The smoothing exponent is set to 0.6, the influence threshold is set to 0.6, and the wind speed range is set to 0-25m / s, referring to the wind rose diagram of Yunnan Province.
[0164] from Figure 10 It is not difficult to observe that, regardless of whether it is real or synthetic data, under the same impact threshold, the frequency of pollutant patterns gradually increases with increasing wind speed. This is mainly because pollution sources gradually shift northeastward under the influence of southwesterly winds, leading to an increase in pollutants in areas that were originally sparsely polluted. For the real data distribution, the increase is relatively slow from wind speeds of 0-15 m / s, but the frequency of patterns increases sharply after 15 m / s. This indicates that when the wind speed is greater than 15 m / s, carcinogenic pollutants that were originally far from human settlements will move closer to them due to wind activity. The higher the wind speed, the more pollutants accumulate, and the greater the risk of cancer in humans, which aligns with the hypothesis of this invention. In contrast, in the synthetic data experiment, because the data is uniformly distributed, the impact of pollutant shifts on the spatial distribution of the data is relatively small. Although the frequency of patterns increases slowly after wind activity, the increase is very small, indicating that wind has a relatively small impact on the carcinogenicity of uniformly distributed pollution sources. Nevertheless, this does not mean that enterprises should be established evenly throughout Yunnan Province, because wind speeds in Yunnan Province are distributed in the range of 0-15 m / s for about 90 percent of the time. Within this range, synthetic data generates more frequent patterns than real data.
[0165] Analysis of the effect of concentration:
[0166] The impact of differences in the radius of influence gradient corresponding to pollution source concentration on model mining results:
[0167] Based on the relationship between the local average concentration and the global average concentration of pollution sources, the radius of influence can be divided into three levels. The radius of influence varies with gradient, and the pattern mining results corresponding to different gradients are also different. Figure 11 The effect of three levels of influence radius on frequent pattern mining results under different gradient variations is shown. The selection of influence radius is made as realistic as possible; when the gradient is 3km, r... min =4km, r mid =7km, r max =10km, which just reaches the maximum value within a small scale. Therefore, we choose r. min The range is between 0.5km and 4km, and the activity radius of cancer cases is r. c The distance is fixed at 1km. The carcinogenicity coefficients for Group I, Group II, and Group III are set to 1, 0.7, and 0.3, respectively. The smoothness exponent is set to 0.6, and the impact threshold is set to 0.3.
[0168] For both real and synthetic datasets, under the same conditions, the larger the gradient affecting the radius change, the more frequent patterns are mined, which is as expected. This is because r min Equal, the larger the gradient r mid and r max The larger the value, the more pairs of instances with proximity relationships there are, and therefore the greater the calculated influence.
[0169] The impact of varying proportions of high-concentration instances from pollution sources on pattern mining results:
[0170] This embodiment also analyzes the impact of the proportion of high-concentration instances of pollution sources on the results of frequent pattern mining. Figure 12 The results are from experiments conducted under different wind conditions, increasing the proportion of high-concentration instances of pollution sources from 10% to 100% on both the real dataset and synthetic dataset 1. The classification of instances into low, medium, and high concentrations was done randomly; excluding high-concentration instances, the remaining instances were randomly assigned low and medium concentrations in a 1:1 ratio. The influence radii for low to high pollutant concentrations were 5 km, 6 km, and 7 km, respectively, while the activity radius for cancer cases was fixed at 2 km. The carcinogenicity coefficients for Group I, II, and III were set to 1, 0.7, and 0.3, respectively, with a smoothing exponent of 0.6 and an influence threshold of 0.6.
[0171] Under the same conditions, regardless of whether there was wind or not, the proportion of high-concentration instances from pollution sources was positively correlated with the number of frequent patterns. The more high-concentration instances there were, the more pollution source instances with a wider impact range, thus increasing the number of frequent patterns. It is noteworthy that under wind conditions, both real and composite data showed more frequent patterns than in windless conditions. In real data, the number of frequent patterns increased significantly with the increase in the proportion of high-concentration instances, exceeding normal levels. This indicates that when there is wind, the more high-concentration instances from pollution sources there are, the significantly increased risk of cancer in humans.
[0172] Analysis of the impact of carcinogenicity coefficient:
[0173] The carcinogenicity coefficient is mainly used to calculate the weighted impact rate. Given different carcinogenicity coefficients, they will eventually be converted into weights. Therefore, in this embodiment, the carcinogenicity coefficient is directly set as the weight for the experiment. Figure 13 This demonstrates the impact of changes in the carcinogenicity coefficient weights of the three types of carcinogens on the pattern mining results. Experiments were conducted on a real dataset, with the pollution source's influence radius r... min The range is 1 to 5 km, with an influence radius gradient of 1 km, and the activity radius r of cancer cases. c With a fixed distance of 1km, a smoothing exponent of 0.6, and an influence threshold of 0.35, the system is designed to be 1km long.
[0174] Overall, regardless of windy or calm conditions, when a Group 1 carcinogen has a higher carcinogenicity factor weight, more frequent patterns are generated. Conversely, when the carcinogenicity factor weights of Group 1 carcinogens are the same, a higher weight for Group 2 carcinogens also results in more frequent patterns. This indicates that the higher the weight of a highly carcinogenic pollutant, the more frequent the patterns generated, and the greater the risk of cancer in humans.
[0175] Comparative analysis:
[0176] This embodiment compares the method of the present invention with the PSSOPP_OA algorithm proposed in “Xie Wang, Wang Lizhen, Chen Hongmei, et al. Identifying Relationship Between Pollution Sources and Cancer Cases with Spatial Ordered Pair Patterns[J]. Data Analysis and Knowledge Discovery, 2021, 5(02):14-31.]” on a real dataset. Although the input parameters of the two algorithms are different, the parameters of the PSSOPP_OA algorithm can still be controlled to make it have the same distance threshold when generating star-shaped neighbor sets of cancer instances. Table 4 shows the distribution of cancer features and the number of instances in the consequents of the Top 100 patterns mined by the algorithm Kde_SOPPMA and PSSOPP_OA of the present invention when the distance threshold is the same.
[0177] Table 4. Distribution of cancer characteristics and number of cases in the Top 100 patterns.
[0178] Sort Kde_SOPPMA Instance count percentage PSSOPP_OA Instance count 1 o 1592 5.5% w 104 0.3% 2 w 104 0.3% t 362 1.2% 3 d 1637 5.6% p 291 1.0% 4 g 449 1.5% y 127 0.4% 5 x 470 1.6% u 132 0.4% 6 alpha 308 1.0% f 359 1.2% 7 e 2032 7.0% 8 k 1065 3.7%
[0179] In this embodiment, multiple experiments were conducted on real datasets. In the top 100 patterns of Kde_SOPPMA, the average number of cancer features was approximately 11, accounting for about 40% of the total cancer features. Furthermore, the number of cancer instances varied, ranging from a large to a small percentage. For example, in Table 4, the number of instances in the top 100 patterns of Kde_SOPPM ranged from as high as 2032 to as low as 104. In contrast, in the top 100 patterns of the PSSOPP_OA algorithm, the average number of cancer features was approximately 6, accounting for about 20% of the total cancer features. The number of cancer instances was almost always small, concentrated in features with relatively few instances in the dataset. This is mainly because the PSSOPP_OA algorithm uses the total number of instances corresponding to cancer features as the denominator when calculating the pattern influence rate. This means that features with the fewest cancer instances often have the largest influence rate, resulting in the top k patterns almost always being cancers with a small percentage of instances. In reality, cancers can be frequent or rare, yet the PSSOPP_OA algorithm performs poorly when mining frequent cancers. The proposed Kde_SOPPMA algorithm improves upon this problem by utilizing a smoothing factor, resulting in greater stability when the number of cancer cases is unevenly distributed.
[0180] Analysis of multiple experimental results revealed a link between pollution source A (particulate matter) and many types of cancer, with lung cancer, liver cancer, and colorectal cancer showing a particularly strong correlation. This indicates that particulate matter in the air poses a significant health risk, and the government should strengthen regulations on particulate matter emissions from factories to minimize harm. Furthermore, while some pollutants alone pose little risk of cancer, their influence increases dramatically when combined with other pollutants. Some pollutants appear to act as catalysts; for example, acid mist has negligible influence when occurring alone, but its influence increases significantly when combined with smoke and particulate matter, potentially linking it to the development of many cancers. This study does not aim to establish a causal relationship but rather to identify the correlation between pollutants and cancer, hoping to provide some insights for cancer prevention and treatment.
[0181] In summary, traditional spatial isomorphic pattern mining algorithms have many limitations when mining spatial ordered even patterns. This invention combines a kernel density estimation model with a spatial isomorphic pattern mining algorithm and proposes a frequency measurement method based on distance decay effect and influence weighting. This not only solves the limitations of traditional methods, but also improves the mining anomalies caused by uneven distribution of cancer cases by using a smoothing factor. Finally, it takes into account the influence of wind environment and pollutant concentration on pollutant diffusion as much as possible, making the mining results more reasonable and accurate, and better reflecting the influence of pollution sources on cancer.
[0182] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0183] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A method of mining the effect of industrial pollution on cancer based on improved spatial co-location patterns, characterized in that, The method comprises the following steps: Step S1, calculate the offset coordinates of the pollution source instance set according to the wind direction and wind speed, calculate the cancer instance c s from the Euclidean distance dist(c s .t, p i .j) between the pollution source instance p i and the cancer instance c s . Step S2, performing calculation processing on the concentration of the input pollution source instance set to obtain a local average concentration and a global average concentration of the pollution source; Step S3, judging the size of the influence radius of the pollution source according to the relationship between the local average concentration and the global average concentration of the pollution source, and further judging the cancer instance c s whether there is a spatial adjacent relationship between the cancer instance c i and the pollution source instance p Step S4, generating a star neighbor set of the cancer instance set; Step S5, generating a star influence instance table of the candidate pattern based on the star neighbor set; Step S6, based on the KDE kernel density estimation model, calculate the improved influence rate of the candidate mode The improved impact rate in the step S6 is calculated as follows: where n represents the number of all instances of cancer characteristics c s n(c _ave ) represents the average of the total sum of all instances of cancer characteristics, δ is called a smoothing factor, and -1 < δ < 0; the impact rate before improvement is calculated as follows: wherein InsT(SOPP_c) is a star-influenced instance table, cancer instance c s is influenced by the star-influenced instance set is calculated as follows: where n(p i ) denotes the total number of instances of pollution source p i , denotes the number of instances of cancer c s .t corresponding to the star-shaped influence of pollution source feature p i , is the influence radius of pollution source instance p i .j, is the activity radius of cancer instance c s .t, and exp(·) is the exponential function with base the natural constant e; Step S7, calculating a weighted impact rate of the candidate pattern based on the carcinogenic category to which the pollution source belongs Step S8, calculating the influence degree of the candidate pattern, and judging the relationship between the influence degree WEI(SOPP_c) of the candidate pattern and the influence degree threshold min_pii, if WEI(SOPP_c)≥min_pii, the frequent pattern is output.
2. The method of claim 1, wherein the method is based on improved spatial co- pattern mining to mine the effect of industrial pollution on cancer. The step S2 is specifically: Step S21, Example of a pollution source p i .j Concentration values at different time periods Local average concentration for The average value is calculated as follows: wherein, conc t denotes the concentration value for the time period t, t e [1, q], Step S22, global average concentration For pollution source characteristic p i of the example set The local average concentration and the average value of are calculated as follows: wherein, represents an instance set of pollution source characteristics p i of an instance number.
3. The method of claim 1, wherein the method is based on improved spatial co- pattern mining to mine the effect of industrial pollution on cancer. The step S3 is specifically: Step S31, according to the relationship between the local average concentration and the global average concentration of the pollution source, the pollution source influence radius is divided into three levels, respectively denoted as r min , r mid , r max , and r min < r mid < r max , and r min , r mid , r max respectively correspond to the low, medium and high levels of the concentration of the pollution source; Step S32, pollution source instance p i Influence radius of j The following is determined: Step S33, judging spatial proximity relation R between instances c_p If then cancer instance c s .t is said to be in proximity to pollution source instance p i .j c_p , denoted as where is the activity radius of cancer instance c s .t.
4. The method of claim 1, wherein the method is based on improved spatial co- pattern mining to mine the effect of industrial pollution on cancer. The star neighbor set in the step S4 refers to the cancer instance c s A set of pollution source instances that satisfy the spatial proximity relationship.
5. The method of claim 1, wherein the method is based on improved spatial co- pattern mining to mine the effect of industrial pollution on cancer. The step S5 is specifically: Step S51, in cancer instance c s Single pollution source feature p in the star neighbors of.t i Instance set of p called p i To c s Star influence instance set of.t, denoted by Step S52, the star influence instance table is defined as the cancer characteristics c s of the example subset and its corresponding pollution source characteristic set F p The collection of star influence instance sets of the example subset, denoted as wherein, is a set of instances of source features p i is a set of instances of source features p s is a set of instances of source features p s is a set of instances of source features p is a set of instances of source features p p is a set of instances of source features p 6. The method of claim 1, wherein, The weighting influence rate in the step S7 is calculated as follows: wherein, is the improved impact rate, c s represents the cancer characteristics, v(p i ) represents the pollution source characteristics p i corresponding carcinogenic coefficients, ε1, ε2, ε3 are carcinogenic coefficients and 0 < ε3 < ε2 < ε1 ≤ 1.
7. The method of claim 6, wherein the method is based on improved spatial co- pattern mining to mine the effect of industrial pollution on cancer. The carcinogenicity coefficient v(p i ) is determined as follows: Wherein, category(1) represents a class of carcinogens, category(2) represents a class of carcinogens, and category(3) represents a class of carcinogens.
8. The method of claim 1, wherein the method is based on improved spatial co- pattern mining to mine the effect of industrial pollution on cancer. The step S8 further comprises: based on the star neighbor set, a k-order candidate pattern is generated in a loop, and the relationship between the influence degree WEI(SOPP_c) of the candidate pattern and the influence degree threshold min_pii is constantly judged until the k+1-order candidate pattern is empty.
9. A method for mining the effect of industrial pollution on cancer based on improved spatial co-location patterns according to claim 1 or 8, characterized in that, The influence degree WEI(SOPP_c) is calculated as follows: wherein, is the weighted influence rate obtained at step S7.
Citation Information
Patent Citations
Cancer gene mining analysis method based on ceRNA network
CN115631172A
Systems, methods, and computer program products for analysis of vessel attributes for diagnosis, disease staging, and surfical planning
US20070019846A1