Association analysis method and device based on weighted count, equipment and storage medium
Patent Information
- Application Number
- CN202611089613.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-22
- Publication Date
- 2026-09-29
AI Technical Summary
相关技术如FP-Growth算法,采用等权计数机制,完全忽略了环境数据的时空衰减特性
本发明实施例提供了一种基于加权计数的关联分析方法、装置、设备及存储介质,通过
Smart Images

Figure CN122838418A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of association analysis technology, and in particular to an association analysis method, apparatus, device and storage medium based on weighted counting. Background Technology
[0002] Environmental monitoring data exhibits spatiotemporal heterogeneity and varying risk intensity. Traditional association rule mining employs equal-weighted counting, assigning equal weight to all events. Related techniques, such as the FP-Growth algorithm, completely ignore the spatiotemporal decay characteristics of environmental data. This spatiotemporal blindness results in mining results filled with numerous outdated, spatially unrelated "pseudo-strong rules," obscuring truly urgent near-term risk signals. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a method, apparatus, device and storage medium for association analysis based on weighted counting, which effectively filters out false strong rules and improves the accuracy of risk identification by introducing a three-dimensional weighted counting mechanism of time, space and risk intensity.
[0004] In a first aspect, embodiments of the present invention provide a weighted counting-based association analysis method, comprising: acquiring environmental monitoring data and preprocessing the environmental monitoring data to generate standardized transactions; determining the time decay factor, spatial distance factor, and risk intensity factor of each risk item in the standardized transactions; dynamically determining the weights of each dimension based on the time decay factor, spatial distance factor, and risk intensity factor using the entropy weight method, and calculating the comprehensive weight of each transaction as the transaction weighted counting weight; inputting the standardized transactions and their corresponding transaction weighted counting weights into an empty weighted frequent pattern tree; during the sequential insertion of each standardized transaction, for each tree node on the path, adding the corresponding transaction weighted counting weight to the current count value of the tree node to update the count value of the tree node, thereby constructing a complete weighted frequent pattern tree; and performing association rule mining based on the weighted frequent pattern tree to output weighted association rules.
[0005] In a preferred embodiment of the present invention, the above-mentioned dynamic determination of the weights of each dimension based on the time decay factor, spatial distance factor, and risk intensity factor using the entropy weight method, and the calculation of the comprehensive weight of each transaction as the transaction weighted counting weight, includes: calculating a basic objective weight vector for multiple dimensions using the entropy weight method; obtaining a preset medium characteristic correction coefficient vector based on the type of the current environmental medium; multiplying the basic objective weight vector and the medium characteristic correction coefficient vector element by element to obtain a corrected weight vector; normalizing the corrected weight vector to obtain a scenario-based weight vector; using the scenario-based weight vector to perform a weighted summation of the time decay factor, spatial distance factor, and risk intensity factor to obtain a comprehensive weight; and using the comprehensive weight as the transaction weighted counting weight.
[0006] In a preferred embodiment of the present invention, the above-mentioned preprocessing of environmental monitoring data to generate standardized transactions includes: cleaning the environmental monitoring data, removing outliers, and filling missing values using interpolation; unifying environmental monitoring data from different monitoring sources into the smallest common time granularity, and establishing a unique monitoring point identifier and a standardized timestamp; determining the unidirectional influence relationship between monitoring points based on the physical topology matrix, and determining the optimal time lag through cross-correlation analysis; extracting driving factors backtracking from the abnormal moment of the response factor, and recombining the driving factors and the response factor into logically synchronized transaction units to generate standardized transactions.
[0007] In a preferred embodiment of the present invention, the above-mentioned construction of a complete weighted frequent pattern tree includes: a first scan of the weighted transaction database to count the weighted support of each risk item; a construction of an item header table based on the weighted support and sorting it in descending order of weighted support; a second scan of the weighted transaction database to rearrange the items within each transaction according to the priority order of the item header table; and inserting the rearranged transactions into the weighted frequent pattern tree.
[0008] In a preferred embodiment of the present invention, the method further includes a multi-dimensional rule evaluation step: calculating the weighted support, weighted confidence, weighted lift, and weighted certainty of the mined association rules; normalizing the weighted support, weighted confidence, weighted lift, and weighted certainty; weighting and summing the normalized weighted support, weighted confidence, weighted lift, and weighted certainty based on a preset AHP weight vector to generate a comprehensive rule score; and filtering and outputting association rules with a comprehensive rule score higher than a preset score value based on preset differential confidence thresholds for different environmental media.
[0009] In a preferred embodiment of the present invention, the above method further includes a step of dynamic threshold determination: determining a statistical threshold based on statistical distribution characteristics and a regulatory threshold based on environmental regulatory standards; for positive risk indicators, taking the minimum value between the statistical threshold and the regulatory threshold; for negative risk indicators, taking the maximum value between the statistical threshold and the regulatory threshold; setting a three-level risk level boundary point based on the positive risk indicators and the negative risk indicators; and mapping the monitoring values of continuous indicators to multiple discrete risk levels.
[0010] In a preferred embodiment of the present invention, the method further includes an incremental update step: when new monitoring data flows in, the data growth rate is calculated; if the data growth rate is less than a preset growth rate threshold, the new monitoring data is converted into a new transaction and directly inserted into the existing frequent pattern tree; if the data growth rate is greater than or equal to the growth rate threshold or the number of tree nodes exceeds a preset threshold, the weighted frequent pattern tree is reconstructed in a background thread based on all historical data and new monitoring data, and switched by atomic pointers.
[0011] Secondly, embodiments of the present invention also provide a weighted counting-based association analysis device, comprising: a standardized transaction generation module for acquiring environmental monitoring data and preprocessing the environmental monitoring data to generate standardized transactions; a determination module for determining the time decay factor, spatial distance factor, and risk intensity factor of each risk item in the standardized transaction; a transaction weighted counting weight calculation module for dynamically determining the weights of each dimension based on the time decay factor, spatial distance factor, and risk intensity factor using the entropy weight method, and calculating the comprehensive weight of each transaction as the transaction weighted counting weight; a data input module for inputting the standardized transactions and their corresponding transaction weighted counting weights into an empty weighted frequent pattern tree; a weighted frequent pattern tree construction module for, during the sequential insertion of each standardized transaction, adding the corresponding transaction weighted counting weight to the current count value of each tree node on the path to update the count value of the tree node, thereby constructing a complete weighted frequent pattern tree; and a weighted association rule output module for performing association rule mining based on the weighted frequent pattern tree and outputting weighted association rules.
[0012] Thirdly, embodiments of the present invention also provide an electronic device, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the weighted counting-based association analysis method of the first aspect described above.
[0013] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are invoked and executed by a processor, the computer-executable instructions cause the processor to implement the weighted counting-based association analysis method of the first aspect described above.
[0014] The embodiments of the present invention bring the following beneficial effects: This invention provides a method, apparatus, device, and storage medium for association analysis based on weighted counting, through... Other features and advantages of this disclosure will be set forth in the following description, or some features and advantages may be inferred from the description or determined without doubt, or may be learned by practicing the techniques described above.
[0015] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 A flowchart illustrating a weighted counting-based association analysis method provided in an embodiment of the present invention; Figure 2 A flowchart illustrating another association analysis method based on weighted counting provided in an embodiment of the present invention; Figure 3 A flowchart illustrating another association analysis method based on weighted counting provided in an embodiment of the present invention; Figure 4 A schematic diagram of a correlation analysis device based on weighted counting provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] Environmental monitoring data exhibits spatiotemporal heterogeneity and varying risk intensity. Traditional association rule mining employs equal-weighted counting, assigning equal weight to all events. Related techniques, such as the FP-Growth algorithm, completely ignore the spatiotemporal decay characteristics of environmental data. This spatiotemporal blindness results in mining results filled with numerous outdated, spatially unrelated "pseudo-strong rules," obscuring truly urgent near-term risk signals.
[0020] Based on this, the present invention provides a method, apparatus, device, and storage medium for association analysis based on weighted counting. This method can acquire environmental monitoring data, preprocess the data to generate standardized transactions, determine the time decay factor, spatial distance factor, and risk intensity factor for each risk item in the standardized transactions, dynamically determine the weights of each dimension based on these factors using the entropy weight method, and calculate the comprehensive weight of each transaction as the transaction weighted counting weight. The standardized transactions and their corresponding transaction weighted counting weights are input into an empty weighted frequent pattern tree. During the sequential insertion of each standardized transaction, for each tree node along the path, the current count value of the tree node is added to the corresponding transaction weighted counting weight to update the tree node's count value, thus constructing a complete weighted frequent pattern tree. Association rule mining is then performed based on the weighted frequent pattern tree to output weighted association rules. This method, by introducing a three-dimensional weighted counting mechanism of time, space, and risk intensity, effectively filters out false strong rules and improves the accuracy of risk identification.
[0021] To facilitate understanding of this embodiment, a weighted counting-based association analysis method disclosed in this embodiment of the invention will first be described in detail.
[0022] Example 1 This invention provides a correlation analysis method based on weighted counting. Figure 1 This is a flowchart illustrating a weighted counting-based association analysis method provided in an embodiment of the present invention. Figure 1 As shown, this association analysis method based on weighted counting may include the following steps: Step S101: Obtain environmental monitoring data and preprocess the environmental monitoring data to generate standardized transactions.
[0023] Specifically, preprocessing environmental monitoring data to generate standardized transactions can include: cleaning the environmental monitoring data, removing outliers, and filling in missing values using interpolation; unifying environmental monitoring data from different monitoring sources into a minimum common time granularity, and establishing unique monitoring point identifiers and standardized timestamps; determining the unidirectional influence relationship between monitoring points based on the physical topology matrix, and determining the optimal time lag through cross-correlation analysis; and backtracking to extract driving factors based on the abnormal moments of response factors, recombining the driving factors and response factors into logically synchronized transaction units to generate standardized transactions.
[0024] First, a rationality review can be conducted based on professional knowledge to identify and remove abnormal records that do not conform to physical and chemical laws. For missing values, different interpolation strategies are adopted according to the data characteristics: cubic spline interpolation is used to process smooth time series data, and kriging interpolation is used to process missing spatial dimensions.
[0025] To eliminate discrepancies in sampling frequencies from different monitoring sources, a minimum common time unit can be used as a benchmark to align the data with a unified time granularity. For high-frequency monitoring data, such as online monitoring, mean pooling or max pooling strategies are used to aggregate it into a standard time granularity, such as hourly. For low-frequency data, such as manually sampled data, interpolation algorithms are used to fill in the gaps. Simultaneously, by establishing a unique monitoring point identifier (SiteID) and a standardized timestamp, it is ensured that each transaction has an equal statistical benchmark in both spatiotemporal dimensions.
[0026] Furthermore, in order to eliminate the influence of dimensions and meet the requirement of non-negativity of data for subsequent entropy weight method calculation, the spatiotemporally aligned multi-source data is subjected to Min-Max normalization, as shown in the following formula (1): (1) in, For the normalized data, The original data, and These are the minimum and maximum values of the indicator within the statistical period, respectively. It should be noted that, in order to maintain the intuitiveness of the physical semantics of the environmental standard, the subsequent "dynamic threshold determination" (Formulas 2 and 3) and "risk level mapping" (Formula 4) are still based on the raw physical quantity data and do not use the normalized data here.
[0027] In order to fully explore the driving role of meteorological and hydrological conditions on receptor response and capture causal relationships without explicitly modeling complex transport paths, a sliding window technique and a statistically based spatiotemporal backtracking mechanism can be introduced to construct composite features: Extracting statistical features using the sliding window technique can include: Cumulative characteristics: such as cumulative rainfall in the past 24 hours and duration of continuous rainfall, are used to capture the cumulative effect of environmental stress.
[0028] Statistical features, such as daily maximum temperature and daily temperature range, are used to capture the impact of extreme weather conditions.
[0029] To address the complexity of environmental risk transmission, a two-level mapping mechanism combining spatial topological constraints and temporal cross-correlation optimization can be constructed, achieving precise end-to-end association between the source set and the receptor set. The first step is spatial mapping based on physical topology: The system pre-configures a physical topology matrix based on GIS geographic information system and hydrodynamic / gas field data. This matrix defines the unidirectional influence relationships between monitoring points. Let the set of points be... Matrix elements If the point lie in If the upstream catchment area or the upwind direction of the prevailing wind direction is within the effective influence threshold, then it is marked. ,Will Included The set of candidate driving sources. Conversely, if the physical path is unreachable, such as if located in different watersheds or downstream, then Through this step, the system uses domain knowledge to filter out more than 90% of invalid combinations before data mining, such as using downstream data to explain upstream anomalies, ensuring the physical compliance of the causal direction.
[0030] The second step is time lag optimization based on cross-correlation analysis: Within a defined set of candidate driving sources, for the upstream-to-downstream time lag problem, pure data-driven cross-correlation analysis can replace the traditional estimation method based on physical flow velocity, automatically finding the optimal lag amount. Lag Calculation: For any candidate driving factor (e.g., the discharge sequence of an upstream chemical plant), and receptor response. Cross-correlation analysis was performed on historical time series data (such as downstream cross-sectional water quality series). The cross-correlation function was calculated. The system searches within a preset time window (e.g., 0 to 72 hours) and retrieves... When the maximum value is reached, the corresponding Optimal transport lag time as the driving factor .
[0031] Feature backtracking and virtual transaction construction: based on the moment when the receptor anomaly occurs Using the reference anchor point, the calculated Targeted backtracking extraction The upstream driving data at any given moment. Through this end-to-end mapping mechanism that determines direction based on spatial topology and step size based on temporal correlation, the system reorganizes the originally discrete and asynchronous multi-point monitoring data into logically synchronous virtual transactions.
[0032] For example: {Upstream heavy rain (T-4h) → Chemical plant wastewater discharge (T-2h)} → {Downstream fish mortality (T)}. This method avoids complex physical modeling errors, greatly simplifies model dimensions, and improves the efficiency of association rule mining.
[0033] Step S102: Determine the time decay factor, spatial distance factor, and risk intensity factor for each risk item in the standardized transaction.
[0034] The method may further include a dynamic threshold determination step: determining a statistical threshold based on statistical distribution characteristics and a regulatory threshold based on environmental regulations and standards; for positive risk indicators, taking the minimum value between the statistical threshold and the regulatory threshold; for negative risk indicators, taking the maximum value between the statistical threshold and the regulatory threshold; setting three-level risk level boundary points based on positive and negative risk indicators; and mapping the monitoring values of continuous indicators to multiple discrete risk levels.
[0035] It should be noted that, in order to maintain the intuitiveness of the physical semantics of environmental standards, the dynamic threshold determination is still based on the original physical quantity data, and the normalized data mentioned above is not used.
[0036] Among them, by calculating statistical thresholds and regulatory thresholds The strictest value of both will be used as the final risk assessment benchmark. .
[0037] Statistical threshold Calculated based on a moving time window, such as raw monitoring data from the past 3 months. Where μ is the moving average, σ is the standard deviation, and k is the statistical multiple, typically k=3.0. Regulatory threshold. Based on the national / local standard limit (S) corresponding to the environmental functional zoning of the monitoring point, and introducing the regional safety factor η.
[0038] This application's embodiments can automatically identify the risk direction of indicators and calculate the final risk assessment benchmark. : Positive risk indicators, the higher the value, the worse, such as COD and PM2.5: adopt the strict principle and take the minimum value between the statistical threshold and the regulatory threshold, as shown in the following formula (2): (2) If the original monitoring value x> If so, it is considered a risk.
[0039] The lower the value of the adverse risk indicator, the worse it is. For example, dissolved oxygen (DO): adopt the strict principle and take the maximum value of the statistical threshold and the regulatory threshold, as shown in the following formula (3): (3) If the original monitoring value x < If so, it is considered a risk.
[0040] This mechanism ensures dual security: in areas with high environmental background values, regulatory thresholds serve as mandatory constraints; in areas with excellent environmental quality, statistical thresholds can sensitively detect early signals of water quality deterioration (or abnormal decline in dissolved oxygen). The embodiments of this application, through the aforementioned adaptive logic, transform complex and ever-changing environmental monitoring data into a structured set of risk transactions.
[0041] Among them, the setting of the three-level risk level dividing point: basic threshold (Risk Initiation Line), Early Warning Threshold (Warning line) and high-risk threshold (High-risk line), thereby constructing a continuous hierarchical spectrum covering no risk to extremely high risk. Based on this, the mapping function... Data transformation is achieved through dual-channel processing: input variables The aim is to handle environmental characteristics with varying dimensions through a unified interface. In practical implementation, it can be based on the indicator attributes... The construction can be divided into the following two cases: Direct mapping to instantaneous values of instantaneous physicochemical indicators: For water quality parameters such as COD and ammonia nitrogen, Directly take the physical monitoring value at the current sampling time, such as At this point, the aforementioned basic threshold... Warning threshold and high-risk threshold The corresponding water quality concentration standard limit for this functional area is loaded.
[0042] Statistical mapping of time-series characteristics for cumulative statistical indicators: For environmental factors with cumulative effects, a statistical counting mode can be enabled. Taking the construction of composite characteristics of high-temperature heat accumulation risk as an example, the specific execution process is as follows: Feature definition and window calculation: First, the indicator definition (RuleID:TEMP_ACC_HIGH) is loaded according to the preset rule base, and the sliding window is dynamically configured according to the characteristic action cycle of environmental pressure. Regarding this indicator, considering the temperature rise lag (thermal inertia) caused by the specific heat capacity of water and the tolerance limit of aquatic organisms, the following parameters are set: The criterion is set as a daily maximum temperature > 30℃. The preprocessing module scans the data of the past 7 days and calculates the number of consecutive days that meet the criterion up to the current time as the input variable. Assuming monitoring data shows that the highest temperature has exceeded 30℃ for the past three consecutive days, then x=3, unit: days.
[0043] Threshold system adaptive switching: If the indicator is cumulative, it automatically loads the corresponding time-sensitive threshold set from the database, rather than the concentration threshold. (Settings) Day (the beginning of accumulated risk, i.e., the thermal stress begins to manifest), Day (Severe cumulative warning) Day (extremely high risk, i.e., reaching the tolerance threshold).
[0044] Mapping execution: Substitute x=3 into the following formula (4) for judgment. Since 2<3≤4 (i.e., satisfies the condition) The system determines this record to be of high risk level (Code=3) and assigns an intensity factor I(x) = 0.8, ultimately generating a semantic discrete identifier DRIVER_TEMP_ACC_LEVEL3. This function makes the determination based on the adapted threshold system described above, simultaneously outputting the discrete level code Code and intensity factor. .
[0045] (4) The application logic of the mapping results is clearly layered: First, the system uses a prefix-based semantic naming convention to generate unique identifiers (ItemIDs) to automatically distinguish between driver factors and receptor responses. The specific format follows "[ROLE]_[INDICATOR]_[LEVEL]:".
[0046] [ROLE] (role prefix): DRIVER_ represents historical / real-time driving factor, RC_ represents receptor response, and PRED_ represents virtual driving factor generated by the prediction model (used for branch prediction).
[0047] [INDICATOR] (Indicator Code): A standardized and unique code representing the monitoring parameter (e.g., RAIN represents rainfall, TEMP represents air temperature, and COD represents chemical oxygen demand).
[0048] [LEVEL] (Level suffix): A standardized suffix generated based on the Code value output by the above formula (4), in the format of LEVEL plus the corresponding number (such as LEVEL0 to LEVEL4).
[0049] The parsing logic is as follows: Driver Factor: When the system resolves to the prefix DRIVER_ (e.g., DRIVER_RAIN_ACC_LEVEL4), it automatically categorizes it into the historical / real-time driver predecessor set. .
[0050] Virtual driver factor: When the system resolves to the prefix PRED_ (e.g., PRED_RAIN_ACC_LEVEL4), it automatically classifies it into the predictive driver antecedent set. And apply the confidence decay coefficient.
[0051] Receptor response: When the prefix RC_ is parsed (e.g., RC_COD_LEVEL3), it is automatically classified into the consequent set. .
[0052] Furthermore, intensity factor numerical set Employing a nonlinear heuristic mapping, its design strictly adheres to the response prioritization principle in environmental management practice: low risk (0.2) corresponds to controllable impact, requiring only routine monitoring; medium risk (0.5) corresponds to 1.0–1.5 times the exceedance (i.e., to (Range) requires enhanced monitoring and intervention; high risk (0.8) corresponds to 1.5–2.0 times the limit (i.e., to (range), triggering special governance; extremely high risk (1.0) corresponds to >2.0 times the standard (i.e., greater than) This requires the activation of emergency plans to avoid irreversible ecological damage. Therefore, this mechanism not only highlights the decision-making weight of high-level events through non-linear amplification of risk intensity, but also effectively suppresses low-risk noise interference at the algorithmic level, significantly improving the detection accuracy and explanatory power of key risk patterns in association rule mining.
[0053] Step S103: Based on the time decay factor, spatial distance factor, and risk intensity factor, the weights of each dimension are dynamically determined using the entropy weight method, and the comprehensive weight of each transaction is calculated as the transaction weighted counting weight.
[0054] Step S104: Input the standardized transactions and their corresponding transaction weighted count weights into an empty weighted frequent pattern tree.
[0055] Before constructing the FP-tree, the spatiotemporally weighted environmental risk data needs to be converted into a standard transaction format. This process not only requires discretization of continuous numerical values but also necessitates the fusion of the causal chain between driving factors and response results through a spatiotemporal alignment mechanism, laying the data foundation for association rule mining. Its core lies in a four-stage collaborative algorithm of "discretization-alignment-assembly-sorting." First, based on the threshold system determined by formulas (2) and (3) above, continuous monitoring values are mapped to discrete item identifiers (ItemID), and the time factor T(t) and spatial factor S(s) are transformed into semantic context labels (such as TIME_RECENT, GEO_NEAR), fully preserving the characteristics of risk metadata. On this basis, the system abandons the traditional time-based packaging strategy and instead uses the abnormal time of the receptor. As a spatiotemporal anchor point, the time lag is estimated based on flow velocity / wind speed. Backtracking and crawling The upstream driving state within the window. This process internalizes the physical transport process into data preprocessing logic, accurately preserving cross-scale causal temporality while avoiding the complexity of hydrodynamic models.
[0056] Furthermore, the transaction structure adopts a set-based representation. In order to adapt to the requirements of the weighted FP-Tree algorithm for flattened input and to clearly distinguish between logical itemsets and numerical weights, the embodiment of this application defines the standard transaction as the following binary tuple structure, as shown in equation (9): (9) in, A flattened itemset is a set of driving factors. Receptor response A list of unique identifiers is generated by taking the union of the three spatiotemporal context labels and flattening them.
[0057] For example: {DRIVER_RAIN_ACC_LEVEL4, RC_FISH_DO_LEVEL4, TIME_RECENT}. These labels are mixed together and will be rearranged uniformly based on global support in subsequent steps to construct the path of the FP-tree. Transaction weight: This is a floating-point value calculated using formula (5) or formula (10) above. It is attached to the transaction as an independent attribute and is only used for the weighted accumulation of node counts (i.e., Node.count += ). ), does not participate in the sorting and topology construction of itemsets.
[0058] The system automatically removes redundant intermediate features and assembles the above identifiers into an itemset portion of a standard transaction: ["DRIVER_RAIN_ACC_LEVEL4", "DRIVER_TEMP_ACC_LEVEL3", "DRIVER_DISCHARGE_LEVEL3", "RC_FISH_DO_LEVEL4", "TIME_RECENT"].
[0059] Subsequently, the system performs a sorting operation based on global support in descending order. Since context labels such as TIME_RECENT or GEO_BASIN_A typically cover a large number of records and have extremely high support, they are automatically sorted to the head of the transaction (e.g., TIME_RECENT→DRIVER_...). This sorting mechanism allows the FP-tree to naturally form a topology guided by the "spatiotemporal context" as the root, maximizing prefix sharing rate (i.e., compression efficiency) and supporting fast slice mining along the spatiotemporal dimension.
[0060] Taking a typical risk event as an example: When the dissolved oxygen at the downstream fish breeding station is abnormal (RC_FISH_DO_LEVEL4), the system traces back the upstream 24-hour window and assembles the accumulated rainfall exceeding 200mm (DRIVER_RAIN_ACC_LEVEL4), the high temperature exceeding 30℃ (DRIVER_TEMP_ACC_LEVEL3), and the sewage discharge event (DRIVER_DISCHARGE_LEVEL3) with the time tag (TIME_RECENT) into a standard transaction: ["DRIVER_RAIN_ACC_LEVEL4", "DRIVER_TEMP_ACC_LEVEL3", "DRIVER_DISCHARGE_LEVEL3", "RC_FISH_DO_LEVEL4", "TIME_RECENT"].
[0061] The resulting data-driven implicit attribution mechanism replaces explicit physical modeling with spatiotemporal alignment logic, transforming complex transmission constraints into data preprocessing rules.
[0062] In step S105, during the sequential insertion of each standardized transaction, for each tree node on the path, the current count value of the tree node is added to the corresponding transaction weighted count weight to update the count value of the tree node, so as to construct a complete weighted frequent pattern tree.
[0063] In this embodiment, a weighted FP-Growth algorithm is used to construct the tree structure. The core of this approach lies in employing the aforementioned weighted counting mechanism, introducing a comprehensive transaction weight to correct the tree node count. Input: Transaction sequence and their corresponding comprehensive weights Process: When inserting a node along the path, the node count is no longer simply incremented, but rather the weight of the transaction is accumulated, as shown in the following formula (14): (14) This mechanism preserves both the categorical characteristics of risk indicators (identified by node IDs) and the timeliness and intensity characteristics of risks at the statistical level (retained by floating-point values of node counts).
[0064] Specifically, constructing a complete weighted frequent pattern tree may include: first scanning the weighted transaction database to count the weighted support of each risk item; constructing an item header table based on the weighted support and sorting it in descending order of weighted support; second scanning the weighted transaction database to rearrange the items within each transaction according to the priority order of the item header table; and inserting the rearranged transactions into the weighted frequent pattern tree.
[0065] Among them, for any risk indicator Its weighted support is defined as the sum of the weights of all transactions that include the indicator, as shown in the following formula (15): (15) in, This represents the total number of transactions in the transaction database. Indicates the first Transaction record; For logical indicator functions, when the condition Establishment (i.e., indicator) Existing in transactions The value is 1 when the condition is met, and 0 otherwise. For the first The overall weight of each transaction (i.e., the weight calculated above) The relative weighted support is then normalized, as shown in the following formula (16): (16) The resulting logical consistency mechanism ensures that the screening of frequent items takes into account both the frequency of occurrence and the risk contribution. For example, a single extremely high-risk event ( In terms of statistical power, it is equivalent to 5 low-risk events. This effectively avoids the omission of high-risk sparse patterns.
[0066] Based on this, a dynamic threshold strategy is adopted when constructing the item header table. The theoretical basis is that the inherent sparsity of environmental risk events requires lowering the support threshold to retain key patterns. Accordingly, this framework sets thresholds based on the characteristics of the medium: 0.05 for aquatic environment (high mobility leads to pattern dispersion), 0.03 for atmospheric environment (high volatility), 0.04 for ecological environment (moderate complexity), and 0.08 for soil environment (low mobility).
[0067] DRIVER_RAIN_ACC_LEVEL4 (weighted support 0.30) is preferentially selected as the root node, while RC_SOIL_PH_LEVEL1 (0.06) is located as a leaf node. The node linked list structure (e.g., Node1→Node4→Node7) explicitly maintains the topological association of high-frequency risk patterns. Finally, through a global support descending order sorting mechanism, high-frequency basic items (such as season and geographical basis) are placed at the front of the transaction sequence.
[0068] Table 2 below shows an example of the weighted support statistics for environmental risk indicators and the construction of the FP tree header table: Table 2:
[0069] Among them, the system calculates the global weighted support based on the first scan ( The item header table is constructed in descending order, establishing the global priority of all risk indicators. Before inserting any transaction (such as T={RC_FISH_DO_LEVEL4,DRIVER_RAIN_ACC_LEVEL4,TIME_RECENT}) into the FP-tree, the items within the transaction must be rearranged according to the priority order of the item header table.
[0070] Assume the priority is: If TIME_RECENT > DRIVER_RAIN_ACC_LEVEL4 > RC_FISH_DO_LEVEL4, then the transaction is reorganized into an ordered sequence. T'=[TIME_RECENT,DRIVER_RAIN_ACC_LEVEL4,RC_FISH_DO_LEVEL4].
[0071] This step is crucial: it ensures that high-frequency items (such as spatiotemporal labels and primary driving factors) are always located at the root or shallow level of the tree, thus forcing different transactions to share the same path prefix.
[0072] Among them, the FP-tree construction process dynamically achieves three optimizations based on the rearranged ordered sequence: Path sharing compression: When inserting an ordered sequence T', the algorithm prioritizes searching for the shared prefix starting from the root node. If the path Root→TIME_RECENT→DRIVER_RAIN_ACC_LEVEL4 already exists in the tree, the system does not create a new node, but directly applies the transaction's comprehensive weight. The increments are added to the counter of each node on the existing path (i.e. A new child node is created only when the path branches at a point (such as RC_FISH_DO_LEVEL4). This mechanism compresses massive amounts of redundant data into a compact tree topology.
[0073] Node-linked list acceleration: The item header table maintains linked list pointers to items with the same risk level, reducing the query complexity of conditional patterns for specific risk items from... Down to ( (where the length is the linked list length), significantly improving the efficiency of high-frequency traversal.
[0074] Furthermore, the association rules mined in the embodiments of this application strictly follow the standard paradigm of driver factor set → receptor response, which logically clearly defines the causal chain between environmental drivers and responses. Specifically, the antecedent of a rule is characterized as a combination of environmental driving factors, such as the synergistic effect of upstream continuous cumulative rainfall exceeding 200mm (DRIVER_RAIN_ACC_LEVEL4) and temperature above 30℃ (DRIVER_TEMP_ACC_LEVEL3). In the physical structure of the FP tree, such "parallel combination relationships" are not presented as parallel directed arrows, but are mapped as a vertical sequence path of parent node → child node (e.g., Root → spatiotemporal context → rainfall node → high temperature node).
[0075] Furthermore, node positions are optimized using a descending support ranking strategy: frequently occurring common driving factors are typically located in internal nodes near the root, thus playing a core role in path guidance. Based on this, the consequent of a rule corresponds to the receptor response state, such as extremely low dissolved oxygen at the cross-section (RC_FISH_DO_LEVEL4). In the FP-tree, such responses are all located at the end of the aforementioned driving path (i.e., leaf nodes or deep nodes), serving as the ultimate environmental result led by that path. It is important to note that NodeLinks in this structure only act as a horizontal indexing mechanism, connecting nodes with the same name in different branches to accelerate algorithm computation; they themselves do not carry any causal logical meaning. In summary, this embodiment, through the above structured design, ensures that all mining rules conform to the rigorous academic norms of driving factor set → receptor response in both form and content.
[0076] Step S106: Perform association rule mining based on the weighted frequent pattern tree and output weighted association rules.
[0077] Furthermore, the method may also include an incremental update step: when new monitoring data flows in, the data growth rate is calculated; if the data growth rate is less than a preset growth rate threshold, the new monitoring data is converted into a new transaction and directly inserted into the existing frequent pattern tree; if the data growth rate is greater than or equal to the growth rate threshold or the number of tree nodes exceeds a preset threshold, the weighted frequent pattern tree is reconstructed in a background thread based on all historical data and new monitoring data, and switched via atomic pointers.
[0078] In response to the continuously growing nature of environmental monitoring data, this application proposes an incremental update framework based on a dynamic growth rate threshold to overcome the efficiency bottleneck of traditional full recalculation. Its core lies in first defining a data growth rate indicator. To quantify the scale of the new data, see formula (17): (17) in, For the amount of new transactions within the update cycle, This represents the total number of historical transactions. Based on this, the system abandons the static update mode and instead adopts a dual-track adaptive strategy: when... (In daily real-time monitoring scenarios), activate the append-based lightweight update mechanism. This mechanism directly embeds new transactions into the existing FP-tree: if the new path (e.g., A→B→C) matches the existing structural prefix, the transaction weight of the corresponding node is accumulated. Create a new branch only at the fork point. Conversely, when (Quarterly archive / history import) or the number of tree nodes exceeds the preset balance threshold When the system triggers a full, refactoring update mechanism, it merges all historical and new data in a separate background thread, recalculates weighted support, and optimizes the item header table sorting. Once the new tree is built, seamless replacement is achieved through atomic pointer switching, instantly releasing resources from the old structure. The key value of this design lies in avoiding tree imbalance caused by long-term appending while maintaining globally optimal mining efficiency. A 20% threshold is used as a baseline reference value, and in practical applications, it is adaptively adjusted according to the environmental medium type.
[0079] The weighted counting-based association analysis method provided in this invention can acquire environmental monitoring data, preprocess the data to generate standardized transactions, determine the time decay factor, spatial distance factor, and risk intensity factor for each risk item in the standardized transactions, dynamically determine the weights of each dimension based on the time decay factor, spatial distance factor, and risk intensity factor using the entropy weight method, and calculate the comprehensive weight of each transaction as the transaction weighted counting weight. The standardized transactions and their corresponding transaction weighted counting weights are input into an empty weighted frequent pattern tree. During the sequential insertion of each standardized transaction, for each tree node on the path, the current count value of the tree node is added to the corresponding transaction weighted counting weight to update the tree node's count value, thus constructing a complete weighted frequent pattern tree. Association rule mining is then performed based on the weighted frequent pattern tree to output weighted association rules. This method, by introducing a three-dimensional weighted counting mechanism of time, space, and risk intensity, effectively filters out false strong rules and improves the accuracy of risk identification.
[0080] Example 2 This invention also provides another association analysis method based on weighted counting; this method is implemented on the basis of the method in the above embodiments; this method focuses on describing the specific implementation of dynamically determining the weights of each dimension based on the time decay factor, spatial distance factor and risk intensity factor through the entropy weight method, and calculating the comprehensive weight of each transaction as the transaction weighted counting weight.
[0081] Figure 2A flowchart of another association analysis method based on weighted counting provided in an embodiment of the present invention is shown below. Figure 2 As shown, this method, which dynamically determines the weights of each dimension based on the time decay factor, spatial distance factor, and risk intensity factor using the entropy weight method, and calculates the comprehensive weight of each transaction as the transaction weighted counting weight, may include the following steps: Step S201: The entropy weight method is used to calculate the basic objective weight vector of multiple dimensions.
[0082] Addressing the core challenge of unifying the quantification of multi-source heterogeneous environmental data (including historical monitoring data and predictive simulation data) in the construction of frequent-pattern FP-trees, a general transaction weight calculation model is proposed. This model introduces a dynamic confidence decay mechanism to achieve adaptive weighting of the confidence of different data sources. Specifically, for any transaction... Update count weight Defined as follows (5): (5) in, For comprehensive risk scoring, To predict the confidence decay coefficient. Its core mechanism lies in: assigning a certain value to historical / real-time monitoring data (hard data). The full confidence weight, where the transaction weight is determined solely by the risk score. The decision is to ensure that highly reliable data dominates pattern mining; for intelligent branch prediction data (soft data), the following settings are made. The dynamic decay range is determined by a coefficient based on the quantitative assessment of the uncertainty of the prediction model, ensuring that the weight of the predicted data is strictly lower than that of the measured data (1.0). As a result, low-confidence prediction results will automatically reduce their transaction weight through the multiplier effect, thereby significantly weakening their contribution to the support count during the FP-tree construction process.
[0083] The theoretical value of this design lies in its ability to retain the supplementary value of predictive data for risk forecasting through a confidence-weight coupling mechanism, while effectively avoiding the interference of uncertain data on association rule mining. (Predictive confidence coefficients of historical / real-time monitoring data) The value is always 1.0.
[0084] Among them, dynamic coding can be carried out by comprehensively considering spatiotemporal characteristics and risk intensity, as shown in the following formula (6): (6) Among them, weight parameters Based on the dynamic determination of entropy weight, it can adaptively adjust according to historical risk data, ensuring that the coding results are more objective and accurate.
[0085] Time decay factor: This reflects the timeliness of the data. This is the attenuation coefficient. For example, sudden risk... (Half-life 7 days), ongoing risk (Half-life 23 days).
[0086] Spatial distance factor: This reflects the distance attenuation effect. The impact thresholds are set (500m for aquatic environment and 2000m for atmospheric environment). The attenuation index is 1.5-2.0 for aquatic environments and 0.5-1.0 for atmospheric environments.
[0087] Intensity factor: A non-linear mapping based on risk level. (Using...) The numerical set reflects the response priorities for different risk levels.
[0088] Table 1 below is a risk indicator discretization and intensity mapping table (based on a unified five-level classification system).
[0089] Table 1:
[0090] Note: In the architectural design of this application, the functional boundaries of topology construction and weight calculation are strictly distinguished. The discrete identifier (ItemID) is specifically used to construct the topology node structure of the FP-tree; while the strength factor... It only participates in the calculation of the comprehensive risk score in the above formula (6), thereby affecting the transaction weight in the above formula (5). This separation design not only ensures the simplicity of the FP tree structure, but also ensures that the risk intensity information can be accurately quantified in the mining process.
[0091] Step S202: Based on the type of the current environmental medium, obtain the preset medium feature correction coefficient vector.
[0092] To address the limitations of single weighting methods in simultaneously considering both statistical data patterns and environmental physical characteristics, this application constructs a three-level dynamic correction algorithm for attribute weights to generate a scenario-based weight vector that ultimately participates in risk scoring calculations. Its core design follows the principle of incremental optimization: First, based on the entropy weighting method, the information entropy of historical samples in the three dimensions of time, space, and intensity is calculated to generate a basic objective weight vector. As shown in formula (6) above, this vector exhibits an equilibrium distribution in the absence of prior knowledge. This provides a statistical benchmark for subsequent corrections. However, the basic weights fail to reflect the physical transport characteristics of the environmental medium. To address this limitation, a medium characteristic correction mechanism is introduced in the second stage: a preset environmental medium correction coefficient matrix is used. This implicitly incorporates physical laws such as hydrological diffusion and atmospheric transport into the weight allocation. For example, due to the highly mobile nature of the water environment, the spatial weight coefficient is increased to 1.3. This automatically suppresses interference in non-hydraulic connected areas; the atmospheric environment, due to its strong time sensitivity, has a time weighting coefficient set to 1.4. The system prioritizes capturing meteorologically driven pollution transport paths. Therefore, it can dynamically select driving factors based on physical accessibility. For example, in river monitoring, when a water conservancy project blocks the hydraulic connection between the upstream pollution source and the receptor, its spatial weight will be significantly reduced, ensuring the algorithm focuses on real and effective causal chains. Furthermore, a third-level risk situation dynamic optimization mechanism (…) In real-time data mining, the system responds to changes in risk type: when identifying highly toxic emergencies, the intensity dimension weighting coefficient is automatically increased to 1.5. This is to prevent high-risk signals from being drowned out by distance attenuation.
[0093] Step S203: Multiply the basic objective weight vector and the medium characteristic correction coefficient vector element by element to obtain the corrected weight vector.
[0094] Step S204: Normalize the modified weight vector to obtain the scenario-based weight vector.
[0095] The three-level weights are fused through element-wise multiplication, as shown in formula (7) below, and then normalized to generate a scenario-specific weight vector, as shown in formula (8) below: (7) (8) Where ⊙ represents the Hadamard product, which is the element-wise multiplication of vectors; The three components of the vector correspond to the formula (6) above. .
[0096] Step S205: Use the scenario-based weight vector to perform a weighted summation of the time decay factor, spatial distance factor, and risk intensity factor to obtain the comprehensive weight.
[0097] Step S206: Use the overall weight as the transaction weighted counting weight.
[0098] To ensure the reliability boundary of the predicted data, the system is designed with a dual weighting control mechanism: First, a confidence decay coefficient is introduced for virtual transactions. Its value is dynamically calibrated by the historical accuracy of the physical model; secondly, the transaction weight is defined by the following formula (10). : (10) in, To predict the confidence decay coefficient, unlike formula (5), the virtual transaction... It needs to be dynamically calculated based on the spatiotemporal decay model (see Formula 11 below), while historical / real-time monitoring data The value is always 1.0. This ensures that the contribution of the predicted data to the support in the FP-tree construction is strictly lower than that of the measured data.
[0099] The system does not consider specific numerical deviations, but only calculates the theoretical confidence decay based on the prediction step size, as shown in the following formula (11): (11) in, This is the baseline confidence level for the model (usually taken as 0.9). To predict the time span, To predict the divergence coefficient. It is important to distinguish that... This describes the process by which the accuracy of a numerical model drops sharply as the prediction step size increases (the "butterfly effect"), and its value is usually much larger than the time decay coefficient of historical data. ; This represents the relative error from the most recent closed-loop verification. The resulting dynamic decay mechanism accurately characterizes the exponential growth of prediction uncertainty over time.
[0100] Formula (11) above only addresses the timeliness issue and cannot eliminate the inherent systematic biases of the physical model in local environments (e.g., always overestimating pollutant diffusion under specific terrain). Therefore, error classification is required: Calculate historical moments Predicted value Compared with measured values relative error ,when If the model is found to have a significant systematic bias, a correction procedure is triggered.
[0101] This also requires dynamic numerical correction: adjusting the predicted values themselves. Based on the most recent... Moving average of error over a time window Using symbolic functions Determine the direction of the deviation and the original predicted value. Reverse compensation is performed using the following formula (12): (12) Here, For the sign function: when historical predicted values are too high ( When the value is positive, it is +1, causing the correction term to decrease the predicted value; conversely, it is -1, causing the predicted value to increase. The corrected value is... The original value will be replaced by the discretization encoding of the above formula (4).
[0102] This requires a confidence level closed-loop update: while numerical correction reduces bias, it introduces new correction uncertainties. Therefore, the system will update the corrected residuals... Feedback is sent to the weighting model to make a final correction to the initial confidence level calculated by formula (11), as shown in formula (13) below: (13) It should be noted that this final attenuation coefficient This is the final coefficient that is substituted into formula (10) to calculate the transaction weight.
[0103] The weighted counting-based association analysis method provided in this invention constructs a closed-loop strategy of error identification, numerical correction, and weight update: First, the relative error between the predicted and measured values is calculated in real time, and correction is triggered when the error exceeds a preset threshold; second, dynamic numerical correction is performed, using the sign function direction of historical errors to compensate the original predicted values in reverse, eliminating systematic overestimation or underestimation; finally, a closed-loop confidence update is performed, feeding the corrected residuals back to the weight model to further reduce the confidence coefficient of the corrected data. This mechanism ensures that even if the physical model has biases, the virtual transactions input into the FP-Tree still possess statistical reliability, effectively preventing erroneous predictions from misleading the mining results.
[0104] Example 3 This invention also provides another association analysis method based on weighted counting; this method is implemented based on the method in the above embodiments.
[0105] Figure 3 A flowchart of another association analysis method based on weighted counting provided in an embodiment of the present invention is shown below. Figure 3 As shown, the method may also include the following multidimensional rule evaluation steps: Step S301: Calculate the weighted support, weighted confidence, weighted lift, and weighted certainty of the mined association rules.
[0106] It should be noted that, to maintain consistency in the algorithm logic, all the following metric calculations are based on the weighted count (Sum of Weights) defined in PHASE III, rather than the traditional frequency count: Weighted Support: Measures the prevalence (i.e., frequency) of a rule. It reflects the proportion of a specific environmental risk pattern in the historical dataset, as shown in the following formula (18). (18) The numerator is the sum of the weights of all transactions that simultaneously include the predecessor X and the successor Y, and the denominator is the sum of the weights of all transactions in the database.
[0107] Weighted confidence reflects the reliability of a rule (i.e., conditional probability). It is characterized by the driving factors. Under the condition that the (antecedent) occurs, the receptor response The probability of the (subsequent) event occurring.
[0108] (19) Weighted Lift: Reflects the relevance of a rule. Used to determine whether a rule is better than random guessing. When... Time indicates and A positive correlation is observed, with larger values indicating a stronger causal relationship; if This indicates that the two are independent (i.e., pseudo-association) and should be eliminated, as shown in the following formula (20): (20) Weighted Conviction: As a quality auxiliary verification metric, it measures the independence between the antecedent and consequent of a rule, and is used to exclude rules that are ruled out due to the consequent itself occurring with an extremely high frequency (i.e., The spurious strong association caused by the large number of large numbers is shown in the following formula (21): (twenty one) Step S302: Normalize the weighted support, weighted confidence, weighted lift, and weighted certainty.
[0109] Furthermore, to overcome the inherent limitations of single-indicator assessment in environmental risk analysis, this application constructs a multi-dimensional rule scoring system based on the Analytic Hierarchy Process (AHP). This system scientifically weights and integrates four core indicators—support, confidence, lift, and certainty—to generate a comprehensive score, thereby achieving a systematic and quantitative assessment of the quality of associated rules.
[0110] Within this framework, due to the range of values for lift and confidence (Conv) ) and support / confidence ( Inconsistency can lead to scaling bias if directly weighted. Therefore, in this embodiment, before substituting the two indicators into the scoring formula, the arctangent function is used to normalize and map them, compressing them to... Interval, define the normalization operator The following formulas (22) and (23) are: (twenty two) (twenty three) Weight vector The settings are strictly based on the physical meaning and priority of each indicator in environmental decision-making: among which, the support weight ( =0.539) had the highest proportion, aiming to ensure that the rules were statistically significant and effectively exclude spurious associations caused by random perturbations; confidence weight ( =0.297) Secondly, the focus is on ensuring the accuracy of the rule's prediction of the receptor response; the degree weight ( =0.119) is used to identify and remove invalid strong associations (i.e., independent events with a Lift value close to 1); confidence weight ( =0.045) serves as an auxiliary verification indicator to further strengthen the evaluation of the rule's independence.
[0111] Step S303: Based on the preset AHP weight vector, the normalized weighted support, weighted confidence, weighted lift, and weighted certainty are summed to generate a rule-based comprehensive score.
[0112] Based on this weight allocation, any rule The comprehensive score calculation formula can be expressed as the following formula (24): (twenty four) To ensure the reliability and adaptability of the rules, the system first sets a global minimum weighted confidence benchmark. As a basic screening criterion, and considering the significant differences in risk sensitivity and data characteristics of different environmental media (such as the high immediate responsiveness of the water environment to pollution events, while the soil environment has a cumulative lag effect), this system introduces a differentiated confidence threshold mechanism, as shown in Table 3 below. Table 3 is the differentiated confidence threshold table.
[0113] For example, in water environment risk management, given the high sensitivity of drinking water safety, the confidence threshold is strictly set at 0.70; while in the soil environment, due to the cumulative effect of pollution, a certain buffer space is allowed, and the threshold is correspondingly relaxed to 0.55. To intuitively illustrate the operation process of this mechanism, let's take water environment risk mining as an example: Assume that in the weighted transaction set, the total weight contribution of "upstream pollution exceeding standards" (the antecedent, corresponding to DRIVER_DISCHARGE_LEVEL3) is... The sum of the transaction weights that co-occur with "DO (Dissolved Oxygen) Anomaly" (the consequent, corresponding to RC_FISH_DO_LEVEL4) is: The weighted confidence level is then calculated as follows: Because this value reaches the water environment threshold ( The system determines that this rule is a valid strong rule and retains it.
[0114] In summary, this scoring system, through hierarchical weight design and contextualized threshold adaptation, not only overcomes the one-sidedness of single-indicator evaluation, but also realizes a logical closed loop from statistical verification to environmental decision-making, providing a scientific basis for risk warning of complex environmental systems.
[0115] Table 3:
[0116] Step S304: Based on the preset differential confidence thresholds for different environmental media, filter and output association rules whose comprehensive rule score is higher than the preset score value.
[0117] Example 4 Corresponding to the above method embodiments, this invention provides a correlation analysis device based on weighted counting. Figure 4 This is a schematic diagram of a correlation analysis device based on weighted counting, provided in an embodiment of the present invention. Figure 4 As shown, the association analysis device based on weighted counting may include: The standardized transaction generation module 401 is used to acquire environmental monitoring data, preprocess the environmental monitoring data, and generate standardized transactions.
[0118] Module 402 is used to determine the time decay factor, spatial distance factor, and risk intensity factor for each risk item in a standardized transaction.
[0119] The transaction weighted counting weight calculation module 403 is used to dynamically determine the weight of each dimension based on the time decay factor, spatial distance factor and risk intensity factor through the entropy weight method, and calculate the comprehensive weight of each transaction as the transaction weighted counting weight.
[0120] The data input module 404 is used to input standardized transactions and their corresponding transaction weighted count weights into an empty weighted frequent pattern tree.
[0121] The weighted frequent pattern tree construction module 405 is used to add the corresponding transaction weighted count weight to the current count value of each tree node on the path during the sequential insertion of each standardized transaction, so as to update the count value of the tree node and construct a complete weighted frequent pattern tree.
[0122] The weighted association rule output module 406 is used to mine association rules based on a weighted frequent pattern tree and output weighted association rules.
[0123] The weighted counting-based association analysis device provided in this invention can acquire environmental monitoring data, preprocess the data to generate standardized transactions, determine the time decay factor, spatial distance factor, and risk intensity factor for each risk item in the standardized transactions, dynamically determine the weights of each dimension based on the time decay factor, spatial distance factor, and risk intensity factor using the entropy weight method, and calculate the comprehensive weight of each transaction as the transaction weighted counting weight. The standardized transactions and their corresponding transaction weighted counting weights are input into an empty weighted frequent pattern tree. During the sequential insertion of each standardized transaction, for each tree node on the path, the current count value of the tree node is added to the corresponding transaction weighted counting weight to update the tree node's count value, thus constructing a complete weighted frequent pattern tree. Association rule mining is performed based on the weighted frequent pattern tree, and weighted association rules are output. This method, by introducing a three-dimensional weighted counting mechanism of time, space, and risk intensity, effectively filters out false strong rules and improves the accuracy of risk identification.
[0124] In some embodiments, the transaction weighted counting weight calculation module is further configured to calculate a multi-dimensional basic objective weight vector using the entropy weight method; obtain a preset medium characteristic correction coefficient vector based on the type of the current environmental medium; multiply the basic objective weight vector and the medium characteristic correction coefficient vector element by element to obtain a corrected weight vector; normalize the corrected weight vector to obtain a scenario-based weight vector; use the scenario-based weight vector to perform a weighted summation of the time decay factor, spatial distance factor, and risk intensity factor to obtain a comprehensive weight; and use the comprehensive weight as the transaction weighted counting weight.
[0125] In some embodiments, the standardized transaction generation module is also used to perform data cleaning on environmental monitoring data, remove outliers, and fill in missing values using interpolation; unify environmental monitoring data from different monitoring sources into the smallest common time granularity, and establish a unique monitoring point identifier and standardized timestamp; determine the unidirectional influence relationship between monitoring points based on the physical topology matrix, and determine the optimal time lag through cross-correlation analysis; extract driving factors by backtracking from the abnormal moment of the response factor, and reorganize the driving factors and response factors into logically synchronized transaction units to generate standardized transactions.
[0126] In some embodiments, the weighted frequent pattern tree construction module is further used to first scan the weighted transaction database, count the weighted support of each risk item; construct an item header table based on the weighted support and sort it in descending order of weighted support; secondly scan the weighted transaction database, rearrange the items in each transaction according to the priority order of the item header table; and insert the rearranged transaction into the weighted frequent pattern tree.
[0127] In some embodiments, the weighted association rule output module is further configured to calculate the weighted support, weighted confidence, weighted lift, and weighted certainty of the mined association rules; normalize the weighted support, weighted confidence, weighted lift, and weighted certainty; based on a preset AHP weight vector, perform a weighted summation of the normalized weighted support, weighted confidence, weighted lift, and weighted certainty to generate a rule comprehensive score; and based on a preset differential confidence threshold for different environmental media, filter and output association rules whose rule comprehensive scores are higher than the preset score value.
[0128] In some embodiments, the determining module is further configured to determine a statistical threshold based on statistical distribution characteristics and a regulatory threshold based on environmental regulatory standards; for positive risk indicators, the minimum value between the statistical threshold and the regulatory threshold is taken; for negative risk indicators, the maximum value between the statistical threshold and the regulatory threshold is taken; the three-level risk level boundary points are set based on the positive risk indicators and the negative risk indicators; and the monitoring values of continuous indicators are mapped to multiple discrete risk levels.
[0129] In some embodiments, the weighted association rule output module is also used to calculate the data growth rate when new monitoring data flows in; if the data growth rate is less than the preset growth rate threshold, the new monitoring data is converted into a new transaction and directly inserted into the existing frequent pattern tree; if the data growth rate is greater than or equal to the growth rate threshold or the number of tree nodes exceeds the preset threshold, the weighted frequent pattern tree is reconstructed in the background thread based on all historical data and new monitoring data, and switched by atomic pointers.
[0130] The device provided in this embodiment of the invention has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0131] Example 5 This invention also provides an electronic device for running the above-described association analysis method based on weighted counting; see [link to other documentation]. Figure 5 The diagram shows the structure of an electronic device, which includes a memory 500 and a processor 501. The memory 500 stores one or more computer instructions, which are executed by the processor 501 to implement the aforementioned association analysis method based on weighted counting.
[0132] Furthermore, Figure 5 The electronic device shown also includes a bus 502 and a communication interface 503. The processor 501, the communication interface 503 and the memory 500 are connected via the bus 502.
[0133] The memory 500 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 503 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 502 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0134] Processor 501 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 501 or by instructions in software form. Processor 501 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 500, and processor 501 reads information from memory 500 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.
[0135] This invention also provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are called and executed by a processor, they cause the processor to implement the aforementioned association analysis method based on weighted counting. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0136] The computer program product for performing a weighted counting-based association analysis method provided in this embodiment of the invention includes a computer-readable storage medium storing processor-executable non-volatile program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0137] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0138] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0139] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0140] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0141] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0142] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A correlation analysis method based on weighted counting, characterized in that, The method includes: Acquire environmental monitoring data and preprocess the environmental monitoring data to generate standardized transactions; Determine the time decay factor, spatial distance factor, and risk intensity factor for each risk item in the standardized transaction; Based on the time decay factor, the spatial distance factor, and the risk intensity factor, the weights of each dimension are dynamically determined using the entropy weight method, and the comprehensive weight of each transaction is calculated as the transaction weighted counting weight. Input the standardized transaction and its corresponding transaction weighted count weight into an empty weighted frequent pattern tree; During the process of sequentially inserting each standardized transaction, for each tree node on the path, the current count value of the tree node is added to the corresponding transaction weighted count weight to update the count value of the tree node, so as to construct a complete weighted frequent pattern tree. Based on the weighted frequent pattern tree, association rule mining is performed, and weighted association rules are output.
2. The method according to claim 1, characterized in that, The process of dynamically determining the weights of each dimension based on the time decay factor, the spatial distance factor, and the risk intensity factor using the entropy weight method, and calculating the comprehensive weight of each transaction as the transaction weighted counting weight, includes: The entropy weight method is used to calculate the basic objective weight vectors of multiple dimensions. Based on the type of the current environmental medium, obtain the preset medium feature correction coefficient vector; The modified weight vector is obtained by multiplying the basic objective weight vector element by element with the medium feature correction coefficient vector. The modified weight vector is normalized to obtain a scenario-based weight vector; The time decay factor, spatial distance factor, and risk intensity factor are weighted and summed using the scenario-based weight vector to obtain the comprehensive weight. The overall weight is used as the weight for the transaction weighted count.
3. The method according to claim 1, characterized in that, The preprocessing of the environmental monitoring data to generate standardized transactions includes: The environmental monitoring data is cleaned to remove outliers, and interpolation is used to fill in missing values. Environmental monitoring data from different monitoring sources are unified into the smallest common time granularity, and a unique monitoring point identifier and standardized timestamp are established. The unidirectional influence relationship between monitoring points is determined based on the physical topology matrix, and the optimal time lag is determined through cross-correlation analysis. The driving factor is extracted by backtracking based on the abnormal moment of the response factor, and the driving factor and the response factor are recombined into a logically synchronized transaction unit to generate the standardized transaction.
4. The method according to claim 1, characterized in that, The construction of the complete weighted frequent pattern tree includes: The first scan of the weighted transaction database is performed to calculate the weighted support for each risk item; Construct an item header table based on the weighted support, and sort it in descending order of weighted support; The second scan of the weighted transaction database rearranges the items within each transaction according to the priority order of the item header table. Insert the rearranged transactions into the weighted frequent pattern tree.
5. The method according to claim 1, characterized in that, The method also includes a multidimensional rule evaluation step: Calculate the weighted support, weighted confidence, weighted lift, and weighted certainty of the discovered association rules; The weighted support, weighted confidence, weighted lift, and weighted certainty are normalized. Based on the preset AHP weight vector, the normalized weighted support, weighted confidence, weighted lift and weighted certainty are summed to generate a comprehensive rule score. Based on the preset differential confidence thresholds for different environmental media, the association rules with a comprehensive rule score higher than the preset score value are selected and output.
6. The method according to claim 1, characterized in that, The method also includes a dynamic threshold determination step: Statistical thresholds are determined based on statistical distribution characteristics, and regulatory thresholds are determined based on environmental regulations and standards. For positive risk indicators, the minimum value between the statistical threshold and the regulatory threshold is taken; for negative risk indicators, the maximum value between the statistical threshold and the regulatory threshold is taken. Three-level risk level dividing points are set based on the positive risk indicators and the negative risk indicators; The monitoring values of continuous indicators are mapped to multiple discrete risk levels.
7. The method according to claim 1, characterized in that, The method also includes an incremental update step: When new monitoring data is received, calculate the data growth rate. If the data growth rate is less than the preset growth rate threshold, the newly added monitoring data will be converted into a new transaction and directly inserted into the existing frequent pattern tree. If the data growth rate is greater than or equal to the growth rate threshold or the number of tree nodes exceeds the preset threshold, then based on all historical data and newly added monitoring data, the weighted frequent pattern tree is reconstructed in the background thread and switched through atomic pointers.
8. A correlation analysis device based on weighted counting, characterized in that, The device includes: The standardized transaction generation module is used to acquire environmental monitoring data, preprocess the environmental monitoring data, and generate standardized transactions. The determination module is used to determine the time decay factor, spatial distance factor, and risk intensity factor for each risk item in the standardized transaction; The transaction weighted counting weight calculation module is used to dynamically determine the weight of each dimension based on the time decay factor, the spatial distance factor and the risk intensity factor through the entropy weight method, and calculate the comprehensive weight of each transaction as the transaction weighted counting weight. The data input module is used to input the standardized transactions and their corresponding transaction weighted count weights into an empty weighted frequent pattern tree; The weighted frequent pattern tree construction module is used to add the corresponding transaction weighted count weight to the current count value of each tree node on the path during the sequential insertion of each standardized transaction, so as to update the count value of the tree node and construct a complete weighted frequent pattern tree. The weighted association rule output module is used to mine association rules based on the weighted frequent pattern tree and output weighted association rules.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the association analysis method based on weighted counting as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the association analysis method based on weighted counting as described in any one of claims 1 to 7.