Carbon lock evaluation method based on multi-source data projection pursuit and frequent item set identification

Through the methods of multi-source data projection pursuit and frequent item set identification, the problem of identifying carbon lock-in in household energy consumption behavior was solved, nonlinear dimensionality reduction and efficient factor extraction were achieved, the causal path of carbon lock-in was revealed, and the accuracy and systematicness of carbon lock-in assessment were improved.

CN120634038APending Publication Date: 2025-09-12SOUTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510770646.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing carbon emission analysis methods have difficulty revealing the carbon lock-in mechanism in household energy consumption behavior when faced with high-dimensional, multi-source, and heterogeneous data. They lack the ability to model nonlinear coupling relationships, are difficult to adapt to dynamic changes, and lack a systematic identification mechanism. It is difficult to efficiently explore carbon-intensive behavior combinations and construct highly credible behavioral mechanism paths.

Method used

The method of multi-source data projection pursuit and frequent item set identification is adopted. By constructing a household energy consumption factor database, nonlinear dimensionality reduction processing is performed, and the real number coded accelerated genetic algorithm is combined to optimize the projection direction, extract frequent item sets, set support and credibility thresholds, and combine the DNAS behavioral mechanism theory to identify the causal path and mechanism of carbon lock-in.

Benefits of technology

It has achieved accurate assessment of household carbon lock-in, revealed the hidden carbon lock-in path, improved identification efficiency and mechanism analysis capabilities, and provided a scientific basis for energy behavior management intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634038A_ABST
    Figure CN120634038A_ABST
Patent Text Reader

Abstract

The invention discloses a carbon lock evaluation method based on multi-source data projection pursuit and frequent item set identification. The method comprises the following steps: S1, constructing a household energy consumption factor database comprising a plurality of dimensions; s2, performing nonlinear dimension reduction processing on the data by using a projection pursuit model, and optimizing a projection direction in combination with a real number coding acceleration genetic algorithm; s3, based on the factor vector after dimension reduction, extracting a frequent item set, setting support degree and credibility thresholds, and identifying a carbon locking behavior path; and S4, in combination with a DNAS behavior mechanism theory, identifying a causal path and mechanism of carbon locking. The method can integrate multi-source data, has nonlinear dimension reduction and efficient factor extraction capabilities, and realizes accurate and effective evaluation of carbon locking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of carbon emission analysis, and in particular to a carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification. Background Art

[0002] Currently, in the context of climate change, household energy consumption, as a significant source of household carbon emissions, is receiving increasing research attention. Traditional carbon emission analysis methods often rely on industry-specific statistics or static feature modeling, which has limited ability to identify carbon lock-in mechanisms in household energy consumption at the micro level. Existing methods, particularly when dealing with high-dimensional, multi-source, and heterogeneous data, suffer from the following shortcomings:

[0003] First, there is a lack of modeling capabilities for the nonlinear coupling relationships between multidimensional influencing factors, making it difficult to reveal the true behavioral mechanisms of carbon lock-in. Second, carbon lock-in assessment methods generally use simplified tools such as linear regression or cluster analysis, which are difficult to adapt to the dynamic changes in complex household behavioral data. Third, there is a lack of a systematic identification mechanism, making it difficult to integrate and identify key carbon lock-in pathways from multiple dimensions, including technology. Furthermore, in the carbon lock-in identification process, how to efficiently mine behavioral types, how to extract frequently occurring carbon-intensive behavior combinations, and how to construct highly credible and explanatory behavioral mechanism pathways remain long-term technical bottlenecks that have not been resolved by existing technologies.

[0004] Therefore, to solve the above problems, a carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification is needed, which can integrate multi-source data, have nonlinear dimensionality reduction and efficient factor extraction capabilities, and realize accurate and effective assessment of carbon lock-in. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to overcome the defects in the prior art and provide a carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification, which can integrate multi-source data, has nonlinear dimensionality reduction and efficient factor extraction capabilities, and realizes accurate and effective assessment of carbon lock-in.

[0006] The carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification of the present invention comprises the following steps:

[0007] S1. Construct a household energy consumption factor database with several dimensions;

[0008] S2. Use the projection pursuit model to perform nonlinear dimensionality reduction on the data and optimize the projection direction by combining it with a real-coded accelerated genetic algorithm.

[0009] S3. Based on the factor vector after dimensionality reduction, extract frequent itemsets, set support and credibility thresholds, and identify carbon lock-in behavior paths;

[0010] S4. Combine DNAS behavioral mechanism theory to identify the causal pathways and mechanisms of carbon lock-in.

[0011] Furthermore, the step S1 specifically includes:

[0012] Integrate data sources to obtain original data;

[0013] The collected raw data are divided into four dimensions according to factor attributes: environmental dimension, social dimension, technical dimension and strategic dimension;

[0014] Clean, normalize, fill missing values ​​and integrate subjective and objective variables on the original data;

[0015] Establish a unified data table structure, where each row corresponds to a family sample and each column is a specific factor variable;

[0016] The database is designed as a dynamic structure that supports periodic updates and multi-source interface access.

[0017] Furthermore, the step S2 specifically includes:

[0018] S21. Standardize the collected household energy consumption factor data from multiple sources;

[0019] S22. Setting a projection pursuit index function to identify projection directions with clustering and outliers in the data;

[0020] S23. Randomly generate an initial projection direction population using real number encoding, where each individual is a real number vector of length d, and perform normalization processing; where d is the data dimension;

[0021] S24. Find the optimal projection direction according to the following method:

[0022] Use roulette or tournament mechanism to select outstanding individuals based on fitness value;

[0023] The selected individuals are recombined using the BLX-α real crossover operator;

[0024] Introducing Gaussian perturbations to the target components in individuals to enhance diversity;

[0025] Keep the individual with the best fitness in the current generation to ensure that the global optimal solution is not lost;

[0026] If the fitness improvement for several consecutive generations is less than the threshold, the algorithm is considered to have converged and the optimal projection direction is output;

[0027] S25. Use the optimal projection direction obtained in step S24 to perform projection transformation on the standardized high-dimensional samples to obtain a one-dimensional dimensionality reduction result.

[0028] Furthermore, the step S3 specifically includes:

[0029] Perform interval partitioning on the continuous eigenvector after dimensionality reduction;

[0030] Each household sample is regarded as a transaction, which is composed of its discretized labels on multiple behavioral variables. By integrating all sample transactions, a complete energy consumption behavior transaction database is constructed.

[0031] Run the Apriori algorithm on the transactional database;

[0032] Extract strong association rules that meet the set threshold and filter rule sets with high energy consumption and high emission characteristics;

[0033] Graphically display frequent item sets and their corresponding support and credibility.

[0034] Furthermore, the step S4 specifically includes:

[0035] For each type of behavioral variable appearing in the frequent item set, it is mapped one by one to the four subsystems of the DNAS model according to the variable semantics and behavioral attributes;

[0036] Connect the mapped variables in the logical order of drive, demand, action, and system to construct multiple candidate behavior chains;

[0037] By counting the frequency of occurrence of variable combinations in each chain in the sample, typical paths with a high probability of occurrence are screened as carbon lock-in causal paths. For each path, its representativeness and credibility are evaluated by combining the support and credibility indicators of previous frequent item sets.

[0038] Based on typical paths, summarize the common mechanisms behind different paths;

[0039] The typical DNAS chain is visualized in the form of a flow chart or directed graph, with the direction of causal relationships and the frequency of occurrence between nodes indicated.

[0040] Furthermore, the ratio of the inter-class dispersion to the intra-class aggregation is used as the projection pursuit index function J(a):

[0041] J(a)=S B / S w ; Where a is the projection direction vector; S B is the inter-class dispersion; S w is the intra-class aggregation degree.

[0042] Furthermore, the Apriori algorithm is run on the transaction database, specifically including:

[0043] Count the support of all 1-item sets;

[0044] Iteratively generate 2-item sets, 3-item sets, etc., and filter out item sets that do not meet the support threshold;

[0045] The candidate itemset is pruned using the Apriori property.

[0046] The present invention has the following beneficial effects: A carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification, disclosed herein, effectively improves the accuracy of carbon lock-in identification and causal analysis capabilities for urban households by integrating multi-source energy behavior data, projection pursuit nonlinear dimensionality reduction, and frequent item set identification. Compared to traditional regression models, this method can uncover deeper behavioral patterns and reveal hidden carbon lock-in pathways. It boasts high identification efficiency, robust mechanism analysis, and broad applicability, providing a scientific basis for household energy behavior management interventions. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The present invention will be further described below in conjunction with the accompanying drawings and embodiments:

[0048] Figure 1 Schematic diagram of the carbon lock-in assessment method of the present invention. DETAILED DESCRIPTION

[0049] The present invention is further described below with reference to the accompanying drawings, as shown in the drawings:

[0050] This embodiment discloses a carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification, comprising the following steps:

[0051] S1. Construct a household energy consumption factor database with several dimensions;

[0052] S2. Use the projection pursuit model to perform nonlinear dimensionality reduction on the data and optimize the projection direction by combining it with a real-coded accelerated genetic algorithm.

[0053] S3. Based on the factor vector after dimensionality reduction, extract frequent itemsets, set support and credibility thresholds, and identify carbon lock-in behavior paths;

[0054] S4. Combine DNAS behavioral mechanism theory to identify the causal pathways and mechanisms of carbon lock-in.

[0055] In this embodiment, in step S1, to achieve high-precision modeling and causal mechanism identification of household carbon lock-in behavior in urban residential buildings, a household energy factor database containing multi-dimensional factors must first be constructed. This database uses the household as the analysis unit and integrates information variables from multiple dimensions, such as environmental, social, technological, and strategic dimensions, for subsequent dimensionality reduction, behavior identification, and mechanism modeling. Specifically, it includes:

[0056] Integrate data sources to obtain raw data; obtain multi-source heterogeneous data through the following channels: obtain subjective information such as user behavior patterns, energy-saving awareness, strategic cognition, and self-control preferences; collect high-frequency energy consumption data such as household electricity, gas, and heating; introduce structural data such as building energy efficiency rating, climate zone classification, and regional strategy; and obtain external environmental variables such as building orientation, greening rate, and ventilation conditions;

[0057] The collected raw data was divided into four dimensions based on factor attributes: environmental dimension, social dimension, technical dimension, and strategic dimension. The environmental dimension includes factors such as the city's climate zone, annual average temperature change, air humidity, lighting conditions, and building orientation; the social dimension includes factors such as annual household income, family composition, education level, occupation type, and energy price sensitivity; the technical dimension includes factors such as the type and energy efficiency rating of household appliances, whether there is an intelligent temperature control system, insulation structure type, and equipment maintenance frequency; and the strategic dimension includes factors such as whether energy-saving subsidies are available, awareness of existing policies, community energy feedback mechanisms, and green travel incentive policies.

[0058] To ensure data consistency and model input compatibility, the raw data was cleaned, normalized, missing values ​​imputed, and subjective and objective variables fused. Categorical variables were encoded using one-hot encoding, and continuous variables were normalized using the Z-score.

[0059] A unified data table structure is established, with each row corresponding to a family sample and each column being a specific factor variable. A total of no less than 80 variables are set, covering the core influencing factors of the above four dimensions.

[0060] The database is designed as a dynamic structure, supporting periodic updates and multi-source interface access. It is compatible and portable with samples from other regions, allowing for cross-regional comparative analysis of carbon lock-in behavior.

[0061] The construction of the household energy consumption factor database provides a solid data foundation and multidimensional semantic support for the projection pursuit dimensionality reduction analysis, frequent behavior item set identification and causal path modeling in the subsequent implementation steps.

[0062] In this embodiment, in step S2, to effectively extract the main variation structure among household energy behavior factors in urban residential buildings and identify potential carbon lock-in behavior paths, a projection pursuit (PP) model is used to perform nonlinear dimensionality reduction on multidimensional energy behavior data. This method can map high-dimensional data into a one-dimensional or low-dimensional subspace to reveal the nonlinear structure and distribution characteristics in the data. Specifically, it includes:

[0063] S21. First, standardize the collected multi-source household energy factor data to ensure that each dimension has the same scale and eliminate the dimensionality effect. The standardization method can be Z-score or Min-Max normalization;

[0064] S22. To effectively identify the projection directions with clustering and outliers in the data, the ratio of the inter-class dispersion to the intra-class aggregation is used as the projection pursuit index function J(a):

[0065] J(a)=S B / S w ; Where a is the projection direction vector; S B is the inter-class dispersion; S w is the intra-class aggregation degree.

[0066] S23. Randomly generate an initial projection direction population using real number encoding, where each individual is a real vector of length d and normalize it to ensure it is a unit vector; where d is the data dimension;

[0067] S24. To find the optimal projection direction, the real-coded accelerated genetic algorithm is used for optimization:

[0068] Using a roulette or tournament mechanism, excellent individuals are selected based on their fitness values ​​(i.e., the projection pursuit index function J(a));

[0069] The selected individuals are recombined using the BLX-α real crossover operator;

[0070] Introducing Gaussian perturbations to the target components in individuals to enhance diversity;

[0071] Keep the individual with the best fitness in the current generation to ensure that the global optimal solution is not lost;

[0072] If the fitness improvement for several consecutive generations is less than the threshold, the algorithm is considered to have converged and the optimal projection direction is output;

[0073] S25. Use the optimal projection direction a obtained in step S24 * , for the standardized high-dimensional sample x i Perform projection transformation to obtain the one-dimensional dimensionality reduction result y i :

[0074] y i =a * ·x i ;

[0075] The resulting projections form a new set of reduced-dimensionality feature vectors, which are subsequently used for frequent item set extraction and carbon lock-in behavior identification. If the dimensionality is reduced to two dimensions, clustering trends of different categories or carbon lock-in behaviors can be displayed through scatter plots, enhancing model interpretability.

[0076] The above steps achieve the extraction of the most discriminative feature dimensions from high-dimensional complex behavioral factors, providing dimensionality reduction support and feature optimization basis for subsequent frequent item set identification and causal mechanism modeling.

[0077] In this embodiment, in step S3, after completing the nonlinear dimensionality reduction process of the projection pursuit model, a set of low-dimensional feature vectors representing household energy consumption behavior patterns is obtained. To further explore the potential combined patterns of carbon lock-in behavior, the Apriori association rule mining algorithm is used to identify frequent item sets from the dimensionality reduction results. Specifically, the following steps are performed:

[0078] To meet the Apriori algorithm's requirement for discrete input, the continuous feature vectors after dimensionality reduction are first partitioned into intervals. Equal-frequency or equal-interval binning is then used to discretize each variable into several categories, such as behavioral labels like "low-frequency air conditioning use," "moderate heating load," and "high-frequency bathing."

[0079] Each household sample is considered a transaction, consisting of its discretized labels on multiple behavioral variables, such as {low-frequency window opening, old building, high income}. By integrating all sample transactions, a complete energy consumption behavior transaction database is constructed.

[0080] Running the Apriori algorithm on a transactional database involves:

[0081] Initial item set count: count the support of all 1 item sets;

[0082] Frequent item set generation: iteratively generate 2-item sets, 3-item sets, etc., and filter out item sets that do not meet the support threshold;

[0083] Pruning optimization: Utilize the Apriori property to prune the candidate item set to improve algorithm efficiency.

[0084] The support threshold is set to 0.3, that is, a certain item set appears in at least 30% of the samples, and the credibility threshold is set to 0.7, that is, if the rule When established, at least 70% of the samples containing A also contain B. This parameter can be adjusted by region and sample size;

[0085] Extract strong association rules that meet the set threshold and filter out rule sets with high energy consumption and high emissions characteristics; for example, mine the following typical carbon lock-in paths: {old buildings, large households, high comfort needs} Excessive energy use; the above rules reveal the implicit relationship between energy use behavior combinations and carbon emissions, which can be used to further construct a causal mechanism for carbon lock-in;

[0086] Frequent item sets and their corresponding support and credibility are displayed graphically. Bar charts, matrix charts, or network diagrams can be used to represent the degree of coupling between behavioral variables, enhancing the readability of model results and decision support functions.

[0087] The above steps effectively identify the combination of carbon lock-in patterns in household energy consumption behavior through frequent item set mining technology, which can provide a strong behavioral data basis for subsequent path causal modeling and carbon unlocking strategy formulation.

[0088] In this embodiment, after mining typical frequent item sets in household energy consumption behavior in step S4, the DNAS behavioral mechanism theoretical framework is further introduced to identify and model the causal path of the formation mechanism of carbon lock-in behavior. The DNAS theory (Drive-Need-Action-System) emphasizes that the evolution of individual behavior is driven by driving factors, which in turn trigger specific actions and are ultimately subject to the influence of external system conditions. It has a strong behavioral causal chain logic. Specifically, it includes:

[0089] For various behavioral variables that appear in the frequent item set, such as "high frequency of air conditioning use," "low window opening rate," and "weak awareness of energy-saving strategies," we map them one by one to the four subsystems of the DNAS model based on the variable semantics and behavioral attributes:

[0090] Drivers: external triggers such as climate, income, availability of energy-saving information, etc.

[0091] Demand: psychological or physical needs of users, such as comfort, convenience, and health;

[0092] Actions: Specific behavioral manifestations, such as device usage frequency and energy usage habits;

[0093] System: Structural conditions that constrain behavior, such as equipment energy efficiency, energy prices, feedback mechanisms, etc.

[0094] Connect the mapped variables in the logical order of drive, demand, action, and system to construct multiple candidate behavior chains; for example: lack of information channel → increased comfort preference → high-frequency air conditioning turned on → lack of energy efficiency feedback mechanism.

[0095] By counting the frequency of occurrence of variable combinations in each chain in the sample, typical paths with a high probability of occurrence are screened as carbon lock-in causal paths. For each path, its representativeness and credibility are evaluated by combining the support and credibility indicators of previous frequent item sets.

[0096] Based on typical paths, the common mechanisms behind different paths are summarized; for example, the starting points of most high-carbon lock-in paths are concentrated on the driving factors of "weak strategic awareness" or "high climate adaptation pressure"; most action nodes are reflected in "high frequency use of air conditioners / water heaters"; and system nodes are mostly concentrated on "poor building envelope structure" and "lack of energy consumption feedback".

[0097] The typical DNAS chain is visualized in the form of a flow chart or directed graph, with the direction of causal relationships and the frequency of occurrence between nodes indicated, to assist strategy designers in clearly identifying the causes of carbon lock-in and its dynamic evolution path.

[0098] Through the above steps, the present invention not only statically identifies carbon lock-in behavior, but also systematically reveals the dynamic relationship path between its behavior driver-demand motivation-energy consumption action-system support through the DNAS causal framework, providing a technical basis and model foundation for target identification of intervention measures and decarbonization path design.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification, characterized by: The steps include: S1. Construct a household energy consumption factor database with several dimensions; S2. Use the projection pursuit model to perform nonlinear dimensionality reduction on the data and optimize the projection direction by combining it with a real-coded accelerated genetic algorithm. S3. Based on the factor vector after dimensionality reduction, extract frequent itemsets, set support and credibility thresholds, and identify carbon lock-in behavior paths; S4. Combine DNAS behavioral mechanism theory to identify the causal pathways and mechanisms of carbon lock-in.

2. The carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification according to claim 1 is characterized by: The step S1 specifically includes: Integrate data sources to obtain original data; The collected raw data are divided into four dimensions according to factor attributes: environmental dimension, social dimension, technical dimension and strategic dimension; Clean, normalize, fill missing values ​​and integrate subjective and objective variables on the original data; Establish a unified data table structure, where each row corresponds to a family sample and each column is a specific factor variable; The database is designed as a dynamic structure that supports periodic updates and multi-source interface access.

3. The carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification according to claim 1 is characterized by: The step S2 specifically includes: S21. Standardize the collected household energy consumption factor data from multiple sources; S22. Setting a projection pursuit index function to identify projection directions with clustering and outliers in the data; S23. Randomly generate an initial projection direction population using real number encoding, where each individual is a real number vector of length d, and perform normalization processing; where d is the data dimension; S24. Find the optimal projection direction according to the following method: Use roulette or tournament mechanism to select outstanding individuals based on fitness value; The selected individuals are recombined using the BLX-α real crossover operator; Introducing Gaussian perturbations to the target components in individuals to enhance diversity; Keep the individual with the best fitness in the current generation to ensure that the global optimal solution is not lost; If the fitness improvement for several consecutive generations is less than the threshold, the algorithm is considered to have converged and the optimal projection direction is output; S25. Use the optimal projection direction obtained in step S24 to perform projection transformation on the standardized high-dimensional samples to obtain a one-dimensional dimensionality reduction result.

4. The carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification according to claim 1 is characterized by: The step S3 specifically includes: Perform interval partitioning on the continuous eigenvector after dimensionality reduction; Each household sample is regarded as a transaction, which is composed of its discretized labels on multiple behavioral variables. By integrating all sample transactions, a complete energy consumption behavior transaction database is constructed. Run the Apriori algorithm on the transactional database; Extract strong association rules that meet the set threshold and filter rule sets with high energy consumption and high emission characteristics; Graphically display frequent item sets and their corresponding support and credibility.

5. The carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification according to claim 1 is characterized by: The step S4 specifically includes: For each type of behavioral variable appearing in the frequent item set, it is mapped one by one to the four subsystems of the DNAS model according to the variable semantics and behavioral attributes; Connect the mapped variables in the logical order of drive, demand, action, and system to construct multiple candidate behavior chains; By counting the frequency of occurrence of variable combinations in each chain in the sample, typical paths with a high probability of occurrence are screened as carbon lock-in causal paths. For each path, its representativeness and credibility are evaluated by combining the support and credibility indicators of previous frequent item sets. Based on typical paths, summarize the common mechanisms behind different paths; The typical DNAS chain is visualized in the form of a flow chart or directed graph, with the direction of causal relationships and the frequency of occurrence between nodes indicated.

6. The carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification according to claim 3 is characterized by: The ratio of the inter-class dispersion to the intra-class aggregation is used as the projection pursuit index function J(a): J(a)=S B / S w ; Where a is the projection direction vector; S B is the inter-class dispersion; S w is the intra-class aggregation degree.

7. The carbon lock-in assessment method based on multi-source data projection pursuit and frequent item set identification according to claim 4 is characterized by: Running the Apriori algorithm on a transactional database involves: Count the support of all 1-item sets; Iteratively generate 2-item sets, 3-item sets, etc., and filter out item sets that do not meet the support threshold; The candidate itemset is pruned using the Apriori property.