Energy efficiency carbon asset integrated fusion analysis method and system based on multi-source data lake

CN122797945APending Publication Date: 2026-09-22SHAANXI KUNLEI DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611037712.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]根据本发明的一方面,提出一种基于多源数据湖的能效碳资产一体化融合分析方法,用于解决现有方法无法满足能效碳资产一体化融合分析对于高效能、低延迟及资源集约化的要求的问题,包括如下步骤:

Benefits of technology

[0006]本发明通过将融合分析过程分解为数据处理节点集合,并为各节点生成静态标识,实现了对不同计算步骤的区分、追踪与管理。所建立的结果缓存机制结合多维状态特征向量和效用值对中间结果进行精细化存储,其中效用值综合反映节点计算成本与下游影响范围,并随时间衰减、随复用增强,从而有助于优先保留复用收益较高的计算结果,减少低价值缓存数据占用空间,提升缓存资源利用效率并降低重复计算量。在处理新的分析请求时,系统不仅能够复用静态标识和状态特征相匹配的历史缓存结果,还能够在命中近邻缓存点时,基于近邻输入数据与缓存结果构建代理模型,对当前节点计算结果进行快速预测,扩展了历史计算结果的可复用范围。在满足能效与碳资产数据融合分析需求的基础上,有利于缩短全流程计算响应时间,降低底层算力资源消耗,并提升高频分析场景下的数据处理效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797945A_ABST
    Figure CN122797945A_ABST
Patent Text Reader

Abstract

The application provides a kind of energy efficiency carbon asset integrated fusion analysis method and system based on multi-source data lake, including collecting energy efficiency, carbon asset data, disassembling fusion process into data processing node, generating the static identification of each node according to calculation logic, parameter, upstream dependence hash;Build cache mechanism, extract the dimensionless processed multi-dimensional state feature vector of input data after calculation, give time decay utility value positively correlated with calculation cost and number of downstream nodes, store data, results and utility value with static identification and feature vector index, enhance utility when reusing cache;New request disassembles node, then search cache first, if there is no match, complete operation;When there is candidate set, search high-utility neighbor cache, if hit, reduce dimension on data and build proxy model to predict results, improve the utility value of used neighbor cache point, integrate all node output integrated fusion analysis results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of integrated energy efficiency and carbon assets, and in particular relates to an integrated analysis method and system for energy efficiency and carbon assets based on a multi-source data lake. Background Technology

[0002] With the continuous advancement of the "dual carbon" goals, the synergistic integration of energy efficiency management and carbon asset management has become a crucial support for various enterprises to achieve green and low-carbon transformation. Energy efficiency data and carbon asset data are closely linked in terms of business and computation. Integrated analysis of these two data points can provide reliable data for energy conservation and emission reduction strategy formulation, carbon quota management, and energy operation optimization. However, the data generated during energy efficiency monitoring and carbon emission accounting are typically diverse in origin, massive in scale, and highly dimensional, often requiring complex data processing workflows and computationally intensive models for fusion analysis. In practical applications, due to the continuous operation of energy systems and the dynamic changes in carbon asset operations, analysis systems typically need to handle high-frequency, continuous analysis requests. Under high-concurrency, high-load data processing conditions, how to reduce the repetitive computational pressure caused by massive amounts of data and improve system computational efficiency and response speed has become a pressing technical problem to be solved in the fields of energy digitalization and carbon asset management.

[0003] To address the aforementioned high-frequency analysis needs, traditional analysis systems typically employ a per-request independent computation processing model. This means that each time a data fusion analysis request is received, the entire computation process is re-executed from the source data. This approach easily leads to a large amount of redundant computation, resulting in high computational resource consumption and significant system response latency. While some systems alleviate some computational pressure through basic caching mechanisms, in real-world scenarios, energy efficiency data and carbon asset data input at different times often exhibit only local fluctuations or high similarity in multidimensional state characteristics. Traditional processing mechanisms can usually only reuse historical results with completely identical input conditions, lacking the ability to intelligently reuse similar input data and failing to fully utilize the large amount of historical computation results accumulated in the data lake. Furthermore, in complex data processing chains, existing systems often lack a scientific cache value assessment mechanism, making it difficult to comprehensively consider factors such as computational costs, upstream and downstream node dependencies, and time variations to manage the full lifecycle utility of a large number of intermediate results. This easily leads to the premature obsolescence of high-value computation results and the long-term accumulation of low-value cached data, thus limiting the system's ability to respond quickly to similar analysis requests and failing to meet the requirements of high efficiency, low latency, and resource intensification for integrated energy efficiency and carbon asset fusion analysis. Summary of the Invention

[0004] According to one aspect of the present invention, a method for integrated analysis of energy efficiency and carbon assets based on multi-source data lakes is proposed to address the problem that existing methods cannot meet the requirements of high efficiency, low latency, and resource intensification for integrated analysis of energy efficiency and carbon assets, comprising the following steps: Acquire target energy efficiency data and carbon asset data, decompose the fusion analysis process into a set of data processing nodes, and generate a static identifier for each node consisting of calculation logic, parameters and upstream dependency hashes; A result caching mechanism is established. After the data processing node completes the calculation, the multidimensional state feature vector of the input data after dimensionless processing is extracted. A utility value that is positively correlated with the calculation cost and the number of downstream nodes is assigned. The input data, calculation results and utility value are stored in the result caching mechanism using the static identifier and the multidimensional state feature vector after dimensionless processing as indexes. The utility value decays over time and is enhanced when reused. The system receives new analysis requests and decomposes nodes. It searches for candidate sets in the result caching mechanism using the static identifier of each node. If the candidate set is empty, it performs calculations and stores the results in the result caching mechanism. If the candidate set is not empty, it searches for nearest neighbor cache points with a utility value higher than a set threshold, using the dimensionless multidimensional state feature vector of the current input data as the center. If no nearest neighbor cache point is found, it performs calculations and stores the results in the result caching mechanism. If a nearest neighbor cache point is found, it performs dimensionality reduction processing on the input data of the nearest neighbor cache point and the current input data. It constructs a proxy model using the dimensionality-reduced input data of the nearest neighbor cache point as input and the cached result as output. It inputs the dimensionality-reduced current input data into the proxy model to predict the current calculation result as the node output, thereby improving the utility value of the used nearest neighbor cache points. It integrates the calculation results of each node and outputs a unified fusion analysis result.

[0005] According to another aspect of the present invention, an integrated fusion analysis system for energy efficiency and carbon assets based on a multi-source data lake is proposed, comprising the following modules: The generation module is used to acquire target energy efficiency data and carbon asset data, decompose the fusion analysis process into a set of data processing nodes, and generate a static identifier for each node consisting of calculation logic, parameters and upstream dependency hashes. The extraction module is used to establish a result caching mechanism. After the data processing node completes the calculation, it extracts the multidimensional state feature vector of the input data after dimensionless processing, assigns a utility value that is positively correlated with the calculation cost and the number of downstream nodes, and stores the input data, calculation results and utility value into the result caching mechanism using the static identifier and the multidimensional state feature vector after dimensionless processing as indexes. The utility value decays over time and is enhanced when reused. The output module receives new analysis requests and decomposes nodes. It searches for candidate sets in the result caching mechanism using the static identifier of each node. If the candidate set is empty, it performs calculations and stores the results in the result caching mechanism. If the candidate set is not empty, it searches for nearest neighbor cache points with a utility value higher than a set threshold, centered on the dimensionless multidimensional state feature vector of the current input data. If no nearest neighbor cache point is found, it performs calculations and stores the results in the result caching mechanism. If a nearest neighbor cache point is found, it performs dimensionality reduction processing on the input data of the nearest neighbor cache point and the current input data. It constructs a proxy model using the dimensionality-reduced input data of the nearest neighbor cache point as input and the cached results as output. It inputs the dimensionality-reduced current input data into the proxy model to predict the current calculation result as the node output, thereby improving the utility value of the used nearest neighbor cache points. It integrates the calculation results of each node and outputs a unified fusion analysis result.

[0006] This invention decomposes the fusion analysis process into a set of data processing nodes and generates static identifiers for each node, enabling the differentiation, tracking, and management of different computational steps. The established result caching mechanism combines multi-dimensional state feature vectors and utility values ​​to store intermediate results in a refined manner. The utility value comprehensively reflects the node's computational cost and downstream impact range, decaying over time and increasing with reuse. This helps to prioritize the retention of computational results with higher reuse benefits, reduce the space occupied by low-value cached data, improve cache resource utilization efficiency, and reduce redundant computation. When processing new analysis requests, the system can not only reuse historical cached results matching static identifiers and state features, but also, when a nearest neighbor cache point is hit, construct a proxy model based on the nearest neighbor input data and cached results to quickly predict the current node's computational results, expanding the reusable range of historical computational results. While meeting the needs of energy efficiency and carbon asset data fusion analysis, this approach helps to shorten the overall computational response time, reduce the consumption of underlying computing resources, and improve data processing efficiency in high-frequency analysis scenarios. Attached Figure Description

[0007] Figure 1 A flowchart of the first embodiment; Figure 2 This is a schematic diagram illustrating the trend of utility value before truncation during the decay and reuse enhancement process; Figure 3 A bar chart comparing the state characteristics of time series data for four types of samples. Detailed Implementation

[0008] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0009] In the first embodiment, the present invention proposes an integrated analysis method for energy efficiency and carbon assets based on a multi-source data lake, such as... Figure 1 It includes the following steps: S1 acquires target energy efficiency data and carbon asset data, decomposes the fusion analysis process into a set of data processing nodes, and generates a static identifier for each node consisting of calculation logic, parameters, and upstream dependency hashes.

[0010] A unified metadata catalog from a multi-source data lake is accessed. Multimodal target energy efficiency data and carbon asset data, including smart meter time-series data, carbon trading system relationship tables, and unstructured environmental monitoring logs, are extracted from the data lake object storage in read-time mode across sources. The fusion analysis process is constructed as a directed acyclic graph (DAG). Each data processing node in the DAG is traversed, and the computational logic (i.e., the Python source code string) of the current node is extracted. Invalid control characters, whitespace characters, and comment text are cleaned to obtain computational logic characteristics. The `dumps` function is used to serialize the current node's runtime parameters into a byte stream, and the hash values ​​of all upstream nodes are extracted as upstream dependencies. The SHA256 algorithm is used to hash the string concatenated from the computational logic characteristics, runtime parameter byte stream, and upstream dependencies. The hexadecimal digest output by the SHA256 algorithm is used as the static identifier of the data processing node. The SHA256 algorithm output is a 256-bit binary digest, typically represented by 64 hexadecimal characters, thus avoiding confusion between "bit length" and "hexadecimal character count."

[0011] In an optional embodiment, the acquisition of target energy efficiency data and carbon asset data involves decomposing the fusion analysis process into a set of data processing nodes, generating a static identifier for each node consisting of computational logic, parameters, and upstream dependency hashes, including: The workflow of the analysis process is analyzed, the data mapping relationship between tasks of each data processing node is extracted, and a directed acyclic graph is constructed. Extract the source code content corresponding to the data processing node, clean up invalid control characters and comment text through regular expression matching, and retain the core continuous operation instruction string as the computational logic feature; Read all preset parameter key-value pairs in the data processing node's runtime environment, sort them alphabetically in ascending order, and then concatenate them to form control parameter features; Locate all upstream parent nodes that transmit data to the current node, obtain the retained static identifiers of each parent node and combine them to generate upstream dependency features; The computational logic features, control parameter features, and upstream dependency features are sequentially concatenated and input into a one-way hash algorithm to generate a fixed-length hash value, which serves as the unique static identifier of the data processing node.

[0012] During the workflow parsing phase, the workflow configuration file in JSON or XML format is read to identify the input and output data dependencies between nodes, generating a directed acyclic graph (DAG) in the form of an adjacency matrix or adjacency list. When extracting the source code, regular expressions are used to filter whitespace and newline characters, and line-level or block-level comments and their contents are removed. The cleaned data processing instruction fragments are concatenated into a continuous core string, typically between 100 and 1000 characters long, serving as the computational logic feature. For control parameters, dictionary-style key-value pairs of parameters configured at node runtime, such as thresholds and learning rates (floating-point variables), are read, sorted in ascending order by the ASCII code of the parameter keys, and converted into strings using a fixed format, such as lr=0.01&threshold=0.85, generating control parameter features. If the current node has no parent node, a preset empty dependency character is used as the upstream dependency feature. If multiple parent nodes exist, the static identifiers represented by 64 hexadecimal characters, calculated and stored by each parent node, are retrieved in ascending order of their internal system IDs and combined using a specific delimiter to generate the upstream dependency feature. The cleaned computational logic features, formatted control parameter features, and concatenated upstream dependency features are joined into a complete character stream using delimiters such as ||, and then the standard SHA-256 one-way hash algorithm is called for computation.

[0013] This process outputs a static identifier of 64 hexadecimal characters in the form of e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855, corresponding to a 256-bit digest. This ensures that during energy efficiency analysis and modeling, any fine-tuning of parameters or addition / deletion of code will trigger a change in the static identifier, achieving collision prevention and unique identification for each processing node.

[0014] S2. Establish a result caching mechanism. After the data processing node completes the calculation, extract the multidimensional state feature vector of the input data after dimensionless processing. Assign a utility value that is positively correlated with the calculation cost and the number of downstream nodes. Store the input data, calculation results and utility value in the result caching mechanism using the static identifier and the multidimensional state feature vector after dimensionless processing as indexes. The utility value decays over time and is enhanced when reused.

[0015] The underlying storage engine of the result caching mechanism is Redis distributed in-memory database. After the data processing node completes its computation, the input data is standardized. A unified standardizer and a unified PCA transformation model, pre-trained and solidified based on historical samples, are invoked. Principal Component Analysis (PCA) is used to extract the first K principal components of the standardized data to form a multi-dimensional state feature vector, where K is no greater than 10 and no greater than the smaller of the original feature dimension and the number of samples used to train the PCA model. The unified standardizer and unified PCA transformation model maintain the same version under the same node static identifier or the same input structure version during the cache writing stage and the subsequent reuse retrieval stage, and the model version number is saved as cache index metadata, so that the feature vectors of historical cached items and current input data are in the same coordinate system. The time module is used to record the computation time of the data processing node, and the resource monitoring interface is used to record the maximum memory peak during the operation. The computation time and the maximum memory peak are dimensionless and then merged according to preset weights to form the computation cost. The descendants function or a downstream traversal function based on the topology graph is used to count the total number of downstream nodes that are directly or indirectly affected by the data processing node in the directed acyclic graph. The utility value is calculated using the formula: Utility Value = Dimensionless Computational Cost × (1 + Number of Downstream Nodes) or an equivalent forward weighting formula. Higher computational costs and a wider downstream impact result in a higher utility value for the cached result. A hash table structure is used, with a static identifier as the outer key and the hash value of the multidimensional state feature vector as the inner key. The input data, computational results, and utility value are stored via the HSET command. The original numerical array of the multidimensional state feature vector is synchronously saved as a cache record field, or a vector index table that can be retrieved by the static identifier is maintained for subsequent nearest neighbor distance calculations.

[0016] In one specific implementation, when a subsequent read call is triggered, or when a background daemon triggers a utility value update according to a preset period, the system reads the lifecycle baseline point recorded when each cached item was written to the cache system and the current system timestamp, calculates the time difference between the current timestamp and the lifecycle baseline point, divides the time difference by a preset baseline time interval to obtain a dimensionless time ratio, substitutes the dimensionless time ratio (plus one) into a preset natural logarithmic decay function to obtain the decay amount, and subtracts the decay amount from the current utility value. If the utility value after deduction is lower than a preset lower limit of utility value, it is then adjusted. The system uses the preset lower limit of utility value to achieve timeliness degradation and updates the lifecycle benchmark point to the current system timestamp. When a cached item is hit in the nearest neighbor search and used to build a local proxy model, a first fixed additional value is added to the degraded utility value for this level of enhancement. When a downstream node that depends on the current node is successfully hit and reused in the cache, a second fixed additional value is added to the current node's utility value for association enhancement, where the first fixed additional value is greater than the second fixed additional value. If the utility value after this level of enhancement or association enhancement is higher than the preset upper limit of utility value, it is corrected to the preset upper limit of utility value. Figure 2 As shown, this illustrates the natural decay of cached entries over time before being truncated to a preset utility value upper limit, and the increase in value after a reuse hit. When actually writing to or updating the cache, if the enhanced utility value exceeds the preset utility value upper limit, it is truncated and corrected according to the preset utility value upper limit.

[0017] In an optional embodiment, the establishment of a result caching mechanism, which extracts the dimensionless multidimensional state feature vector of the input data after the data processing node has completed its calculation, includes: Data is read from the main data collection entry point that receives local data, and the data is divided into time series numerical data and non-numerical discrete classification data. For time series numerical data, the absolute difference between adjacent data points is calculated by difference operation to form a sequence, and the overall variance and absolute average deviation of the sequence are calculated as the first feature subset. For non-numerical discrete classification data, the frequency of each enumeration value in the global data is counted, and the uncertainty of the discrete distribution is calculated by the information entropy algorithm, which is used as the second feature subset. The first feature subset and the second feature subset are respectively processed into dimensionless form according to the corresponding preset scale or historical statistical range to obtain a first standardized feature subset and a second standardized feature subset with a unified scale. The first standardized feature subset and the second standardized feature subset are concatenated and linked according to a preset queue to generate a continuous vector whole, which serves as the dimensionless multidimensional state feature vector of the node input data.

[0018] When the original multimodal target energy efficiency and carbon emission data existing in the source-attached layer of the multi-source data lake are scheduled to the fusion computing node, the field type tags are parsed by calling the unified metadata service of the data lake. Sample sequences such as historical electricity consumption curves and greenhouse gas concentration waveforms are categorized into time-series numerical data. The length of the numerical data array is preferably controlled between 100 and 10000, while equipment models and regional operation mode tags are categorized into non-numerical discrete classification data. For a time-series numerical sequence containing n samples, the absolute value of the difference between adjacent time points is calculated using first-order differencing to obtain a difference sequence. Then, a statistical algorithm is executed on this new sequence to obtain the overall variance and absolute mean deviation. For example, if the difference variance of a certain carbon emission sequence is calculated to be 15.2 and the absolute mean deviation is 3.4, these two indicators constitute a 2-dimensional floating-point first feature subset, used to represent the degree of sequence fluctuation in the current input time data window. Figure 3 The diagram shows the degree of data fluctuation dispersion in each group of time-series samples. For non-numerical discrete classification data such as device modes, the frequency of occurrence of each type of state enumeration value is statistically analyzed using a hash table. The entropy value is obtained by applying the Shannon information entropy formula based on the statistically derived distribution frequency sequence. For example, if the entropy value is 1.15 derived from the frequency distribution, this single floating-point number or a one-dimensional floating-point array aggregated from multiple field entropy values ​​is used as the second feature subset to evaluate the data distribution characteristics. In the feature integration stage, the first feature subset and the second feature subset are concatenated in the array using floating-point operations according to the established data structure column specifications. The above example will be transformed into a multi-dimensional floating-point vector of [15.2, 3.4, 1.15]. Normalization is performed according to the historical minimum and maximum values ​​corresponding to each feature dimension, and the feature mapping of each dimension is limited to the feature closed interval of [0, 1]. When further PCA dimensionality reduction is required, the standardized parameters and PCA transformation matrix consistent with the system call and cache write stages are used to uniformly transform the vector, prohibiting independent fitting of a new PCA coordinate system for a single request. The result is persisted into the system's computation cache and used in high-dimensional spatial clustering matching engines such as KNN (K nearest neighbors) or other Euclidean distance metrics during reuse verification.

[0019] In an optional embodiment, assigning a utility value that is positively correlated with computational cost and the number of downstream nodes includes: Monitor the total processing latency of the host component currently performing the operation, and detect the maximum peak memory usage during the operation. After the total processing latency and the maximum memory peak are respectively dimensionless, they are summed using a preset weighted formula to obtain a first feature value representing the absolute cost of the node's computation. In the global processing topology graph, all downstream nodes that have a dependency relationship with the current node are investigated, and the total number of downstream nodes affected by the current node is counted as the second feature value. Add the second feature value to 1 to obtain the downstream correction factor; multiply the first feature value by the downstream correction factor, or perform a weighted summation of the first feature value and the downstream correction factor according to preset weights that are both positive numbers to obtain the fusion value; The fusion value increases as the first feature value increases, and also increases as the second feature value increases; The fusion value is input into a monotonically increasing nonlinear transformation function and mapped to a preset specified numerical range to obtain the utility value.

[0020] For the underlying execution component of data fusion operations, a built-in probe captures the start and end timestamps of thread execution to calculate the total processing latency, for example, a total processing latency of 5400 milliseconds. Simultaneously, high-frequency sampling every 50 milliseconds identifies the maximum memory peak used by the process within this task cycle, such as recording a maximum memory peak of 2048MB. Then, using system configuration constants such as a latency limit of 10000 milliseconds and a memory peak limit of 8192MB, dimensionless processing is performed using division. Subsequently, the time and space consumption are multiplied by preset coefficients, such as a time weight parameter of 0.6 and a memory weight parameter of 0.4, and then summed. In this example, the first characteristic value generated is 0.424. This first characteristic value represents the underlying computational resource burden. In the graph theory spatial dimension processing, a depth-first traversal mechanism is triggered to perform topology tracing on the node network graph, accumulating the scale of the terminal links affected by this analysis step, and obtaining the number of affected downstream nodes without duplicate statistics, for example, 15 including intermediate conversion layers and report merging nodes. This value is set as the second characteristic value. The process then proceeds to multi-factor transformation, directly fusing the first and second eigenvalues ​​(both greater than 0) in a positive direction. For example, multiplying the first eigenvalue (0.424) by the second eigenvalue (plus one, 16) yields an initial base of 6.784. Alternatively, a weighted sum is performed according to positive weights, ensuring that the fused value increases synchronously with the computational cost. For standardized mapping, a Sigmoid nonlinear transformation model based on a smoothing adjustment factor is applied to scale this base. This base is transformed and mapped within a set scoring interval (0, 100) through a Sigmoid nonlinear transformation with a specified stretching coefficient of 0.1. The corresponding utility value obtained in this way is ultimately solidified as the utility value of the result node, serving as a reference weight for the lifespan of this result node in the memory eviction algorithm.

[0021] In an optional embodiment, the step of storing input data, calculation results, and utility values ​​in a result caching mechanism using static identifiers and dimensionless multidimensional state feature vectors as indices, where the utility value decays over time and is enhanced during reuse, includes: Record the absolute system timestamp when the calculation results are written to the cache system, as a lifecycle benchmark point; When a subsequent instruction initiates a read call, or when the utility value is updated according to a preset period, the current timestamp is obtained, and the time difference between the current timestamp and the life cycle benchmark point is calculated. Divide the time difference by a preset baseline time interval to obtain a dimensionless time ratio. Substitute the dimensionless time ratio after adding one into a preset natural logarithmic decay function to obtain the decay amount. Subtract the decay amount from the current utility value to achieve timeliness downgrading. If the downgraded utility value is lower than the preset lower limit of utility value, it is corrected to the preset lower limit of utility value, and the life cycle baseline point is updated to the current timestamp. When the current cache node is hit in the nearest neighbor search and used to build a local proxy model, a first fixed additional value is added to the degraded utility value for this level of enhancement; When a downstream node that depends on the current node is successfully cached and reused, a second fixed additional value is added to the utility value of the current node to enhance the association, where the first fixed additional value is greater than the second fixed additional value. If the utility value after enhancement at this level or by association is higher than the preset upper limit of utility value, then it will be corrected to the preset upper limit of utility value.

[0022] Each time the calculated target energy efficiency data is written to the memory cache, the engine extracts the system timestamp as a persistent lifecycle baseline, such as timestamp 1690000000. When a subsequent request triggers the comparison process and requires verification of this record, the read interface retrieves the current system timestamp, such as 1690003600, and calculates the difference in seconds between the two, converting it into a time difference on a specified scale. For example, the time difference is converted to a 1-hour time unit, and a +1 operation is performed on the time difference to avoid the input value of the natural logarithm function being 0. The 2 is then imported into the natural logarithmic decay function to calculate a decay of 0.69. After extracting the original utility value of 65.38, this decay is deducted to update it to 64.69, and the compared system time overwrites the old timestamp as the new lifecycle baseline. After a query verification hit occurs, the system's enhancement module will activate: If the query hits the cache block during the filtering process of the proxy prediction model reconstruction, a first fixed additional value, set in the range of 3.0 to 5.0 (e.g., 5.0), will be added to the current reduced utility value, raising the utility value to 69.69. If a leaf-level derived cache unit in a subsequent inheritance chain of the topology graph achieves a hit, even if the current node is not hit, the engine will still, through cascading backtracking, add a second fixed additional value, with a parameter range limited to 1.0 to 2.0 (e.g., adding 1.5), to the original utility value. The first fixed additional value is greater than the second fixed additional value. This utility value upgrade adjusts the lifecycle of different cache datasets, and the enhanced utility value does not exceed the preset utility value upper limit.

[0023] S3 receives new analysis requests and decomposes nodes. It searches for candidate sets in the result caching mechanism using the static identifier of the nodes. If the candidate set is empty, it performs calculations and stores the results in the result caching mechanism. If the candidate set is not empty, it searches for nearest neighbor cache points with a utility value higher than a set threshold, using the dimensionless multidimensional state feature vector of the current input data as the center. If no nearest neighbor cache point is found, it performs calculations and stores the results in the result caching mechanism. If a nearest neighbor cache point is found, it performs dimensionality reduction processing on the input data of the nearest neighbor cache point and the current input data. It constructs a proxy model using the dimensionality-reduced input data of the nearest neighbor cache point as input and the cached result as output. It inputs the dimensionality-reduced current input data into the proxy model to predict the current calculation result as the node output, improves the utility value of the used nearest neighbor cache points, and integrates the calculation results of each node to output a unified fusion analysis result.

[0024] Upon receiving a new energy efficiency carbon asset fusion analysis request, a directed acyclic graph (DAG) decomposition method is used to generate the computational logic and static identifiers of each data processing node. Nodes are then processed one by one according to the topological sorting algorithm of the DAG. For the currently processed node, the Redis HGETALL command is used to query all corresponding cache entries to form a candidate set based on the node's static identifier. If HGETALL returns an empty result, the computational logic function of the current node is called to perform data operations. The dimensionless multidimensional state feature vector is extracted, and the utility value is calculated before being stored in the Redis distributed in-memory database via the HSET command. The nearest neighbor cache point can represent a single cache point or a set of nearest neighbor cache points that satisfy both distance and utility thresholds. When using a proxy model requiring multi-sample training, this set of nearest neighbor cache points serves as the modeling sample source. If HGETALL returns a non-empty result, it calls the normalizer and PCA transformation model that match the current node's static identifier and input structure version, and are the same version as those in the S2 cache write phase. It extracts the current multidimensional state feature vector from the current input data, traverses the candidate set, and filters out cache items with a utility value greater than a set threshold. It calculates the Euclidean distance between the current multidimensional state feature vector and the filtered cache item's multidimensional state feature vector, selecting several cache items whose Euclidean distance does not exceed a set distance threshold and whose number is not less than the preset minimum modeling sample number, as the nearest neighbor cache point set. When the number of cache items within the distance threshold is less than the preset minimum modeling sample number, it is determined that no nearest neighbor cache point satisfying the surrogate modeling condition has been found, and the current node's calculation logic is called to perform native operations before storing the data in the Redis distributed memory database. If a nearest neighbor cache point set satisfying the surrogate modeling condition is found, a unified dimensionality reduction model or a PCA-type dimensionality reduction model that can perform transform operations on new samples is used to perform dimensionality reduction processing on the input data matrix of the nearest neighbor cache point set and the current input data matrix in the same coordinate system.

[0025] A proxy model is constructed using the RandomForestRegressor algorithm. When the cached result is structured data, a data frame, or a report fragment, the predictable numerical fields are first expanded into target vectors according to a preset field order. Non-numerical fields or template fields do not participate in regression training. If there are no predictable numerical fields in the cached result, or the numerical target vector is empty, or the set of predictable numerical fields for each cached result in the nearest cache point set is inconsistent, then the nearest cache point set is determined not to meet the proxy modeling conditions. The original calculation logic of the current node is executed, and the original calculation result is output as the calculation result of the current node. If there are predictable numerical fields, and the set of predictable numerical fields and the field order of each cached result in the nearest cache point set are consistent, then the dimensionality-reduced input data of multiple cached samples in the nearest cache point set is used as feature variables, and the numerically expanded target vector in the corresponding cached result is used as the target variable to train the proxy model. The dimensionality-reduced current input data is then used for inference and prediction. After the proxy model completes inference and prediction, the predicted target vector is backfilled into the original output structure according to the preset field order. The reconstructed result is then output as the calculation result of the current node. When performing utility enhancement on cached items in the set of used nearest neighbor cache points, the current utility value is first read using the Redis HGET command, and a first fixed additional value is added to the current utility value to calculate the enhanced target utility value. If the enhanced target utility value is higher than the preset utility value upper limit, it is corrected to the preset utility value upper limit and then written back using the HSET command; or the first fixed additional value is used as an increment to call the HINCRBYFLOAT command for addition accumulation, and after accumulation, it is truncated and updated according to the preset utility value upper limit, thereby achieving hit enhancement at this level. After all nodes in the directed acyclic graph have been processed, the calculation result data frame of the tail node is converted into JSON format and output as the integrated fusion analysis result.

[0026] In an optional embodiment, the step of receiving a new analysis request and performing node decomposition, searching for a candidate set in the result caching mechanism using the static identifier of the node, and performing calculations and storing the results in the result caching mechanism if the candidate set is empty, includes: Extract the static identifier of the target data processing node to be verified in the new analysis request and use it as the target object for global search; The system calls the full cache database index table and uses all historical static identifier records stored in the table as the comparison range. Within the comparison range, each historical static identifier is verified to be absolutely equal to the string of the target object; All historical cached record entries that meet the absolute equality condition are selected, and the extracted entries are aggregated and combined to form a candidate set for subsequent nearest neighbor search.

[0027] After receiving the analysis request and executing the workflow schedule, extract the SHA-256 static identifier of the data node, which is a 64-hexadecimal character generated by the code attributes and parameters.

[0028] For example, d8e4a1c5f0b2a93e6c7841d5b9f0234a7e2c0b6d89f1a0c35e74b8f6a2d9c113 is used as a static identifier for retrieval and comparison. A read operation is initiated to the NoSQL architecture or caching environment, retrieving the cached database index table containing the primary key directory prefix mapping structure of historical runtime snapshots from the system cache to determine the scope of the comparison operation. Then, the comparison verification function is invoked to compare each static identifier in the candidate set with the static identifier extracted from the query. When the search determines that the static identifiers are completely identical, it means that a task of the same level with the same computational logic, runtime parameters, and upstream dependencies has been previously executed. However, this does not mean that the current input data is completely identical to the historical input data; whether the input data is similar is still determined by subsequent multi-dimensional state feature vector nearest neighbor comparison. The engine will capture and organize the relevant feature space vector group sequence entities that match the static identifiers, along with the stored specific calculation result values ​​and other attributes, into a list. This list will be combined into a record set object array and handed over to the nearest neighbor query calculation KNN, i.e., the K nearest neighbor proxy reconstruction module. The engine will then perform subsequent nearest neighbor matching and proxy reconstruction operations. If no matching result is found, the engine will send a notification back to the main control engine to execute the native direct calculation process.

[0029] In an optional embodiment, the integration of the calculation results from each node to output a unified fusion analysis result includes: Pre-allocate and establish a structured virtual template location within the system's global storage area to accommodate the final output results; Traverse the system topology graph from the front end node to the back end node, and extract all back end computation results that are not called by other nodes as upstream dependencies; Obtain the identity name tags carried by the calculation results of each end, and fill the identity name tags one by one into the corresponding slots of the structured virtual template according to the preset output mapping rules; Once all slots are filled, the overall data within the virtual template location is serialized to generate and output a complete integrated analysis report containing energy efficiency indicators and carbon asset data.

[0030] For data fusion operations, the control node allocates a storage area in memory or disk space, typically between 1MB and 5MB. Basic field structures and empty slots are injected in JSON format, generating a structure template. For example, the initialization text embeds marker dictionaries such as Total_Emission:null and Efficiency_Index:null to define the slot system to be assembled. When the worker engine detects that all processes in the topology have completed, it scans the DAG dependencies to find all worker branch endpoints with an out-degree of 0. It automatically collects the data dictionary and scalar file arrays associated with these end nodes, forming a complete set of end-point computation results. Using the feature identification names and IDs carried in the read end-point result set, such as Carbon_Model_v3 or Node_112_Output, and mapping them to the XML parameters configured in the system, for example, the XML configuration item specifies that the identifier name Node_112_Output needs to be filled into the Total_Emission dictionary slot in the template. Based on these rules, the end-point statistical results are filled into the corresponding slots. After the engine monitoring confirms that all blank identifiers have been replaced with entity integers or strings, it triggers the encapsulation module, calls the data packaging function for serialization processing, and obtains a document object in the form of a single data stream. This file set covers the energy efficiency indicators and carbon asset data extracted from the analysis, and pushes the generated integrated report back to the client.

[0031] In the second embodiment, the present invention also proposes an integrated analysis system for energy efficiency and carbon assets based on a multi-source data lake, comprising the following modules: The generation module is used to acquire target energy efficiency data and carbon asset data, decompose the fusion analysis process into a set of data processing nodes, and generate a static identifier for each node consisting of calculation logic, parameters and upstream dependency hashes. The extraction module is used to establish a result caching mechanism. After the data processing node completes the calculation, it extracts the multidimensional state feature vector of the input data after dimensionless processing, assigns a utility value that is positively correlated with the calculation cost and the number of downstream nodes, and stores the input data, calculation results and utility value into the result caching mechanism using the static identifier and the multidimensional state feature vector after dimensionless processing as indexes. The utility value decays over time and is enhanced when reused. The output module receives new analysis requests and decomposes nodes. It searches for candidate sets in the result caching mechanism using the static identifier of each node. If the candidate set is empty, it performs calculations and stores the results in the result caching mechanism. If the candidate set is not empty, it searches for nearest neighbor cache points with a utility value higher than a set threshold, centered on the dimensionless multidimensional state feature vector of the current input data. If no nearest neighbor cache point is found, it performs calculations and stores the results in the result caching mechanism. If a nearest neighbor cache point is found, it performs dimensionality reduction processing on the input data of the nearest neighbor cache point and the current input data. It constructs a proxy model using the dimensionality-reduced input data of the nearest neighbor cache point as input and the cached results as output. It inputs the dimensionality-reduced current input data into the proxy model to predict the current calculation result as the node output, thereby improving the utility value of the used nearest neighbor cache points. It integrates the calculation results of each node and outputs a unified fusion analysis result.

[0032] The above description represents the preferred embodiments of the present invention. It should be noted that, for those skilled in the art, various improvements and modifications can be made without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for integrated analysis of energy efficiency and carbon assets based on multi-source data lakes, characterized in that, include: Acquire target energy efficiency data and carbon asset data, decompose the fusion analysis process into a set of data processing nodes, and generate a static identifier for each node consisting of calculation logic, parameters and upstream dependency hashes; A result caching mechanism is established. After the data processing node completes the calculation, the multidimensional state feature vector of the input data after dimensionless processing is extracted. A utility value that is positively correlated with the calculation cost and the number of downstream nodes is assigned. The input data, calculation results and utility value are stored in the result caching mechanism using the static identifier and the multidimensional state feature vector after dimensionless processing as indexes. The utility value decays over time and is enhanced when reused. The system receives new analysis requests and decomposes nodes. It searches for candidate sets in the result caching mechanism using the static identifier of each node. If the candidate set is empty, it performs calculations and stores the results in the result caching mechanism. If the candidate set is not empty, it searches for nearest neighbor cache points with a utility value higher than a set threshold, using the dimensionless multidimensional state feature vector of the current input data as the center. If no nearest neighbor cache point is found, it performs calculations and stores the results in the result caching mechanism. If a nearest neighbor cache point is found, it performs dimensionality reduction processing on the input data of the nearest neighbor cache point and the current input data. It constructs a proxy model using the dimensionality-reduced input data of the nearest neighbor cache point as input and the cached result as output. It inputs the dimensionality-reduced current input data into the proxy model to predict the current calculation result as the node output, thereby improving the utility value of the used nearest neighbor cache points. It integrates the calculation results of each node and outputs a unified fusion analysis result.

2. The method according to claim 1, characterized in that, The acquisition of target energy efficiency data and carbon asset data involves decomposing the fusion analysis process into a set of data processing nodes. A static identifier is generated for each node, consisting of computational logic, parameters, and upstream dependency hashes, including: The workflow of the analysis process is analyzed, the data mapping relationship between tasks of each data processing node is extracted, and a directed acyclic graph is constructed. Extract the source code content corresponding to the data processing node, clean up invalid control characters and comment text through regular expression matching, and retain the core continuous operation instruction string as the computational logic feature; Read all preset parameter key-value pairs in the data processing node's runtime environment, sort them alphabetically in ascending order, and then concatenate them to form control parameter features; Locate all upstream parent nodes that transmit data to the current node, obtain the retained static identifiers of each parent node and combine them to generate upstream dependency features; The computational logic features, control parameter features, and upstream dependency features are sequentially concatenated and input into a one-way hash algorithm to generate a fixed-length hash value, which serves as the unique static identifier of the data processing node.

3. The method according to claim 2, characterized in that, The established result caching mechanism extracts the dimensionless multidimensional state feature vector of the input data after the data processing node has completed its calculations, including: Data is read from the main data collection entry point that receives local data, and the data is divided into time series numerical data and non-numerical discrete classification data. For time series numerical data, the absolute difference between adjacent data points is calculated by difference operation to form a sequence, and the overall variance and absolute average deviation of the sequence are calculated as the first feature subset. For non-numerical discrete classification data, the frequency of each enumeration value in the global data is counted, and the uncertainty of the discrete distribution is calculated by the information entropy algorithm, which is used as the second feature subset. The first feature subset and the second feature subset are respectively processed into dimensionless form according to the corresponding preset scale or historical statistical range to obtain a first standardized feature subset and a second standardized feature subset with a unified scale. The first standardized feature subset and the second standardized feature subset are concatenated and linked according to a preset queue to generate a continuous vector whole, which serves as the dimensionless multidimensional state feature vector of the node input data.

4. The method according to claim 1 or 2, characterized in that, The utility value assigned, which is positively correlated with computational cost and the number of downstream nodes, includes: Monitor the total processing latency of the host component currently performing the operation, and detect the maximum peak memory usage during the operation. After the total processing latency and the maximum memory peak are respectively dimensionless, they are summed using a preset weighted formula to obtain a first feature value representing the absolute cost of the node's computation. In the global processing topology graph, all downstream nodes that have a dependency relationship with the current node are investigated, and the total number of downstream nodes affected by the current node is counted as the second feature value. Add the second feature value to 1 to obtain the downstream correction factor; multiply the first feature value by the downstream correction factor, or perform a weighted summation of the first feature value and the downstream correction factor according to preset weights that are both positive numbers to obtain the fusion value; The fusion value increases as the first feature value increases, and also increases as the second feature value increases; The fusion value is input into a monotonically increasing nonlinear transformation function and mapped to a preset specified numerical range to obtain the utility value.

5. The method according to claim 1, characterized in that, The mechanism that stores input data, calculation results, and utility values ​​in a result cache using static identifiers and dimensionless multidimensional state feature vectors as indices, where the utility value decays over time and is enhanced during reuse, includes: Record the absolute system timestamp when the calculation results are written to the cache system, as a lifecycle benchmark point; When a subsequent instruction initiates a read call, or when the utility value is updated according to a preset period, the current timestamp is obtained, and the time difference between the current timestamp and the life cycle benchmark point is calculated. Divide the time difference by a preset baseline time interval to obtain a dimensionless time ratio. Substitute the dimensionless time ratio after adding one into a preset natural logarithmic decay function to obtain the decay amount. Subtract the decay amount from the current utility value to achieve timeliness downgrading. If the downgraded utility value is lower than the preset lower limit of utility value, it is corrected to the preset lower limit of utility value, and the life cycle baseline point is updated to the current timestamp. When the current cache node is hit in the nearest neighbor search and used to build a local proxy model, a first fixed additional value is added to the degraded utility value for this level of enhancement; When a downstream node that depends on the current node is successfully cached and reused, a second fixed additional value is added to the utility value of the current node to enhance the association, where the first fixed additional value is greater than the second fixed additional value. If the utility value after enhancement at this level or by association is higher than the preset upper limit of utility value, then it will be corrected to the preset upper limit of utility value.

6. The method according to claim 1, characterized in that, The process of receiving new analysis requests and performing node decomposition, searching for candidate sets in the result caching mechanism using static node identifiers, and performing calculations and storing the results in the caching mechanism if the candidate set is empty, includes: Extract the static identifier of the target data processing node to be verified in the new analysis request and use it as the target object for global search; The system calls the full cache database index table and uses all historical static identifier records stored in the table as the comparison range. Within the comparison range, each historical static identifier is verified to be absolutely equal to the string of the target object; All historical cached record entries that meet the absolute equality condition are selected, and the extracted entries are aggregated and combined to form a candidate set for subsequent nearest neighbor search.

7. The method according to claim 1, characterized in that, The integrated calculation results from each node are output as a unified fusion analysis result, including: Pre-allocate and establish a structured virtual template location within the system's global storage area to accommodate the final output results; Traverse the system topology graph from the front end node to the back end node, and extract all back end computation results that are not called by other nodes as upstream dependencies; Obtain the identity name tags carried by the calculation results of each end, and fill the identity name tags one by one into the corresponding slots of the structured virtual template according to the preset output mapping rules; Once all slots are filled, the overall data within the virtual template location is serialized to generate and output a complete integrated analysis report containing energy efficiency indicators and carbon asset data.

8. A fusion analysis system for energy efficiency and carbon assets based on a multi-source data lake, characterized in that, include: The generation module is used to acquire target energy efficiency data and carbon asset data, decompose the fusion analysis process into a set of data processing nodes, and generate a static identifier for each node consisting of calculation logic, parameters and upstream dependency hashes. The extraction module is used to establish a result caching mechanism. After the data processing node completes the calculation, it extracts the multidimensional state feature vector of the input data after dimensionless processing, assigns a utility value that is positively correlated with the calculation cost and the number of downstream nodes, and stores the input data, calculation results and utility value into the result caching mechanism using the static identifier and the multidimensional state feature vector after dimensionless processing as indexes. The utility value decays over time and is enhanced when reused. The output module receives new analysis requests and decomposes nodes. It searches for candidate sets in the result caching mechanism using the static identifier of each node. If the candidate set is empty, it performs calculations and stores the results in the result caching mechanism. If the candidate set is not empty, it searches for nearest neighbor cache points with a utility value higher than a set threshold, centered on the dimensionless multidimensional state feature vector of the current input data. If no nearest neighbor cache point is found, it performs calculations and stores the results in the result caching mechanism. If a nearest neighbor cache point is found, it performs dimensionality reduction processing on the input data of the nearest neighbor cache point and the current input data. It constructs a proxy model using the dimensionality-reduced input data of the nearest neighbor cache point as input and the cached results as output. It inputs the dimensionality-reduced current input data into the proxy model to predict the current calculation result as the node output, thereby improving the utility value of the used nearest neighbor cache points. It integrates the calculation results of each node and outputs a unified fusion analysis result.

9. The system according to claim 8, characterized in that, The acquisition of target energy efficiency data and carbon asset data involves decomposing the fusion analysis process into a set of data processing nodes. A static identifier is generated for each node, consisting of computational logic, parameters, and upstream dependency hashes, including: The workflow of the analysis process is analyzed, the data mapping relationship between tasks of each data processing node is extracted, and a directed acyclic graph is constructed. Extract the source code content corresponding to the data processing node, clean up invalid control characters and comment text through regular expression matching, and retain the core continuous operation instruction string as the computational logic feature; Read all preset parameter key-value pairs in the data processing node's runtime environment, sort them alphabetically in ascending order, and then concatenate them to form control parameter features; Locate all upstream parent nodes that transmit data to the current node, obtain the retained static identifiers of each parent node and combine them to generate upstream dependency features; The computational logic features, control parameter features, and upstream dependency features are sequentially concatenated and input into a one-way hash algorithm to generate a fixed-length hash value, which serves as the unique static identifier of the data processing node.

10. The system according to claim 8, characterized in that, The established result caching mechanism extracts the dimensionless multidimensional state feature vector of the input data after the data processing node has completed its calculations, including: Data is read from the main data collection entry point that receives local data, and the data is divided into time series numerical data and non-numerical discrete classification data. For time series numerical data, the absolute difference between adjacent data points is calculated by difference operation to form a sequence, and the overall variance and absolute average deviation of the sequence are calculated as the first feature subset. For non-numerical discrete classification data, the frequency of each enumeration value in the global data is counted, and the uncertainty of the discrete distribution is calculated by the information entropy algorithm, which is used as the second feature subset. The first feature subset and the second feature subset are respectively processed into dimensionless form according to the corresponding preset scale or historical statistical range to obtain a first standardized feature subset and a second standardized feature subset with a unified scale. The first standardized feature subset and the second standardized feature subset are concatenated and linked according to a preset queue to generate a continuous vector whole, which serves as the dimensionless multidimensional state feature vector of the node input data.