Interpolation completion method and device for carbon emission factor data missing, equipment, storage medium and program product

Through the deep learning model of self-attention mechanism and diagonal mask mechanism, combined with the joint training of observation value reconstruction and mask filling tasks, the problem of low accuracy in missing data completion of carbon emission factors is solved, higher-precision interpolation completion and dynamic characteristic adaptation are achieved, and the robustness of the model is enhanced.

CN120744334APending Publication Date: 2025-10-03SHENZHEN POWER SUPPLY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511193458.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing methods for completing missing carbon emission factor data have low accuracy when processing complex and nonlinear data, and cannot accurately reflect the dynamic changes of carbon emission factors, resulting in inaccurate carbon emission reduction strategies.

Method used

A deep learning model based on self-attention mechanism and diagonal mask mechanism is adopted to optimize the interpolation and completion of missing data through joint training of observation value reconstruction task and mask filling task. Combined with the joint training of observation value reconstruction and mask filling tasks, the model parameters are optimized to improve the interpolation and completion accuracy.

Benefits of technology

The accuracy of interpolation and completion of carbon emission factor data has been improved, which can better capture the temporal nature and complex patterns of data, enhance the robustness of the model, adapt to the dynamic characteristics of carbon emission factors, and improve the accuracy of interpolation and completion results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744334A_ABST
    Figure CN120744334A_ABST
Patent Text Reader

Abstract

The invention relates to an interpolation completion method and device for carbon emission factor data missing, equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence. According to the invention, the precision of the interpolation completion result can be improved. The method comprises the following steps: performing data segmentation, normalization processing and missing data simulation on a carbon emission factor data set to obtain a missing mask, an indication mask and an incomplete data set containing missing values; determining a total loss function based on the weight coefficient, the first loss function and the second loss function; training and verifying the initial deep learning model by using the training set and the verification set, and optimizing model parameters of the initial deep learning model by minimizing a total loss function until the interpolation completion precision of the deep learning model after parameter optimization on the test set meets a threshold condition, obtaining a target deep learning model based on a self-attention mechanism and a diagonal mask mechanism; the target deep learning model is used for performing interpolation completion on the carbon emission factor data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for interpolation and completion of missing carbon emission factor data. Background Art

[0002] With growing global concern about climate change, carbon emission reduction has become a key focus. Against this backdrop, the accurate estimation and monitoring of carbon emission factors has become fundamental to promoting carbon reduction policies. Carbon emission factors are important indicators for measuring the contribution of energy consumption to CO2 emissions and are widely used in various fields, including energy production, industrial manufacturing, and construction. However, in actual production processes, carbon emission factor data is often missing. This is especially true in large-scale industrial applications. Incomplete data makes it difficult to accurately estimate carbon emissions, which in turn impacts the formulation of emission reduction strategies. Missing data often occurs for a variety of reasons, such as equipment failure, sensor failure, and data transmission interruptions. These issues not only reduce data quality but also negatively impact subsequent carbon emission monitoring and optimization decisions.

[0003] Currently, research on missing data for carbon emission factors primarily focuses on traditional interpolation methods, such as linear interpolation and spline interpolation. Missing data are typically filled with simple constants, or by extrapolating known data around the missing data. However, these methods often assume that data trends are relatively stable over time, ignoring the temporal nature and dynamic changes of the data. These methods are not suitable for datasets with highly nonlinear or complex patterns. In particular, for data such as carbon emission factors, which are influenced by multiple factors, the interpolation results suffer from low accuracy and fail to accurately reflect the actual changes in the carbon emission factors. Summary of the Invention

[0004] Based on this, it is necessary to provide an interpolation and completion method, device, computer equipment, computer-readable storage medium and computer program product for missing carbon emission factor data in response to the above technical problems.

[0005] In the first aspect, the present application provides a method for interpolation and completion of missing carbon emission factor data, including:

[0006] Obtaining a carbon emission factor dataset to be processed, performing data segmentation, normalization, and missing data simulation on the carbon emission factor dataset to obtain a missing mask, an indicator mask, and an incomplete dataset containing missing values; the incomplete dataset includes a training set, a validation set, and a test set;

[0007] Determine a first loss function for an observation reconstruction task and a second loss function for a mask filling task according to the missing mask, the indicator mask, and the incomplete data set, and determine a total loss function based on a weight coefficient, the first loss function, and the second loss function;

[0008] The initial deep learning model is trained and verified using the training set and the validation set, and the model parameters of the initial deep learning model are optimized by minimizing the total loss function until the interpolation and completion accuracy of the deep learning model after parameter optimization on the test set meets the threshold condition, thereby obtaining a target deep learning model based on the self-attention mechanism and the diagonal mask mechanism; the target deep learning model is used to interpolate and complete missing carbon emission factor data.

[0009] In one embodiment, determining a first loss function for an observation reconstruction task and a second loss function for a mask filling task based on the missing mask, the indicator mask, and the incomplete data set includes:

[0010] Determine the real missing data based on the missing mask and the incomplete data set, and generate simulated completion data based on the real missing data; determine the first loss function of the observation value reconstruction task based on the real missing data, the simulated completion data and the missing mask; determine the second loss function of the mask filling task based on the real missing data, the simulated completion data and the indication mask.

[0011] In one embodiment, before determining the total loss function based on the weight coefficient, the first loss function, and the second loss function, the method further includes: obtaining relative importance information between the observation reconstruction task and the mask filling task, and determining the weight coefficient according to the relative importance information;

[0012] The determining of the total loss function based on the weight coefficient, the first loss function and the second loss function includes: based on the weight coefficient, fusing the first loss function and the second loss function to obtain the total loss function.

[0013] In one embodiment, the step of performing data segmentation, normalization, and missing data simulation on the carbon emission factor dataset to obtain an indicator mask, a missing mask, and multiple incomplete datasets includes:

[0014] The carbon emission factor data set is segmented according to a preset ratio to obtain an initial training set, an initial validation set, and an initial test set; the initial training set, the initial validation set, and the initial test set are normalized using the same normalizer; and missing data simulation is performed on the normalized training set, the normalized validation set, and the normalized test set according to a preset missing rate to obtain the training set, the validation set, and the test set.

[0015] In one embodiment, the method further comprises:

[0016] Obtain interpolation and completion data of the parameter-optimized deep learning model on the test set, and calculate the mean absolute error between the interpolation and completion data and the test set; determine the interpolation and completion accuracy based on the mean absolute error; if the interpolation and completion accuracy does not meet the threshold condition, estimate the time required to complete the training of the target deep learning model based on the difference between the interpolation and completion accuracy and the threshold condition.

[0017] In one embodiment, the method further comprises:

[0018] The target deep learning model after training is saved to the data storage system; in response to a secondary training instruction for the target deep learning model, the target deep learning model is reloaded from the data storage system according to the number of the target deep learning model, and the target deep learning model is trained for the second time.

[0019] Secondly, the present application also provides an interpolation and completion device for missing carbon emission factor data, comprising:

[0020] A data processing module is used to obtain a carbon emission factor dataset to be processed, perform data segmentation, normalization, and missing data simulation on the carbon emission factor dataset, and obtain a missing mask, an indicator mask, and an incomplete dataset containing missing values; the incomplete dataset includes a training set, a validation set, and a test set;

[0021] a function determination module, configured to determine a first loss function for an observation reconstruction task and a second loss function for a mask filling task based on the missing mask, the indicator mask, and the incomplete data set, and determine a total loss function based on a weight coefficient, the first loss function, and the second loss function;

[0022] A model training module is used to train and verify the initial deep learning model using the training set and the validation set, and optimize the model parameters of the initial deep learning model by minimizing the total loss function until the interpolation and completion accuracy of the deep learning model after parameter optimization on the test set meets the threshold condition, thereby obtaining a target deep learning model based on the self-attention mechanism and the diagonal mask mechanism; the target deep learning model is used to interpolate and complete missing carbon emission factor data.

[0023] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0024] Obtain a carbon emission factor dataset to be processed, perform data segmentation, normalization processing, and missing data simulation on the carbon emission factor dataset to obtain a missing mask, an indicator mask, and an incomplete dataset containing missing values; the incomplete dataset includes a training set, a validation set, and a test set; according to the missing mask, the indicator mask, and the incomplete dataset, determine a first loss function for the observation value reconstruction task and a second loss function for the mask filling task, and determine a total loss function based on the weight coefficient, the first loss function, and the second loss function; use the training set and the validation set to train and verify the initial deep learning model, and optimize the model parameters of the initial deep learning model by minimizing the total loss function until the interpolation and completion accuracy of the deep learning model after parameter optimization on the test set meets the threshold condition, thereby obtaining a target deep learning model based on the self-attention mechanism and the diagonal mask mechanism; the target deep learning model is used to interpolate and complete missing carbon emission factor data.

[0025] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0026] Obtain a carbon emission factor dataset to be processed, perform data segmentation, normalization processing, and missing data simulation on the carbon emission factor dataset to obtain a missing mask, an indicator mask, and an incomplete dataset containing missing values; the incomplete dataset includes a training set, a validation set, and a test set; according to the missing mask, the indicator mask, and the incomplete dataset, determine a first loss function for the observation value reconstruction task and a second loss function for the mask filling task, and determine a total loss function based on the weight coefficient, the first loss function, and the second loss function; use the training set and the validation set to train and verify the initial deep learning model, and optimize the model parameters of the initial deep learning model by minimizing the total loss function until the interpolation and completion accuracy of the deep learning model after parameter optimization on the test set meets the threshold condition, thereby obtaining a target deep learning model based on the self-attention mechanism and the diagonal mask mechanism; the target deep learning model is used to interpolate and complete missing carbon emission factor data.

[0027] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:

[0028] Obtain a carbon emission factor dataset to be processed, perform data segmentation, normalization processing, and missing data simulation on the carbon emission factor dataset to obtain a missing mask, an indicator mask, and an incomplete dataset containing missing values; the incomplete dataset includes a training set, a validation set, and a test set; according to the missing mask, the indicator mask, and the incomplete dataset, determine a first loss function for the observation value reconstruction task and a second loss function for the mask filling task, and determine a total loss function based on the weight coefficient, the first loss function, and the second loss function; use the training set and the validation set to train and verify the initial deep learning model, and optimize the model parameters of the initial deep learning model by minimizing the total loss function until the interpolation and completion accuracy of the deep learning model after parameter optimization on the test set meets the threshold condition, thereby obtaining a target deep learning model based on the self-attention mechanism and the diagonal mask mechanism; the target deep learning model is used to interpolate and complete missing carbon emission factor data.

[0029] The above-mentioned interpolation and completion method, device, computer equipment, computer-readable storage medium and computer program product for missing carbon emission factor data, by combining the observation value reconstruction task and the mask filling task for collaborative optimization, enables the initial deep learning model to simultaneously optimize the interpolation of missing data and the reconstruction of observed data during the training process, avoiding the problem of traditional methods only focusing on the missing part and ignoring the overall data quality, and can accurately restore the carbon emission factor in large-scale data and accurately reflect its temporal changes, thereby improving the accuracy of the interpolation and completion results. In addition, the present application adopts a deep learning model based on the self-attention mechanism and the diagonal mask mechanism, which can better capture the temporal nature and complex laws in the carbon emission factor data, avoid the stationary assumption of the data during the interpolation process, and thus better adapt to the dynamic characteristics of the carbon emission factor data. Secondly, the present application not only improves the model's ability to handle missing data through the joint training of the observation value reconstruction task and the mask filling task, but also enhances the robustness of the model by reconstructing the observed data. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0031] Figure 1 This is a diagram of an application environment for an interpolation method for completing missing carbon emission factor data in one embodiment;

[0032] Figure 2 1 is a flow chart of a method for interpolation and completion of missing carbon emission factor data in one embodiment;

[0033] Figure 3 A schematic diagram of a method for joint training of an observation reconstruction task and a mask filling task in one embodiment;

[0034] Figure 4 Schematic diagram of the algorithm framework of the SAITS (Self-Attention-based Imputation for Time Series) model in one embodiment;

[0035] Figure 5 A schematic diagram of the interpolation and completion effect of missing data of a carbon emission factor for one day in one embodiment;

[0036] Figure 6 A schematic diagram of a process for determining a loss function in one embodiment;

[0037] Figure 7 Schematic diagram of a flow chart of a method for interpolation and completion of missing carbon emission factor data in a specific embodiment;

[0038] Figure 8 This is a structural block diagram of a device for interpolating and completing missing carbon emission factor data in one embodiment;

[0039] Figure 9 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0041] The interpolation method for missing carbon emission factor data provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown in FIG. , the terminal can communicate with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be integrated on the server or placed on the cloud or other network servers. Figure 1 In the application environment shown, the terminal can be, but is not limited to, various personal computers, laptops, smart phones, and tablet computers. The server can be implemented as an independent server or a server cluster consisting of multiple servers.

[0042] In one embodiment, Figure 2 As shown in the figure, an interpolation method for missing carbon emission factor data is provided, which can be applied to Figure 1 The terminal in the method may include the following steps:

[0043] Step S201 : obtaining a carbon emission factor dataset to be processed, performing data segmentation, normalization, and missing data simulation on the carbon emission factor dataset, and obtaining a missing mask, an indicator mask, and an incomplete dataset containing missing values.

[0044] Specifically, the terminal first processes the carbon emission factor dataset, which is divided into a training set, a validation set, and a test set. The training set is used for model training, the validation set is used to adjust hyperparameters, and the test set is used for final model evaluation. During the data processing process, in order to avoid information leakage, the training set is first normalized to 0~1, and then the same normalizer is applied to the validation set and the test set. This ensures that no test set features are leaked during the data normalization process, avoiding the model's prior knowledge of the test set data. Next, the data is processed into a standard incomplete dataset format. During the simulation of missing data, some values ​​in the data are randomly replaced with missing values ​​according to a certain missing ratio to generate an incomplete carbon emission factor dataset. These missing values ​​are then completed through subsequent model training.

[0045] Step S202: Determine the first loss function of the observation reconstruction task and the second loss function of the mask filling task based on the missing mask, the indication mask, and the incomplete data set, and determine the total loss function based on the weight coefficient, the first loss function, and the second loss function.

[0046] It should be noted that in order to effectively solve the problem of data interpolation where large areas of carbon emission factor data are missing, such as Figure 3 As shown in Figure 2, this example proposes a joint training method that combines the observed reconstruction task (ORT) with the masked imputation task (MIT). The core idea of ​​this method is to enhance the robustness of the model by simultaneously optimizing the interpolation of missing data (MIT task) and the reconstruction of observed data (ORT task), thereby improving the accuracy of the interpolation results.

[0047] Specifically, the terminal first determines the first loss function of the observation value reconstruction task and the second loss function of the mask filling task based on the missing mask, the indication mask and the incomplete data set, and then determines the weight coefficient based on the relative importance information between the observation value reconstruction task and the mask filling task, and finally determines the total loss function based on the weight coefficient, the first loss function and the second loss function.

[0048] In step S203, the initial deep learning model is trained and verified using the training set and the validation set, and the model parameters of the initial deep learning model are optimized by minimizing the total loss function until the interpolation completion accuracy of the deep learning model after parameter optimization on the test set meets the threshold condition, thereby obtaining the target deep learning model based on the self-attention mechanism and the diagonal mask mechanism.

[0049] Among them, the target deep learning model based on the self-attention mechanism and the diagonal mask mechanism can be a SAITS model, such as Figure 4 As shown in the figure, the SAITS model consists of two DMSA (Diagonally-Masked Self-Attention) modules and a weighted combination block. Each DMSA block uses the self-attention mechanism to capture the dependencies of time series data and avoids self-interference between time steps by introducing diagonal masks. The specific steps are as follows:

[0050] ①Introducing the Self-Attention Mechanism:

[0051] First, the SAITS model uses the self-attention mechanism to process time series data to capture the dependencies between different time steps. The conventional self-attention mechanism uses the following formula to calculate the attention weight:

[0052]

[0053] in, 、 and represent query, key, and value vectors respectively, is the dimension of the vector, Represents the similarity between the query and the key. The self-attention mechanism generates attention weights by calculating the similarity between each pair of time steps and applies them to the corresponding values. Because self-attention is non-autoregressive, it overcomes the speed bottleneck of recurrent neural networks (RNNs) in processing long sequences through parallel processing.

[0054] ②Introducing the diagonal mask mechanism (DiagMasked Self-Attention, DMSA):

[0055] To enhance the time series interpolation capability, SAITS introduced a diagonal masking mechanism. DMSA improves upon the standard self-attention mechanism by restricting each time step from seeing itself through diagonal masking, allowing predictions to rely solely on data from other time steps. Specifically, DMSA is calculated as follows:

[0056]

[0057]

[0058] Through this mechanism, the model can only use past time step data to predict the data of the current time step, ensuring that the temporal dependency during interpolation and completion is better captured.

[0059] ③Weighted Combination Block:

[0060] In the SAITS model, the outputs of the two DMSA modules are weightedly combined using a weighted combination block to generate better learning representations. , the weighted combination block dynamically weights the synthesis based on the dependencies and missing information of the time steps and , generate the final learning representation, the specific process is as follows:

[0061] First, the attention weight (output from the second DMSA block) is obtained by averaging multiple heads :

[0062]

[0063] Then, combined with the missing mask Calculate weighting coefficients ,in is a Values ​​in range:

[0064]

[0065] then, and By weighting coefficient synthesis :

[0066]

[0067] Finally, by replacing missing values, the completed data is generated :

[0068]

[0069] in, Represents the element-by-element product (Hadamard product). This weighted combination mechanism can effectively combine missing data and observed data in the interpolation process, thereby improving the interpolation accuracy of the model until the threshold condition is met. Finally, by using the target deep learning model to interpolate the missing carbon emission factor data, we can obtain the following: Figure 5 The interpolation effect of missing data for carbon emission factors is shown.

[0070] In this embodiment, by combining the observation reconstruction task and the mask filling task for collaborative optimization, the initial deep learning model can simultaneously optimize the interpolation of missing data and the reconstruction of observed data during the training process, avoiding the problem of traditional methods that only focus on the missing part and ignore the overall data quality, and can accurately restore the carbon emission factor in large-scale data and accurately reflect its temporal changes, thereby improving the accuracy of the interpolation and completion results. In addition, the present application adopts a deep learning model based on the self-attention mechanism and the diagonal mask mechanism, which can better capture the temporal nature and complex laws in the carbon emission factor data, avoid the stationary assumption of the data during the interpolation process, and thus better adapt to the dynamic characteristics of the carbon emission factor data. Secondly, through the joint training of the observation reconstruction task and the mask filling task, the present application not only improves the model's ability to handle missing data, but also enhances the robustness of the model by reconstructing the observed data.

[0071] In one embodiment, Figure 6 As shown, in the above step S202, determining the first loss function of the observation reconstruction task and the second loss function of the mask filling task according to the missing mask, the indicator mask and the incomplete data set may include the following steps:

[0072] Step S301 : determining the actual missing data according to the missing mask and the incomplete data set, and generating simulated completion data according to the actual missing data.

[0073] Step S302: Determine a first loss function for the observation reconstruction task based on the real missing data, the simulated completed data, and the missing mask.

[0074] Step S303: Determine a second loss function for the mask filling task based on the real missing data, the simulated completed data, and the indication mask.

[0075] It should be noted that the joint training method in this embodiment includes two main tasks: the mask filling task (MIT task) and the observation value reconstruction task (ORT task), and these two tasks are jointly optimized through the loss function.

[0076] Among them, the goal of the MIT task is to effectively fill in missing values ​​based on known data and solve the problem of low interpolation accuracy for large-scale missing data. In traditional interpolation methods, missing data are often filled with simple constants or filled by interpolation of neighboring data, ignoring the temporal nature and dynamic changes of the data. Under the MIT framework, the temporal learning ability of the model is utilized to predict appropriate supplementary values ​​through the contextual information of the missing data. The core idea of ​​the ORT task is to enhance the learning ability of the model and avoid overfitting by constructing a task to reconstruct the observed data. Specifically, when training the model, it not only focuses on the interpolation of missing data, but also requires the model to be able to reconstruct data on the observed data, thereby optimizing the model performance globally. The key to the combination of ORT and MIT lies in the complementarity of the two tasks: the ORT task focuses on ensuring the reconstruction accuracy of the observed data, while the MIT task focuses on the interpolation of missing data.

[0077] During the training process, the original carbon emission factor dataset is first randomly masked to generate an incomplete dataset containing missing values. The mask matrix Used to mark which data are missing and which are observed:

[0078] Representation data is missing, Representation data It has been observed.

[0079] At this time, input data For incomplete data sets, is an incomplete dataset containing artificially randomly masked missing data.

[0080] (Missing mask) is used to indicate which data is missing (including missing data and missing data that are artificially masked). During training, the model learns to fill in missing values ​​through this mask, which is defined as:

[0081]

[0082] (Indicator mask) is used to indicate which data are artificially missing. During training, the indicator mask helps the model distinguish between missing parts in the original data and missing parts that are artificially generated. It is defined as:

[0083]

[0084] Specifically, the goal of the mask filling task is to fill the missing parts with the observed data. In each iteration, the model learns the temporal features of the missing data through the known data and the partially missing data. The loss function of the mask filling task is as follows:

[0085]

[0086]

[0087] in, Indicates the completion data output by the model, is the real missing data, Indicates the indication mask.

[0088] The goal of the observation reconstruction task is to reconstruct the observed data to ensure the model's processing accuracy of the observed data. The loss function of the observation reconstruction task is as follows:

[0089]

[0090] in, Indicates missing mask.

[0091] This loss function focuses on the reconstruction of the observed data and enhances the robustness of the model by minimizing the error of the observed data.

[0092] In one embodiment, before determining the total loss function based on the weight coefficient, the first loss function, and the second loss function, the method of the present application further includes the following steps:

[0093] Obtain the relative importance information between the observation reconstruction task and the mask filling task, and determine the weight coefficient based on the relative importance information;

[0094] In the above step S202, determining the total loss function based on the weight coefficient, the first loss function, and the second loss function may include the following steps:

[0095] Based on the weight coefficient, the first loss function and the second loss function are fused to obtain the total loss function.

[0096] Specifically, during the training process, the loss functions of the MIT task and the ORT task are optimized simultaneously, and the total loss function can be expressed as:

[0097]

[0098] in, is the weight coefficient, which is used to adjust the relative importance between MIT and ORT tasks.

[0099] In one embodiment, in step S201, the carbon emission factor dataset is segmented, normalized, and subjected to missing data simulation to obtain an indicator mask, a missing mask, and multiple incomplete datasets, which may include the following steps:

[0100] The carbon emission factor data set is split according to a preset ratio to obtain an initial training set, an initial validation set, and an initial test set; the initial training set, the initial validation set, and the initial test set are normalized using the same normalizer; and missing data simulation is performed on the normalized training set, the normalized validation set, and the normalized test set according to a preset missing rate to obtain a training set, a validation set, and a test set.

[0101] Specifically, the detailed steps are as follows:

[0102] ① Dataset segmentation:

[0103] First, the carbon emission factor dataset is divided into training set, validation set and test set. Assume that the original carbon emission factor dataset is a dataset containing A sequence of samples , the original carbon emission factor dataset is divided into the following proportions:

[0104] Training set: 70%, that is ,Include samples;

[0105] Validation set: 15%, i.e. ,Include samples;

[0106] Test set: 15%, i.e. ,Include samples;

[0107] These partitions ensure that unseen samples in the carbon emission factor dataset (test set) do not leak into the training process, thus avoiding the model’s prior knowledge of the test set data.

[0108] ②Data normalization:

[0109] In order to avoid the adverse effects of feature scale differences on model training, the training set, validation set, and test set need to be normalized. The goal of normalization is to scale the data to the range of [0, 1]. The specific operations are as follows:

[0110] For the training set , first normalize it:

[0111]

[0112] here, and The training set Then, the validation set and test set are normalized using the training set normalizer:

[0113]

[0114]

[0115] here, and It is the minimum and maximum values ​​of the training set, not the minimum and maximum values ​​of the validation set and test set. This approach can avoid the leakage of test set information to the training set.

[0116] ③Missing data simulation:

[0117] Before training the model, we need to simulate missing data. Assume that the dataset After normalization, some data in the training set, validation set, and test set need to be randomly missing according to a certain missing rate. For example , i.e. 80% of the data is missing), the missing part of each sample will be randomly masked as follows:

[0118]

[0119] in, is the missing mask, For the The sample in The missing data will be filled with special values ​​(such as NaN) to generate an incomplete data set:

[0120]

[0121] For the validation set and test set, similar operations are performed to generate missing data:

[0122]

[0123]

[0124] ④Data is organized into standard format:

[0125] Finally, the missing data and its corresponding missing mask and indicator mask are organized into a standard input format for subsequent model training and evaluation. The standard format of input data includes:

[0126] Incomplete data ;

[0127] Missing Mask ;

[0128] Indication mask , used to mark those missing values ​​added artificially.

[0129] Indication mask It can be expressed as the following calculation formula:

[0130]

[0131] This data collation provides a data foundation for subsequent joint training of ORT and MIT.

[0132] In one embodiment, the method of the present application may further include the following steps:

[0133] Obtain the interpolation and completion data of the parameter-optimized deep learning model on the test set, and calculate the mean absolute error between the interpolation and completion data and the test set; determine the interpolation and completion accuracy based on the mean absolute error; if the interpolation and completion accuracy does not meet the threshold condition, estimate the time required to complete the training of the target deep learning model based on the difference between the interpolation and completion accuracy and the threshold condition.

[0134] The threshold conditions may be preset by those skilled in the art according to actual needs.

[0135] Specifically, the terminal obtains the interpolation and completion data of the deep learning model after parameter optimization on the test set, and calculates the mean absolute error between the interpolation and completion data and the test set; determines the interpolation and completion accuracy based on the mean absolute error; judges whether the interpolation and completion accuracy meets the preset threshold condition. If the interpolation and completion accuracy does not meet the threshold condition, then estimates and displays the time required to complete the training of the target deep learning model based on the difference between the interpolation and completion accuracy and the threshold condition.

[0136] In one embodiment, the method of the present application may further include the following steps:

[0137] The target deep learning model after training is saved to the data storage system; in response to the secondary training instruction for the target deep learning model, the target deep learning model is reloaded from the data storage system according to the number of the target deep learning model, and the target deep learning model is trained for the second time.

[0138] Specifically, the terminal saves the target deep learning model after training to the data storage system for subsequent use; if the target deep learning model needs to continue training or optimization, the saved target deep learning model is reloaded from the data storage system according to the number of the target deep learning model, the model state is restored, and the target deep learning model is trained again.

[0139] In one embodiment, Figure 7 As shown, a method for interpolation and completion of missing carbon emission factor data in a specific embodiment is provided, which specifically includes the following steps:

[0140] Step S701: obtain the carbon emission factor data set to be processed, split the carbon emission factor data set according to a preset ratio to obtain an initial training set, an initial validation set, and an initial test set; use the same normalizer to normalize the initial training set, the initial validation set, and the initial test set; simulate missing data on the normalized training set, the normalized validation set, and the normalized test set according to a preset missing rate to obtain a training set, a validation set, and a test set.

[0141] Step S702: determine the real missing data based on the missing mask and the incomplete data set, and generate simulated completion data based on the real missing data; determine the first loss function of the observation value reconstruction task based on the real missing data, the simulated completion data and the missing mask; determine the second loss function of the mask filling task based on the real missing data, the simulated completion data and the indication mask.

[0142] Step S703: Obtain the relative importance information between the observation reconstruction task and the mask filling task, and determine the weight coefficient based on the relative importance information; based on the weight coefficient, fuse the first loss function and the second loss function to obtain the total loss function.

[0143] Step S704: Use the training set and validation set to train and validate the initial deep learning model, and optimize the model parameters of the initial deep learning model by minimizing the total loss function until the interpolation and completion accuracy of the deep learning model after parameter optimization on the test set meets the threshold condition, thereby obtaining a target deep learning model based on the self-attention mechanism and the diagonal mask mechanism; the target deep learning model is used to interpolate and complete the missing carbon emission factor data.

[0144] Step S705: Save the trained target deep learning model to the data storage system; in response to a secondary training instruction for the target deep learning model, reload the target deep learning model from the data storage system according to the number of the target deep learning model, and perform secondary training on the target deep learning model.

[0145] The beneficial effects brought about by the above embodiment are as follows:

[0146] 1) Improved accuracy of missing data interpolation: By combining ORT and MIT tasks, this application enables the SAITS model to simultaneously optimize missing data interpolation and observed data reconstruction during training. This joint training strategy not only focuses on filling missing data but also optimizes the reconstruction of observed data, avoiding the problem of traditional methods that only focus on missing data while ignoring overall data quality. Through the coordinated optimization of ORT and MIT, it can accurately recover carbon emission factors in large-scale data and accurately reflect their temporal changes.

[0147] 2) Overcoming the limitations of traditional interpolation methods: Compared to traditional interpolation methods (such as linear interpolation and spline interpolation), this application uses a deep learning model based on self-attention and diagonal masking mechanisms to better capture the temporal and complex patterns in carbon emission factor data. Traditional interpolation methods often assume that data changes smoothly, but carbon emission factor data is affected by multiple factors and its changing trends are relatively complex. The SAITS model, through the self-attention mechanism and the diagonal masked self-attention module (DMSA), avoids the assumption of stationarity during the data interpolation process, making it more adaptable to the dynamic characteristics of carbon emission factor data.

[0148] 3) Improved model robustness and interpretability: This application utilizes joint training on the ORT and MIT tasks to not only enhance the model's ability to handle missing data but also strengthen its robustness by reconstructing observed data. Compared to traditional deep learning models, SAITS can globally optimize both observed and missing data during interpolation, reducing the risk of overfitting. The model's self-attention mechanism provides excellent interpretability, effectively revealing temporal relationships and patterns in the data, helping decision makers better understand the dynamics of carbon emission factors.

[0149] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0150] Based on the same inventive concept, embodiments of the present application also provide an interpolation and completion device for missing carbon emission factor data, which is used to implement the interpolation and completion method for missing carbon emission factor data. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the interpolation and completion device for missing carbon emission factor data provided below can be found in the above-mentioned limitations of the interpolation and completion method for missing carbon emission factor data, and will not be repeated here.

[0151] In an exemplary embodiment, Figure 8 As shown, a device for interpolation and completion of missing carbon emission factor data is provided, which may include:

[0152] The data processing module 801 is used to obtain a carbon emission factor dataset to be processed, perform data segmentation, normalization, and missing data simulation on the carbon emission factor dataset, and obtain a missing mask, an indicator mask, and an incomplete dataset containing missing values; the incomplete dataset includes a training set, a validation set, and a test set;

[0153] A function determination module 802 is configured to determine a first loss function for the observation reconstruction task and a second loss function for the mask filling task based on the missing mask, the indicator mask, and the incomplete data set, and determine a total loss function based on the weight coefficient, the first loss function, and the second loss function;

[0154] The model training module 803 is used to train and verify the initial deep learning model using the training set and the validation set, and optimize the model parameters of the initial deep learning model by minimizing the total loss function until the interpolation and completion accuracy of the deep learning model after parameter optimization on the test set meets the threshold condition, thereby obtaining a target deep learning model based on the self-attention mechanism and the diagonal mask mechanism; the target deep learning model is used to interpolate and complete missing carbon emission factor data.

[0155] In one embodiment, the function determination module 802 is also used to determine the actual missing data based on the missing mask and the incomplete data set, and generate simulated completion data based on the actual missing data; determine the first loss function of the observation reconstruction task based on the actual missing data, the simulated completion data and the missing mask; and determine the second loss function of the mask filling task based on the actual missing data, the simulated completion data and the indication mask.

[0156] In one embodiment, the device may also include: a coefficient determination module, used to obtain the relative importance information between the observation value reconstruction task and the mask filling task, and determine the weight coefficient based on the relative importance information; a function determination module 802, also used to fuse the first loss function and the second loss function based on the weight coefficient to obtain the total loss function.

[0157] In one embodiment, the data processing module 801 is also used to split the carbon emission factor data set according to a preset ratio to obtain an initial training set, an initial validation set, and an initial test set; use the same normalizer to normalize the initial training set, the initial validation set, and the initial test set; and simulate missing data on the normalized training set, the normalized validation set, and the normalized test set according to a preset missing rate to obtain a training set, a validation set, and a test set.

[0158] In one embodiment, the device may also include: an error calculation module, used to obtain the interpolation and completion data of the deep learning model after parameter optimization on the test set, and calculate the mean absolute error between the interpolation and completion data and the test set; determine the interpolation and completion accuracy based on the mean absolute error; if the interpolation and completion accuracy does not meet the threshold condition, then estimate the time required to complete the training of the target deep learning model based on the difference between the interpolation and completion accuracy and the threshold condition.

[0159] In one embodiment, the device may also include: a secondary training module, used to save the target deep learning model after training to a data storage system; in response to a secondary training instruction for the target deep learning model, according to the number of the target deep learning model, the target deep learning model is reloaded from the data storage system, and the target deep learning model is trained for the second time.

[0160] Each module in the aforementioned interpolation and completion device for missing carbon emission factor data may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0161] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 9As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be implemented via Wi-Fi, mobile cellular networks, near-field communication (NFC), or other technologies. When executed by the processor, the computer program implements a method for interpolation and completion of missing carbon emission factor data. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0162] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0163] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0164] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0165] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0166] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0167] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0168] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0169] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for interpolation and completion of missing carbon emission factor data, characterized in that: The method comprises: Obtaining a carbon emission factor dataset to be processed, performing data segmentation, normalization, and missing data simulation on the carbon emission factor dataset to obtain a missing mask, an indicator mask, and an incomplete dataset containing missing values; the incomplete dataset includes a training set, a validation set, and a test set; Determine a first loss function for an observation reconstruction task and a second loss function for a mask filling task according to the missing mask, the indicator mask, and the incomplete data set, and determine a total loss function based on a weight coefficient, the first loss function, and the second loss function; The initial deep learning model is trained and verified using the training set and the validation set, and the model parameters of the initial deep learning model are optimized by minimizing the total loss function until the interpolation and completion accuracy of the deep learning model after parameter optimization on the test set meets the threshold condition, thereby obtaining a target deep learning model based on the self-attention mechanism and the diagonal mask mechanism; the target deep learning model is used to interpolate and complete missing carbon emission factor data.

2. The method according to claim 1, characterized in that The determining, based on the missing mask, the indicator mask, and the incomplete data set, a first loss function for an observation reconstruction task and a second loss function for a mask filling task, comprises: determining real missing data according to the missing mask and the incomplete data set, and generating simulated completed data according to the real missing data; Determining a first loss function for the observation reconstruction task based on the real missing data, the simulated completed data, and the missing mask; A second loss function for the mask filling task is determined based on the real missing data, the simulated completed data, and the indication mask.

3. The method according to claim 2, characterized in that Before determining the total loss function based on the weight coefficient, the first loss function and the second loss function, the method further includes: Obtaining relative importance information between the observation value reconstruction task and the mask filling task, and determining the weight coefficient according to the relative importance information; The determining of the total loss function based on the weight coefficient, the first loss function, and the second loss function includes: Based on the weight coefficient, the first loss function and the second loss function are fused to obtain the total loss function.

4. The method according to claim 1, wherein The carbon emission factor dataset is subjected to data segmentation, normalization, and missing data simulation to obtain an indicator mask, a missing mask, and multiple incomplete datasets, including: Splitting the carbon emission factor dataset according to a preset ratio to obtain an initial training set, an initial validation set, and an initial test set; Normalizing the initial training set, the initial validation set, and the initial test set using the same normalizer; Missing data simulation is performed on the normalized training set, the normalized validation set, and the normalized test set according to a preset missing rate to obtain the training set, the validation set, and the test set.

5. The method according to claim 1, wherein The method further comprises: Obtain interpolated and completed data of the parameter-optimized deep learning model on the test set, and calculate the mean absolute error between the interpolated and completed data and the test set; Determining the interpolation completion accuracy according to the mean absolute error; If the interpolation completion accuracy does not meet the threshold condition, the time required to complete the training of the target deep learning model is estimated based on the difference between the interpolation completion accuracy and the threshold condition.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Saving the trained target deep learning model to a data storage system; In response to a secondary training instruction for the target deep learning model, the target deep learning model is reloaded from the data storage system according to the serial number of the target deep learning model, and the target deep learning model is trained for the second time.

7. An interpolation and completion device for missing carbon emission factor data, characterized in that: The device comprises: A data processing module is used to obtain a carbon emission factor dataset to be processed, perform data segmentation, normalization, and missing data simulation on the carbon emission factor dataset, and obtain a missing mask, an indicator mask, and an incomplete dataset containing missing values; the incomplete dataset includes a training set, a validation set, and a test set; a function determination module, configured to determine a first loss function for an observation reconstruction task and a second loss function for a mask filling task based on the missing mask, the indicator mask, and the incomplete data set, and determine a total loss function based on a weight coefficient, the first loss function, and the second loss function; A model training module is used to train and verify the initial deep learning model using the training set and the validation set, and optimize the model parameters of the initial deep learning model by minimizing the total loss function until the interpolation and completion accuracy of the deep learning model after parameter optimization on the test set meets the threshold condition, thereby obtaining a target deep learning model based on the self-attention mechanism and the diagonal mask mechanism; the target deep learning model is used to interpolate and complete missing carbon emission factor data.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.