Carbon data distributed storage method based on pre-check

By introducing pre-verification and distributed storage technologies into carbon data storage, the problems of data security, processing efficiency and consistency in existing carbon data storage methods are solved, and efficient, secure and accurate carbon data management is achieved.

CN120066404APending Publication Date: 2025-05-30欧冶云商股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411983257.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing carbon data storage methods have problems such as low data security, low concentration, low processing efficiency and difficult to guarantee data consistency.

Method used

A carbon data distributed storage method based on pre-verification is adopted, and a distributed storage is carried out on multiple nodes by collecting carbon emission data, performing pre-verification of data verification algorithms, and according to data retrieval rules.

Benefits of technology

It improves the security and processing efficiency of data, ensures the accuracy and completeness of data, and ensures the security and reliability of data through user authentication and data retrieval rules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066404A_ABST
    Figure CN120066404A_ABST
Patent Text Reader

Abstract

The invention relates to a pre-check-based carbon data distributed storage method. The method comprises the following steps: collecting carbon emission related data; carrying out pre-check on the collected data by utilizing a data check algorithm, and eliminating wrong and incomplete data; storing the checked data on a plurality of nodes in a distributed manner according to a data retrieval rule to form a distributed storage network; and if the distributed storage network receives the user access request, performing user authentication on the user, if the user authentication is not passed, rejecting the user access request, and if the user authentication is passed, analyzing the user access request, and obtaining data requested by the user from a corresponding node in the distributed storage network according to the request content. Compared with the prior art, the distributed storage of the data is realized, and the security and the processing efficiency of the data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of carbon emissions, and in particular to a distributed storage method for carbon data based on pre-verification. Background Art

[0002] Carbon storage data is an important basis for evaluating the global carbon cycle and predicting climate change trends. By accurately measuring and recording carbon storage, we can better understand the Earth's climate system and provide a scientific basis for addressing climate change. Carbon storage data comes from multiple sources, including geological exploration, ecological monitoring, remote sensing satellites, etc. These data need to be processed and analyzed professionally before they can be converted into usable carbon storage information, so there are often some problems with the accuracy of carbon data. With the continuous increase of carbon storage data, the unified storage and management of data has become an important issue, and an efficient data management system needs to be established to ensure the integrity, security, and availability of data.

[0003] In existing carbon data storage methods, a centralized database is mainly used for data storage. This method has the following disadvantages:

[0004] 1. Low data security: The centralized database is vulnerable to attacks. Once the database is damaged, all data will be lost;

[0005] 2. Low data concentration: Most carbon data exists locally on users' devices. Due to privacy and security concerns, it is less likely to be provided to institutions for commercial operation;

[0006] 3. Low data processing efficiency: As the amount of data increases, the read and write speed of the centralized database will decrease significantly;

[0007] 4. Difficulty in ensuring data consistency: In the case of concurrent access by multiple users, it is difficult to maintain data consistency. Summary of the Invention

[0008] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a distributed storage method for carbon data based on pre-verification, which can achieve distributed storage of data while ensuring the accuracy and integrity of carbon emission data.

[0009] The purpose of the present invention can be achieved by the following technical solutions:

[0010] A distributed storage method for carbon data based on pre-verification, the method comprising:

[0011] Collect carbon emission-related data;

[0012] Use a data verification algorithm to perform pre-verification on the collected data, and eliminate incorrect and incomplete data;

[0013] The verified data is distributed and stored on multiple nodes according to the data retrieval rules to form a distributed storage network;

[0014] If the distributed storage network receives a user access request, the user is authenticated. If the user authentication fails, the user access request is rejected. If the authentication passes, the user access request is parsed, and the data requested by the user is obtained from the corresponding nodes in the distributed storage network according to the request content.

[0015] Furthermore, the process of pre-verification includes:

[0016] Perform a preliminary format check on the collected carbon emission-related data to ensure that the format of the data file meets the requirements;

[0017] Process and mark the data that meets the format requirements, including cleaning, verifying data integrity, checking data consistency, verifying data rationality, and detecting errors in the data;

[0018] Classify the processed data, and distinguish the accuracy levels of each data according to the marks on the processed data. The classification levels include accurate, partially accurate, and inaccurate;

[0019] Directly output the accurate data, and output the partially accurate or inaccurate data after supplementing or correcting according to the marks.

[0020] Even further, the cleaning includes:

[0021] Deduplicate the data and delete duplicate data records;

[0022] Eliminate the data with inconsistent formats, including missing key fields and incorrect field types.

[0023] Even further, the verification of data integrity includes:

[0024] Check whether the data is missing key fields, including timestamp, emission amount, and emission source. If the data is missing key fields, mark it as incomplete data. If the data is complete, mark it as complete data.

[0025] Even further, the checking of data consistency includes:

[0026] Compare the data from different sources, check whether there is inconsistent data for the same emission source in the same time period, and check whether the time series of each data is continuous. If the data is inconsistent or discontinuous, mark the data as inconsistent and conduct further investigation and correction. Otherwise, mark the data as consistent data.

[0027] Even further, the verification of data rationality includes:

[0028] Verify whether the carbon emissions in the carbon emission-related data are within a reasonable range, and whether there are negative values or abnormally high values in the carbon emissions. If there are negative values or abnormally high values, it is determined as unreasonable data, mark the data as unreasonable and conduct further investigation and verification. After verification, delete or correct the unreasonable data, otherwise mark the data as reasonable data.

[0029] Furthermore, the errors in the detected data include:

[0030] Check whether the calculation formula of the data is correct and whether the data entry complies with the specifications. If an incorrect calculation formula or non-compliant data entry is detected in the data, it is determined as incorrect data, and the data is corrected or deleted, otherwise the data is marked as correct data.

[0031] Furthermore, the process of user authentication includes:

[0032] Obtain the user's identity credentials, which include the username, password, and digital certificate;

[0033] Verify the user's identity credentials and check whether the user is in the access authorization list. If not, it is determined that the user authentication fails. If so, it is determined that the user authentication passes.

[0034] Furthermore, the data retrieval rules include:

[0035] Establish indexes based on the carbon emission data, including timestamp index and emission source index;

[0036] When carbon emission data needs to be obtained, display the data that meets the acquisition conditions and hide the data that does not meet the acquisition conditions.

[0037] Furthermore, the process of obtaining the data requested by the user includes:

[0038] Convert the user access request into a query statement or instruction for the distributed storage network, and extract the data content and scope that the user wants to query in the request;

[0039] Locate the storage nodes of the distributed storage network according to the data content and scope that the user wants to query in the request, and obtain the node address list of the storage nodes storing the relevant data;

[0040] Read the data that the user wants to query from the storage nodes according to the node address list;

[0041] Integrate the scattered data fragments into a complete data file or dataset according to the structure and order of the data;

[0042] Display the integrated data file or dataset in the form of charts or reports, and at the same time provide users with data analysis functions, including data sorting, filtering, and statistical analysis;

[0043] Record the user's access requests, access times, and accessed data, and store them in a log file. Regularly conduct audit analysis on the log file.

[0044] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0045] 1. The present invention stores data in different nodes according to data retrieval rules, realizing distributed storage of data, improving data security and processing efficiency. At the same time, a pre-verification mechanism is introduced to improve data accuracy and integrity; the user's access requests need to pass user authentication, further ensuring data security.

[0046] 2. The pre-verification mechanism of the present invention performs operations such as cleaning the input data, verifying data integrity, checking data consistency, verifying data rationality, and detecting errors in the data, ensuring data accuracy and integrity, and reducing decision-making mistakes that may be caused by data errors;

[0047] 3. After the present invention performs operations such as cleaning on the data, for the data classified as partially accurate or inaccurate, it is supplemented or corrected again according to the marks and then output. The scattered data is compared, supplemented, and corrected again, further ensuring the overall rationality and security of distributed storage.

[0048] 4. The present invention stores data in different nodes according to data retrieval rules, facilitating quick query of data according to the index, improving data security and reliability. At the same time, the form of the index also improves data processing efficiency, meeting the needs of large-scale carbon data storage and analysis;

[0049] 5. The present invention also records the user's access behavior and user data access situation, facilitating subsequent audit and data management. The subsequent regular audit analysis can ensure the compliance and security of data access. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] Example 1

[0053] This example aims to disclose a distributed carbon data storage method based on pre-verification. The method is as follows Figure 1 shown and includes:

[0054] Step S1, collect carbon emission related data.

[0055] Carbon emission data includes total carbon emission data, industry carbon emission data, carbon emission intensity and efficiency data, carbon trading and carbon market price data, and other carbon related data.

[0056] Total carbon emission data includes:

[0057] Total global and national carbon emissions, including the total annual carbon dioxide emissions globally and in each region, usually measured in hundreds of millions of tons or millions of tons.

[0058] Industry carbon emission data includes:

[0059] Carbon emissions in the energy industry, including carbon emissions in the extraction, transportation, and combustion of fossil fuels such as coal, oil, and natural gas;

[0060] Carbon emissions in the industrial industry, involving carbon emissions in the production processes of industries such as manufacturing, chemical industry, steel, cement, and non-ferrous metals;

[0061] Carbon emissions in the transportation industry, including carbon emissions from vehicles such as cars, airplanes, and ships, as well as carbon emissions in the construction and maintenance of transportation infrastructure;

[0062] Carbon emissions in the construction industry, involving carbon emissions in the production of building materials, building construction, building operation, and demolition;

[0063] Carbon emissions in the agricultural industry, including carbon emissions in agricultural planting, animal husbandry, and fishery, as well as carbon emissions from agricultural production materials such as chemical fertilizers and pesticides.

[0064] Carbon emission intensity and efficiency data includes:

[0065] Carbon emission intensity, which refers to the carbon emissions per unit of GDP or per unit of output, reflecting the carbon emission efficiency of economic activities;

[0066] Carbon emission efficiency, which refers to the ratio between carbon emissions and output or benefits, used to evaluate the utilization efficiency of carbon emissions.

[0067] Carbon trading and carbon market price data includes:

[0068] Carbon trading volume, which refers to the quantity of carbon emission rights traded in the carbon trading market;

[0069] The carbon trading price refers to the trading price of carbon emission rights, which reflects the market demand and supply situation of carbon emission rights.

[0070] Other carbon-related data includes:

[0071] Carbon sink data refers to the amount of carbon dioxide absorbed through measures such as afforestation and forest management, and is an important indicator for evaluating the effect of carbon emission reduction.

[0072] Carbon footprint data refers to the carbon emissions generated by a product or activity throughout its entire life cycle, including direct emissions and indirect emissions.

[0073] Step S2: Use a data verification algorithm to pre-check the collected data and eliminate incorrect and incomplete data.

[0074] In step S2, the process of pre-checking includes:

[0075] Conduct a preliminary format check on the collected carbon emission-related data to ensure that the format of the data file meets the requirements;

[0076] Process and mark the data that meets the format requirements, including cleaning, verifying the integrity of the data, checking the consistency of the data, verifying the rationality of the data, and detecting errors in the data;

[0077] Classify the processed data. According to the marks on the processed data, distinguish the accuracy levels of each data. The classification levels include accurate, partially accurate, and inaccurate;

[0078] Directly output the accurate data, and output the partially accurate or inaccurate data after supplementing or correcting according to the marks.

[0079] Among them, the purpose of cleaning is to clean the imported data and remove invalid, duplicate, or incorrectly formatted data records. The specific steps include:

[0080] Deduplicate the data and delete duplicate data records;

[0081] Eliminate the data with incorrect formats, including missing key fields and incorrect field types.

[0082] In this embodiment, cleaning is specifically manifested as using a data cleaning tool or script to automatically identify and delete invalid data according to the preset cleaning rules.

[0083] The purpose of verifying the integrity of the data is to ensure that each data record contains all necessary fields and information. The specific steps include:

[0084] Check whether the data is missing key fields, including timestamp, emissions, and emission sources. If the data is missing key fields, mark it as incomplete data; if the data is complete, mark it as complete data.

[0085] The purpose of checking data consistency is to avoid data contradictions and conflicts. The specific steps include:

[0086] Compare data from different sources to check if there are inconsistencies in data for the same emission source during the same time period, and check if the time series of each data is continuous. If the data is inconsistent or discontinuous, mark the data as inconsistent and conduct further investigation and correction; otherwise, mark the data as consistent data.

[0087] The purpose of verifying data rationality is to eliminate obviously abnormal or illogical data. The specific steps include:

[0088] Verify whether the carbon emissions in the carbon emission-related data are within a reasonable range, and whether there are negative or extremely high values for carbon emissions;

[0089] If there are negative or extremely high values, determine it as unreasonable data, mark the data as unreasonable and conduct further investigation and verification. After verification, delete or correct the unreasonable data; otherwise, mark the data as reasonable data.

[0090] The purpose of detecting errors in data is to avoid calculation errors or data entry errors. The specific steps include:

[0091] Check whether the calculation formula of the data is correct and whether the data entry complies with the specifications;

[0092] If an error in the data is detected, determine it as error data and correct or delete the data; otherwise, mark the data as correct data.

[0093] The purpose of classifying the processed data is to distinguish the accuracy and integrity levels of the data.

[0094] After performing the above operations of cleaning, verifying data integrity, checking data consistency, verifying data rationality, and detecting errors in data, the data will be marked with corresponding labels. If, after a series of processing, the data labels are complete data, consistent data, reasonable data, and correct data, classify the data into accurate data; if the data labels are incomplete data, inconsistent data, unreasonable data, and error data, classify the data as inaccurate data, and classify the remaining data as partially accurate data.

[0095] Although the data will be verified and corrected according to the actual situation after the above-mentioned cleaning, verifying the integrity of the data, checking the consistency of the data, verifying the rationality of the data, and detecting errors in the data, in order to ensure the overall rationality and security of the distributed storage, the data classified as partially accurate or inaccurate will be supplemented or corrected again according to the marks and then output.

[0096] Step S3: Store the verified data on multiple nodes according to the data retrieval rules to form a distributed storage network.

[0097] In this embodiment, the verified data is distributed and stored on the data owner nodes according to the data retrieval rules. The indexing rules can be constructed based on the primary key, attribute fields, etc. of the data. When the carbon emission data needs to be obtained, the data that meets the acquisition conditions is displayed, and the data that does not meet the acquisition conditions is hidden. For example, a timestamp index, an emission source index, etc. can be established for the carbon emission data to facilitate quick query of the data according to the time range or emission source, improving the security and reliability of the data.

[0098] Step S4: If the distributed storage network receives a user access request, authenticate the user. If the user authentication fails, reject the user access request. If the authentication passes, parse the user access request and obtain the data requested by the user from the corresponding nodes in the distributed storage network according to the request content.

[0099] The process of user authentication includes:

[0100] Obtain the identity credentials of the corresponding user, where the identity credentials include the username, password, and digital certificate;

[0101] Verify the user's identity credentials and check whether the user is in the access authorization list. If not, determine that the user authentication fails. If so, determine that the user authentication passes.

[0102] The process of obtaining the data requested by the user includes:

[0103] Convert the user access request into a query statement or instruction of the distributed storage network, and extract the data content and range that the user wants to query in the request;

[0104] Locate the storage nodes of the distributed storage network according to the data content and range that the user wants to query in the request to obtain a list of node addresses of the storage nodes storing the relevant data;

[0105] Read the data requested by the user from the storage nodes according to the list of node addresses;

[0106] Integrate the scattered data fragments into a complete data file or dataset according to the structure and order of the data;

[0107] Display the integrated data file or dataset in the form of charts or reports, and at the same time provide users with data analysis functions, including data sorting, filtering, and statistical analysis;

[0108] Record the user's access requests, access times, and accessed data, and store them in a log file. Regularly conduct audit analysis on the log file.

[0109] In another embodiment of the present invention, after extracting the data content and scope that the user wants to query in the extraction request, the user's access permission will also be queried. If the user's access permission is not sufficient to access the data content and scope that the user wants to query, the user's access request will be rejected.

[0110] Embodiment 2

[0111] On the basis of Embodiment 1, this embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the carbon data distributed storage method based on pre-verification as described above.

[0112] At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned carbon data distributed storage method based on pre-verification. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.

[0113] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0114] A computer-readable medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0115] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A carbon data distributed storage method based on pre-verification, characterized in that: The method comprises: Collect data related to carbon emissions; Use data verification algorithms to pre-check the collected data and eliminate erroneous and incomplete data; The verified data is distributed and stored on multiple nodes according to data retrieval rules to form a distributed storage network; If the distributed storage network receives a user access request, it will authenticate the user. If the user authentication fails, the user access request will be rejected. If the authentication passes, the user access request will be parsed and the data requested by the user will be obtained from the corresponding node in the distributed storage network according to the request content.

2. A method for distributed storage of carbon data based on pre-verification according to claim 1, characterized in that: The pre-verification process includes: Conduct a preliminary format check on the collected carbon emission-related data to ensure that the format of the data file meets the requirements; Process and mark data that meets format requirements, including cleaning, verifying data integrity, checking data consistency, verifying data rationality, and detecting errors in data; Classify the processed data, and distinguish the accuracy level of each data according to the marks on each data after processing, and the classification levels include accurate, partially accurate and inaccurate; Accurate data is directly output, and partially accurate or inaccurate data is supplemented or corrected according to the marks before being output.

3. A method for distributed storage of carbon data based on pre-verification according to claim 2, characterized in that: The cleaning comprises: De-duplicate the data and delete duplicate data records; Eliminate data that does not meet the required format, including missing key fields and field type errors.

4. A method for distributed storage of carbon data based on pre-verification according to claim 2, characterized in that: The integrity of the verification data includes: Check whether the data is missing key fields, including timestamp, emission amount and emission source. If the data is missing key fields, it will be marked as incomplete data. If the data is complete, it will be marked as complete data.

5. A method for distributed storage of carbon data based on pre-verification according to claim 2, characterized in that: The consistency of the checking data includes: Compare data from different sources to check whether there are inconsistencies in the data for the same emission source in the same time period, and check whether the time series of each data are continuous. If the data are inconsistent or discontinuous, mark the data as inconsistent and conduct further investigation and correction, otherwise mark the data as consistent.

6. A method for distributed storage of carbon data based on pre-verification according to claim 2, characterized in that: The rationality of the verification data includes: Verify whether the carbon emissions in the carbon emission related data are within a reasonable range, and whether there are negative or abnormally high values ​​in the carbon emissions. If there are negative or abnormally high values, it will be determined as unreasonable data, and the data will be marked as unreasonable and further investigated and verified. After verification, the unreasonable data will be deleted or corrected, otherwise the data will be marked as reasonable data.

7. A method for distributed storage of carbon data based on pre-verification according to claim 2, characterized in that: The errors in the detection data include: Check whether the calculation formula of the data is correct and whether the data entry complies with the specifications. If an error is detected in the calculation formula of the data or the data entry does not comply with the specifications, it is determined to be erroneous data and the data is corrected or deleted. Otherwise, the data is marked as correct data.

8. The method for distributed storage of carbon data based on pre-verification according to claim 1, characterized in that: The user authentication process includes: Obtaining the user's identity credentials, which include a user name, password, and digital certificate; Verify the user's identity credentials and check whether the user is in the access authorization list. If not, the user authentication is considered to have failed. If so, the user authentication is considered to have passed.

9. The method for distributed storage of carbon data based on pre-verification according to claim 1, characterized in that: The data retrieval rules include: Establish indexes based on carbon emission data, including timestamp index and emission source index; When carbon emission data needs to be obtained, data that meets the conditions for obtaining is displayed, and data that does not meet the conditions for obtaining is hidden.

10. A method for distributed storage of carbon data based on pre-verification according to claim 1, characterized in that: The process of obtaining the data requested by the user includes: Convert the user access request into a query statement or instruction of the distributed storage network, and extract the data content and scope that the user wants to query in the request; Locate the storage nodes of the distributed storage network according to the data content and range that the user wants to query in the request, and obtain a node address list of the nodes storing the relevant data; According to the node address list, read the data that the user wants to query from the storage node; Integrate scattered data fragments into complete data files or data sets based on the structure and order of the data; Display the integrated data files or data sets in the form of charts or reports, and provide users with data analysis functions, including data sorting, filtering and statistical analysis; Record user access requests, access time, and access data, store them in log files, and audit and analyze the log files regularly.