Enterprise carbon emission accounting system based on TR-OCR
By designing a TR-OCR-based enterprise carbon emission accounting system with integrated multi-module, the processing capability of the TR-OCR identification module is optimized, and the problem of complex background, overlapping text and special symbol picture recognition is solved, and the accurate accounting of enterprise carbon emissions and efficient generation of reports is achieved.
Patent Information
- Application Number
- CN202510239194.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-06
AI Technical Summary
The existing TR-OCR technology is difficult to accurately identify pictures containing complex backgrounds, overlapping texts or special symbols, resulting in data omissions and inconsistencies, affecting the accurate accounting of corporate carbon emissions.
A TR-OCR-based enterprise carbon emission accounting system is designed. Through the integration of data collection, classification, preprocessing, identification, extraction, verification and accounting modules, combined with the optimization of the TR-OCR identification module, it includes identification submodules with complex backgrounds, overlapping texts and special symbols, to improve the recognition accuracy, and ensure the integrity and consistency of data through the information verification module.
It realizes accurate identification of all information in pictures containing complex backgrounds, overlapping texts or special symbols, avoids data omissions, ensures accurate accounting of enterprise carbon emissions, and improves the efficiency and quality of reports.
Smart Images

Figure CN120106384A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of carbon emission accounting, and in particular to an enterprise carbon emission accounting system based on TR-OCR. Background Art
[0002] Corporate carbon emissions accounting refers to the process of statistics, accounting, analysis and evaluation of greenhouse gas emissions (mainly including carbon dioxide, methane, nitrous oxide, etc.) generated by enterprises. The corporate carbon emissions accounting system is a platform that integrates a variety of digital technologies, such as big data, the Internet of Things, artificial intelligence, etc. These technologies jointly support the comprehensive monitoring and management of carbon emissions by enterprises. The core functions of the system include data integration, intelligent analysis, decision support, and report generation. Through this system, enterprises can monitor carbon emissions data in real time, conduct statistical analysis, and generate carbon emissions reports that meet industry standards and policy requirements.
[0003] In the process of carbon emission accounting of enterprises, a large amount of data and information needs to be collected. TR-OCR, as an OCR (optical character recognition) technology, that is, deep learning enhanced optical character recognition technology, can help enterprises quickly and accurately extract these unstructured information (such as paper reports, charts, etc.), thereby accelerating the process of carbon emission accounting. And when compiling carbon emission reports, enterprises need to organize a large amount of data and results into structured ones. TR-OCR can assist enterprises in completing this task and improve the efficiency and quality of report compilation.
[0004] However, in actual use, when the image contains complex backgrounds, overlapping text or special symbols, TR-OCR technology often cannot recognize all the information in the image. For example, when the image background is similar to the text color or there are interfering elements, OCR technology will find it difficult to accurately recognize the text, resulting in data omissions, which will further lead to incomplete carbon emission accounting results and fail to fully reflect the company's carbon emissions. At the same time, due to TR-OCR recognition errors or other problems in the data extraction process, there may be inconsistencies between data from different sources or different time periods, affecting the accurate accounting of corporate carbon emissions. Therefore, it is urgent to improve this shortcoming. The present invention studies and improves the existing technology and its shortcomings, and provides an enterprise carbon emission accounting system based on TR-OCR. Summary of the invention
[0005] The purpose of the present invention is to provide an enterprise carbon emission accounting system based on TR-OCR to solve the problems raised in the above background technology.
[0006] To achieve the above purpose, the present invention provides the following technical solution: a corporate carbon emission accounting system based on TR-OCR, comprising: Data collection module: responsible for collecting various data related to carbon emissions calculations used by enterprises, including but not limited to energy consumption data, production data, and transportation data. The data sources include internal enterprise systems, external databases, and data provided by third-party institutions. Data classification module: responsible for classifying the collected data according to the form of the data, dividing the data into structured, semi-structured, and unstructured, providing a basis for subsequent data preprocessing and identification; Preprocessing module: responsible for targeted preprocessing according to the form of data, so as to better identify information later; TR-OCR recognition module: responsible for text recognition for unstructured data based on TR-OCR technology, used to accurately identify information in pictures containing complex backgrounds, overlapping text or special symbols; Information extraction module: responsible for extracting key data related to corporate carbon emissions from the identified text information, or directly extracting key data from structured and semi-structured data. These data will be used for subsequent carbon emissions accounting; Information verification module: responsible for verifying the extracted information. Once data errors or missing are found, the information will be re-identified and extracted. Carbon emission accounting module: responsible for calculating the carbon emissions of enterprises based on the extracted data using carbon emission accounting methods, and generating carbon emission inventories and accounting reports; Result output module: responsible for using visualization tools such as charts, line charts, pie charts, etc. to display the carbon emission accounting results to users in a visual way and generate carbon emission reports.
[0007] Furthermore, the data collection module connects with the internal enterprise system (such as ERP, MES, etc.) through the API interface to automatically collect data, or obtains relevant data from external databases or third-party organizations, such as climate change data, energy price data, etc., and provides a manual input function for users to upload or input data that cannot be collected automatically.
[0008] Furthermore, the data classification module uses predefined rules and regular expressions to perform automatic classification, and provides a user manual classification function for correcting the automatic classification when it is inaccurate; The predefined rules are used to identify different forms of data by predefining a series of rules. For example, by checking the file extension, content format or specific tags (such as HTML tags, table borders, etc.), it is determined whether the data is text, table or unstructured. The regular expression is a string that matches a specific pattern to identify data in a specific form. For example, a regular expression can be used to identify data containing specific table tags (such as 、etc.), thus classifying it as semi-structured data.
[0009] Further, the preprocessing module includes a first preprocessing unit, a second preprocessing unit and a third preprocessing unit; The first preprocessing unit is used for preprocessing structured data, specifically including: Text cleaning: remove irrelevant characters in the text, such as special symbols, punctuation marks, etc., and convert the text into a unified format, such as removing extra spaces, line breaks, etc., and identifying and correcting spelling errors or abbreviations in the text; Word segmentation and part-of-speech tagging: Perform word segmentation on the text to split it into meaningful words or phrases, and perform part-of-speech tagging on the word segmentation results to facilitate subsequent analysis and processing; Text vectorization: Use methods such as bag-of-words model, TF-IDF, word embedding, etc. to vectorize text and convert text data into numerical data for easy processing by machine learning algorithms; The second preprocessing unit is used for preprocessing the semi-structured data, and specifically includes: Data cleaning and organization: Identify and use interpolation, mean filling and other methods to fill missing values, delete, replace or correct abnormal data, and deduplicate data; Data format conversion: convert table data into a format suitable for subsequent processing, such as CSV, Excel, etc., to ensure that the column names and indexes of the data are clear and accurate; Data standardization and normalization: Use methods such as minimum-maximum normalization and zero-mean normalization to standardize or normalize numerical data to eliminate the impact of dimensional differences on model training; The third preprocessing unit is used for preprocessing unstructured data, specifically including: Grayscale processing: Convert color images into grayscale images to reduce data dimensions and computational complexity; Denoising: Use filters and other algorithms to smooth the image and remove noise from the image; Binarization: Convert the grayscale image into a binary image (black and white image). Specifically, a threshold is set and pixels above the threshold are set to white, and pixels below the threshold are set to black. In a binary image, each pixel has only two possible values: black (0) or white (255). The contrast between the characters and the background is higher, making the characters easier to detect and segment. The binary image has a smaller data volume, which can reduce the computational burden of the OCR algorithm.
[0010] Furthermore, the TR-OCR recognition module uses the open source code of TR-OCR and combines the characteristics of the carbon emission accounting system to carry out targeted secondary development, including adjusting algorithm parameters, optimizing processing procedures, etc. At the same time, by collecting a large amount of image data containing complex backgrounds, overlapping text and special symbols, TR-OCR is trained and optimized to improve its recognition accuracy in specific scenarios.
[0011] Furthermore, the TR-OCR recognition module includes the following submodules: Complex background recognition submodule: It is specially used to recognize and process images with complex backgrounds (such as textured backgrounds, gradient backgrounds, etc.). The complex backgrounds include but are not limited to the following implementation methods: a background segmentation algorithm is used to separate the text area from the background area to reduce the interference of the background on text recognition. Then, TR-OCR technology uses CNN to extract features from the separated images, uses CNN to learn the complex relationship between text and background, and automatically extracts the text area. At the same time, an attention mechanism is introduced to enable CNN to focus on the key text area in the image, further improving the recognition accuracy. Overlapping text recognition submodule: used to recognize and process images with overlapping text (such as intra-line overlap, inter-line overlap, etc.). Implementation method: Use image segmentation technology to separate overlapping text areas so that each character can be recognized separately. For overlapping characters that are difficult to separate, TR-OCR technology uses contextual information or machine learning algorithms to predict and correct them. At the same time, it performs multi-scale analysis on the image and observes the image at different resolutions and scales, so as to more accurately recognize overlapping text. Special symbol recognition submodule: used to recognize and process pictures containing special symbols (such as mathematical symbols, punctuation marks, currency symbols, etc.). Implementation method: Establish a character library containing various special symbols, and use an image matching-based algorithm to match and recognize special symbols in pictures. In addition to introducing existing special symbols, the character library also provides a customized dictionary, including specific symbols and vocabulary, to improve recognition accuracy.
[0012] Furthermore, the information extraction module uses regular expressions or natural language processing technology to extract key data, and provides a function for users to manually correct the extraction results. The data extraction process includes the following steps: S1. Preprocessing: Preprocess the recognized text information to remove irrelevant characters, word segmentation, and part-of-speech tags; S2, matching and extraction: using regular expressions or natural language processing techniques to match and extract strings containing carbon emission related data; S3, data sorting: sorting the extracted data, including removing duplicate data, converting data formats, etc.; S4. Manual review: The user checks and modifies the extracted data based on the actual situation.
[0013] Furthermore, the carbon emission accounting method includes the emission factor method and the mass balance method, which are as follows: 1) For large-scale and wide-ranging carbon emission accounting, such as carbon emission accounting for a country, province, city or large enterprise, the emission factor method is selected; The emission factor method calculates emissions by multiplying activity data (such as fossil fuel consumption) by the corresponding emission factor (greenhouse gas emission coefficient per unit of activity). The calculation formula is as follows: Greenhouse gas emissions (GHG) = activity data (AD) × emission factor (EF); In the formula, activity data (AD) refers to the amount of production or consumption activities that lead to greenhouse gas emissions, and emission factor (EF) is the coefficient corresponding to the activity level data, including carbon content per unit calorific value or elemental carbon content, oxidation rate, etc., which characterizes the greenhouse gas emission coefficient per unit production or consumption activity; 2) For carbon emission calculations of specific facilities and process flows, such as petrochemical enterprises and steel enterprises, the mass balance method is selected; The mass balance method is based on the law of conservation of mass. It calculates carbon emissions by analyzing the mass changes during greenhouse gas emissions. The calculation formula is as follows: Carbon dioxide emissions = (raw material input × raw material carbon content - product output × product carbon content - waste output × waste carbon content) × 44 / 12; Where 44 / 12 is the molecular weight ratio of carbon dioxide to carbon and is used to convert the mass of carbon to the mass of carbon dioxide.
[0014] Furthermore, the output of the carbon emission accounting module includes a carbon emission inventory and an accounting report; The carbon emission inventory lists the carbon emissions of enterprises in detail, including emissions from different emission sources and emission types and the proportion of total emissions, etc., which helps enterprises to clearly understand their own carbon emissions, identify major emitters and emission reduction potential, and generate a carbon emission inventory: the calculated carbon emissions are classified and sorted according to different emission sources, emission types, etc. to generate a carbon emission inventory; The carbon emission accounting report provides a detailed description and explanation of the carbon emission accounting process, including the selection of accounting methods, parameter settings, data sources and processing, etc., and provides accounting results and conclusions. The carbon emission accounting report is generated: based on the carbon emission inventory and accounting results, a carbon emission accounting report is written, including accounting methods, data sources, accounting process, accounting results, etc.
[0015] Furthermore, the carbon emission report includes: Total emissions: the total carbon emissions of an enterprise in a month, which is the basic data for assessing the carbon emission level of an enterprise; Emission intensity: The carbon emission efficiency of an enterprise is evaluated by calculating the carbon emissions per unit of output value, per unit of energy consumption or per unit of area. The lower the emission intensity, the more effective the carbon emission management of the enterprise is and the better its environmental performance is. Suggestions on emission reduction measures: Based on the results of carbon emission accounting, put forward targeted suggestions on emission reduction measures to help enterprises optimize their energy structure, improve energy efficiency, promote clean energy, etc., so as to reduce carbon emissions and achieve sustainable development; Other key information: including comparative analysis of carbon emission data (such as comparison with industry standards, competitors or historical data), carbon emission trend forecasts, carbon asset management recommendations, etc., to provide more comprehensive carbon emission management support.
[0016] The present invention provides a corporate carbon emission accounting system based on TR-OCR, which has the following beneficial effects: The present invention is composed of a data collection module, a data classification module, a preprocessing module, a TR-OCR recognition module, an information extraction module, an information verification module, a carbon emission accounting module and a result output module. The modules work closely together to form a highly integrated system, which can automatically complete the whole process from data collection to result output, reduce manual intervention, and improve work efficiency. The TR-OCR recognition module is optimized for pictures with complex backgrounds, overlapping texts or special symbols, and can accurately recognize such information. With the strict verification of the information verification module, the limitations of traditional OCR technology in processing such pictures are solved, and accurate recognition of all information in pictures containing complex backgrounds, overlapping texts or special symbols is achieved, data omissions are avoided, and accurate accounting of corporate carbon emissions is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a logical architecture diagram of a TR-OCR-based enterprise carbon emission accounting system of the present invention; Figure 2 This is a schematic diagram of a preprocessing module of a corporate carbon emission accounting system based on TR-OCR in the present invention; Figure 3 A schematic diagram of a TR-OCR recognition module of a TR-OCR-based enterprise carbon emission accounting system of the present invention; Figure 4 This is a schematic diagram of the data extraction process of a TR-OCR-based enterprise carbon emission accounting system of the present invention. DETAILED DESCRIPTION
[0018] The following embodiments of the present invention are described in further detail in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0019] like Figure 1-Figure 4 As shown, a carbon emission accounting system for enterprises based on TR-OCR includes: a data collection module, a data classification module, a preprocessing module, a TR-OCR recognition module, an information extraction module, an information verification module, a carbon emission accounting module and a result output module; Data collection module: collects various data related to the calculation of carbon emissions by enterprises, including but not limited to energy consumption data, production data, and transportation data, and the data sources include internal enterprise systems, external databases, data provided by third-party organizations, etc.; this module connects with the internal enterprise systems (such as ERP, MES, etc.) through API interfaces to automatically collect data, or obtain relevant data from external databases or third-party organizations, such as climate change data, energy price data, etc., and also provides manual input functions for users to upload or input data that cannot be collected automatically.
[0020] Data classification module: classifies the collected data according to the form of the data, and divides the data into structured, semi-structured, and unstructured, providing a basis for subsequent data preprocessing and identification; in this embodiment, the data classification module uses predefined rules and regular expressions to perform automatic classification, and provides a user manual classification function for correction when the automatic classification is inaccurate; Predefined rules: A series of rules are predefined to identify different forms of data. For example, by checking the file extension, content format or specific tags (such as HTML tags, table borders, etc.), it can be determined whether the data is text, table or unstructured. Regular expression: A string that matches a specific pattern to identify data in a specific form. For example, a regular expression can be used to identify data that contains specific table tags (such as 、 etc.), thus classifying it as semi-structured data.
[0021] Preprocessing module: performs targeted preprocessing according to the form of data, including a first preprocessing unit, a second preprocessing unit and a third preprocessing unit; The first preprocessing unit is used for preprocessing structured data, specifically including: Text cleaning: remove irrelevant characters in the text, such as special symbols, punctuation marks, etc., and convert the text into a unified format, such as removing extra spaces, line breaks, etc., and identifying and correcting spelling errors or abbreviations in the text; Word segmentation and part-of-speech tagging: Perform word segmentation on the text to split it into meaningful words or phrases, and perform part-of-speech tagging on the word segmentation results to facilitate subsequent analysis and processing; Text vectorization: Use methods such as bag-of-words model, TF-IDF, word embedding, etc. to vectorize text and convert text data into numerical data for easy processing by machine learning algorithms; The second preprocessing unit is used for preprocessing the semi-structured data, specifically including: Data cleaning and organization: Identify and use interpolation, mean filling and other methods to fill missing values, delete, replace or correct abnormal data, and deduplicate data; Data format conversion: convert table data into a format suitable for subsequent processing, such as CSV, Excel, etc., to ensure that the column names and indexes of the data are clear and accurate; Data standardization and normalization: Use methods such as minimum-maximum normalization and zero-mean normalization to standardize or normalize numerical data to eliminate the impact of dimensional differences on model training; The third preprocessing unit is used for preprocessing unstructured data, specifically including: Grayscale processing: converting a color image into a grayscale image to reduce data dimension and computational complexity. In this embodiment, the grayscale processing formula is based on the weighted average method, and the color values of the three channels of red, green, and blue are weighted averaged according to different weights to obtain the grayscale value. The weights in this embodiment are 0.299 (red), 0.587 (green), and 0.114 (blue). Denoising: Use algorithms such as filters to smooth the image and eliminate noise in the image. In this embodiment, the denoising process uses mean filtering, median filtering or Gaussian filtering, wherein mean filtering replaces the original pixel value by calculating the average value of the pixel values in the neighborhood around the pixel point, thereby smoothing the image; median filtering uses the median value of the pixel values in the neighborhood around the pixel point to replace the original pixel value, which is particularly effective in removing salt and pepper noise; Gaussian filtering uses a Gaussian function as a weight to perform a weighted average of the pixel values in the neighborhood around the pixel point, thereby achieving a smoothing effect; Binarization processing: converting the grayscale image into a binary image (black and white image), specifically by setting a threshold, setting the pixels above the threshold in the grayscale image to white, and setting the pixels below the threshold to black. In a binary image, each pixel has only two possible values: black (0) or white (255). The contrast between the characters and the background is higher, making the characters easier to detect and segment, and the data volume of the binary image is smaller, which can reduce the computational burden of the OCR algorithm. In this embodiment, the threshold selection method includes but is not limited to a fixed threshold method, an adaptive threshold method, and a histogram-based method, wherein the fixed threshold method is suitable for images with relatively stable background brightness; the adaptive threshold method dynamically sets the threshold according to the local features of the image; and the histogram-based method determines the optimal threshold by analyzing the grayscale histogram of the image.
[0022] TR-OCR recognition module: Based on TR-OCR technology, it performs text recognition on unstructured data to accurately identify information in pictures containing complex backgrounds, overlapping text or special symbols. In this embodiment, the TR-OCR recognition module uses the open source code of TR-OCR and combines the characteristics of the carbon emission accounting system to carry out targeted secondary development, including adjusting algorithm parameters, optimizing processing procedures, etc. At the same time, by collecting a large amount of picture data containing complex backgrounds, overlapping text and special symbols, TR-OCR is trained and optimized to improve its recognition accuracy in specific scenarios. The TR-OCR recognition module includes the following submodules: Complex background recognition submodule: It is specially used to recognize and process images with complex backgrounds (such as textured backgrounds, gradient backgrounds, etc.). Complex backgrounds include but are not limited to: Implementation method: Use background segmentation algorithm to separate text area from background area to reduce the interference of background on text recognition. Then TR-OCR technology uses CNN to extract features from the separated image, uses CNN to learn the complex relationship between text and background, and automatically extracts text area. At the same time, it introduces attention mechanism to make CNN focus on key text area in the image to further improve recognition accuracy; Overlapping text recognition submodule: used to recognize and process images with overlapping text (such as intra-line overlap, inter-line overlap, etc.). Implementation method: Use image segmentation technology to separate overlapping text areas so that each character can be recognized separately. For overlapping characters that are difficult to separate, TR-OCR technology uses contextual information or machine learning algorithms to predict and correct them. At the same time, it performs multi-scale analysis on the image and observes the image at different resolutions and scales, so as to more accurately recognize overlapping text. Special symbol recognition submodule: used to recognize and process pictures containing special symbols (such as mathematical symbols, punctuation marks, currency symbols, etc.). Implementation method: Establish a character library containing various special symbols, and use an image matching-based algorithm to match and recognize special symbols in pictures. In addition to introducing existing special symbols, the character library also provides a custom dictionary, including specific symbols and vocabulary, to improve recognition accuracy.
[0023] Information extraction module: extract key data related to the enterprise's carbon emissions from the identified text information, or directly extract key data from structured or semi-structured data. These data will be used for subsequent carbon emissions accounting. In this embodiment, the information extraction module uses regular expressions or natural language processing technology to extract key data, and provides the user with the function of manually correcting the extraction results. The data extraction process includes the following steps: S1. Preprocessing: Preprocess the recognized text information to remove irrelevant characters, word segmentation, and part-of-speech tags; S2, matching and extraction: using regular expressions or natural language processing techniques to match and extract strings containing carbon emission related data; S3, data sorting: sorting the extracted data, including removing duplicate data, converting data formats, etc.; S4. Manual review: The user checks and modifies the extracted data based on the actual situation.
[0024] In this embodiment, the regular expression matches a specific pattern in a string by defining a series of rules; in the information extraction module, the regular expression is used to match and extract strings containing carbon emission related data. For example, a regular expression can be used to match strings containing numbers, units (such as tons, kilowatt-hours, etc.) and emission types (such as carbon dioxide, methane, etc.); Natural language processing technology (NLP) can understand and process natural language texts, including word segmentation, part-of-speech tagging, named entity recognition, syntactic analysis and other aspects; in the information extraction module, NLP technology can extract key information from the text, such as company name, date, emissions, etc., and can also be used to understand the contextual relationship in the text, thereby more accurately extracting carbon emission-related data.
[0025] Information verification module: Verify the extracted information, and once data errors or missing are found, re-identify and extract; in this embodiment, the information verification module uses preset rules or algorithms to verify the data, and provides the user with the function of manually verifying and correcting the data; Carbon emission accounting module: Based on the extracted data, the carbon emission accounting method is used to calculate the carbon emissions of the enterprise, and a carbon emission inventory and accounting report are generated; the carbon emission accounting method includes the emission factor method and the mass balance method, as follows: 1) For large-scale and wide-ranging carbon emission accounting, such as carbon emission accounting for a country, province, city or large enterprise, the emission factor method is selected; The emission factor method calculates emissions by multiplying activity data (such as fossil fuel consumption) by the corresponding emission factor (greenhouse gas emission coefficient per unit of activity). The calculation formula is as follows: Greenhouse gas emissions (GHG) = activity data (AD) × emission factor (EF); In the formula, activity data (AD) refers to the amount of production or consumption activities that lead to greenhouse gas emissions, and emission factor (EF) is the coefficient corresponding to the activity level data, including carbon content per unit calorific value or elemental carbon content, oxidation rate, etc., which characterizes the greenhouse gas emission coefficient per unit production or consumption activity; 2) For carbon emission calculations of specific facilities and process flows, such as petrochemical enterprises and steel enterprises, the mass balance method is selected; The mass balance method is based on the law of conservation of mass. It calculates carbon emissions by analyzing the mass changes during greenhouse gas emissions. The calculation formula is as follows: Carbon dioxide emissions = (raw material input × raw material carbon content - product output × product carbon content - waste output × waste carbon content) × 44 / 12; Where 44 / 12 is the molecular weight ratio of carbon dioxide to carbon and is used to convert the mass of carbon to the mass of carbon dioxide.
[0026] In this embodiment, the output of the carbon emission accounting module includes a carbon emission inventory and an accounting report; The carbon emission inventory lists the carbon emissions of enterprises in detail, including emissions from different emission sources, emission types, and the proportion of total emissions, etc., which helps enterprises to clearly understand their own carbon emissions, identify major emitters and emission reduction potential, and generate a carbon emission inventory: the calculated carbon emissions are classified and sorted according to different emission sources, emission types, etc., to generate a carbon emission inventory; The carbon emission accounting report describes and explains the carbon emission accounting process in detail, including the selection of accounting methods, parameter settings, data sources and processing, etc., and provides accounting results and conclusions. The carbon emission accounting report is generated as follows: based on the carbon emission inventory and accounting results, a carbon emission accounting report is compiled, including accounting methods, data sources, accounting processes, accounting results, etc.; in this embodiment, the carbon emission accounting report includes total emissions (giving the total carbon emissions of the enterprise within one month), emission intensity (assessing the carbon emission efficiency of the enterprise by calculating the carbon emissions per unit of output value, unit energy consumption or unit area), emission reduction measures (proposing targeted emission reduction measures based on the carbon emission accounting results), etc., which provides a scientific basis for enterprises to formulate emission reduction measures and conduct carbon trading, and also helps to improve the transparency and credibility of enterprises.
[0027] Result output module: Use visualization tools such as charts, line graphs, and pie charts to display carbon emission accounting results to users in a visual way and generate a carbon emission report; the carbon emission report includes: Total emissions: the total carbon emissions of an enterprise in a month, which is the basic data for assessing the carbon emission level of an enterprise; Emission intensity: The carbon emission efficiency of an enterprise is evaluated by calculating the carbon emissions per unit of output value, per unit of energy consumption or per unit of area. The lower the emission intensity, the more effective the carbon emission management of the enterprise is and the better its environmental performance is. Suggestions on emission reduction measures: Based on the results of carbon emission accounting, put forward targeted suggestions on emission reduction measures to help enterprises optimize their energy structure, improve energy efficiency, promote clean energy, etc., so as to reduce carbon emissions and achieve sustainable development; Other key information: including comparative analysis of carbon emission data (such as comparison with industry standards, competitors or historical data), carbon emission trend forecasts, carbon asset management recommendations, etc., to provide more comprehensive carbon emission management support.
[0028] The embodiments of the present invention are given for the purpose of illustration and description, and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to better illustrate the principles and practical applications of the present invention and to enable those of ordinary skill in the art to understand the present invention and thereby design various embodiments with various modifications suitable for specific uses.
Claims
1. A corporate carbon emissions accounting system based on TR-OCR, including a data collection module, which is responsible for collecting various data related to the enterprise's carbon emissions calculation, and the data sources include the enterprise's internal system, external database, and data provided by a third-party organization, characterized in that: The enterprise carbon emission accounting system also includes a data classification module, a preprocessing module, a TR-OCR recognition module, an information extraction module and a carbon emission accounting module; The data classification module is responsible for classifying the collected data according to the form of the data, and classifying the data into structured, semi-structured, and unstructured; The preprocessing module is responsible for performing targeted preprocessing according to the form of the data; The TR-OCR recognition module is responsible for performing text recognition on unstructured data based on TR-OCR technology, and is used to accurately recognize information in pictures containing complex backgrounds, overlapping text or special symbols; The information extraction module is responsible for extracting key data related to corporate carbon emissions from the identified text information, or directly extracting key data from structured or semi-structured data; The carbon emission accounting module is responsible for calculating the carbon emissions of the enterprise based on the extracted data using the carbon emission accounting method, and generating a carbon emission inventory and accounting report.
2. According to claim 1, a corporate carbon emission accounting system based on TR-OCR is characterized in that: The system also includes an information verification module and a result output module; The information verification module is responsible for verifying the extracted information. Once data errors or missing are found, the information is re-identified and extracted. The result output module is responsible for using visualization tools to display the carbon emission accounting results to users in a visual manner and generate a carbon emission report.
3. According to claim 1, a corporate carbon emission accounting system based on TR-OCR is characterized in that: The data collection module connects with the internal system of the enterprise through an API interface to automatically collect data, or obtain relevant data from an external database or a third-party organization, and also provides a manual input function for users to upload or input data that cannot be automatically collected.
4. According to claim 1, a corporate carbon emission accounting system based on TR-OCR is characterized in that: The data classification module uses predefined rules and regular expressions to perform automatic classification, and provides a user manual classification function for correcting the automatic classification when it is inaccurate; The predefined rules are used to identify data in different forms by predefining a series of rules; The regular expression is a string that matches a specific pattern to identify data in a specific form.
5. According to claim 1, a corporate carbon emission accounting system based on TR-OCR is characterized in that: The preprocessing module includes a first preprocessing unit, a second preprocessing unit and a third preprocessing unit; The first preprocessing unit is used for preprocessing structured data, specifically including: Text cleaning: remove irrelevant characters from the text, convert the text into a unified format, and identify and correct spelling errors or abbreviations in the text; Word segmentation and part-of-speech tagging: Perform word segmentation on the text, split it into meaningful words or phrases, and perform part-of-speech tagging on the word segmentation results; Text vectorization: Use the bag-of-words model, TF-IDF, and word embedding to vectorize text and convert text data into numerical data; The second preprocessing unit is used for preprocessing the semi-structured data, and specifically includes: Data cleaning and organization: Identify and use interpolation and mean filling to fill missing values, delete, replace or correct abnormal data, and deduplicate data; Data format conversion: convert the table data into a format suitable for subsequent processing, ensuring that the column names and indexes of the data are clear and accurate; Data standardization and normalization: Use minimum-maximum normalization and zero-mean normalization to standardize or normalize numerical data to eliminate the impact of dimensional differences on model training; The third preprocessing unit is used for preprocessing unstructured data, specifically including: Grayscale processing: Convert color images into grayscale images to reduce data dimensions and computational complexity; Denoising: Use filter algorithms to smooth the image and remove noise from the image; Binarization: Convert a grayscale image into a binary image by setting a threshold, setting the pixels above the threshold to white and the pixels below the threshold to black.
6. The enterprise carbon emission accounting system based on TR-OCR according to claim 1 is characterized in that: The TR-OCR recognition module uses the open source code of TR-OCR and combines the characteristics of the carbon emission accounting system to carry out targeted secondary development. At the same time, TR-OCR is trained and optimized by collecting a large amount of image data containing complex backgrounds, overlapping text and special symbols; The TR-OCR recognition module includes the following submodules: Complex background recognition submodule: It is specially used to recognize and process images with complex backgrounds, such as textured backgrounds and gradient backgrounds. The implementation method is: the background segmentation algorithm is used to separate the text area from the background area. Then, the TR-OCR technology uses CNN to extract features from the separated image and automatically extract the text area. At the same time, the attention mechanism is introduced to make CNN focus on the key text area in the image. Overlapping text recognition submodule: used to recognize and process images with overlapping text, including intra-line overlap and inter-line overlap. The implementation method is: using image segmentation technology to separate overlapping text areas. For overlapping characters that are difficult to separate, TR-OCR technology uses context information or machine learning algorithms to predict and correct them. At the same time, multi-scale analysis is performed on the image to observe the image at different resolutions and scales. Special symbol recognition submodule: used to recognize and process pictures containing special symbols, including mathematical symbols, punctuation marks, and currency symbols. Implementation method: establish a character library containing various special symbols, and use an image matching-based algorithm to match and recognize special symbols in pictures. In addition to introducing existing special symbols, the character library also provides a customized dictionary, including specific symbols and vocabulary.
7. The enterprise carbon emission accounting system based on TR-OCR according to claim 1 is characterized in that: The information extraction module uses regular expressions or natural language processing technology to extract key data and provides the user with a function to manually modify the extraction results. The data extraction process includes the following steps: S1. Preprocessing: Preprocess the recognized text information to remove irrelevant characters, word segmentation, and part-of-speech tags; S2, matching and extraction: using regular expressions or natural language processing techniques to match and extract strings containing carbon emission related data; S3, data sorting: sorting the extracted data, including removing duplicate data and converting data formats; S4. Manual review: The user checks and modifies the extracted data based on the actual situation.
8. The enterprise carbon emission accounting system based on TR-OCR according to claim 1 is characterized in that: The carbon emission accounting methods include emission factor method and mass balance method, which are as follows: 1) For large-scale and wide-ranging carbon emission accounting, choose the emission factor method; The emission factor method calculates emissions by multiplying activity data by the corresponding emission factor. The calculation formula is as follows: Greenhouse gas emissions (GHG) = activity data (AD) × emission factor (EF); Wherein, activity data (AD) refers to the amount of production or consumption activities that lead to greenhouse gas emissions, and emission factor (EF) is the coefficient corresponding to the activity level data, including carbon content per unit calorific value or elemental carbon content, oxidation rate, and greenhouse gas emission coefficient per unit production or consumption activity; 2) For carbon emission calculations of specific facilities and process flows, select the mass balance method; The mass balance method is based on the law of conservation of mass and estimates carbon emissions by analyzing the mass changes during greenhouse gas emissions. The calculation formula is as follows: Carbon dioxide emissions = (raw material input × raw material carbon content - product output × product carbon content - waste output × waste carbon content) × 44 / 12; Where 44 / 12 is the ratio of the molecular weight of carbon dioxide to carbon and is used to convert the mass of carbon to the mass of carbon dioxide.
9. The enterprise carbon emission accounting system based on TR-OCR according to claim 1 is characterized in that: The output of the carbon emission accounting module includes a carbon emission inventory and an accounting report; The carbon emission inventory lists in detail the carbon emissions of the enterprise, including emissions from different emission sources and emission types and the proportion of total emissions. The carbon emission inventory is generated by classifying and sorting the calculated carbon emissions according to different emission sources and emission types to generate a carbon emission inventory; The carbon emission accounting report provides a detailed description and explanation of the carbon emission accounting process, including the selection of accounting methods, parameter settings, data sources and processing, and provides accounting results and conclusions. The carbon emission accounting report is generated as follows: a carbon emission accounting report is prepared based on the carbon emission inventory and accounting results.
10. The enterprise carbon emission accounting system based on TR-OCR according to claim 2 is characterized in that: The carbon emissions report includes: Total emissions: the total carbon emissions of an enterprise in a month; Emission intensity: The carbon emission efficiency of an enterprise is evaluated by calculating the carbon emissions per unit of output value, per unit of energy consumption or per unit of area. The lower the emission intensity, the more effective the carbon emission management of the enterprise is and the better its environmental performance is. Suggestions on emission reduction measures: Based on the results of carbon emission accounting, put forward targeted suggestions on emission reduction measures to help enterprises optimize their energy structure, improve energy efficiency and promote clean energy; Other key information: including comparative analysis of carbon emission data, carbon emission trend forecast, and carbon asset management recommendations.
Citation Information
Patent Citations
Regional greenhouse gas emission list accounting system and method based on artificial intelligence
CN113849542A
Method and system for measuring and calculating carbon emission of urban sewage treatment plant
CN119358819A
System and method for simulating and predicting forecasts for carbon emissions
EP4307189A1