Intelligent analysis method for carbon emission data
By adopting a method chain of receiving segmentation, element normalization, spatiotemporal positioning, and version locking, the auditability and reproducibility issues of row-level data in enterprise carbon accounting are solved. This enables stable mapping to unique emission factors and the generation of evidence chains, reducing review costs and improving the reliability and consistency of data processing.
Patent Information
- Application Number
- CN202511500236.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-10-21
AI Technical Summary
Existing technologies lack auditability and reproducibility of row-level data processing in enterprise carbon accounting. In particular, it is difficult to reliably map Chinese free-text invoices to a unique emission factor. Furthermore, the lack of row-level parallelism and consistency verification of location-based and market-based approaches in electricity consumption scenarios leads to high verification costs and difficulty in reproducibility.
By using a method chain that includes receiving segmentation, element normalization and dimensional alignment, spatiotemporal positioning and version locking, unique factor determination, and evidence chain solidification, each row is stably mapped to a unique emission factor, and evidence chain objects are generated to support auditable and reproducible carbon emission data analysis.
It achieves stable mapping of each row to a unique emission factor, high completeness of evidence chain fields and high reproducibility of results, significantly reduces review time, allows location method and market method to run in parallel at the row level and undergo consistency verification, naturally aligns report calibers, and controls operational risks.
Smart Images

Figure CN120975409B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of carbon emission data intelligent analysis, in particular to a carbon emission data intelligent analysis method. BACKGROUND
[0002] At present, enterprise carbon accounting relies on a pipeline of "business system export + general OCR / extraction + caliber rule engine + report summary". Data such as purchase invoices, warehouse entry lists and electricity meter readings are usually extracted into table fields first, and then emission factors are selected and evidence is supplemented in the report or audit stage. The emission factor library generally provides tabular data by region and year, and a small number of systems support basic retrieval and alternative factor selection; in the electricity consumption scenario, the calculation of the location method and the market method is usually parallel at the summary level, and the line level rarely maintains both calibrations simultaneously. Evidence materials (original ticket fragments, extraction instructions, conversion basis, factor sources) are usually archived in the form of attachments or notes, and version changes and caliber switching often lack fine-grained traces. Overall, the existing solution is more biased towards report and compliance presentation, and lacks support for "auditable and reproducible line-level data processing".
[0003] In real tickets, the goods description is mainly in Chinese free text, with many aliases and abbreviations, and the caliber of the supplier code is not the same. Key attributes such as material, specification, and measurement basis are often missing or ambiguous; units exist across levels such as "box, roll, kilogram, ton" conversion, density, specific gravity, and temperature and pressure conditions are not uniform; the quality of scanned versions varies, and positioning is unstable. When mapping such line items to "a unique emission factor (including region, year, version, caliber label)", the existing process is prone to problems such as multiple candidates being difficult to choose, insufficient version locking basis, and opaque downgrade replacement path; once a dispute arises, it is often difficult to review at the line level "why this factor was chosen, and where is the conversion basis". In terms of dual-caliber electricity, existing systems often only calculate in parallel at the summary end, lacking parallel and consistency checks for the same line record, resulting in insufficient comparability and reconcilability between subjects and accounts. In addition, line-level evidence is usually not "co-located" with the results, and the extraction traces, conversion tracks, factor sources and version paths lack structured traces, and manual review relies on experience allocation, lacking a shunting mechanism based on uncertainty, with high review costs and difficulty in reproducing in batches.
[0004] Based on the above status quo, there is still a key problem facing engineering landing: under the premise of not modifying the existing business system, how to form a processing method on the line level that can be landed, so that each ticket line can be stably and explainably mapped to a unique emission factor after the Chinese elements are normalized and dimensioned, and the auditable evidence chain (including original text positioning, extraction trace, unit conversion path, factor source and version locking / downgrade path) is synchronized and solidified, while the line level parallelism and consistency of place method and market method are realized in the electricity consumption scene, and the uncertainty indication quantity is used as the basis to trigger hierarchical review and rule trace updating. The existing public scheme generally lacks: version locking and downgrade path record at the line level, objectization of evidence chain and saving in the same domain as the result, step-by-step trajectory and conditional caliber of unit conversion, consistency management of line-level dual-caliber parallelism, and uncertainty-driven shunting and reproducible parameter threshold for batch monthly settlement. SUMMARY
[0005] In view of the deficiencies of the prior art, the present application provides a carbon emission data intelligent analysis method, which forms a line-level method chain by receiving segmentation→element normalization and dimension alignment→spatiotemporal positioning and version locking→unique factor determination→evidence chain solidification and uncertainty shunting, realizes stable mapping of each line to a unique emission factor, and the result is traceable, reproducible and auditable, thereby solving the problems mentioned in the background art.
[0006] To achieve the above object, the present application provides the following technical scheme: a carbon emission data intelligent analysis method, comprising:
[0007] S1, receiving business data and cutting by line, labeling account period, subject and delivery place, retaining original text segment and position index, forming line-level to-be-processed unit, receiving side establishing unified channel, stably accessing tickets from financial system and energy consumption account, receiving according to fixed rhythm and incremental rhythm; mapping ticket number and subject code one by one, unifying time and unit, and de-duplicating according to source priority and arrival order; filling in the missing account period according to the invoice month, mapping the delivery place according to the order of the consignee address, contract address and ticket address, and setting a temporary processing flag when the mapping fails; cutting by line according to the format and generating original text segment and position index, supplementing the account period, subject and delivery place, encapsulating as line-level to-be-processed unit containing batch number, line number and content summary, and delivering to the message queue with the power-free identifier composed of batch number and line number;
[0008] S2, element extraction is performed on Chinese goods description, name and attribute normalization is completed based on alias account book, and unit conversion trajectory is generated by unifying the unit of measurement;
[0009] S3, loading emission factor index containing region, year, version and caliber label, analyzing spatiotemporal information according to account period and delivery place, screening candidate factors and completing version locking;
[0010] S4, calculate matching points according to attribute matching degree, space-time consistency and version priority, output unique emission factor, generate review task and record judgment path when uniqueness is not reached;
[0011] S5, calculate row-level emissions with activity data and unique factors, and generate location method results and market method results in the same row at the same time, perform dimension and caliber consistency check and set flags;
[0012] S6, generate evidence chain object for each row, record text fragments, extraction traces, unit conversion tracks, emission factor sources and version locking path, build uncertainty quantification index, trigger review when threshold is exceeded, and write results back to alias account book and rule snapshot, establish evidence chain object for the row immediately after row-level results are generated, the object includes original text fragments and position index, element extraction traces, unit conversion tracks, emission factor sources and version locking path, person in charge signature and timestamp; generate content summary and write it into an incremental-only log, any modification is appended with a new version and the original version is preserved; set viewing and modification rights and only allow supplementary records; evidence chain and results are persisted in the same domain and partitioned by batch and date, which can be played back and located; batch number and row number are used as idempotent identifiers for transmission and access.
[0013] In a preferred embodiment, S2 includes:
[0014] Element extraction and normalization are performed on Chinese goods descriptions, and brand, material, specification, shape, model, and measurement unit are merged into standard names and attribute groups using an alias account book;
[0015] When normalization is difficult to determine, the context fields are called in order according to adjacent rows of the same ticket, the same batch and the same supplier, and contract materials;
[0016] Measurement units are unified to a reference dimension, unit conversion tracks are gradually generated according to conversion tables, sources and caliber are recorded, and the track number is written back to the row;
[0017] For cases where the model and unit are inconsistent, the material and shape are inconsistent, the specification is missing, and the context is insufficient, a gray marker is set and a review list is formed;
[0018] The alias account book and conversion table have versions and effective intervals, and historical normalization is not rolled back;
[0019] When multiple materials are listed, they are split into sub-items according to the separator and the order and position index are maintained;
[0020] The normalization results are stored in the normalization area for factor screening and version locking calls.
[0021] In a preferred embodiment, S3 includes:
[0022] The index includes regional identification, year, version number, caliber label, source identification and source verification fingerprint;
[0023] According to the account period mapping year and the delivery place mapping region and partition, the region and year are determined first, then the candidate factors are screened and the latest valid version is locked;
[0024] When the year entry is missing, only the latest year is downgraded and the locking path is recorded;
[0025] When the regional entry is missing, the hit is tried item by item according to the recorded partition adjacency sequence, and the locking level is recorded. When not hit, the approved alternative caliber is used and the responsibility information is written;
[0026] The candidate list is written into the factor reference area together with the version locking path, and the list number and path digest are backfilled to the corresponding row-level record.
[0027] In a preferred embodiment, S4 comprises:
[0028] After completing version locking, for each row record, first perform attribute anchor synonym merging, unit reference checking and space-time pre-checking, and eliminate candidates inconsistent with the account period and delivery place;
[0029] For the remaining candidates, the matching score is obtained by comprehensively considering the attribute fit degree, space-time consistency and version priority according to the preset weight;
[0030] When the difference between the highest and the second highest matching score reaches the preset threshold, the unique emission factor is determined, and the judgment path containing the participation field, the score, the weight, the reason for being eliminated and the version information is generated;
[0031] When the threshold is not reached, generate a review task and freeze the candidate list, weight, threshold and rule version, and mark the timestamp for the same version to be consistent when running again;
[0032] The unique emission factor number and the judgment path are written into the factor decision area and backfilled to the row-level record.
[0033] In a preferred embodiment, S5 comprises:
[0034] Based on the activity amount and the unique emission factor, the place method result and the market method result are generated simultaneously in the same row, and the caliber label is marked;
[0035] According to the contract account, the market method coverage amount is mapped by site, time period and calendar in the order of division, and then apportioned;
[0036] The residual amount not covered is apportioned along with the same factor caliber as the place method, and the row record apportionment path and caliber label are marked;
[0037] Cross-dimensional superposition is prohibited within the industry; the unit conversion trajectory that has been saved shall be used as the standard for unit uniformity.
[0038] In a preferred embodiment, a consistency check is performed on the location-based approach and the market-based approach results for the same entity and the same payment period;
[0039] In case of inconsistency, the attribution shall be determined according to the priority order: insufficient contract coverage, insufficient number of vouchers, inconsistent regional mapping, and abnormal measurement conversion. A pending status shall be generated in the bank and the attribution, allocation path and evidence fragment shall be written.
[0040] Green electricity certificates are verified for uniqueness by number and can only be used by a single entity.
[0041] All results and tags are stored in the database with batch number and row number as idempotent keys and can be used for subsequent evidence chain generation.
[0042] In a preferred embodiment, uncertain quantitative indicators are constructed based on the evidence chain object, including activity volume, emission factor, unit conversion, and caliber selection, and thresholds and priorities are set.
[0043] When any uncertain quantification exceeds its threshold, a review task is automatically generated and the main cause is highlighted. The current candidates and parameters are frozen. The review conclusion is written back to the alias ledger and rule snapshot and the version, coverage and effective time are recorded. It is then automatically adopted in subsequent similar rows.
[0044] Review tasks are processed according to grade and time limits and can be upgraded.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] 1. By integrating bill line items onto the same method chain—first, uniformly receiving and segmenting; then, unifying and aligning Chinese elements with dimensions; then, performing spatiotemporal positioning and locking the version based on payment period and delivery location; then, determining a unique factor based on attribute fit and spatiotemporal consistency; and finally, solidifying the original text fragments, location indexes, extraction traces, unit conversion trajectories, and factor source paths into evidence chain objects, and driving hierarchical review and rule reinjection with uncertainty thresholds—this addresses line-level pain points such as high noise in free Chinese text, missing attributes, inconsistent definitions, and difficulty in reviewing factor selection. It achieves the goal of "each line being stably mapped to a unique emission factor that is auditable, reproducible, and traceable": significantly reduced line-level mismatches (target less than one percent), evidence chain field completeness and result reproducibility approaching 100%, human-machine collaboration focusing review on high-uncertainty stages, significantly reducing review time (target no less than half), parallel implementation of the electricity usage scenario location method and market method at the line level with consistency verification, and natural alignment of report definitions.
[0047] 2. By embedding the method chain with engineering governance means such as idempotent identification, at least once arrival semantics, back pressure and elastic expansion, hot cache and sharding loading, version locking and ordered degradation, whitelist and tourniquet line, sampling and acceptance threshold, rule snapshot and effective interval, etc., the operational risks such as unstable delay in bulk processing, version drift, tamperable evidence, cross-currency and cross-border threshold deviation, etc. are solved, and the goals of "stable operation at scale + long-term compliance and auditability" are achieved; the end-to-end delay is controllable within the specified window (no more than ten minutes from receiving to packaging, no more than fifteen minutes for batch threshold verification), the online parallel capability meets the monthly settlement and peak scenario (row-level throughput is in the order of thousands to tens of thousands per minute), the evidence chain only increases traces and ensures that audit replay is available on demand, version withdrawal, annual gaps and regional gaps have clear classification and rollback paths; at the consistency level, the differences between the electricity site method and the market method are governed by subject layer threshold, the target pass rate is not less than 98%, and the dimension alignment accuracy is not less than 99.5%; overall, compliance and audit requirements are converted into executable "evidence object + threshold limit + backfilling closed loop", which not only stabilizes the on-site indicators, but also leaves a replicable foundation for subsequent expansion to more categories and subjects. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 The flowchart of the carbon emission data intelligent analysis method of the present application is given. DETAILED DESCRIPTION
[0049] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0050] Embodiment: Figure 1 The flowchart of the carbon emission data intelligent analysis method of the present application is given. The carbon emission data intelligent analysis method comprises:
[0051] S1, receive service data and cut by line, mark account period, subject and delivery place, keep original text segment and position index, form line-level to-be-processed unit, receive side establishes unified channel, stably accesses bill from financial system and energy consumption account, receives according to fixed rhythm and incremental rhythm; bill number and subject code are one-to-one mapped, unified time and unit, and de-duplicated according to source priority and arrival order; account period is filled according to bill month, delivery place is mapped according to delivery address, contract address and bill address in order, and a temporary processing flag is set when the mapping fails; the original text segment and position index are generated by cutting by line according to the format, the account period, subject and delivery place are supplemented, and the line-level to-be-processed unit containing batch number, line number and content summary is packaged, and the power equivalent identifier composed of batch number and line number is delivered to the message queue;
[0052] S2, factor extraction is performed on Chinese goods description, name and attribute normalization is completed based on alias account book, and unit conversion track is generated by unifying unit of measurement;
[0053] S3, load emission factor index containing region, year, version and caliber label, analyze space-time information according to account period and delivery place, screen candidate factors and complete version locking;
[0054] S4, calculate matching score according to attribute fitting degree, space-time consistency and version priority, output unique emission factor, generate review task and record judgment path when uniqueness is not reached;
[0055] S5, calculate line-level emission quantity by using activity data and unique factor, simultaneously generate site method result and market method result in the same line, perform dimension and caliber consistency check and set flag;
[0056] S6, generate evidence chain object for each line, record text segment, extraction trace, unit conversion track, emission factor source and version locking path, build uncertainty quantification index, trigger review when threshold exceeds, and write results back to alias account book and rule snapshot, establish evidence chain object for the line immediately after line-level result is generated, the object includes original text segment and position index, factor extraction trace, unit conversion track, emission factor source and version locking path, person signature and time stamp; generate content summary and write into only-increase log, any modification is appended with new version and original version is preserved; set viewing and modification division of power and only allow supplement; evidence chain and result are persisted in the same domain and partitioned by batch and date, which can be played back and located; batch number and line number are used as power equivalent identifier for transmission and access.
[0057] The technical connection and implementation logic of the six steps are as follows: first, in S1, the bill is cut by row, the account period, the main body and the delivery place are marked on each row, and the original text fragment and the layout position are stored together, and a "row-level to-be-processed unit" is packaged and put into the queue through batch number and row number; S2 retrieves the row along this identifier, restores the Chinese goods description into standard name and attribute group based on the alias account book, unifies the units according to the reference dimension and generates a traceable conversion track, and writes the two results back to the same row; with the account period and the delivery place, S3 does the space-time positioning in the emission factor index, screens out the candidates of the same region and the same year, and records the version locking and any downgrade attempt as a "locking path"; S4 reads the attribute group, conversion track and candidate list of the row, scores according to the established criteria of "attribute fitting degree + space-time consistency + version priority", selects one, generates a unique emission factor and a replayable "determination path", and if the gap cannot be pulled, generates a review task and freezes the parameters of the time; S5, after aligning the unique factor with the activity quantity, directly produces the place method and the market method in the row in parallel, and does difference threshold checking on the same main body and the same account period, and marks the to-be-checked if the line is touched; finally, S6 merges the original text fragment, the extraction trace, the unit conversion track, the factor source and the version locking / determination path precipitated in the previous five steps into an "evidence chain object", and calculates the activity quantity, the factor, the conversion and the caliber four uncertainty quantification indexes, and if the boundary is crossed, the review is pushed and the review conclusion is written back to the alias account book and the rule snapshot. The whole process uses "batch number + row number" as the idempotent key throughout, and the state and identifier of the previous step are the entrance of the next step; each step only appends a trace without covering the history, which not only ensures the row-level unique mapping and the double-caliber comparability, but also makes the evidence and the rules continuously converge in the closed loop.
[0058] S1, receive service data and cut by row, mark the account period, the main body and the delivery place, keep the original text fragment and the position index, form a row-level to-be-processed unit, and the receiving side establishes a unified channel to stably access bills from the financial system and the energy consumption account, receives according to fixed rhythm and incremental rhythm; the bill number and the main body code are one-to-one mapped, the time and the unit are unified, and the duplicates are removed according to the source priority and the arrival order; the account period is filled in according to the invoice month, the delivery place is mapped according to the delivery address, the contract address and the bill address in order, and a temporary processing flag is set when the mapping fails; cut by row according to the format and generate the original text fragment and the position index, supplement the account period, the main body and the delivery place, package as a row-level to-be-processed unit containing batch number, row number and content summary, and deliver to the message queue with the idempotent identifier composed of batch number and row number, the specific implementation is as follows:
[0059] A unified data access channel is established on the receiving side, and the relevant information of the invoice is stably accessed from the financial system and the energy consumption ledger without changing the rhythm of the existing business system. The parallel access rhythm of daily batch and hourly increment is adopted. Each invoice needs the following fields: invoice number, invoice date, account period, main body name and unified code, delivery place, cargo description text, quantity and unit, amount and tax; all units follow the enterprise benchmark system (quality in kilograms, volume in cubic meters, electricity in kilowatt hours, and number of pieces in pieces), the source is limited to two types of financial system and energy consumption ledger, the invoice date is allowed to differ by at most one day, and the quantity and amount are allowed to have an input error of within one percent, and exceeding it sets a flag to prompt. The format and value requirements of the key fields are as follows: the invoice number consists only of uppercase letters and numbers, the length is not more than thirty characters; the main body unified code is based on the unified social credit code, the length is eighteen characters; the account period uses the format of "year-month" of the Gregorian calendar; the delivery place uses the standard place name and carries the administrative division code; the amount and tax are rounded to the cent.
[0060] After the channel gets a batch, it first maps the invoice number and the main body code one by one, the time is unified to Beijing time, the units are merged according to the equivalence relationship, and the repeated records are kept only one according to "source priority in front, arrival time in back" (the priority of the financial system is higher than that of the energy consumption ledger, and the high priority and the time of the new one are used as the reference in case of conflict between sources). If the account period is missing, it is supplemented according to the month of the invoice date. If the delivery place is missing, it is mapped according to the order of "delivery address → contract address → invoice address", and if the mapping still fails, it is temporarily suspended and a "delivery place missing" is written. To avoid overload or starvation state, the receiving adopts two parallel rules of "whole point trigger" and "cumulative to one hundred rows trigger", and the first one runs; the single batch does not exceed one hundred thousand rows, and the threshold is self-adaptive according to the quantile statistics of the business distribution in the past three months; if the arrival rate drops below the historical lower limit within a ten-minute observation window, it automatically switches to a low-frequency mode (pulls a batch every thirty minutes), and returns to the normal frequency when the continuous two observation windows recover to more than the historical median.
[0061] A batch list is generated each time the processing is completed (including batch number, source mark, row number, time range), and the original content is completely landed as an "original mirror image", which is stored in UTF-8 character set, JSON per row, single row length not exceeding ten thousand characters, exceeding it is split and saved with a mapping table, which is kept for at least ninety days; if it is an electronic invoice with unstructured attachments, it is archived as is and linked with the batch number; if it is a scanned copy, it is first positioned and recognized, and only when the format stability is not less than 95% and the character recognition accuracy is not less than 97% can it be included in the process, otherwise the whole single is temporarily suspended for manual recording.
[0062] Subsequently, according to the bill format, each row is cut off, and the text fragment of each row and the position index in the original format are reserved. The position index is expressed in the coordinate system of "page number from one count + character offset". The mapping table is generated synchronously when the cross-page merging is performed. Immediately after cutting, the row entries are supplemented with the period, the subject and the delivery place. When the period is missing, the billing month is supplemented. When the delivery month is clear in the contract, the contract is used as the reference. The subject is marked according to the unified code. When encountering an alias, it is merged first. When the delivery place is inconsistent in the three sources, the receiving address is used as the reference, and the conflict is recorded in the notes of the batch list. After these are completed, the row number, batch number, period, subject, delivery place, text fragment, quantity and unit, amount, text fragment and position index are packaged into a row-level processing unit, written into the row-level storage area, and the "system source + batch number + row number + content hash" is used as an idempotent identifier to deliver to the message queue, and the downstream pulls according to this. The message channel uses this idempotent identifier on both sides of receiving and delivering to prevent duplicate consumption. Repeat events are recorded in the de-duplication log. The de-duplication window is fixed at seven days.
[0063] The whole link is from receiving to packaging. Under normal circumstances, it does not exceed ten minutes. At the same time, at least ten batches are processed in parallel. The queue back pressure threshold is set to more than twenty processed batches, more than fifteen minutes of average waiting time or more than eighty percent of memory utilization. Any hit will expand by two execution instances at a time. If the threshold is lower for two consecutive observation windows, it will be recycled. Any step fails and retries three times at intervals of one, three and five minutes. If it still fails, the batch is directly sent to the manual channel and is prohibited from being automatically triggered again until the manual confirmation is restored.
[0064] The time dimension is uniformly recorded in Beijing time. Cross-border bills will also retain the original time zone and conversion timestamp. If summer time is involved, it will be executed according to local rules on the conversion day. The amount is recorded with tax by default, and the tax rate and currency are also written. If conversion is involved, the exchange rate is taken from the enterprise financial main number warehouse at the closing price of the day and the version number is locked. The judgment of "one percent error" is calculated based on the original precision. If the relative error exceeds the threshold, the row is marked as "precision in doubt" and entered into the manual checking list. The page number and character offset of the position index require playback positioning. If playback fails, a flag is set and must be completed before being disclosed externally.
[0065] All abnormal conditions in the whole process are uniformly coded. The reason code is fixed as six categories of missing subject, missing delivery place, duplicate number, date out of bounds, precision in doubt, and unstable format, with an extension bit reserved. The first discovery time and the person in charge are written for each record. If the proportion of single batch temporary processing exceeds five percent, an alarm will be sent immediately. If any of the indicators such as temporary processing proportion, back pressure or time delay exceeds the threshold for two consecutive observation windows, it will be upgraded to the on-duty person in charge through enterprise instant messaging and email channels until the processing loop is closed.
[0066] All batch lists, temporary processing reasons, warnings and manual processing records are retained for at least two years, supporting retrieval by batch number, line number, source system, reason code, person in charge, and timestamp. Access follows the principle of minimum necessary authority and records access audit. To ensure that the range can be checked, no less than 10,000 lines are monthly sampled using stratified random sampling (stratified by source system, ticket type, and amount interval), and the coverage rate, field completeness rate, de-duplication accuracy rate, and end-to-end delay compliance rate of the sampled samples are calculated. The rectification list is formed within ten working days and fed back to the rule library; the coverage rate target is not less than 99%, the field completeness rate is not less than 95%, the de-duplication accuracy rate is not less than 99.5%, and the end-to-end delay compliance rate is not less than 95%.
[0067] In practical applications, the following scenarios can be referred to: On a certain day, a forty-line steel invoice is received in addition to the monthly electricity list, and the system combines the two into a batch into the receiving process, and falls into the original mirror according to the above de-duplication, merging, and completion rules, and cuts each line and records the text segment and character position, extends the account period, subject, and delivery place, and encapsulates more than forty line-level small units to deliver to the downstream; one line is missing the delivery place and is supplemented once according to the contract account and marked "supplemented once", and two lines are suspended for temporary processing due to low scanning quality below the threshold value, and manual recording is required; the whole process delay of this batch falls within ten minutes, and during this period, elastic expansion is triggered once due to the queue reaching the high threshold value; after the batch is completed, it participates in monthly sampling, and the indicators are above the red line.
[0068] So far, the three main lines of fields, time, and region have been aligned since receiving, and the original text evidence and positioning information have been completely deposited; once the name is normalized, the factor is selected or the uniqueness is determined, and the batch list, the original mirror, and the line-level position index can be found back along the identifier, which quickly locates and corrects the problem in a small range, preserves the time limit, and meets the requirements of traceability, reproducibility, and checkability.
[0069] S2, element extraction is performed on Chinese goods descriptions, name and attribute normalization is completed based on the alias account book, unit of measurement is unified, and unit conversion track is generated, and the specific implementation is:
[0070] In the process of element extraction and name and attribute normalization of Chinese goods description, this step aims to restore the description in colloquial, abbreviated, and even mixed misspelled form in the first-line ticket to a comparable and reviewable standard item, which is convenient for subsequent factor screening, version locking, and evidence chain generation. The scope covers material, equipment, and service entries. The text is received on-site according to the rhythm consistent with the upstream line-level record, and six types of elements, including brand, material, specification, shape, model, and measurement unit, are fixedly collected. The unit is always in the enterprise standard system (mass in kilograms, volume in cubic meters, electric quantity in kilowatt-hours, and piece in pieces), the source is mainly based on financial document text, supplier commodity catalog, and enterprise commonly used list, and is cross-verified with historical procurement records. Missing of a single row key field is allowed once, but continuous missing is not allowed.
[0071] Standardization before going online: unified coding, cleaning of non-printable characters and extra spaces, handling of mixed full and half-width characters, equivalent replacement of common oral writing (e.g., "φ two" as "diameter two millimeters", "stainless steel 316" as "316 stainless steel"), and alignment of unit tables. The unit conversion table fixedly records the conversion starting point and ending point, reference source, conversion accuracy and tolerance, and applicable product category. The accuracy is generally retained to three decimal places, and the tolerance is one percent by default. Preferably, supplier specifications and contract provisions are used, and for lack of evidence, general conversion is confirmed by sampling weighing and reconciliation, and once confirmed, the sample source is written into the conversion table.
[0072] Normalization action is promoted by "alias account book": the alias account book is used to merge colloquial Chinese goods description into a standard name and attribute group of word entry collection, including alias, standard name, material / specification points, common units, version number, effective interval, and source record. The account book is maintained weekly, and the fields fixedly include alias, standard name, material points, specification points, shape category, common units, measurement reference, common misspellings, applicable scope, source record, version number, reviewer, and effective time. Any new or modified record is used for on-site analysis after the effective time, and the in-stock line before the effective time is not recalculated to avoid historical result drift. On-site analysis prioritizes hitting the account book; if not enough, use the context field linkage to supplement the judgment in the order of "same ticket adjacent row → same batch same supplier → contract material column". If any level clue is not enough to form a stable conclusion, stop immediately, and do not make cross-layer guesses.
[0073] To ensure pacing, a one-second limit is set for single-line parsing. An interpretable threshold is set to measure normalization reliability, using a percentage scale ranging from 60% to 85%. The initial value is taken from the quantiles of stable samples over the past quarter, and is automatically fine-tuned using rolling samples on the first working day of each month. If a supplier's writing style changes abruptly within the same batch, a local adjustment is performed on the judgment threshold for that batch without altering the global threshold. This adjustment only takes effect within that batch, and reverts to the monthly baseline after the batch ends. If three consecutive lines fail to reach the threshold, a manual assistance and rule convergence task is automatically triggered, generating a list of supplementary terms and example sentences.
[0074] Unit conversions require step-by-step recording of the path and basis, such as "box → roll → kilogram," with each step clearly stating the source, precision, and tolerance. These paths, along with standard names and attribute groups, are stored in a unified area, and the trajectory number and attribute group are written back to the row-level record for direct reference later, eliminating the need for repeated reasoning. Throughput boundaries are clearly defined: no less than 3,000 lines per minute per machine; automatic expansion occurs when backlog exceeds the threshold; single-line parsing failures can be retried twice, and those still failing to meet the standard are added to the to-do list.
[0075] When obvious conflicts occur, they are uniformly grayed out. Graying out is divided into three categories: contradiction between model and unit, contradiction between material and shape, and missing specifications with insufficient context. For each category, the trigger words, the location of occurrence, and suggested supplementary items are recorded. When the proportion of grayed-out items in a single batch exceeds 5%, the system automatically generates a rule adjustment work order, which includes a detailed list of new aliases, corrected rules, and supplementary conversion items. After review, a new version is formed. Issues found in monthly spot checks are included in the monthly convergence task. The generated new version takes effect on the first day of the following month and the effective time is written in. On-site switching is based on this.
[0076] The normalization results are stored in the normalization area using key-value pairs. Fields include the standard name, attribute group, and unit conversion trajectory number. Quick retrieval by batch number and row number is supported. The unit conversion trajectory unifies activity volume into a step-by-step conversion path to the baseline unit, recording the source / target unit, basis, precision and tolerance, confirmer, and time. Downstream users can retrieve the standard item and trajectory number for a row using only the batch number and row number, avoiding duplicate parsing. Service items use lightweight normalization, extracting three fixed items: service type, billing unit, and billing quantity, aligned with the unit thesaurus. When multiple units are listed in the same row, the measurement baseline is used as the primary unit, and other units are converted step-by-step according to the conversion table order, with the complete path recorded. If the path is broken at an intermediate node, the conversion stops and is marked as pending confirmation.
[0077] To avoid caliber discrepancy, when the same expression appears in multiple available records in the alias account book, the supplier-specific item is preferred to the industry-wide item, which is preferred to the enterprise-wide item; the one with a more recent effective time is preferred, and if the priority and the effective time are the same, the source credibility is determined according to the order of the contract appendix, the supplier specification, and the historical reconciliation. When involving cross-dimension conversion such as length and mass, volume and mass, the material density and specific gravity library is called, the density value is subdivided according to the material and specification and the source and temperature condition is marked; when the density is missing, it is degraded according to the order of "historical weighing of the same material and specification → interpolation of the same material and adjacent specification → industry public reference value", and the degradation level is marked in the conversion path.
[0078] When encountering multiple materials in the same row (using the hyphen, semicolon, Chinese comma, and hyphen as separators), first split them into sub-items according to the separators, then perform normalization and conversion respectively, and keep the original order number and position index, and append a merged row marker at the end of the row for tracing. The numerical writing is all standardized to the enterprise standard: Chinese numerals, full-width numerals, and thousandth place and multiplier words are uniformly converted to Arabic numerals and base units; if obvious misrecords are found (such as decimal point position exceeding the limit, quantity and unit logic not matching), they are temporarily suspended for processing and recording of rectification suggestions, and are written back after manual confirmation.
[0079] The alias account book and unit conversion table support the invalidation mark and the effective interval, and the invalidation does not affect the historical records that have been normalized before the effective time; any addition, modification, and invalidation generates a change record and keeps the auditor, time, and effective range, and the minimum retention period is not less than five years. At the same time, the running hard boundary is clearly defined: the default temperature of the density and specific gravity library is 25 degrees Celsius, the default error upper limit is 5%, the update rhythm is once a month, and the new value only works for newly added rows after the effective time; the release time point of the alias account book and unit conversion table is fixed at 0 o'clock on Tuesday and 0 o'clock on the first day of each month, both with invalidation mark and effective interval, and historical normalized records are not rolled back; the separator priority of multiple material splitting is in the order of hyphen, semicolon, Chinese comma, and hyphen; the fixed enumeration of service items is installation, maintenance, detection, and transportation, the charging unit is based on times, man-hours, sets, and kilometers, and the supplier-defined unit is converted according to the conversion table path and leaves a trace. In this way, the on-site personnel can clearly know which information to take, how to read according to what caliber, how to clean and normalize, when to hand over to human review, where to place the normalized records, and how to directly use them by the downstream. The threshold range, update rhythm, version effectiveness, gray dictionary, multiple units, and cross-dimension boundaries are also clearly defined, and the same row can be independently implemented and stably reproduced based on this.
[0080] S3, load the emission factor index with region, year, version, and caliber tags, analyze the space-time information according to the account period and delivery location, filter the candidate factors and complete version locking, the specific implementation is:
[0081] In loading the emission factor index containing regional, annual, version and caliber labels, the available range is first compressed to the correct spatio-temporal coordinates, and then the specific version is completely locked, which facilitates subsequent row-level unique determination and clean landing. The index is maintained according to "geographical identification, year, version number, caliber label, source identification, source verification fingerprint". The source is simultaneously connected to the enterprise factor library and the authoritative release library, and is synchronized normally according to year or season; in order to ensure clear evolution, new and old versions are kept in parallel and the inheritance relationship is indicated. The geographical name is unified to the standard of the national administrative region, and then mapped to the partition identification according to the power grid partition table. Common aliases are first merged into the same entry. The emission factor refers to the coefficient for converting activity into greenhouse gas emissions, which has attributes such as dimension, caliber label, geographical identification and year. In this application, the emission factor has the dimension of "emissions per mass / energy / electricity / other reference unit", and the caliber label is limited to place method, market method, supplier measurement, industry average, and expenditure method. Each factor is accompanied by source identification, version number and source verification fingerprint. The place method refers to a method for calculating energy consumption emissions according to the regional average emission factor of the energy supply geographical area and the corresponding year, without considering contract reduction and green certificate offset. In this application, the place method result is consistent with the geographical mapping and year locking. The market method refers to a method for calculating emissions according to the factor corresponding to the remaining electricity after the activity is reduced based on the contract terms and green power / certificate market tools; contract coverage and certificate uniqueness must be recorded at the row level. After the account period is mapped to the year and the delivery place is mapped to the region / partition, the same coordinate set is filtered from the factor index.
[0082] Before going online, do a full de-duplication and source identification integrity check; use event triggering and timing inspection dual mechanism during operation: once there is an update in the external library, the hot cache is invalidated and rebuilt, and if there is no update, the hot cache remains valid for two hours, and expires automatically. The effective period is automatically refreshed, which can be shortened to one hour during peak hours; the cache is managed according to sixteen fixed shards, and the source retrieval failure follows the rhythm of one second, two seconds and four seconds for a maximum of three times, and if it fails again, it is recorded as a failure, and a merge is performed for concurrent requests with the same coordinates to limit the flow. Each row-level record only depends on the account period and delivery place marked in the previous step as anchor points: the account period is first mapped to the Gregorian year, and the delivery place is positioned according to the standard place name and zone code, and then mapped to the power grid partition; based on these two coordinates, the region and year are determined, and then the candidate entries are filtered from the index.
[0083] If there are multiple versions under the same coordinate, the latest valid version that is not revoked is selected by default. If the annual entry is missing, only one-year forward push is allowed and it only happens once. The time, reason and basis of the "annual downgrade" in the lock path record are recorded. If the regional entry is missing, the adjacent partition is tried one by one according to the enterprise's recorded partition adjacency table (the execution version number is TAB-NEIGH-202509-V0.1). The first hit stops and cannot cross more than two levels of adjacency. If it still cannot be hit, the recorded alternative range can be enabled, and the approval ticket number and responsible person code are written into the lock path. The range label is limited to five types: place law, market law, supplier measurement, industry average, and expenditure method. Place law and market law are only used in scenarios where energy can be purchased and settled based on contracts. Supplier measurement requires verifiable third-party or metered manufacturer reports. Industry average and expenditure method are only used as downgrade ranges when the above evidence is not available.
[0084] The source verification fingerprint is generated by concatenating the source agency name, version number, official release date, acquisition channel identifier, and the normalized byte sequence of the first segment of the original document (only visible characters and line breaks are retained, headers and footers are removed, and excess white space is removed). The entry is stored together for easy spot checks at any time. The latest valid version is determined by the "latest release date" that is not revoked. If there is a correction, the release date after the correction is used as the basis. If a new version is released and then a withdrawal statement is released, the system immediately reverts to the previous valid version, and the time and basis of "withdrawal rollback" are written in the lock path. The regional mapping and power grid partition both indicate the version number used. If the delivery location has a boundary adjustment during the accounting period, the version of the accounting period is used as the basis and a note is recorded. Cross-border and Hong Kong, Macau and Taiwan deliveries are mapped to available overseas or regional partitions according to the enterprise's equivalent range table, and the basis and responsible side of "cross-domain mapping" are marked in the lock path.
[0085] On-site execution sets clear failure criteria: single screening exceeds two seconds, candidate list is empty, regional annual conflict cannot be resolved, source verification fingerprint does not pass, candidate version has a withdrawal mark and there is no previous valid version. Any of the above is considered a failure. Two consecutive failures of the same record will be transferred to manual confirmation. To balance throughput and latency, local hot cache and shard pull are used for index loading, with a single load limited to within thirty seconds, and concurrent screening capacity not less than three thousand rows per minute. The factor reference area provides candidate lists and lock paths in key-value format to the outside. The minimum set of fields includes candidate number set, regional identifier, annual, version number, range label, source identifier, source verification fingerprint, lock level, path text and generation time. The interface uses an authenticated internal network channel. The expected response time is two hundred milliseconds, the ninety-fifth percentile is not more than one second, the timeout threshold is two seconds, automatic retries are three times and are idempotent, the idempotent key is composed of batch number, row number and second-level timestamp, repeated delivery is considered successful when the first successful write is written and a "already exists" flag is returned.
[0086] Loading completion will generate a candidate list, along with the version lock path written in the factor reference area, and the list number and path digest are backfilled to the corresponding row-level record, and the next step is directly determined by the unique judgment. The version lock path determines the latest effective version from the candidate factor set or the complete step record of the designated order degradation alternative, including lock level, time, and basis. On the quality control side, unknown regions, version missing pages, and source verification failures are all recorded as "degradation reason + timestamp + responsible side" and transferred to the review channel; if the proportion of entries that have been degraded in the same batch exceeds one-tenth, the system will automatically initiate a management task, prioritizing the inspection of whether the regional mapping table is too old, whether the partition table has changed, and whether the external library has been synchronized. The sampling method uses a hierarchical equidistant approach, with levels divided by region and year. Each layer has at least fifty entries, and the total sample size is at least one thousand entries; if a layer does not meet the standards, the sampling rate for that layer is immediately doubled, and a re-inspection is completed within seven days. The log and evidence retention period is consistent with the accounting period and is not less than five years, covering lock paths, degradation reasons, approval records, and source verification fingerprints.
[0087] For enterprises with higher time requirements, an alternative approach can be used: maintain a local mirror in the data warehouse, synchronize incrementally at 02:00 every week (Beijing time), allow a delay of up to 24 hours, and perform mirror health checks every four hours. If a lag is detected, immediately switch back to remote pulling and record the reason to prevent mirror lag from causing systemic bias. Through the clear definition of the above boundaries, order, thresholds, and trace requirements, loading and locking not only restrict the candidate range to the correct coordinates, but also allow for a one-step review of the reason for each selection when disputes arise. Interface performance, cache strategies, failure stop-loss, and sampling closed loops all have specific numerical values that can be executed on the table.
[0088] S4, according to the attribute matching degree, spatio-temporal consistency, and version priority, calculate the matching score, output the unique emission factor, and generate a review task and record the judgment path if uniqueness is not achieved, the specific implementation is:
[0089] In the determination of the unique emission factor, the goal is to converge the candidate given by the previous step to a specific value, while also leaving a complete trace of the judgment process for future review and audit. The information relied upon includes: the attribute groups (material, specification, shape, brand, measurement benchmark, etc.) and unit conversion trajectories (the source and accuracy of each step are recorded) sorted in the previous step, the candidate list formed by factor index loading (including region, year, version, aperture label, and source identification), and the version lock path (if the region or year is degraded, the complete degradation order is recorded).
[0090] First, let's clarify the terminology: an attribute anchor point is a field in the enterprise directory that is marked as a strong constraint, at least including material and model, and optionally including brand and shape; the minimum required set of attribute anchor points is material and model, and when both are missing, the review is directly entered without attempting to tighten the boundaries. Temporal and spatial consistency means that the annual mapping of the account period and the regional mapping of the delivery location are completely consistent with the candidate record, and no fuzzy matching is adopted. The version priority follows the double-dimensional sorting of the issuing agency and the release date, and the level order of the issuing agency is national, industry, local, and enterprise self-built, and the same level takes the later release date. At the same time, the measurement caliber is hammered: the enterprise reference unit is fixed as mass kilogram, volume cubic meter, energy kilowatt hour, length meter, area square meter, and quantity piece; non-reference units cannot directly participate in uniqueness judgment and must first register a complete conversion track before entering the scoring.
[0091] The preparation actions are attribute synonym merging, unit reference checking, and temporal and spatial pre-checking; any candidate that is inconsistent with the regional or annual mapping of the account period and delivery location is first removed and the reason is recorded. The subsequent scoring focuses on three factors: attribute fit, temporal and spatial consistency, and version priority; to ensure reproducibility, the weights are given a clear default band, with attribute factors accounting for 50%, temporal and spatial factors accounting for 40%, and version factors accounting for 1%. Before going online, a historical acceptance batch is used for calibration, and then a slight self-adaptation is made every month based on real-time operation data, with a single adjustment amplitude not exceeding 1%. Any changes to weights, thresholds, and white lists require double confirmation and take effect at the same time, and a change summary is automatically generated and written into the judgment path.
[0092] The threshold uses the difference method, comparing the difference between the highest score and the second highest score, and only when it reaches the specified threshold is it considered unique; the default band of the threshold is set between the 10th and 20th percentiles of the distribution of the last season, with an initial value of 10%, adjusted by 3% per month; the sample source for weights and thresholds is the acceptance batch of the last three months, covering materials, equipment, and services, with an equal proportion of each subject, with at least 200 lines per subject. To facilitate on-site understanding and spot checks, confidence levels are divided into three categories: high, medium, and low. The difference is not less than 0.15 and the key field completeness is not less than 95%, the difference is between 0.08 and 0.15, or the completeness is between 85% and 95%, and the difference is less than 0.08 or the completeness is less than 85%.
[0093] The decision order when parallel is also revealed: if the comprehensive score is tied, first check if the attribute anchor point is completely matched, then check the spatiotemporal matching degree, still tied, then choose the version with higher publishing agency level and later publishing date; if the enterprise white list has covered the corresponding supplier and category, and the aforementioned rules still cannot distinguish the difference, then the white list is preferred, and the entry number and applicable period are written into the track; if the comprehensive score has been significantly distinguished, but the white list points to different factors, then the score result is still the final answer, and the white list difference is recorded for audit interpretation.
[0094] To balance stability and timeliness, single row limit time is two seconds; the first calculation allows a small boundary tightening once, for example, the synonym mapping similarity threshold is adjusted by a small step to the robust side, the step length does not exceed 2% and is only allowed once, and if it still does not meet the standard, it is considered that it does not have uniqueness for the time being and generates a review task, while freezing the candidate list, weight value, threshold parameter, rule version number and timestamp at this time, ensuring that the same version is run at any point in time. The same version can get consistent conclusions.
[0095] Each judgment will deposit three blocks of content: unique emission factor number and confidence level, complete judgment path, and necessary prompt information; candidate list retains regional code, year, version number, caliber label, publishing agency number, and source review summary; the field set of the judgment path contains participating field list, factor score, weight, difference, rejected candidates and reasons, if downgrading, list the downgrade order and time point, rule version number and source review summary. The above content is written into the factor decision area, and the number and confidence level are backfilled in the current row record; if a review task is generated, a card with the original text fragment positioning and suggested attention points is pushed to the collaborative channel, and the review time limit within working days is four hours, and it is automatically upgraded after the deadline; the same entry enters the review for the second time within thirty days, triggering a review and freezing adaptive adjustment of this type of target, so as to prevent indicators from being biased by short-term samples. The connection with upstream and downstream is completed through an ordered queue: the previous link throws the candidate list and attribute group according to batch and line number, and this link completes the uniqueness and throws the number and judgment path to the downstream to do line-level numerical and caliber consistency verification.
[0096] To ensure throughput and stability, high water level expansion and queue protection are set. The throughput of the line level determination under normal circumstances is not less than 2000 rows per minute, and the queue waiting time exceeds 30 seconds to trigger expansion. If a single row appears timeout, it will be recorded as a failure, and it will try at most twice, with a retry interval of one second and two seconds. When the platform is unavailable for more than 15 minutes, the low confidence level result is suspended, and only the accumulated pending queue is accumulated. After the platform is restored, it will be disposed of. Three common accidents are also set up for disposal: when the same factor is repeated across regions and the source is not completely consistent, use the newer version as a temporary value and force review; if the time and space pre-check and dimension check are removed after the candidate is removed, it means that there may be a mismatch between the account period and the delivery location. Immediately suspend processing the line and send a governance request to the upstream. Before the governance is completed, the line conclusion is not output externally; if the unit conversion track is found to be inconsistent with the candidate dimension, request the upstream to correct it. If the same track node needs to be paid attention to in the determination path, it will be highlighted and reviewed.
[0097] To prove that the mechanism can run and be checked on site, no less than 2000 lines are checked every month, stratified by subject and category. The success rate of uniqueness is calculated by the proportion of lines that form unique factors, with a target of not less than 95%. The first review pass rate is calculated by the proportion of lines whose first review conclusion and temporary determination are consistent, with a target of not less than 90%. High-confidence decisions are further reviewed on site, with a sampling ratio of not less than 10% and a target pass rate of not less than 98%. All determination paths support real-time playback, using the batch number, line number, rule version number, and timestamp four-tuple as the idempotent key. If the difference between the playback result and the original record exceeds 0.02, an alarm will be raised immediately. The determination path is retained in hot storage for three years and in cold backup for seven years. At the end of the period, it is desensitized and destroyed, and the destruction action is bound to the batch number and generates a receipt for annual compliance audit.
[0098] Run on the spot, with a ground example to illustrate the logic: under the steel category, the default weight is adopted, and the two candidates in a certain row only differ in material matching degree on the premise of consistent region and year, and the difference in comprehensive score reaches 0.16, directly determining the high-score item as the only emission factor; the other row in the same batch is automatically sent to the collaborative platform due to incomplete material field, with a difference of only 0.06, and the card is highlighted with two concerns of "material missing" and "specification incomplete", which can be re-triggered for judgment after being completed by the reviewer; if the situation of parallelism is encountered, and the enterprise whitelist covers the supplier and the category, the whitelist rule is used to determine the unique emission factor when the candidates exist in parallel or the matching difference does not reach the preset threshold, and the item number and applicable period are written into the track; if the whitelist and the score conclusion are different, the score is used as the final result and the difference is recorded for explanation. Through the above criteria and traces, this link explains the "why choose it, how to review after choosing it, what to do if not chosen, what to pay attention to, how to set and change parameters" in place, and the boundaries, thresholds, samples, permissions, time sequences and disclosure gates are also implemented in executable scales.
[0099] S5, calculate the row-level emission amount based on activity data and unique factors, and simultaneously generate location method results and market method results in the same row, perform dimension and caliber consistency check and set a flag, and the specific implementation is as follows:
[0100] On the basis of the preface, the unique emission factor has been determined for each row record, and the version and caliber label is locked, while the unit conversion track, account period and subject information are preserved, this link is based on daily batches as the rhythm to process row by row within the account period window, the goal is to output location method and market method results in the same row and to check each other, to ensure that the subsequent statistical caliber is comparable and the reconciliation is smooth. The caliber of activity quantity is unified as enterprise standard unit: kilogram, cubic meter or piece for materials, kilowatt-hour for electricity, cubic meter for gas, and liter or kilogram for liquid fuel; the conversion of volume and mass is based on 20 degrees Celsius and one standard atmosphere, if the temperature or pressure on the bill is different from the above-mentioned standard, the bill conditions will be used as the reference, and the deviation source and parameter source will be recorded in the evidence set.
[0101] In the electricity scenario, the reading is based on the active power meter installed at the subject metering point, the metering level is not less than the second level specified by the national standard, and the verification period is not more than 24 months; if there are double tables or compound rate tables, the main table used for financial settlement is used as the reference, and the vice table is preserved as evidence. For non-electricity metering such as material weighing and gas flow, the national verification regulation corresponding to the enterprise metering account is executed, and the equipment level and verification period shall not be lower than the minimum requirement of internal control. The latest verification certificate number is retained on site and associated in the record.
[0102] When encountering returns or red tickets, the quantity is negative and is considered as a cancellation of the original record in principle: if the cancelled record falls within the previous accounting period and is no more than three months away from the current period, establish cross-period correlation and make a negative adjustment in the original accounting period; red tickets exceeding three months are recalculated in the current accounting period, and the reason and original ticket number are noted in the record. When the unit of the line record does not match the required dimension of the factor, it is converted step by step according to the saved unit conversion track; when the conversion is missing key parameters (such as density, moisture content or packing coefficient), it is first supplemented once according to the supplier's specifications or the enterprise's commonly used table, and if it still cannot be determined, it is marked as "conversion missing parameters", and this line is temporarily not involved in the consistency check and main level summary.
[0103] To prevent the caliber from being adjusted subjectively, the same subject is prohibited from directly stacking across dimensions in the same accounting period, and the system only enters the main level check after completing the conversion within the line. Market law related contract information (covering varieties, covering power or proportion, start and end date, settlement subject, traceable number and signing party) is pulled once by the contract account according to the accounting period; when there are multiple valid contracts for the same subject, the one with the earliest start time and still within the valid period is given priority, and the rest are allocated in descending order of coverage according to the coverage; if the contract only covers part of the days, it is preferentially mapped by site, then by time period, and finally by calendar equal allocation, and the allocation path is written into the evidence set; site mapping takes any one of the three as the matching success condition: power supply household number, metering point number, and contract site code, and the default granularity is aggregated by hour.
[0104] Green power certificates and renewable power deductions strictly implement uniqueness and validity management: the same certificate number cannot be used repeatedly in multiple subjects within the same accounting period, and if duplication is detected, the first record is confirmed according to the first-come-first-served principle, and the remaining entries are transferred to the investigation and the source is marked; the certificate number needs to meet the length and check digit rules, and all failed checks are considered as not covered and included in the investigation. The region and year used by the location method are based on the previous locking results, and if missing, they are mapped according to the power supply unit on the enterprise electricity settlement sheet; the residual power formed after market law reduction continues to use the location method factor of the same region and year.
[0105] Generate two results of location method and market method in line, and perform consistency check on the same subject in the same accounting period: the difference threshold is set in layers, the default is 20% for regular subjects, 15% for large energy-using subjects, and special industry subjects can be fixed after annual review in the range of 10% to 25%; the determination caliber of large energy-using subjects is that the annual electricity consumption reaches or exceeds 5 million kilowatt-hours, and the threshold is fixed by annual review; it can be adjusted slightly by quarter, and the single adjustment range is not more than plus or minus 5%; once the threshold is exceeded, it is marked as under investigation, and the difference is attributed to the contract coverage deficiency, certificate quantity deficiency, regional mapping inconsistency, and measurement conversion anomaly in the order of priority; if multiple items are hit at the same time, the primary cause is determined in this order, and the secondary causes and evidence fragments are listed in the notes.
[0106] To meet the monthly and large-scale batch processing, the end-to-end time limit of the whole batch verification is controlled within 15 minutes, and the parallel scale is not less than 200,000 rows; when the queue level approaches the high threshold, it is cut into multiple fragments according to the main body and account period and executed in parallel; the batch number and row number are used as the unique identifier for idempotent writing in each fragment, and the repeated rows are refreshed only when the state bit and timestamp are refreshed, and the repeated measurement is not repeated; two retries are performed for the row-level number failure, the interval is gradually lengthened, and the third time still fails to enter the waiting for verification and freeze the subsequent operation related to the row; when the same main body and the same account period fail to exceed 5 / 1000, an alarm is triggered and a management task is generated.
[0107] After the calculation is completed, three core records will be output for each row and synchronized: first, the row-level value saved according to the enterprise benchmark unit, internally retained to the fifth decimal place, and displayed uniformly to the third decimal place; when the factor release precision is lower than three, the factor precision is used as the upper limit and the source precision is marked in the record; second, the caliber label, using a fixed coding set, place method is marked as A and market method is marked as B; when both are parallel, they must appear at the same time and cannot be split; third, consistency state, divided into normal, waiting for verification and temporary processing. The three records are written into the row-level value area, and the state bit is refreshed on the row record; all "waiting for verification" are aggregated into a list according to the main body and account period and pushed to the on-site inspection channel, the list is accompanied by the primary cause, secondary cause and necessary context fragments, reducing back-and-forth communication.
[0108] In terms of time specification, all time fields use Beijing time; if the source system is in other time zones, time zone conversion is completed on the receiving side and the original time zone is kept as a note field. To ensure institutionalized operation, row-level records are kept for at least five years; monthly closing should be completed within five working days after the end of the account period, and quarterly closing should be completed within ten working days after the end of the quarter and archived; if the contract and voucher related to the market method have expired, duplicated or failed in number verification, they are considered as not covered and put into the waiting for verification; after subsequent supplementation, they are backfilled according to the original account period and the correction trace is kept. The process effectiveness is measured by two hard indicators as the common board: the consistency verification pass rate of place method and market method is not less than 98%, and the dimension alignment accuracy is not less than 99.5%; sampling uses a proportionate strategy by main body and account period, with a confidence level of not less than 95%, and the sample size is not less than 10,000 rows per month; the details that do not pass must be manually reviewed and written back within five working days.
[0109] A typical electricity consumption scenario in operation is: a manufacturing base records a total electricity of 100,000 kWh at the metering point in April, the green electricity contract takes effect on the 15th and covers 30,000 kWh, the system calculates the total amount according to the factor of the East China division in the corresponding year under the site method, and the residual part is still calculated according to the site method under the market method, and the difference between the two results falls within the threshold value and is marked as normal; in the same batch, a rental site is not recorded in the contract, and the market method corresponds to zero electricity, and the difference between the site method and the market method exceeds the threshold value, and the system lists it as to be checked and gives the primary attribution prompt of "contract coverage deficiency".
[0110] The material scenario also has clear boundaries: a batch of lubricating oil is measured by volume and the bill indicates 15 degrees Celsius, and the system converts it to kilograms according to the conversion coefficient at that temperature; another one in the same batch is "box", and the catalog lacks the box packing coefficient, which is marked as "conversion parameter missing" and enters the to-be-checked state, and does not participate in the main body aggregation. After the supplier provides the box packing coefficient, it will be backfilled. If the business side is limited in resources or requires higher timeliness, a conservative path can be adopted: first, quickly output the site method and complete the dimension check, then batch generate the market method and backfill the same row after the contract account is filled in asynchronously, and supplement a "reason for late arrival and allocation path" in the evidence set to ensure that the audit is traceable and the caliber is consistent. The whole process connects the front and back of the unique factor determination and the evidence chain and uncertainty evaluation, not only nails the enterprise's common dimension and contract boundaries, but also writes the threshold, stratification and bottom path, ensuring that personnel in this field can reproduce it stably under general system conditions.
[0111] S6, generate an evidence chain object for each row, record the text segment, extract the trace, unit conversion track, emission factor source and version locking path, build uncertainty quantification index, trigger review when threshold exceeds and write results back to alias account and rule snapshot, establish evidence chain object for the row immediately after the row-level result is generated, the object includes original text segment and position index, element extraction trace, unit conversion track, emission factor source and version locking path, person signature and timestamp; generate content summary and write to incremental only log, any modification is appended with new version and original version is preserved; set separate viewing and modifying rights and only allow supplementary records; evidence chain and results are persisted in the same domain and partitioned by batch and date, which can be played back and located; use batch number and row number as idempotent identifiers for transmission and access, specific implementation is:
[0112] After the row-level result output, the system immediately establishes a directly comparable evidence chain for each row, aiming to fix where the row information comes from, how it is identified, how unit conversion is completed, which version of emission factors is used, and when and by whom it is confirmed, to meet the traceability requirements of the audit level, while making uncertainty explicit, and triggering review when the boundary is crossed, and sinking the review conclusion into the enterprise rule library for subsequent automatic adoption. The evidence chain elements include: the original text segment and its positioning in the ticket format (annotated with page number, in-line offset, column offset, and character range, with character encoding unified to enterprise general encoding, and time unified to Beijing time, allowing clock drift not exceeding one minute, and exceeding recording clock adjustment events); the word matched when the element is extracted, the corresponding directory source, the identified field and time, and the current rule version number; the complete track of unit conversion (source unit, target unit, conversion basis, precision and tolerance, confirmation person and time of each step); the source identification, region and year of the emission factor, version number and scope label; if replacement occurs, the downgrade order and reason are recorded as they are; and the signature and timestamp of the person handling.
[0113] The evidence chain object refers to the traceable record coexisting with a certain row result, including original text segment and location index, element extraction trace, unit conversion track, factor source and version locking path, person handling signature and timestamp, and only-increase log summary. After generation, the content summary is written into the only-increase log, and any change is appended in the form of a new version, with the original version remaining. The log records batch number, row number, summary value, operator and timestamp, supporting cross-database collation and spot check comparison; at the same time, access control marks are set, with viewing and modifying belonging to different roles, with permissions configured as minimal, and modification only allowed to be supplemented, with original records not allowed to be overwritten. In order to make the evidence long-term verifiable, the evidence chain and the row-level result are saved in the same domain, managed by batch and date partition, with single-row writing limited to within one second; commonly used fields reside in online hot storage, original long text and bitmap snapshots are placed in cold backup, with a retention period not less than ten years; message delivery uses at least once arrival semantics, consumers use "batch number + row number" as idempotent key to deduplicate, queues automatically expand and throttle when reaching high water level and lasting for fifteen minutes, emptying in-transit tasks before restoring standard rhythm.
[0114] Before the evidence falls, a boundary check is made: the original text fragment should be able to play back in the original format, the position index should not exceed the boundary, and the source identification must be able to trace back to the corresponding entry in the authoritative library; the format derived from the scanned copy also needs to be checked for recognition quality score, and if it is less than 80%, it is determined to be unstable and a line bitmap snapshot must be saved simultaneously; if it is less than 70%, it is directly put into the manual channel. If the proportion of evidence field missing, source check failure or signature abnormality in the same batch exceeds 1%, the system automatically generates a compliance check list; if it exceeds 3%, it immediately stops external disclosure, only the lines that have passed the review are released, and the remaining ones are processed according to the emergency review process. To avoid the currency drift caused by currency crossing, if the bill contains multiple currencies, it will be converted according to the enterprise settlement exchange rate table on the same day, and the exchange rate source and time will be written into the evidence object; suppliers and bill formats that repeatedly appear problems are included in the quality control list, and the main body that enters the list for three consecutive times enters the blacklist and triggers the scene review.
[0115] Around uncertainty, the system sets interpretable indicators from activity, emission factor, unit conversion, and caliber selection, and gives priority: if the measurement chain difference exceeds 3%, it is considered to be out of bounds on the activity side; if the account period and version deviate more than a year, or the factor source cannot be traced back, it is considered to be out of bounds on the factor side; if the unit conversion cascades more than twice, or any step tolerance is higher than 0.005%, it is considered to be out of bounds on the conversion side; if the difference between place law and market law for the same subject in the same account period exceeds 10%, or is higher than the upper limit of the contract, it is considered to be out of bounds on the caliber side. If any of the above is out of bounds, the system immediately generates a review task, and highlights the main causes and recommended check paths on the task sheet; the task is classified into high, medium and general, the high level is processed within four hours, the medium level is processed within 24 hours, and the general is completed within three working days, and the time is automatically upgraded and the responsible person is notified. The threshold value is rolled over daily and adjusted, and when the rolling result fluctuates more than 2% of the original set value within three days, it is automatically rolled back to the stable segment of the previous week and frozen for seven days, and then unfrozen after review by the rules administrator.
[0116] Once the review conclusion is confirmed, it is written back to the alias account book and rule snapshot, and the version number, coverage and effective time are recorded simultaneously, to ensure that similar lines in the future can directly use the improved caliber, avoiding repeated warnings; for example, the two-stage conversion of "box -> kg -> ton" in a certain steel line makes the conversion side indicator out of bounds, the system automatically generates a review task, highlights "unit conversion" and provides a playback link, after the reviewer checks according to the prompt, a more stringent conversion caliber is adopted according to the enterprise's local standard, and the signature is confirmed and the rule is backfilled, so that the same path will not exceed the boundary from the next batch. The concurrent capability of the whole link is configured to be no less than 5,000 lines per minute, with cold standby asynchronous landing, and online writing priority is given during business peak; if a single line writing cannot meet the time limit, it is allowed to retry three times with increasing intervals, and if it still fails, it is transferred to the manual list and the original evidence reference is preserved, to avoid link interruption; sensitive fields such as tax identification number are stored in face mask to support spot checks and reduce the risk of leakage.
[0117] To measure the degree of implementation and on-site performance, monthly random checks are carried out with sufficient coverage: no less than 1,000 rows are selected, at least three subjects, three categories, and three account periods are covered, double independent review is adopted and consistency is used as the criterion, the completeness rate and reproducibility rate of evidence elements are counted, and the average manual review time, concurrent time delay, and peak stability are continuously monitored, and if the indicators do not meet the standards, the rule optimization work order is triggered.
[0118] Through the above arrangement, evidence and results coexist, uncertainty is no longer "hidden", review has clear priority and time limit, and the improvement formed after review will be remembered by the system for a long time and automatically applied, combined with the non-tamperable increasing only log, long-term storage and hierarchical storage, arrival and idempotent constraints, compliance stop line and blacklist mechanism, the whole chain achieves a verifiable balance between time, resources, compliance and auditability, and on-site personnel can quickly locate problems, close loop processing and directly reuse the confirmed criteria in the next same situation with a piece of evidence object and a task link, reducing disputes and rework.
[0119] The calculation involved in the embodiments is a dimensionless calculation of the numerical value, and the preset parameters and threshold values in the calculation are set by a person skilled in the art according to the actual situation.
[0120] It should be noted that the present application can be deployed in the device itself to realize embedded application, or run on PC or other terminal with user interface, thereby meeting various hardware environments and use requirements.
[0121] The above-described embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented by software, the above-described embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through wireless or wired transmission. The wired transmission includes optical fiber, twisted pair, coaxial cable, etc. The wireless transmission includes infrared, microwave, etc. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.
[0122] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and module can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0123] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the above-described device embodiments are only schematic, for example, the division of the modules is only a logical function division, and actual implementation can have another division manner, for example, a plurality of modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0124] The modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical modules, which can be located in one place or distributed on a plurality of network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.
[0125] In addition, each functional module in the various embodiments of the present application can be integrated in one processing module, or each module can exist physically separately, or two or more modules can be integrated in one module.
[0126] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0127] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0128] Finally: the above is only the preferred embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A method for intelligent analysis of carbon emission data, characterized in that, include: S1. Receive business data and segment it by line, marking the payment period, subject, and delivery location. Retain the original text fragments and location indexes to form line-level pending units. Establish a unified channel on the receiving side to stably access invoices from the financial system and energy consumption ledger, receiving them according to fixed and incremental rhythms. Map the invoice number and subject code one-to-one, unify the time and unit, and remove duplicates according to source priority and arrival order. Fill in missing payment periods according to the invoice month, and map the delivery location according to the order of receiving address, contract address, and invoice address. Set a temporary processing mark when mapping fails. Segment the data line by line according to the format and generate original text fragments and location indexes, supplement the payment period, subject, and delivery location, encapsulate them into line-level pending units containing batch number, line number, and content summary, and deliver them to the message queue with an idempotent identifier composed of batch number and line number. S2. Extract elements from Chinese cargo descriptions, unify names and attributes based on alias ledgers, unify units of measurement, and generate unit conversion trajectories. S3. Load the emission factor index containing regional, annual, version and caliber labels, parse the spatiotemporal information based on the payment period and delivery location, screen out candidate factors and complete version locking; S4. Calculate the matching score based on attribute fit, spatiotemporal consistency and version priority, output a unique emission factor, and generate a review task and record the judgment path if uniqueness is not achieved. S5. Calculate row-level emissions using activity data and a unique factor, and simultaneously generate location-based and market-based results in parallel on the same row, perform dimension and caliber consistency checks and set markers; S6. Generate an evidence chain object for each row, recording text fragments, extraction traces, unit conversion trajectories, emission factor sources, and version locking paths. Construct uncertain quantitative indicators, triggering a review when thresholds are exceeded and writing the results back to the alias ledger and rule snapshot. Immediately after the row-level results are generated, establish an evidence chain object for that row. The object includes the original text fragment and location index, element extraction traces, unit conversion trajectories, emission factor sources and version locking paths, handler signatures, and timestamps. Generate a content summary and write it to the incremental-only log. Any modifications are appended to the new version while retaining the original version. Set viewing and modification permissions and only allow supplementary entries. The evidence chain and results are persisted in the same domain and partitioned by batch and date, allowing for playback and location; batch number and line number are used as idempotent identifiers for transmission and retrieval.
2. The intelligent analysis method for carbon emission data according to claim 1, characterized in that, S2 include: The elements of the Chinese product description are extracted and normalized, and the brand, material, specifications, shape, model and unit of measurement are grouped into standard names and attribute groups using an alias ledger; When the normalization is difficult to determine, the context fields are called sequentially based on adjacent rows of the same invoice, the same supplier in the same batch, and the contract material strip. Units of measurement are unified to the standard units. Unit conversion trajectories are generated step by step according to the conversion table. The source and scope are recorded and the trajectory number is written back to the line. For cases where the model and unit are inconsistent, the material and shape are inconsistent, the specifications are missing and the context is insufficient, a gray mark is set and a checklist is generated; The alias ledger and conversion table have versions and effective ranges, and their history is unified and not rolled back; When multiple materials are listed side-by-side, they are split into sub-items by a separator while maintaining their order and position index; The normalization results are stored in the normalization region for factor selection and version locking.
3. The intelligent analysis method for carbon emission data according to claim 1, characterized in that, S3 include: When loading the emission factor index, the index includes the region identifier, year, version number, caliber label, source identifier, and source verification fingerprint; Based on the payment term mapping year and the delivery location mapping region and zone, first determine the region and year, then filter candidate factors and lock the latest valid version that has not been withdrawn; When an annual entry is missing, only the most recent year is downgraded and the locked path is recorded. When a region entry is missing, the system will attempt to match it item by item according to the order of adjacent partitions in the filing and record the locking level. If no match is found, the approved alternative will be used and the responsibility information will be written. Write the candidate list along with the version lock path into the factor reference area, and backfill the list number and path summary into the corresponding row-level record.
4. The intelligent analysis method for carbon emission data according to claim 1, characterized in that, S4 include: After version locking is completed, attribute anchor synonym merging, unit benchmark verification and spatiotemporal pre-check are performed on each record to eliminate candidates that are inconsistent with the payment period and delivery location. The remaining candidates are matched based on a combination of attribute fit, spatiotemporal consistency, and version priority according to preset weights to obtain a matching score. When the difference between the highest and second-highest matching scores reaches a preset threshold, a unique emission factor is determined, and a decision path containing participating fields, scores, weights, reasons for removal, and version information is generated. If the threshold is not reached, a review task is generated and the candidate list, weights, thresholds and rule versions are frozen, and timestamps are added to ensure consistency when rerunning with the same version. The unique emission factor number and determination path are written into the factor decision area and backfilled to the row-level record.
5. The intelligent analysis method for carbon emission data according to claim 1, characterized in that, S5 include: At the row level, based on activity volume and unique emission factor, both location-based and market-based results are generated on the same row and labeled with caliber labels. The market coverage is allocated according to the contract ledger in the order of site mapping, time period mapping, and calendar equalization. Uncovered residual amounts are accounted for using the same factor caliber as the location method, and the allocation path and caliber label are recorded in this line; Cross-dimensional superposition is prohibited within the industry; the unit conversion trajectory that has been saved shall be used as the standard for unit uniformity.
6. The intelligent analysis method for carbon emission data according to claim 5, characterized in that: For the same entity and the same payment period, a consistency check should be performed between the location-based approach and the market-based approach. In case of inconsistency, the attribution shall be determined according to the priority order: insufficient contract coverage, insufficient number of vouchers, inconsistent regional mapping, and abnormal measurement conversion. A pending status shall be generated in the bank and the attribution, allocation path and evidence fragment shall be written. Green electricity certificates are verified for uniqueness by number and can only be used by a single entity. All results and tags are stored in the database with batch number and row number as idempotent keys and can be used for subsequent evidence chain generation.
7. The intelligent analysis method for carbon emission data according to claim 1, characterized in that: Based on the evidence chain object, uncertain quantitative indicators are constructed for activity volume, emission factor, unit conversion, and caliber selection, and thresholds and priorities are set; When any uncertain quantification exceeds its threshold, a review task is automatically generated and the main cause is highlighted. The current candidates and parameters are frozen. The review conclusion is written back to the alias ledger and rule snapshot and the version, coverage and effective time are recorded. It is then automatically adopted in subsequent similar rows. Review tasks are processed according to grade and time limits and can be upgraded.
Citation Information
Patent Citations
Supplier contract safety management system and method based on block chain
CN120258843A
Business data security protection method and system for digital enterprise management
CN120567444A