Green carbon data transaction and standardization conversion method adapting to international carbon accounting standards

CN122798533APending Publication Date: 2026-09-22FANGYUANBIAOZHIRENZHENG GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610811320.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

现阶段绿碳交易业态复杂,参与主体类型多样,各主体的业务执行流程存在明显差异,且行业内交易数据格式缺乏统一规范,造成不同交易类型、不同流程节点产生的绿碳交易数据难以适配国际碳核算标准,严重阻碍了碳数据的跨主体、跨区域互认与流通

Benefits of technology

双维度聚类分析:首次提出单节点格式聚类与全链格式链聚类相结合的双维度聚类方法,突破了现有技术单一维度对齐的局限,场景覆盖率有效提升;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122798533A_ABST
    Figure CN122798533A_ABST
Patent Text Reader

Abstract

The application provides a green carbon data transaction and standardization conversion method suitable for international carbon accounting standards, and belongs to the technical field of accounting standardization, and comprises the following steps: the method mines historical green carbon transaction data to construct transaction threads, divides the threads into thread sets according to transaction types, and carries out single-set format clustering and full-chain format clustering; candidate threads are obtained through secondary division, matching and screening, and the candidate threads are inducted for secondary clustering and calculation of standard format transition probability; the format stability is determined through a sequence analysis model, and reference results are obtained through fusion or screening; and then, a blank thread is generated through optimal thread recommendation, and a unique standard thread is formed by extracting a high-confidence format. The application solves the problems of non-uniform green carbon transaction data format, difficulty in adapting to international carbon accounting standards, and difficulty in cross-subject and cross-regional mutual recognition, realizes data format standardization and multi-scene adaptation, and improves carbon data circulation and accounting efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of accounting standardization technology, and in particular to a green carbon data trading and standardization conversion method adapted to international carbon accounting standards. Background Technology

[0002] As the global carbon market gradually integrates and becomes more interconnected, international carbon accounting standards are also converging. The standardized conversion of green carbon trading data has become a crucial foundation for cross-border carbon emission reduction accounting, carbon asset trading, and compliance certification. Currently, the green carbon trading landscape is complex, with diverse participating entities and significantly different business execution processes. Furthermore, the lack of unified data formats within the industry makes it difficult for green carbon trading data generated from different trading types and process stages to adapt to international carbon accounting standards. This severely hinders the cross-entity and cross-regional mutual recognition and circulation of carbon data.

[0003] Current mainstream green carbon data standardization and conversion solutions in the industry mostly adopt static template matching or single-dimensional format alignment, which have obvious technical limitations. These solutions rely entirely on preset fixed transaction templates to complete data conversion, and cannot deeply explore the differences in format characteristics of different transaction types and transaction nodes in massive historical transaction data. At the same time, they cannot combine the complete format chain of the transaction thread to carry out multi-dimensional and multi-granular cluster analysis, and have insufficient coverage of various differentiated business scenarios. Overall, they are difficult to achieve universal adaptation with international carbon accounting standards.

[0004] Therefore, this invention proposes a green carbon data trading and standardization conversion method that is compatible with international carbon accounting standards. Summary of the Invention

[0005] This invention provides a green carbon data trading and standardization conversion method that is compatible with international carbon accounting standards, in order to solve the aforementioned technical problems.

[0006] This invention provides a method for green carbon data trading and standardization conversion adapted to international carbon accounting standards, including: Step 1: Based on the historical database, mine each historical green carbon transaction data to obtain the transaction thread, and obtain the historical transaction standard format of each transaction node in the transaction thread. Perform initial division of all transaction threads according to the transaction type to obtain several thread division sets. Step 2: Perform a format clustering analysis on each partitioned thread set according to the historical transaction standard format to obtain several small clustering results for the corresponding partitioned thread set. At the same time, perform chain clustering analysis on all transaction threads according to the historical transaction standard format chain of each transaction thread to obtain the large clustering result. Step 3: Perform secondary partitioning on each large cluster result according to transaction type. Based on the matching relationship with the transaction type of the corresponding partition thread set, sort the secondary partitioning results by size to obtain an initial list; Step 4: Determine the transaction format matching set, matching probability, and partitioning level confidence of each sub-partition result in the initial list of each large cluster result for the corresponding small cluster result, and select candidate threads that meet the candidate criteria of the corresponding small cluster result from each large cluster result; Step 5: All candidate threads obtained based on the small clustering results under the same transaction type are summarized into the corresponding partitioned thread set, and secondary format clustering analysis is performed to obtain several new clustering results for the corresponding summarized thread set; Step 6: Compare and analyze the new clustering results of the corresponding inductive thread set with the original small clustering results, determine the standard format transition probability of each small clustering result, and construct the probability transition sequence of the corresponding inductive thread set. Step 7: Input the probability transition sequence into the pre-trained sequence analysis model to determine the format fixing coefficients of the corresponding small clustering results; If the fixed coefficient of the format is greater than or equal to the preset coefficient, the clustering results with high confidence of the format standard are selected from the corresponding small clustering results and the matched new clustering results as reference results; If the fixed coefficient of the format is less than the preset coefficient, the corresponding small clustering result and the matching new clustering result will be fused and analyzed as a reference result. Step 8: For each reference result under each partitioned thread set, perform optimal thread recommendation processing according to transaction attributes and transaction feedback to obtain a unique blank thread for the corresponding reference result; Step 9: Extract the historical transaction standard format with the highest confidence level for each transaction node based on the unique blank thread from each reference result corresponding to each partitioned thread set, and use it as the current format to obtain the unique standard thread. Then, take all the unique standard threads under each transaction type as the standard threads for green carbon trading.

[0007] Preferably, the transaction thread is obtained by mining each historical green carbon transaction data from the historical database, including: For each historical green carbon transaction data, feature extraction is performed to obtain transaction features and green carbon features, and a transaction template is selected from the dual-feature database based on the transaction features and green carbon features; Simultaneously, the flow process of corresponding historical green carbon trading data and the time-sequential operation behaviors involved in the flow process are mined from the historical database, and the start and end points of each operation behavior are extracted to construct the initial thread; The transaction template is matched with the initial thread to obtain the transaction thread.

[0008] Preferably, the transaction template is matched with the initial thread to obtain a transaction thread, including: Based on the predefined transaction description set of each template node in the transaction template, semantic matching and window association analysis are performed between the operation start point and operation end point of each operation behavior in the initial thread; When the semantic matching degree is greater than the preset matching degree, the template node is retained unchanged; Otherwise, based on the window association analysis results, determine the variable group of the corresponding operation behavior, and compare the variable types involved in the variable group with the transaction description features of the predefined transaction description set of the adjacent nodes of the extracted template node to determine the node addition position, and determine and retain the addition transaction description set of the corresponding addition node based on the variable group. The transaction thread is obtained based on the sequential positional relationship of the retained template nodes.

[0009] Preferably, the initial list is obtained by sorting the results of the secondary partitioning by size, including: The corresponding large clustering results are further divided according to the transaction type to obtain several first partition sets of the corresponding large clustering results, and each first partition set corresponds to a transaction type. The transaction types of each first partition set are matched with the transaction types of the corresponding partition thread set to obtain the matching relationship; At the same time, determine the first number of transaction data that is completely consistent with the corresponding partition thread set and the second number of similar transaction data in each first partition set, and obtain the statistical array of the corresponding first partition set; The matching relationship is optimized based on the statistical array to obtain the sorting value of each first partition set; Sort all the first partition sets corresponding to the large clustering results according to the size of the sorting value to obtain an initial list, where the first partition set is the sub-partition result.

[0010] Preferably, candidate threads that meet the candidate criteria for the corresponding smaller clustering results are selected from each large clustering result, including: Construct a first matrix of corresponding small clustering results and corresponding initial lists based on the transaction format matching set, matching proportion probability and division level confidence; Based on each candidate index in the candidate criteria of the corresponding small clustering results, traverse each element in the first matrix in turn, and mark the matching elements as first; The importance of each row vector in the first matrix is ​​obtained based on the number of times each labeled element is labeled, and the number of sub-partition results to be filtered for each row vector is determined. Candidate threads are obtained by filtering from the corresponding sub-partition results.

[0011] Preferably, the standard format transition probability of each sub-clustering result is determined, including: Construct a first clustering graph corresponding to several new clustering results of the thread set before induction, and at the same time, construct a second clustering graph of several original small clustering results of the thread set before induction. By comparing and analyzing the first clustering graph and the second clustering graph, the first distribution of threads that originally belonged to each small clustering result in the first clustering graph and the second distribution in the second clustering graph are determined, and the change point of the second distribution based on the first distribution is determined, wherein the point is each transaction data in several new clustering results of the corresponding inductive thread set; The performance coefficient of each change point is obtained based on the nearest distance between each change point and the corresponding small cluster result, the cluster distance between each change point and the central cluster of the corresponding small cluster result, and the first distance between the central cluster of the corresponding small cluster result and the point corresponding to the nearest distance. Based on the variation distribution of the variation points and the performance coefficient of each variation point, the variation trend of the variation distribution is determined, and the standard format transition probability of the corresponding small clustering results is determined.

[0012] Preferably, the corresponding small clustering results and the matching new clustering results are fused and analyzed, including: Using the second distance between the cluster center of the corresponding small clustering result and the cluster center of the matched new clustering result as the radius, draw a circle with the cluster center of the corresponding small clustering result as the center, and retain the transaction data within the drawn circle as the first retention; All identical transaction data in the corresponding small clustering results and the matching new clustering results are globally and individually retained; Transaction data with performance coefficients greater than the performance threshold in the newly matched clustering results are retained as a second group. The first retention result, the global single retention result, and the second retention result are all processed by single retention to obtain a reference result. Among them, the completely identical transaction data in the first retention result are locally single retained.

[0013] Preferably, the unique blank thread that obtains the corresponding reference result includes: For each transaction data in the reference results, static features of the transaction attribute dimension and dynamic features of the transaction feedback dimension are extracted, and the correlation mapping relationship between static features and dynamic features is constructed to form a transaction feature matrix. The static features include the format definition of the transaction node, data field constraints and carbon accounting rule parameters, and the dynamic features include historical execution compliance rate, format matching pass rate and accounting feedback correction items. Based on the transaction feature matrix, the transaction scenarios corresponding to the reference results are analyzed in multiple dimensions to identify the scenario type, compliance level and format adaptation requirements, and generate a scenario adaptation tag set. Based on the historical feedback data of the transaction feature matrix, the adaptation value of different thread configurations is evaluated, and high-value thread configurations with adaptation values ​​higher than the preset value are selected. At the same time, based on the scenario adaptation tag set, underutilized adaptation rules are mined and supplementary thread configurations are generated. The high-value thread configuration and the supplementary thread configuration are input into a pre-trained policy fusion module to generate a unique thread optimization policy. Based on the thread optimization strategy and the pre-built scenario adaptation rule library, the transaction nodes in the reference results are formatted and constrained to generate an initial blank thread. The initial blank thread is matched with the dynamic feedback features in the transaction feature matrix in multiple rounds. When the matching degree reaches the preset convergence threshold, the final unique blank thread is output.

[0014] Compared with the prior art, the beneficial effects of this application are as follows: Two-dimensional clustering analysis: This paper proposes a two-dimensional clustering method that combines single-node format clustering and full-chain format chain clustering for the first time. This method breaks through the limitation of single-dimensional alignment in existing technologies and effectively improves the scenario coverage. Quantification of format evolution trends: The introduction of format transition probability modeling and sequence analysis model to determine format stability has enabled accurate prediction of format evolution trends, and the standard update response cycle has been shortened from 3 months to 7 days; Dynamic blank thread generation: A unique blank thread is generated by combining static format requirements and dynamic transaction feedback, instead of directly reusing historical templates, which effectively improves the actual compliance rate. Multi-dimensional fusion filtering: By using a fusion method that combines layered filtering, global deduplication, and retention of high-performance data, it balances format stability and scenario adaptability, and solves the problem of rigidity in existing technical templates.

[0015] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.

[0016] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a green carbon data trading and standardization conversion method adapted to international carbon accounting standards in an embodiment of the present invention. Detailed Implementation

[0018] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0019] This invention provides a method for green carbon data trading and standardization conversion that is compatible with international carbon accounting standards, such as... Figure 1 As shown, it includes: Step 1: Based on the historical database, mine each historical green carbon transaction data to obtain the transaction thread, and obtain the historical transaction standard format of each transaction node in the transaction thread. Perform initial division of all transaction threads according to the transaction type to obtain several thread division sets. Step 2: Perform a format clustering analysis on each partitioned thread set according to the historical transaction standard format to obtain several small clustering results for the corresponding partitioned thread set. At the same time, perform chain clustering analysis on all transaction threads according to the historical transaction standard format chain of each transaction thread to obtain the large clustering result. Step 3: Perform secondary partitioning on each large cluster result according to transaction type. Based on the matching relationship with the transaction type of the corresponding partition thread set, sort the secondary partitioning results by size to obtain an initial list; Step 4: Determine the transaction format matching set, matching probability, and partitioning level confidence of each sub-partition result in the initial list of each large cluster result for the corresponding small cluster result, and select candidate threads that meet the candidate criteria of the corresponding small cluster result from each large cluster result; Step 5: All candidate threads obtained based on the small clustering results under the same transaction type are summarized into the corresponding partitioned thread set, and secondary format clustering analysis is performed to obtain several new clustering results for the corresponding summarized thread set; Step 6: Compare and analyze the new clustering results of the corresponding inductive thread set with the original small clustering results, determine the standard format transition probability of each small clustering result, and construct the probability transition sequence of the corresponding inductive thread set. Step 7: Input the probability transition sequence into the pre-trained sequence analysis model to determine the format fixing coefficients of the corresponding small clustering results; If the fixed coefficient of the format is greater than or equal to the preset coefficient, the clustering results with high confidence of the format standard are selected from the corresponding small clustering results and the matched new clustering results as reference results; If the fixed coefficient of the format is less than the preset coefficient, the corresponding small clustering result and the matching new clustering result will be fused and analyzed as a reference result. Step 8: For each reference result under each partitioned thread set, perform optimal thread recommendation processing according to transaction attributes and transaction feedback to obtain a unique blank thread for the corresponding reference result; Step 9: Extract the historical transaction standard format with the highest confidence level for each transaction node based on the unique blank thread from each reference result corresponding to each partitioned thread set, and use it as the current format to obtain the unique standard thread. Then, take all the unique standard threads under each transaction type as the standard threads for green carbon trading.

[0020] In this embodiment, the historical database is a collection of data related to previously completed green carbon transactions, including structured / unstructured data such as transaction process, node information, format specifications, and transaction type.

[0021] Historical green carbon trading data refers to information related to past actual transactions of green carbon assets (such as forestry carbon sinks and photovoltaic emission reductions), including the trading entities, targets, process nodes, submission document formats, and review standards.

[0022] A transaction thread is the complete process link of a green carbon transaction from initiation, review, registration, delivery to settlement, including transaction nodes and their flow relationships arranged in chronological order.

[0023] A transaction node is a single process step in the transaction thread, such as submitting a transaction application to a third-party verification agency for review, carbon asset registration, and transfer.

[0024] Historical transaction standard formats are the file / data format specifications that each transaction node has used in the past and that conform to carbon accounting standards, such as the field format, data unit, and verification rules of the verification report.

[0025] The transaction type is a category of green carbon transactions classified according to the transaction target, mode, and applicable standards, such as the primary transaction of forestry carbon sink projects and the international VERS standard cross-border carbon transaction.

[0026] A thread set is a collection of threads of the same type after being categorized by transaction type, such as the collection of all CCER forestry carbon sink primary transaction threads.

[0027] For example, using ETL (Extract-Transform-Load) tools, transaction time-series data can be extracted from historical databases, and data from each stage can be linked by serial number to construct a single transaction thread. The historical file metadata of each transaction node can be parsed to extract field definitions, data types, verification rules, and accounting standard versions, forming a node format specification. Based on three dimensions—transaction target, transaction market, and applicable standards—threads with the same dimension combination are grouped into the same thread set. For instance, a 2023 forestry carbon sink primary transaction in the historical database follows the process: project filing application → third-party verification → listing and trading → delivery registration → settlement. This is then linked into transaction threads through data mining. The format for the node "third-party verification" is a PDF report submitted according to the CCER standard, including project boundaries and emission reduction calculations. Fields include project number, monitoring period, and emission reduction amount (…). (Keep 2 decimal places); classify transactions into the CCER Forestry Carbon Sink Primary Transaction Thread Set according to transaction type.

[0028] Format clustering analysis is a statistical method based on the similarity of transaction node formats to group threads. Threads with highly similar format requirements are grouped into one class. The small clustering result is the grouping obtained by clustering the nodes within a single partitioned thread set according to the similarity of the node formats, such as the CCER V3.0 certification format group.

[0029] The historical transaction standard format chain is a complete format link formed by connecting the historical transaction standard formats of all nodes in a transaction thread in the order of the process.

[0030] Chain clustering analysis is an analysis method that groups all transaction threads based on the similarity of the entire thread format chain. Threads with similar overall format chains are grouped into one category. The large clustering result is a grouping obtained by clustering all transaction threads according to the format chain similarity, which may include threads of different transaction types but with similar format chains.

[0031] For small clusters: Extract the format feature vector of each transaction node, with 12 dimensions, including: number of fields, percentage of required fields, data type distribution (numerical / text / date type percentage), file format (PDF / Word / Excel encoding), number of validation rules, applicable standard version, unit specifications, accuracy requirements, attachment requirements, signature requirements, encoding format, and data length limit; Calculate the formatted feature vectors of two nodes using cosine similarity. , The similarity is calculated using the following formula: ; The K-Means++ algorithm is used for clustering, and the K value is determined by the elbow method: the silhouette coefficients under different K values ​​(2-10) are calculated, and the K value with the largest silhouette coefficient is selected as the final number of clusters; Through orthogonal experiments, it was determined that when the threshold is 0.85, the clustering accuracy reaches 94.3% and the misclassification rate is the lowest. Therefore, the preset single-node clustering similarity threshold is 0.85.

[0032] For large clustering: concatenate the format feature vectors of all nodes in the transaction thread in the order of the process to form a format chain feature vector of length 12×N (N is the number of nodes in the transaction thread). The Dynamic Time Warping (DTW) algorithm is used to calculate the sequence similarity between two format chains, which solves the problem of inconsistent node numbers in different transaction threads. The smaller the DTW distance, the higher the similarity. Hierarchical clustering algorithm was adopted, with DTW distance as the distance metric and a clustering threshold of 0.8 (i.e., threads with DTW distance ≤ 0.8 are grouped into the same class). This threshold was determined experimentally. At this point, the purity of the large cluster reached 92.7%. All format chain feature vectors were normalized in length, and the DTW distance ranged from 0 to 1.

[0033] For example: Small clusters: In the CCER forestry carbon sink primary trading thread cluster, 80 threads have third-party verification nodes in CCERV3.0 format and 20 in V2.0 format. After clustering, two small clusters are obtained: CCER V3.0 verification format group and CCER V2.0 verification format group. Large clusters: The format chains of some CCER forestry carbon sink and CCER photovoltaic emission reduction threads are all filing application → verification (PDF) → listing (platform template) → delivery (registration interface) → settlement (transfer voucher). The similarity is ≥0.8 and they are classified into the same large cluster.

[0034] The secondary partitioning involves further splitting the threads in the large clustering result according to transaction type, resulting in subgroups of different transaction types within the same large cluster.

[0035] The matching relationship is the correspondence between the subgroup transaction types after the secondary partitioning and the partitioned thread set.

[0036] The size sort is based on the number of threads contained in the subgroups after the secondary partitioning, sorting them from most to least.

[0037] The initial list is an ordered list formed by sorting the subgroups after the secondary partitioning based on factors such as the number of threads, and is used for subsequent matching and filtering.

[0038] Matching statistics involves matching subgroups with the partitioned thread sets and counting the number of threads in each subgroup. For example, a large cluster result contains 200 threads, of which 120 are CCER forestry carbon sink primary transactions, 50 are CCER photovoltaic emission reduction primary transactions, and 30 are VER cross-border carbon transactions. After a second partition, three subgroups are obtained, sorted by the number of threads, and the initial list is: [CCER forestry carbon sink primary transactions (120), CCER photovoltaic emission reduction primary transactions (50), VER cross-border carbon transactions (30)].

[0039] In this embodiment, the transaction format matching set is the set of threads in which the small clustering results and the initial list subgroups completely match the format of all nodes.

[0040] The matching probability is the proportion of the number of threads in the transaction format matching set to the total number of threads in the subgroup, reflecting the prevalence of format matching.

[0041] Hierarchical confidence is an indicator of the reliability of the matching between small clusters and subgrouping formats, and can be calculated by the number of matching nodes and the matching degree threshold.

[0042] In this embodiment, the matching set calculation is to compare the format requirements of small clusters and subgroup threads node by node, and count the threads that match all nodes to form a transaction format matching set; Matching probability = (Number of threads in the matching set / Total number of threads in the subgroup) × 100%; The confidence score for tiered classification is calculated as follows: (Number of matching nodes / Total number of nodes) × 0.5 + Probability of matching percentage × 0.5. For example, matching the CCER V3.0 certification format group (80 items) with the CCER Forestry Carbon Sink Primary Trading (120 items) subgroup yields 90 matching threads; the matching percentage = 90 / 120 × 100% = 75%; the confidence level = (5 / 5) × 0.5 + 0.75 × 0.5 = 0.875 ≥ 0.8, which meets the candidate criteria, and the 90 threads are candidate threads.

[0043] The inductive thread set is a set of threads containing candidate threads formed by grouping candidate threads of the same transaction type into the corresponding partitioned thread set.

[0044] Secondary clustering uses the same format clustering method as step 2 (K-Means algorithm, based on node format feature vector similarity) to cluster the threads in the inductive thread set again. That is, the threads in the inductive thread set are clustered again according to node format similarity, and the format grouping is updated.

[0045] The output results in several new clustering results, each representing a group of candidate threads with similar formats.

[0046] For example, the 90 CCER forestry carbon sink primary trading candidate threads selected in step 4 are categorized into the original thread set (originally 100 threads, now the summarized thread set contains 190 threads). After secondary clustering, three new clustering results are obtained: V3.0 verification + platform delivery format group (120 threads), V3.0 verification + offline delivery format group (50 threads), and V2.0 verification format group (20 threads).

[0047] In this embodiment, the new clustering results are matched with the original small clustering results to analyze the distribution changes of threads in different clusters, and then the distribution of threads in the original small clusters in the new clustering results is statistically analyzed. The transition probability calculation involves calculating the proportion of threads that transition to each new cluster for each original small cluster. Sequence construction involves piecing together the transition probabilities of each node into a probability transition sequence, following the order of the transaction node process.

[0048] For example, the threads in the original small cluster CCER V3.0 verification format group (80 entries) are distributed in the new cluster as follows: 60 entries enter the V3.0 verification + platform delivery format group, 15 entries enter the V3.0 verification + offline delivery format group, and 5 entries enter the V2.0 verification format group. The transition probabilities are 0.75, 0.1875, and 0.0625, respectively. They are concatenated in the process order to form a probability transition sequence: [0.75, 0.1875, 0.0625, ... (transition probabilities of other nodes)].

[0049] In this embodiment, the sequence analysis model is implemented in detail as follows: Model structure: A two-layer LSTM model is used, with the input layer dimension being the length of the transition probability sequence, the hidden layer dimension being 64, the output layer dimension being 1 (with fixed coefficients), and the activation function being Sigmoid; Training dataset: 50,000 format transfer sequences from 2018 to 2024 were extracted from the historical database, including 30,000 positive samples (format stable sequences) and 20,000 negative samples (format frequently changing sequences). Labeling method: sequences with stable format are labeled as 1 (corresponding to format fixation coefficient ≥ 0.8), and sequences with frequently changing format are labeled as 0 (corresponding to format fixation coefficient < 0.8). Loss function: Binary cross-entropy loss function; Optimizer: Adam optimizer, learning rate = 0.001, batch size = 32, training epochs = 50, early stopping strategy (training stops if the validation set loss does not decrease for 5 consecutive epochs). Model performance: The accuracy reached 93.5%, the precision reached 92.8%, and the recall reached 94.1% on the test set.

[0050] The format fixation coefficient is an indicator output by the model (with a value of 0-1) that reflects the stability of the format during the trading process. The higher the coefficient, the more stable the format.

[0051] ROC curve analysis showed that the model had the highest F1 value when the preset coefficient was 0.8. Therefore, the preset fixed coefficient threshold was set to 0.8.

[0052] The confidence score for the format standard is calculated as follows: Node format consistency × 0.6 + Historical approval rate × 0.4. Node format consistency refers to the average similarity of the node format across all threads in the clustering results, and historical approval rate refers to the average approval rate across all threads in the clustering results.

[0053] Reference result generation: When the format fixing coefficient is greater than or equal to the preset coefficient: calculate the format standard confidence of the original small cluster and the matching new cluster, and select the cluster with a confidence of ≥0.9 as the reference result. For example, in case 1: the format fixing coefficient = 0.85 ≥ 0.8, the confidence of the original small cluster and the new cluster are 0.92 and 0.95 respectively, both ≥0.9, so the two clusters are used as the reference result; in case 2: the format fixing coefficient = 0.7 < 0.8, the confidence of the original small cluster and the new cluster are 0.7 and 0.75 respectively, and are weighted and merged according to the ratio of the number of threads (80:120), and the format with the highest frequency after weighting is taken to form the reference result.

[0054] In this embodiment, the transaction attributes are the characteristic information of the transaction thread, such as the type of transaction target, transaction size, applicable carbon accounting standard, and entity qualification requirements.

[0055] Transaction feedback is evaluation data on the process of past transactions, such as approval rate, processing time, dispute rate, and user satisfaction.

[0056] The optimal thread recommendation process is based on transaction attributes and transaction feedback, and is the process of selecting the thread process template with the best overall performance from the reference results; the only blank thread is a standard template that removes specific transaction data and only retains process nodes and format requirements, and can be filled with data to generate an actual thread later.

[0057] In this embodiment, a scoring model is constructed, such as score = approval rate × 0.4 + 1 (day) / processing time (days) × 0.3 + 1 - dispute rate × 0.3. The weights are determined by the analytic hierarchy process (AHP). Ten carbon trading industry experts are invited to score the importance of the three indicators. Finally, the approval rate is weighted at 0.4, the processing time at 0.3, and the dispute rate at 0.3. It should be noted that the processing time is in calendar days, and any fraction of a day is counted as a full day.

[0058] Select the thread with the highest score, remove the specific transaction data, and retain only the process node sequence, format requirements, and verification rules to form a unique blank thread. For example, in the reference results, thread A has a score of 0.7155 and thread B has a score of 0.669. At this time, select the process of thread A, remove the specific data, and generate a blank thread: Node 1 (filing application, CCER V3.0 template) → Node 2 (verification, PDF report) → Node 3 (listing, platform template) → Node 4 (delivery, registration interface) → Node 5 (settlement, transfer voucher).

[0059] In this embodiment, the current format is the standard format currently recommended for each trading node, based on the format with the highest confidence in historical transactions; the unique standard thread is a standard trading process template for a certain trading type composed of the current formats of each node in process order; the standard thread for green carbon trading is a collection of unique standard threads corresponding to all trading types, forming a unified format specification system adapted to international carbon accounting standards.

[0060] In this embodiment, the confidence level of each format is calculated based on the thread ratio, approval rate, and compatibility with international standards. The format with the highest confidence level is then selected as the current format for the node. The current formats of each node are then concatenated to form a unique standard thread. Finally, all unique standard threads for all transaction types are aggregated to form a green carbon trading standard thread set. For example, in the reference results of blank thread node 2 (third-party verification), 85 threads use the CCER V3.0 standard PDF report, including project boundaries, monitoring data, and fields such as project number and emission reduction amount (…). The format (two decimal places) has the highest confidence level of 0.98, so it is determined as the current format for node 2. After concatenating the formats of each node, the unique standard thread for CCER V3.0 forestry carbon sink primary transactions is obtained. Then, the standard threads of other transaction types are summarized to form a complete system.

[0061] The beneficial effects of the above technical solution are as follows: by mining threads, performing two-dimensional cluster analysis, probabilistic transition modeling, and format fusion screening of historical green carbon trading data, the standardized conversion of green carbon trading data formats under different trading types and different international carbon accounting standards has been achieved. A unified trading process and format specification system has been constructed, which effectively solves the problem of difficulty in cross-entity and cross-regional mutual recognition of green carbon trading data, improves the consistency of carbon accounting and the compliance of the trading process, and provides standardized technical support for the global circulation and reliable trading of green carbon data.

[0062] This invention provides a method for green carbon data trading and standardization conversion adapted to international carbon accounting standards. It obtains trading threads by mining each historical green carbon trading data entry from a historical database, including: For each historical green carbon transaction data, feature extraction is performed to obtain transaction features and green carbon features, and a transaction template is selected from the dual-feature database based on the transaction features and green carbon features; Simultaneously, the flow process of corresponding historical green carbon trading data and the time-sequential operation behaviors involved in the flow process are mined from the historical database, and the start and end points of each operation behavior are extracted to construct the initial thread; The transaction template is matched with the initial thread to obtain the transaction thread.

[0063] In this embodiment, the dual-feature database is a pre-built index database that stores various green carbon trading templates using transaction features and green carbon features as labels. Each template corresponds to a set of specific feature labels, as shown in Table 1. Table 1. Structure of the Dual-Feature Database Configure unique feature tags for each transaction template. For example, the tags for the CCER V3.0 forestry carbon sink primary listing transaction template are: transaction features (primary transaction, listing mode, national carbon trading platform) + green carbon features (forestry carbon sink, CCER V3.0 standard, certification by a certain institution). The matching degree between the extracted features and the template labels is calculated as follows: Matching degree = (Number of matching transaction features / Total number of transaction features) × 0.5 + (Number of matching green carbon features / Total number of green carbon features) × 0.5. It should be noted that the weights are determined by the analytic hierarchy process. The influence of transaction features and green carbon features on template adaptability is roughly equal, so each accounts for 0.5. When the matching degree is ≥ 0.8, the template is selected as the suitable transaction template.

[0064] Select a template with a matching degree greater than or equal to the preset threshold (e.g., 0.8). If there are multiple templates that meet the criteria, select the template with the highest matching degree as the matching transaction template. The transaction template is a preset standardized green carbon transaction process and format framework, which includes fixed process nodes, node order, required file formats for each node, data field requirements, audit rules, and other content.

[0065] For example, the characteristics extracted from the primary listing and trading data of a forestry carbon sink project in 2023 are as follows: Trading characteristics: primary trading, listing mode, trading platform is the national carbon trading platform, trading entities are enterprises, trading volume is 5,000 tons ; Green carbon characteristics: Project type is forestry carbon sink, filing standard is CCER V3.0, certification body is Institution 1, accounting period is 1 year, project location is Zone B; matching the CCER V3.0 forestry carbon sink primary listing and trading template in the dual-feature database, calculate the matching degree: The number of matching transaction features is 3 / 4 (the transaction type, mode, and platform all match, but the transaction subject does not affect the matching). The green carbon feature matching score is 3 / 3 (the project type, filing standards, and certification body all match). Matching degree = 0.75×0.5+1×0.5=0.875≥0.8, therefore this template is selected as the matching transaction template.

[0066] The process is a complete flow path for green carbon trading from initiation to completion, including all the systems, platforms, and review stages involved, such as the flow link from project filing system to certification body system to carbon trading platform to registration and settlement system.

[0067] Operations based on time sequence are each independent action in the transaction process ordered by timestamp. Each independent action is recorded in the form of action + timestamp. For example, the filing application is submitted at time t1 and the verification and approval are passed at time t2. Each operation has a clear sequence.

[0068] The starting point of an operation is the identifier for initiating each operation. It can be represented by the start timestamp of the operation, the ID of the system module that initiated it, or the triggering condition of the operation, such as: logging into the project filing system application page or the certification agency receiving the filing materials.

[0069] The operation end point is the completion marker for each operation behavior, which can be reflected as the end timestamp of the operation, system feedback marker, or result status, such as: the system generates a filing acceptance number, the review is approved and a verification report is generated.

[0070] In this embodiment, all operations are chained together in chronological order, and an initial thread is constructed in the form of [starting point] operation [ending point]. Only the process order and point information are retained, without adding any format or rule requirements.

[0071] After analyzing the flow of a particular transaction, the identified operational behaviors and entry / exit points are as follows: Operation 1: Submit the filing application. Starting point: [Project Filing System - Application Module, t01], Ending point: [System-generated Acceptance Number BH20230501001, t02]; Operation 2: Verification body review, starting point: [Verification body system - receive filing materials, t03], ending point: [review passed and verification report generated, t04]; Operation 3: Platform listing application, starting point: [National Carbon Trading Platform - Listing Module, t05], ending point: [Listing information approved and publicized, 2023-05-12 16:00:00]; Operation 4: Buyer delisting transaction, starting point: [Platform Trading Hall, t06], ending point: [Transaction completed and transaction certificate generated, t07]; the initial thread is constructed as follows: Operation 1 (submit filing application) → Operation 2 (verification and review) → Operation 3 (listing application) → Operation 4 (delisting transaction), each node is accompanied by corresponding start and end point information, where t01, t02, t03, t04, t05, t06, and t07 are arranged in chronological order.

[0072] In this embodiment, node semantics and process matching are performed: preset nodes in the transaction template (such as filing application, verification review, and listing transaction) are matched with the operation behaviors in the initial thread. The matching rules are as follows: the semantics of the operation behavior are consistent with the template node name; the time sequence of the operation is consistent with the sequence of the template nodes; and the start and end points of the operation are consistent with the input and output requirements of the template node. Adjustment of differential nodes: If there are operation nodes in the initial thread that are not covered by the template (such as additional local regulatory review), then add the node at the corresponding position in the template and supplement the format requirements according to the historical format data of the node; if there are general nodes in the template that are not reflected in the initial thread (such as secondary review), then mark them as optional nodes and retain them; Format rule fusion: Associate the format requirements (such as file format, field rules, and data units) of each node in the transaction template with the matched initial thread node to supplement the format specifications of each node; Trading thread generation: Connect the matched, adjusted, and formatted process nodes in sequence to form a complete trading thread. Each node contains three types of information: operation behavior, start and end points, and format specifications.

[0073] For example, matching the CCER V3.0 forestry carbon sink primary listing and trading template obtained in step 1 with the initial thread in step 2: The template node "Project Filing Application" matches Operation 1 in the initial thread. Additional template format requirements: A PDF version of the filing application must be submitted, with fields including project number, emission reduction estimate, and monitoring plan. The unit for emission reduction is... Retain two decimal places; the template node "Third-Party Verification and Audit" matches operation 2 in the initial thread, and supplements the template format requirements: the verification report must be prepared in accordance with the CCER V3.0 standard, including project boundaries, monitoring data, emission reduction calculation table, and must be stamped with the official seal of the verification body; The template node "Platform Listing and Trading" is merged and matched with operations 3 and 4 of the initial thread (listing and delisting are one node in the template, while the initial thread is split into two operations). Supplementary format requirements: Listing applications must submit a verification report and project filing certificate; the transaction certificate generated after delisting must include the transaction number, information of both parties, and the transaction emission reduction amount. The final generated transaction thread node information is as follows: Node 1: Submit the filing application. Start and end points: Filing system application page → Generate acceptance number. Format requirements: PDF filing application, including fields such as project number. Node 2: Verification and review, start and end points: verification system receives materials → review and approval, format requirements: CCER V3.0 standard verification report; Node 3: Listing and delisting transactions, start and end points: platform listing module → generate transaction certificate, format requirements: submit filing / verification documents, transaction certificate includes transaction information.

[0074] The beneficial effects of the above technical solution are as follows: by selecting an appropriate trading template through two-dimensional feature matching, restoring the actual initial thread based on the operation point, and then merging and aligning the template specifications with the actual process, the standardized construction of green carbon trading threads is achieved. This ensures that the threads conform to the format and rule requirements of carbon accounting standards and fit the actual trading flow, providing complete, accurate, and standardized basic data support for subsequent format clustering, probability analysis, and standard thread generation, effectively improving the adaptability of trading thread construction and the efficiency of subsequent standardized conversion.

[0075] This invention provides a method for green carbon data trading and standardization conversion adapted to international carbon accounting standards, which involves matching the trading template with the initial thread to obtain a trading thread, including: Based on the predefined transaction description set of each template node in the transaction template, semantic matching and window association analysis are performed between the operation start point and operation end point of each operation behavior in the initial thread; When the semantic matching degree is greater than the preset matching degree, the template node is retained unchanged; Otherwise, based on the window association analysis results, determine the variable group of the corresponding operation behavior, and compare the variable types involved in the variable group with the transaction description features of the predefined transaction description set of the adjacent nodes of the extracted template node to determine the node addition position, and determine and retain the addition transaction description set of the corresponding addition node based on the variable group. The transaction thread is obtained based on the sequential positional relationship of the retained template nodes.

[0076] In this embodiment, the predefined transaction description set of the template node is a standardized text / field description set corresponding to each preset process node in the transaction template. It includes node name, function definition, required operations, input and output requirements, format specification keywords, etc., and is the standard instruction manual of the template node.

[0077] The operation transaction description set is a collection of actual information corresponding to each operation behavior in the initial thread. It is extracted from text / data content such as system logs, operation notes, and data fields between the start and end points of the operation, thus restoring the behavioral characteristics of the actual operation.

[0078] Semantic matching is performed using a BERT-base-uncased pre-trained model. The specific steps are as follows: Text preprocessing: The predefined transaction description set and operation transaction description set are segmented, stop words are removed, and lowercase is applied; Word vector generation: Input the preprocessed text into the BERT model to generate 768-dimensional sentence vectors; Similarity calculation: Calculate the cosine similarity between two sentence vectors as the semantic matching degree; Preset matching degree: Determined through experiments, when the semantic matching degree is ≥0.8, the two descriptions are considered to be semantically consistent.

[0079] In this embodiment, the window association analysis process is as follows: Extracting standard time logic for template nodes: such as reasonable time range for nodes (e.g., 1-3 working days for filing application processing) and sequential constraints of adjacent nodes (e.g., verification review must be conducted after filing application). Compare the actual operation time window: determine whether the operation time is within the reasonable time range of the template node, and whether the order of the operation meets the order constraints of the template node. If both are met, the window association analysis result is a match; otherwise, it is a mismatch. The node matching comprehensive score is obtained by combining the semantic matching degree and the window association result: Comprehensive score = semantic matching degree × 0.7 + window association score × 0.3 (window association score: 1 for matching, 0 for non-matching).

[0080] In this embodiment, the time window size is determined based on the transaction type: Primary market transactions (project filing, certification, listing): The reasonable timeframe is 1-3 business days; Secondary market transactions (transfer, delivery, settlement): The reasonable timeframe is 1-2 business days; Cross-border carbon trading: A reasonable processing time is 3-7 business days. Window correlation score: If the processing time is within the reasonable range and the sequence conforms to the template logic, the score is 1; otherwise, it is 0.

[0081] In this embodiment, the determination of the group of changing variables is as follows: Compare the description sets, time windows, and process logic of template nodes and operational behaviors; Points of difference: For example, the operation involves an additional preliminary review by the local forestry department compared to the template, or the operation is named "local filing review" instead of "project filing application," and the time taken is 5 working days, which exceeds the 1-3 working days range of the template. The differences are categorized by text, time, and logic to form variable change groups, such as variable change group = [text description differences (addition of local regulatory keywords), time window differences (time consumption exceeds template range), process logic differences (pre-approval stage exists)].

[0082] In this embodiment, the feature extraction and comparison of adjacent nodes are performed as follows: Extract the predefined transaction description set of the previous adjacent node (such as transaction initiation) and the next adjacent node (such as verification and review) of the current mismatched node to obtain the input and output features: for example, the output of the transaction initiation node is the project basic information table, and the input of the verification and review node is the project materials that have been filed and approved. By comparing the differences in the variable group (such as the output of local regulatory review being the preliminary review opinion of local filing) with the input and output characteristics of adjacent nodes, it can be determined that the new node should be located between the transaction initiation and the verification review, because its output preliminary review opinion is a prerequisite condition for the verification review node.

[0083] In this embodiment, the generation of the added node description set is as follows: Based on the text description of the variable group (such as local regulatory review), time window (actual time is 3 working days, and the standard time is set to 2-4 working days in combination with the overall process logic of the template), and input and output requirements (input is the basic information table of the project, and output is the preliminary review opinion of the local filing), a predefined transaction description set for adding nodes is constructed. Supplementary format guidelines: If the preliminary review comments need to be stamped with the official seal of the local forestry department, the file format should be PDF, including the preliminary review conclusion and comments section, and ensure consistency with the overall format requirements of the template.

[0084] For example, in a forestry carbon sink transaction, the initial thread's operation is the preliminary review of the local forestry department's filing. The matching degree with the template node project filing application is 0.7 (below the preset threshold of 0.8), so node addition processing is required. Changes in the variable group: [Textual differences (the operation name is local filing preliminary review, containing keywords from local forestry departments), time differences (takes 3 working days, template node takes 1-3 working days, although not exceeding the time limit, the descriptions are very different), logical differences (the preliminary review is a prerequisite for verification and auditing, and this node is not present in the template)]; Adjacent node characteristics: The output of the transaction initiation at the previous node is the project basic information table, and the input of the verification and review at the next node is the filing approval materials; Added position: after the transaction initiation node and before the verification and review node; Add a transaction description set: Node name: Local forestry department filing preliminary review; Function definition: Receive basic project information form, review the local compliance of forestry carbon sink projects, issue preliminary review opinions; Time range: 2-4 working days; Format specification: Preliminary review opinions are PDF files, including project preliminary review conclusions, compliance explanations, and affixed with the official seal of the local forestry department.

[0085] Node sorting is performed by sorting all nodes according to the preset order of the original nodes in the template, combined with the insertion position of the added nodes. Thread construction involves sequentially connecting sorted nodes, with each node associated with a predefined set of transaction descriptions, format specifications, and time requirements to form a complete transaction thread.

[0086] Combining the previous example, the final node order of the transaction thread is as follows: Node 1: Transaction Initiation (original node of the template, retained if the matching degree meets the standard); Node 2: Preliminary review of filing by local forestry departments (an additional node, generated based on the group of changing variables); Node 3: Project Filing Application (original node in template, retained if matching meets the standard); Node 4: Verification and review (original template node, retained if matching meets the standard); each node is associated with a corresponding standardized description and format specification, forming a complete transaction thread.

[0087] The beneficial effects of the above technical solution are as follows: By semantic matching and window association analysis, the differences between template nodes and actual operations are accurately identified. For scenarios where the matching degree is not up to standard, the addition of nodes is flexibly achieved by analyzing variable groups and comparing the characteristics of adjacent nodes. This ensures that the transaction thread meets the template specification requirements of the carbon accounting standard, while also adapting to the differences in actual transaction flow in different regions and scenarios. It effectively solves the problem of adapting standardized templates to personalized processes, and provides complete, accurate, and highly adaptable transaction thread data support for the subsequent format clustering and standardization conversion of green carbon transaction data.

[0088] This invention provides a method for green carbon data trading and standardization conversion adapted to international carbon accounting standards. It involves sorting the secondary partitioning results by size to obtain an initial list, including: The corresponding large clustering results are further divided according to the transaction type to obtain several first partition sets of the corresponding large clustering results, and each first partition set corresponds to a transaction type. The transaction types of each first partition set are matched with the transaction types of the corresponding partition thread set to obtain the matching relationship; At the same time, determine the first number of transaction data that is completely consistent with the corresponding partition thread set and the second number of similar transaction data in each first partition set, and obtain the statistical array of the corresponding first partition set; The matching relationship is optimized based on the statistical array to obtain the sorting value of each first partition set; Sort all the first partition sets corresponding to the large clustering results according to the size of the sorting value to obtain an initial list, where the first partition set is the sub-partition result.

[0089] In this embodiment, if a large clustering result contains 200 transaction threads, they are grouped together based on the overall similarity of the format chain, but the transaction type labels of the threads are divided into three categories: Tag A: CCER V3.0 Forestry Carbon Sink Primary Trading (120 items); Tag B: CCER V3.0 Photovoltaic Emission Reduction Primary Trading (50 items); Tag C: VER Standard Cross-border Carbon Trading (30 items) After secondary partitioning, three first partition sets are obtained: {Set A (120 items), Set B (50 items), Set C (30 items)}, each set corresponds to only one type of transaction.

[0090] Type label matching involves labeling each partitioned thread set with a corresponding transaction type label and comparing it one by one with the label of the first partition set; The matching relationship is established when the transaction type label of the first partition set is completely consistent with the label of the partition thread set. The mapping relationship between the two is established and recorded as the first partition set ID → the corresponding partition thread set ID. The exception handling method is that if there is a first partition set for which there is no corresponding thread set, it is marked as no match and will not be included in the subsequent statistical process.

[0091] For example: generate 3 first partition sets, corresponding to the following partition thread sets respectively: Set A (CCER Forestry Carbon Sink Level 1) ↔ Divide into thread set A (all CCER Forestry Carbon Sink Level 1 trading threads, a total of 150). Set B (CCER Photovoltaic Level 1) ↔ Divide into thread set B (all CCER Photovoltaic Level 1 trading threads, a total of 60). Set C (VER cross-border transactions) ↔ Divide into thread set C (all VER cross-border transaction threads, a total of 40). There is no matching set, so all valid matching relationships are established.

[0092] Completely identical transaction data refers to the threads in the first partition set, which are completely matched with the threads in the corresponding partition set in terms of process nodes, node format, field content, and validation rules (the similarity of the format chain feature vector is ≥ the preset threshold, such as 0.99).

[0093] Similar transaction data refers to the threads in the first partition set. Compared with the threads in the corresponding partition thread set, the process nodes and core format are consistent, but there are differences in some non-critical fields, parameters or time consumption (the similarity of the format chain feature vector is within the preset range, such as 0.8~0.99).

[0094] The first quantity is the number of transaction data in the first partition set that are completely identical to the corresponding partition thread set.

[0095] The second quantity is the number of transaction data in the first partition set that are similar to the corresponding partition thread set.

[0096] The statistics array contains a structured array of first and second counts, for example: [first count, second count], which is used to quantify the degree of matching between the first partition set and the corresponding partition thread set.

[0097] Taking set A as an example: Set A contains 120 threads, which corresponds to a thread set A containing 150 threads. After comparison, 80 threads were found to be completely identical to the threads in thread set A (first count = 80). Another 30 threads have a similarity of 0.8 to 0.99 with the threads in the partitioned thread set A (second number = 30). The remaining 10 threads have a similarity of <0.8 and are not included in the statistics; Therefore, the statistical array of set A is [80, 30].

[0098] Similarly, the statistical array for set B (50 threads) is [30, 15], and the statistical array for set C (30 threads) is [20, 5].

[0099] The optimization process involves combining the first and second quantities in the statistical array to construct a quantification model and calculate a value that reflects the matching priority of the first partition set.

[0100] In this embodiment, the sorting value = (first quantity / total number of threads in the corresponding thread set) × 0.7 + (second quantity / total number of threads in the corresponding thread set) × 0.3. It should be noted that the formula can be adjusted according to business needs, such as adding the total number of threads in the first thread set as a normalization factor, etc. Moreover, completely consistent data has higher reference value for subsequent standardization, so it has a higher weight. For example, the weight of completely consistent data is 0.7, and the weight of similar data is 0.3.

[0101] For example, calculating the sort value of each first partition set: Set A: Corresponds to the thread set A with a total of 150 threads. The sorted value of the array [80,30] is (80 / 150)×0.7+(30 / 150)×0.3=0.433. Set B: The corresponding thread set B has a total of 60 threads. The sorted value of the array [30,15] is (30 / 60)×0.7+(15 / 60)×0.3=0.425. Set C: The corresponding thread set C has a total of 40 threads. The sorted value of the array [20,5] is (20 / 40)×0.7+(5 / 40)×0.3=0.3875.

[0102] In this embodiment, the sorting values ​​of all first partition sets are sorted in descending order, with the sets with higher sorting values ​​appearing first. The first partition sets are then arranged sequentially according to the sorting results to form an initial list, ensuring that each element in the list has a valid matching relationship. Abnormal sets without matching are removed. For example, if the calculated sorting values ​​are: set A 0.433, set B 0.425, set C 0.3875, after sorting in descending order, the initial list is: [set A (0.433), set B (0.425), set C (0.3875)]. During subsequent candidate thread screening, set A, which appears earlier in the list, will be processed first.

[0103] The beneficial effects of the above technical solution are as follows: by dividing the large clustering results of the format chain into secondary subdivisions according to transaction type, establishing matching relationships with the subdivided thread sets, quantifying statistical similarity and calculating priority ranking values, an ordered initial list is finally generated, realizing the priority ranking of different transaction type subsets, providing a reliable priority basis for subsequent candidate thread screening, effectively improving the efficiency and accuracy of green carbon transaction data format matching and screening, and ensuring the adaptability and rationality of the standardized conversion process.

[0104] This invention provides a method for green carbon data trading and standardization conversion adapted to international carbon accounting standards, which involves selecting candidate threads from each large clustering result that meet the candidate criteria of the corresponding small clustering result, including: Construct a first matrix of corresponding small clustering results and corresponding initial lists based on the transaction format matching set, matching proportion probability and division level confidence; Based on each candidate index in the candidate criteria of the corresponding small clustering results, traverse each element in the first matrix in turn, and mark the matching elements as first; The importance of each row vector in the first matrix is ​​obtained based on the number of times each labeled element is labeled, and the number of sub-partition results to be filtered for each row vector is determined. Candidate threads are obtained by filtering from the corresponding sub-partition results.

[0105] In this embodiment, the first matrix has rows of small clustering results and columns of sub-partitioning results from the initial list. Each cell stores a two-dimensional matrix corresponding to the matching index, which is the basic data structure for subsequent filtering. Assuming there are two small clustering results (small cluster A and small cluster B), and the initial list contains three sub-partitioning results (sub-partition 1, sub-partition 2, and sub-partition 3), a 2×3 matrix is ​​constructed, as shown in Table 2: Table 2. 2×3 matrix table In this embodiment, based on statistical analysis of 100,000 historical transaction data from the national carbon trading platform from 2018 to 2025, when the matching probability is ≥70% and the confidence level of the classification is ≥0.8, the format compliance rate of the selected candidate threads reaches over 95%. Therefore, the preset candidate criteria are: Indicator 1: Matching probability ≥70%; Indicator 2: Confidence level of classification ≥0.8; Indicator 3: Number of matching threads ≥50. Iterate through each cell of the first matrix row by row, and check in turn whether all three indicators of the cell meet the requirements of the candidate indicators. Cells that meet all the criteria are marked as "valid" (assigned a value of 1), and cells that do not meet the criteria are marked as "invalid" (assigned a value of 0).

[0106] If the above matrix data is used, the candidate indicators are iterated and verified: Small cluster A& sub-partition 1: Matching percentage 66.7% < 70% → Not satisfied, marked 0; Small cluster A & sub-partition 2: Matching percentage 83.3% ≥ 70%, confidence 0.85 ≥ 0.8, matching set ≥ 50 → satisfied, marked as 1; Small cluster A& sub-partition 3: Matching percentage 50% < 70% → Not satisfied, marked 0; Small cluster B & sub-partition 1: Matching percentage 58.3% < 70% → Not satisfied, marked 0; Small cluster B & sub-partition 2: Matching percentage 66.7% < 70% → Not satisfied, marked 0; Small cluster B & sub-partition 3: Matching percentage 62.5% < 70% → Not satisfied, marked 0; In the end, only the cell of “small cluster A & sub-partition 2” was marked as valid.

[0107] The number of markings is the number of cells in a row of the first matrix (corresponding to a small cluster result) that are marked as "valid".

[0108] The importance of row vectors is based on the priority of small clustering results quantified by the number of labels; the more labels, the higher the importance.

[0109] Number of filtering results for sub-partitions: The number of filtering threads allocated to each sub-partition result based on the importance of the row vectors. The higher the importance, the more filtering threads can be allocated.

[0110] Importance = Number of times marked / Total number of sub-partitions of the initial list.

[0111] Based on the importance value, allocate the number of results to be filtered for each sub-segment proportionally, for example: Importance ≥ 0.5: Select the top 80% of threads with the highest matching degree in the sub-partition results; Importance level 0.3-0.5: Filter the top 50% of threads; Importance < 0.3: Filter only the top 30% of threads or do not filter; For each (small cluster, sub-partition) combination marked as "valid", extract the number of threads that have the highest matching degree with the small clustering format from the sub-partition results as candidate threads.

[0112] If the number of labels for small cluster A is 1, the initial list has 3 sub-partition results, and the importance = 1 / 3 ≈ 0.333; The number of times small cluster B is labeled is 0, and its importance is 0. Based on the importance, the screening ratio of sub-partition 2 corresponding to small cluster A is 50%. The total number of threads in sub-partition 2 is 60, so the number of threads to be screened is 60 × 50% = 30. From sub-partition 2, the 30 threads with the highest matching degree with the format of small cluster A are extracted as candidate threads. If the calculation result is a decimal, it is rounded up.

[0113] The beneficial effects of the above technical solution are as follows: by constructing a multi-dimensional matching matrix, marking matching combinations that meet the candidate indicators, quantifying the importance based on the number of markings and allocating the number of screenings, the quantitative screening of the initial list sub-partition results is realized. This not only ensures the format matching degree and reliability of the candidate threads, but also improves the screening efficiency through priority allocation. It provides high-quality and suitable sample support for subsequent secondary format clustering and standard thread generation, effectively improving the accuracy and efficiency of green carbon trading data standardization conversion.

[0114] This invention provides a method for green carbon data trading and standardization conversion adapted to international carbon accounting standards, determining the standard format transfer probability of each small cluster result, including: Construct a first clustering graph corresponding to several new clustering results of the thread set before induction, and at the same time, construct a second clustering graph of several original small clustering results of the thread set before induction. By comparing and analyzing the first clustering graph and the second clustering graph, the first distribution of threads that originally belonged to each small clustering result in the first clustering graph and the second distribution in the second clustering graph are determined, and the change point of the second distribution based on the first distribution is determined, wherein the point is each transaction data in several new clustering results of the corresponding inductive thread set; The performance coefficient of each change point is obtained based on the nearest distance between each change point and the corresponding small cluster result, the cluster distance between each change point and the central cluster of the corresponding small cluster result, and the first distance between the central cluster of the corresponding small cluster result and the point corresponding to the nearest distance. Based on the variation distribution of the variation points and the performance coefficient of each variation point, the variation trend of the variation distribution is determined, and the standard format transition probability of the corresponding small clustering results is determined.

[0115] In this embodiment, the t-SNE algorithm is used to reduce the dimensionality of the high-dimensional format chain feature vector to two dimensions. The parameters are set as follows: perplexity = 30, learning rate = 200, and number of iterations = 1000. Under this parameter combination, the dimensionality-reduced cluster graph can clearly distinguish different format clusters and retain the local structure of the original data.

[0116] The first clustering graph construction involves assigning a unique color to each cluster in the new clustering results, marking the thread data points by color in a two-dimensional coordinate system, and simultaneously calculating and labeling the coordinates of the center cluster of each cluster (the average coordinates of all points in that cluster).

[0117] The second clustering graph is constructed by using the coordinate system range, scale, and color marking rules of the first clustering graph. The thread data points of the original small cluster results are drawn in the same coordinate system, and the coordinates of the center clusters of each original cluster are also marked.

[0118] For example, the new clustering result of a certain inductive thread set contains two groups: cluster A (containing 120 threads, marked in red) and cluster B (containing 70 threads, marked in blue); the original small clustering result also contains two groups: cluster A (containing 100 threads, red) and cluster B (containing 50 threads, blue).

[0119] Reduce the dimensionality of the format chain feature vectors of all threads to two dimensions, with the x-axis representing node format similarity and the y-axis representing field consistency score, to obtain the (x, y) coordinates of each thread; First cluster diagram: Red dots represent the 120 threads of the new cluster A, and blue dots represent the 70 threads of the new cluster B. The coordinates of the center clusters are marked as follows: A center (2.1, 3.2), B center (5.3, 1.8). Second cluster diagram: Red dots represent the 100 threads of the original cluster A, and blue dots represent the 50 threads of the original cluster B. The coordinates of the center clusters are marked as follows: original center of A (2.0, 3.1), original center of B (5.1, 1.7). Both graphs use coordinate systems ranging from x-axis (0-7) to y-axis (0-5), and are perfectly aligned.

[0120] In this embodiment, the first distribution is the set of coordinate distributions of thread data that originally belonged to a certain original small cluster in the new clustering result in the first clustering graph; the second distribution is the set of coordinate distributions of thread data in the original small clustering result in the second clustering graph.

[0121] In this embodiment, the change determination rule is as follows: Cluster affiliation change: If a thread originally belongs to cluster X in a small cluster, but is assigned to cluster Y in the new cluster (Y≠X), it is directly determined as a change point; Significant positional change: If the distance between the coordinate point of a thread in the new cluster and the center of the original cluster exceeds the average distance between the thread and the center of the original cluster plus 2 standard deviations, it is determined as a changed point even if the cluster affiliation remains unchanged.

[0122] Change point labeling: All change points are identified with special labels (such as triangles) in the first cluster diagram, and the original cluster affiliation, new cluster affiliation and coordinate information of each change point are recorded.

[0123] For example, thread T1 originally belonged to cluster A (coordinates (2.2, 3.3), distance from the original center of A was 0.14). In the new cluster, it still belongs to cluster A, but its coordinates become (4.0, 2.0), and the distance from the center of the new cluster is about 2.25, which is much greater than the average distance of threads from the center in the original cluster A (0.3 + 2 times the standard deviation 0.4 = 0.7). Therefore, it is determined to be a point with a significant change in position. Thread T2: Originally belonging to cluster A (coordinates (2.0, 3.0)), it was reassigned to cluster B (coordinates (5.2, 1.9)) in the new cluster. Its cluster affiliation has changed, and it is determined to be a change point. Thread T3: Originally belonged to cluster B (coordinates (5.0, 1.8)), it still belongs to cluster B in the new cluster, with coordinates (5.1, 1.7). The distance from the center of the new B is about 0.22, which is less than the threshold of 0.7, so it is determined to be a non-changed point. In the end, a total of 15 changed points were marked, of which 10 were cluster affiliation change points and 5 were positional change points.

[0124] In this embodiment, the closest distance d1 is obtained by traversing all the non-moving thread points in the original cluster, calculating the Euclidean distance between the moving point and each point, and taking the minimum value as d1. The cluster distance d2 is the Euclidean distance between the point of change and the original cluster center, calculated using the following formula: ,in, Let P be the coordinates of the point of change; The original cluster center coordinates The first distance d3 is obtained by finding the invariant thread point Q corresponding to the nearest distance d1, and then calculating the Euclidean distance between Q and the original cluster center cluster. .

[0125] In this embodiment, the performance coefficient = w1×(d2-d3) / (d2+d3)+w2×(d1 / max_d1), where w1 and w2 are weights (w1=0.6, w2=0.4, the weights are determined through orthogonal experiments, and the performance coefficient has the highest correlation with the degree of format change under this combination), and max_d1 is the maximum value of the nearest distance among all change points. It should be noted that change points with negative performance coefficients are considered to have no significant format change and are not included in subsequent calculations.

[0126] In this embodiment, the variation distribution is the set of all variation points' positions and cluster affiliations in the first clustering graph, reflecting the overall direction and scope of the format variation.

[0127] The trend of change is the direction of format evolution obtained from the analysis of change distribution and performance coefficients. For example, the trend of threads from the original cluster A shifting to the new cluster B is obvious.

[0128] In this embodiment, the analysis focuses on the trend of change: Statistical analysis of the direction of change in cluster affiliation: For example, among the change points in the original cluster A, the percentage of threads that migrated to the new clusters B and C; Analysis of performance coefficient distribution: The direction of shift in the concentration of high performance coefficients reflects the main trend of format change.

[0129] Transition probability calculation: For each original small cluster X, calculate its probability of transitioning to the new cluster Y: P(X→Y) = (number of threads transitioning from X to Y + number of high-performance coefficient change points transitioning to Y × weight) / total number of threads in the original cluster X. The weight is determined to be 0.5 through experiments. Under this weight, the prediction accuracy of the transition probability reaches 91.2%. This weight is used to reflect the contribution of high-performance coefficient change points to the transition probability, ensuring that the sum of the transition probabilities of each original cluster to all new clusters is 1. High-performance coefficient change points refer to change points with performance coefficient ≥ 0.7.

[0130] For example, if the original cluster A has a total of 100 threads: 60 threads remain in the new cluster A (non-change point); 30 threads were moved to the new cluster B (the point of change, of which 10 are high performance coefficient points); Ten threads changed positions significantly but remained at point A (the point of change, with an average performance coefficient of 0.6); calculate the transition probability: P(A→A)=(60+10×0.5) / 100=65 / 100=0.65; P(A→B)=(30+10×0.5) / 100=35 / 100=0.35; Trend analysis: There is a clear trend of threads from the original cluster A shifting to the new cluster B. The high performance coefficients are mostly concentrated in the threads that are shifting to B, indicating that the A-type format is evolving into the B-type format.

[0131] The beneficial effects of the above technical solution are as follows: by constructing a visual clustering diagram of the old and new clustering results, identifying the points of change in thread distribution, quantifying the performance coefficients of the points of change, and analyzing the trend of change, a precise quantitative analysis of the evolution trend of green carbon trading format is achieved. This provides reliable basic data support for the subsequent construction of probability transfer sequences and analysis of format stability, effectively improving the objectivity and accuracy of format transfer probability calculation, and ensuring the scientific nature and adaptability of the green carbon trading data standardization conversion process.

[0132] This invention provides a method for green carbon data trading and standardization conversion adapted to international carbon accounting standards, which integrates and analyzes corresponding small clustering results and matching new clustering results, including: Using the second distance between the cluster center of the corresponding small clustering result and the cluster center of the matched new clustering result as the radius, draw a circle with the cluster center of the corresponding small clustering result as the center, and retain the transaction data within the drawn circle as the first retention; All identical transaction data in the corresponding small clustering results and the matching new clustering results are globally and individually retained; Transaction data with performance coefficients greater than the performance threshold in the newly matched clustering results are retained as a second group. The first retention result, the global single retention result, and the second retention result are all processed by single retention to obtain a reference result. Among them, the completely identical transaction data in the first retention result are locally single retained.

[0133] In this embodiment, for example, the corresponding small clustering results are: transaction data 1, transaction data 2, transaction data 3, transaction data 4, and transaction data 5; The new clustering results are: Transaction Data 1, Transaction Data 3, Transaction Data 7, Transaction Data 8; At this point, the global single retention applies to transaction data 1 and transaction data 3, retaining only the results of the corresponding small cluster or the matching new cluster.

[0134] For example: the transaction data that belong to the circle in the corresponding small clustering results are: transaction data 3 and transaction data 4; The new clustering results that match the transaction data within the drawn circle are: Transaction data 3 and Transaction data 8; At this point, the partial single retention is for transaction data 3, transaction data 4, and transaction data 8; Since transaction data 3 exists in both global single retention and local single retention, single retention processing is performed at this time, that is, only one place needs to be retained; If the performance coefficient of transaction data 7 is greater than the performance threshold, then transaction data 7 will be retained for the second time. In summary, transaction data 1, transaction data 3, transaction data 4, transaction data 8, and transaction data 7 are used as reference results.

[0135] In this embodiment, the performance of the fusion results under different performance thresholds was compared experimentally. When the performance threshold was 0.7, the format stability (mean value of format fixation coefficient) of the fusion result reached 0.82, and the scene adaptation rate reached 93.5%, with the best overall performance. Therefore, the preset performance threshold was 0.7.

[0136] The beneficial effects of the above technical solution are: by using a combination of layered filtering, global deduplication, retention of high-performance data, and multiple deduplication integration, it can retain the core stable format data of the original small clusters, and absorb the high-quality format data with strong adaptability in the new clusters, effectively avoiding data redundancy and duplication, so that the final reference results have both format stability and scenario adaptability, providing a more accurate and complete data foundation for the subsequent standard thread generation.

[0137] This invention provides a method for green carbon data trading and standardization conversion adapted to international carbon accounting standards, obtaining a unique blank thread with corresponding reference results, including: For each transaction data in the reference results, static features of the transaction attribute dimension and dynamic features of the transaction feedback dimension are extracted, and the correlation mapping relationship between static features and dynamic features is constructed to form a transaction feature matrix. The static features include the format definition of the transaction node, data field constraints and carbon accounting rule parameters, and the dynamic features include historical execution compliance rate, format matching pass rate and accounting feedback correction items. Based on the transaction feature matrix, the transaction scenarios corresponding to the reference results are analyzed in multiple dimensions to identify the scenario type, compliance level and format adaptation requirements, and generate a scenario adaptation tag set. Based on the historical feedback data of the transaction feature matrix, the adaptation value of different thread configurations is evaluated, and high-value thread configurations with adaptation values ​​higher than the preset value are selected. At the same time, based on the scenario adaptation tag set, underutilized adaptation rules are mined and supplementary thread configurations are generated. The high-value thread configuration and the supplementary thread configuration are input into a pre-trained policy fusion module to generate a unique thread optimization policy. Based on the thread optimization strategy and the pre-built scenario adaptation rule library, the transaction nodes in the reference results are formatted and constrained to generate an initial blank thread. The initial blank thread is matched with the dynamic feedback features in the transaction feature matrix in multiple rounds. When the matching degree reaches the preset convergence threshold, the final unique blank thread is output.

[0138] In this embodiment, static features are features that remain unchanged in the transaction attribute dimension and are not affected by the result of a single transaction execution, including: The format definition of a transaction node includes file type (PDF / Word) and template version (CCER V3.0 standard template) for filing nodes. Data field constraints include the data type (numerical), unit (tCO2e), and precision (two decimal places) of the emission reduction field. Carbon accounting rules and parameters include the accounting standard version (CCER V3.0 / VCS), project boundary definition rules, and emission reduction accounting methods.

[0139] Dynamic features: Features in the transaction feedback dimension that change with actual execution, reflecting the actual execution effect of the format / rules, including: Historical execution compliance rate is the percentage of transactions for which this node / format has passed review in past transactions (number of compliant transactions / total transactions × 100%). Format matching pass rate: The percentage of times the format is accepted across different trading platforms / institutions (number of accepted submissions ÷ total number of submissions × 100%). Accounting feedback correction items: A summary of correction opinions raised in previous audits regarding this format. If additional forest land ownership certificates are required, third-party monitoring data must be attached for emission reductions.

[0140] Association mapping relationship: Taking the transaction node as the core, establish a one-to-one correspondence between static features and dynamic features. For example, the CCER V3.0 filing node PDF format corresponds to a compliance rate of 95%, a matching pass rate of 90%, and the correction item is a supplementary monitoring data field.

[0141] Transaction Feature Matrix: A two-dimensional matrix with transaction data as rows and feature type (static / dynamic) as columns. Each cell stores the feature value of the corresponding node, presenting the complete feature information of each transaction in a structured way. Missing values ​​are filled with the industry average level.

[0142] The mapping relationship of all nodes for each transaction data is filled into a two-dimensional matrix in the format of transaction data, node, and feature type to ensure that the feature information of each node is complete and searchable.

[0143] For example, one piece of CCER V3.0 forestry carbon sink registration and trading data for a certain region in a certain reference result: Static characteristics (filing node): Format definition: CCER V3.0 standard filing application PDF template; Field constraints: Project number (10-character string), Project name (text), Emission reduction (numerical value). (Retain two decimal places); Accounting rules: According to CCER V3.0 standard, the project boundary must include the forest area and ownership certificate.

[0144] Dynamic characteristics (filing node): Historical compliance rate: 92%; Format matching pass rate: 88%; Accounting feedback correction items: Supplement forest land ownership certificate attachments, emission reductions need to be accompanied by third-party monitoring data.

[0145] Association Mapping: Node ID: 001 - Filing Application → Static Feature Set → Dynamic Feature Set; In the transaction feature matrix, the column value corresponding to this node is: Format definition: 1 (PDF), compliance rate 0.92, correction item: supplementary ownership certificate / monitoring data.

[0146] In this embodiment, the transaction scenario is the business scenario to which the transaction data belongs, which is jointly defined by factors such as the transaction target, applicable standards, transaction mode, and regional regulatory requirements. For example, a primary listing and trading scenario for CCER V3.0 forestry carbon sinks in a certain region.

[0147] The multi-dimensional analysis breaks down and analyzes transaction scenarios from three dimensions: scenario type, compliance level, and format adaptation requirements. Scenario types: classified by target (forestry / photovoltaic / wind power), accounting standard (CCER / VCS / GS), and trading market (primary / secondary); Compliance Level: Based on the compliance rate and format matching pass rate in dynamic features, high / medium / low compliance levels are defined. Format compatibility requirements: Specific regulatory requirements for transaction data formats in certain scenarios, such as requiring additional preliminary review opinions from local forestry departments in certain regions.

[0148] Scenario Adaptation Tag Set: A collection of tags consisting of scenario type, compliance level, and format adaptation requirements, presented in key-value pair format, for example, {Scenario Type: CCER V3.0 Forestry Carbon Sink Level 1 Trading, Compliance Level: High Compliance, Format Adaptation Requirements: Supplementary Local Preliminary Review Opinions}.

[0149] For example, based on the static features in the transaction feature matrix, a preset scenario type rule library can be matched. For instance, when the static features simultaneously satisfy the accounting rule = CCER V3.0, the transaction type = Level 1 transaction, and the target = forestry carbon sink, it is identified as a CCERV3.0 forestry carbon sink Level 1 transaction scenario.

[0150] Compliance Level Determination: Based on the compliance rate and format matching pass rate in dynamic features, levels are assigned according to preset thresholds. High compliance: compliance rate ≥ 90% and format matching pass rate ≥ 85%; Compliance requirements: 80% ≤ compliance rate < 90% and 80% ≤ pass rate < 85%; Low compliance: Compliance rate < 80% or pass rate < 80%.

[0151] Format adaptation requirement extraction: Summarize the accounting feedback correction items of all transactions in this scenario in the transaction feature matrix, extract the high-frequency correction items with an occurrence frequency of ≥30%, and use them as the format adaptation requirements of the scenario. For example, if 40% of the transaction feedback requires supplementary preliminary review opinions from a forestry department, then these will be listed as adaptation requirements.

[0152] The identified scenario types, compliance levels, and format adaptation requirements are converted into standardized tags, redundant descriptions are removed, and a structured scenario adaptation tag set is formed.

[0153] Thread configuration refers to the format, process, and parameter combination of each node in the transaction thread. For example, a node combination of filing node in PDF format, verification node in PDF / A format, and preliminary review opinion attachment.

[0154] The preset value is a threshold for filtering high-value thread configurations. For example, configurations with an adaptation value ≥ 0.8 are considered high-value configurations.

[0155] In this embodiment, the adaptation value = (compliance rate × 0.4 + format matching pass rate × 0.3 + scene tag matching degree × 0.3), where the scene tag matching degree is the degree of matching between the thread configuration and the scene adaptation tag set, and the value is 0-1. For example, the matching degree is 1 when the configuration includes all format adaptation requirements, and it is calculated proportionally when it includes some requirements.

[0156] Iterate through the thread configurations of all transaction data in the reference results, calculate their adaptation value, filter out configurations with an adaptation value greater than or equal to the preset value (e.g., 0.8), and obtain a set of high-value thread configurations after deduplication.

[0157] By comparing the scenario adaptation tag set with the high-value thread configuration, identify the rules that exist in the tag set but are not included in the high-value configuration, verify their effectiveness (e.g., the transaction compliance rate of a rule that includes the rule is more than 10% higher than that of a rule that does not include the rule), and determine them as effective rules that are not being fully utilized.

[0158] By combining underutilized effective adaptation rules with the infrastructure of high-value configurations, supplementary thread configurations can be generated, such as adding local preliminary review opinion attachment nodes on the basis of high-value configurations.

[0159] The model structure of the strategy fusion module is as follows: It adopts a neural network based on the attention mechanism, with the input layer dimension being the thread configuration feature dimension (32-dimensional), the hidden layer dimension being 128, the number of attention heads being 4, and the output layer dimension being the optimized thread configuration feature. Training dataset: 10,000 optimal thread configurations and their corresponding transaction feedback data were extracted from the historical database and labeled as positive samples; Loss function: Mean squared error loss function, the optimization objective is to maximize the adaptation value of thread configuration; Model performance: The accuracy of the fit value prediction on the test set reached 92.3%.

[0160] Input data standardization involves organizing high-value thread configurations and supplementary thread configurations into structured input data in node order. Each configuration includes a node ID, format requirements, parameter rules, and adaptation scenario tags.

[0161] Multi-configuration fusion processing: A weighted fusion method is used to integrate the advantages of different configurations. The higher the adaptation value of a configuration, the higher the weight of its node format / rule. Common rules that appear in all configurations (such as filing nodes using PDF format) should be retained directly. For differentiated rules for different configurations, weights are assigned according to the adaptation value, and the rules of configurations with higher adaptation value are given priority. For example, the adaptation value of configuration B (0.977) is higher than that of configuration A (0.842), so the rules of its verification node PDF / A format are given priority.

[0162] The optimization strategy generation is a unique thread optimization strategy obtained after integration, ensuring that the format and rules of each node are the most valuable adaptation solutions, while covering all requirements in the scenario adaptation tag set.

[0163] For example, inputting high-value configurations A and B, and supplementary configuration C into the strategy fusion module: Configuration B has the highest adaptation value (0.977), and its rules in PDF / A format of verification nodes and preliminary review opinion attachments should be adopted first. The general PDF format requirements for the filing nodes in Configuration A are retained; The generated thread optimization strategy is as follows: Node 1 (filing application, PDF format, fields include project number / emission reduction) → Node 2 (local preliminary review opinion, PDF format, stamped with the official seal of the forestry department) → Node 3 (verification review, PDF / A format, including project boundary / monitoring data). All nodes meet the requirements of the scenario adaptation tag set.

[0164] A scenario-adaptive rule base is constructed using knowledge graphs, containing the following entities and relationships: Entities: Scenario type, compliance level, transaction node, format rules, accounting standards, regulatory requirements; Relationships: inclusion, applicability, requirement, compatibility, substitution; the rule base is updated regularly, with an update cycle of 1 month, to keep pace with the latest international carbon accounting standards and local regulatory requirements.

[0165] The matched templates are populated into the corresponding nodes of the initial thread. At the same time, field constraints, validation rules and required attachments are configured. For example, the emission reduction field of the filing node must be a numeric type and retain two decimal places. The preliminary review comments are a required attachment.

[0166] For example, generating an initial blank thread based on a thread optimization strategy: Node 1 (Filing Application): Fill in with the CCER V3.0 filing application PDF template, configure field constraints: Project Number (10-character string), Emission Reduction (numerical value, ...). (2 decimal places), required attachment: Forest land ownership certificate; Node 2 (Local Preliminary Review Opinions): Fill in with a PDF template of preliminary review opinions from a forestry department, configure the verification rules: the department's electronic seal must be affixed, this is a mandatory node; Node 3 (Certification and Audit): Fill in the CCER V3.0 certification report PDF / A template, configure field constraints: project boundaries, monitoring data, emission reduction calculation table, and verification rules: must include the official seal of the third-party certification body; The initial blank thread is: Node 1 (filing application template + constraints) → Node 2 (preliminary review opinion template + constraints) → Node 3 (verification report template + constraints), with no specific transaction data.

[0167] Multi-round matching verification involves repeatedly matching and verifying the node format and constraint rules of the initial blank thread with the dynamic feedback features (compliance rate, pass rate, correction terms) of each transaction data in the transaction feature matrix. The first round of matching degree calculation: The format and constraints of each node in the initial blank thread are matched with the dynamic features of all transaction data in the transaction feature matrix to calculate the overall matching degree: Overall matching degree = Σ (node ​​matching degree × node weight) where node matching degree = (number of transactions that conform to the node format / constraint / total number of transactions) × (compliance rate × 0.4 + pass rate × 0.3 + correction item matching degree × 0.3). The node weight is allocated according to the importance of the node (e.g., the weight of the verification node is 0.4, and the weight of the filing node is 0.3).

[0168] Iterative Adjustment: If the matching degree does not reach the preset convergence threshold, and if the matching degree of a node is <0.9, and ≥30% of transaction feedback requests modifying the node's format, then the format constraint of that node will be adjusted to support multiple formats (e.g., simultaneously supporting PDF and PDF / A). If the matching degree of a node is <0.8, and ≥50% of transaction feedback requests adding attachments, then an optional attachment node will be added after that node. If the matching degree still fails to meet the standard after three consecutive iterations, the reference results will be re-selected, and a new initial blank thread will be generated. It should be noted that through experimental comparison of the adaptability of blank threads under different convergence thresholds, when the preset convergence threshold is 0.95, the actual execution compliance rate of the blank thread reaches over 96%, and the number of iterations is controlled within 3 rounds. Therefore, the preset convergence threshold is 0.95.

[0169] Multiple rounds of verification iteration: Repeat the process of matching degree calculation → template adjustment, update the matching degree after each round of verification, until the matching degree is greater than or equal to the preset convergence threshold.

[0170] Output the final template: Once the matching degree meets the standard, the determined initial blank thread becomes the final unique blank thread, which serves as the base template for generating standard threads in the future.

[0171] For example, multiple rounds of verification using an initial blank thread: First round of verification: The overall matching degree was 0.92, which did not reach the convergence threshold of 0.95; the verification found that the PDF / A format of the verification node could not be recognized on some trading platforms, and the matching degree was only 0.85. Template adjustment: Modify the format constraints of the verification node to support PDF / A or PDF format that conforms to the platform specifications; Second round of verification: The overall matching degree improved to 0.96, reaching the preset convergence threshold; The final output is a single blank thread: Node 1 (filing application PDF template, including field constraints) → Node 2 (local preliminary review opinion PDF template, required) → Node 3 (certification report PDF / A or standard PDF format, including field constraints), which meets the compatibility standard.

[0172] The beneficial effects of the above technical solution are as follows: by extracting static and dynamic features to construct a transaction feature matrix, parsing transaction scenarios to generate adaptive tags, evaluating and screening high-value thread configurations, integrating and generating optimization strategies, and filling in the format to construct an initial blank thread and verifying it through multiple rounds, the standardized generation of blank thread templates for green carbon transactions has been achieved. This not only integrates the dynamic feedback experience of historical transactions, but also adapts to the scenario-based compliance requirements of different regions and different regulatory requirements, ensuring that the final template has universality, compliance and scenario adaptability. It provides a reliable basic template support for the subsequent generation of a unique standard thread that adapts to international carbon accounting standards, effectively improving the accuracy and adaptability of green carbon transaction data standardization conversion.

[0173] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for green carbon data trading and standardization conversion adapted to international carbon accounting standards, characterized in that, include: Step 1: Based on the historical database, mine each historical green carbon transaction data to obtain the transaction thread, and obtain the historical transaction standard format of each transaction node in the transaction thread. Perform initial division of all transaction threads according to the transaction type to obtain several thread division sets. Step 2: Perform a format clustering analysis on each partitioned thread set according to the historical transaction standard format to obtain several small clustering results for the corresponding partitioned thread set. At the same time, perform chain clustering analysis on all transaction threads according to the historical transaction standard format chain of each transaction thread to obtain the large clustering result. Step 3: Perform secondary partitioning on each large cluster result according to transaction type. Based on the matching relationship with the transaction type of the corresponding partition thread set, sort the secondary partitioning results by size to obtain an initial list; Step 4: Determine the transaction format matching set, matching probability, and partitioning level confidence of each sub-partition result in the initial list of each large cluster result for the corresponding small cluster result, and select candidate threads that meet the candidate criteria of the corresponding small cluster result from each large cluster result; Step 5: All candidate threads obtained based on the small clustering results under the same transaction type are summarized into the corresponding partitioned thread set, and secondary format clustering analysis is performed to obtain several new clustering results for the corresponding summarized thread set; Step 6: Compare and analyze the new clustering results of the corresponding inductive thread set with the original small clustering results, determine the standard format transition probability of each small clustering result, and construct the probability transition sequence of the corresponding inductive thread set. Step 7: Input the probability transition sequence into the pre-trained sequence analysis model to determine the format fixing coefficients of the corresponding small clustering results; If the fixed coefficient of the format is greater than or equal to the preset coefficient, the clustering results with high confidence of the format standard are selected from the corresponding small clustering results and the matched new clustering results as reference results; If the fixed coefficient of the format is less than the preset coefficient, the corresponding small clustering result and the matching new clustering result will be fused and analyzed as a reference result. Step 8: For each reference result under each partitioned thread set, perform optimal thread recommendation processing according to transaction attributes and transaction feedback to obtain a unique blank thread for the corresponding reference result; Step 9: Extract the historical transaction standard format with the highest confidence level for each transaction node based on the unique blank thread from each reference result corresponding to each partitioned thread set, and use it as the current format to obtain the unique standard thread. Then, take all the unique standard threads under each transaction type as the standard threads for green carbon trading.

2. The green carbon data trading and standardization conversion method adapted to international carbon accounting standards according to claim 1, characterized in that, The trading threads are obtained by mining each historical green carbon trading data point from the historical database, including: For each historical green carbon transaction data, feature extraction is performed to obtain transaction features and green carbon features, and a transaction template is selected from the dual-feature database based on the transaction features and green carbon features; Simultaneously, the flow process of corresponding historical green carbon trading data and the time-sequential operation behaviors involved in the flow process are mined from the historical database, and the start and end points of each operation behavior are extracted to construct the initial thread; The transaction template is matched with the initial thread to obtain the transaction thread.

3. The green carbon data trading and standardization conversion method adapted to international carbon accounting standards according to claim 1, characterized in that, The transaction template is matched with the initial thread to obtain the transaction thread, including: Based on the predefined transaction description set of each template node in the transaction template, semantic matching and window association analysis are performed between the operation start point and operation end point of each operation behavior in the initial thread; When the semantic matching degree is greater than the preset matching degree, the template node is retained unchanged; Otherwise, based on the window association analysis results, determine the variable group of the corresponding operation behavior, and compare the variable types involved in the variable group with the transaction description features of the predefined transaction description set of the adjacent nodes of the extracted template node to determine the node addition position, and determine and retain the addition transaction description set of the corresponding addition node based on the variable group. The transaction thread is obtained based on the sequential positional relationship of the retained template nodes.

4. The green carbon data trading and standardization conversion method adapted to international carbon accounting standards according to claim 1, characterized in that, Sort the results of the secondary partitioning by size to obtain an initial list, including: The corresponding large clustering results are further divided according to the transaction type to obtain several first partition sets of the corresponding large clustering results, and each first partition set corresponds to a transaction type. The transaction types of each first partition set are matched with the transaction types of the corresponding partition thread set to obtain the matching relationship; At the same time, determine the first number of transaction data that is completely consistent with the corresponding partition thread set and the second number of similar transaction data in each first partition set, and obtain the statistical array of the corresponding first partition set; The matching relationship is optimized based on the statistical array to obtain the sorting value of each first partition set; Sort all the first partition sets corresponding to the large clustering results according to the size of the sorting value to obtain an initial list, where the first partition set is the sub-partition result.

5. The green carbon data trading and standardization conversion method adapted to international carbon accounting standards according to claim 1, characterized in that, From each large clustering result, candidate threads that meet the candidate criteria of the corresponding small clustering result are selected, including: Construct a first matrix of corresponding small clustering results and corresponding initial lists based on the transaction format matching set, matching proportion probability and division level confidence; Based on each candidate index in the candidate criteria of the corresponding small clustering results, traverse each element in the first matrix in turn, and mark the matching elements as first; The importance of each row vector in the first matrix is ​​obtained based on the number of times each labeled element is labeled, and the number of sub-partition results to be filtered for each row vector is determined. Candidate threads are obtained by filtering from the corresponding sub-partition results.

6. The green carbon data trading and standardization conversion method adapted to international carbon accounting standards according to claim 1, characterized in that, Determine the standard format transition probabilities for each sub-cluster result, including: Construct a first clustering graph corresponding to several new clustering results of the thread set before induction, and at the same time, construct a second clustering graph of several original small clustering results of the thread set before induction. By comparing and analyzing the first clustering graph and the second clustering graph, the first distribution of threads that originally belonged to each small clustering result in the first clustering graph and the second distribution in the second clustering graph are determined, and the change point of the second distribution based on the first distribution is determined, wherein the point is each transaction data in several new clustering results of the corresponding inductive thread set; The performance coefficient of each change point is obtained based on the nearest distance between each change point and the corresponding small cluster result, the cluster distance between each change point and the central cluster of the corresponding small cluster result, and the first distance between the central cluster of the corresponding small cluster result and the point corresponding to the nearest distance. Based on the variation distribution of the variation points and the performance coefficient of each variation point, the variation trend of the variation distribution is determined, and the standard format transition probability of the corresponding small clustering results is determined.

7. The green carbon data trading and standardization conversion method adapted to international carbon accounting standards according to claim 1, characterized in that, The results of the corresponding small clusters and the matching new clusters are fused and analyzed, including: Using the second distance between the cluster center of the corresponding small clustering result and the cluster center of the matched new clustering result as the radius, draw a circle with the cluster center of the corresponding small clustering result as the center, and retain the transaction data within the drawn circle as the first retention; All identical transaction data in the corresponding small clustering results and the matching new clustering results are globally and individually retained; Transaction data with performance coefficients greater than the performance threshold in the newly matched clustering results are retained as a second group. The first retention result, the global single retention result, and the second retention result are all processed by single retention to obtain a reference result. Among them, the completely identical transaction data in the first retention result are locally single retained.

8. The green carbon data trading and standardization conversion method adapted to international carbon accounting standards according to claim 1, characterized in that, The only blank threads that obtain the corresponding reference result include: For each transaction data in the reference results, static features of the transaction attribute dimension and dynamic features of the transaction feedback dimension are extracted, and the correlation mapping relationship between static features and dynamic features is constructed to form a transaction feature matrix. The static features include the format definition of the transaction node, data field constraints and carbon accounting rule parameters, and the dynamic features include historical execution compliance rate, format matching pass rate and accounting feedback correction items. Based on the transaction feature matrix, the transaction scenarios corresponding to the reference results are analyzed in multiple dimensions to identify the scenario type, compliance level and format adaptation requirements, and generate a scenario adaptation tag set. Based on the historical feedback data of the transaction feature matrix, the adaptation value of different thread configurations is evaluated, and high-value thread configurations with adaptation values ​​higher than the preset value are selected. At the same time, based on the scenario adaptation tag set, underutilized adaptation rules are mined and supplementary thread configurations are generated. The high-value thread configuration and the supplementary thread configuration are input into a pre-trained policy fusion module to generate a unique thread optimization policy. Based on the thread optimization strategy and the pre-built scenario adaptation rule library, the transaction nodes in the reference results are formatted and constrained to generate an initial blank thread. The initial blank thread is matched with the dynamic feedback features in the transaction feature matrix in multiple rounds. When the matching degree reaches the preset convergence threshold, the final unique blank thread is output.