A fingerprinting method based on coal consumption process

By constructing carbon fingerprint and waste fingerprint maps of the coal consumption process, and combining machine learning and Bayesian inversion algorithms, the problem of coal waste traceability has been solved, enabling accurate identification of coal type and source, and improving the efficiency of resource management and environmental monitoring.

CN120853744BActive Publication Date: 2026-02-03ANHUI UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510942538.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2026-02-03
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

Existing technologies cannot effectively track various wastes and residues generated during coal consumption, nor can they accurately identify their sources and types, thus limiting the accuracy and efficiency of resource management and environmental monitoring.

Method used

By constructing carbon fingerprint maps of coal and coal combustion waste fingerprint maps, and combining the TabPFN machine learning model and Bayesian inversion algorithm, a mapping relationship from coal fingerprint to waste fingerprint is established to achieve accurate traceability of waste in the coal consumption process.

Benefits of technology

It enables precise identification of coal types and sources, monitors carbon emissions and waste generation during coal combustion, provides a scientific basis for the efficient use of coal, and improves the effectiveness of environmental monitoring and resource optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853744B_ABST
    Figure CN120853744B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on coal consumption process's fingerprint tracing method, including the construction of coal carbon fingerprint atlas: by detecting the carbon content of organic matter component in coal sample, form coal carbon fingerprint vector;Build coal combustion waste fingerprint atlas: detect solid waste, and in liquid waste, form waste fingerprint vector;Establish forward mapping model: with coal carbon fingerprint vector as input matrix, waste fingerprint vector as output matrix, adopt TabPFN machine learning model to carry out multiple output regression, by quantile transformation and field threshold bucketing preprocessing data, combine physical constraint check post-processing, train and obtain the mapping relationship of coal fingerprint to waste fingerprint;Reverse deduce coal source: based on measured waste fingerprint, utilize Bayes inversion algorithm to solve posterior distribution, by MAP estimation or MCMC sampling reverse deduce original coal carbon fingerprint.The application can provide effective technical support in the identification of coal combustion process monitoring and waste management and the like, and help to achieve carbon emission reduction target.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of tracing technology in coal consumption process, and particularly relates to a fingerprint tracing method based on coal consumption process. BACKGROUND

[0002] In the current technical background, tracing technology has been widely applied and continuously innovated, especially in the fields of energy, medicine, industrial detection and scientific research. These technologies can achieve real-time monitoring and accurate analysis of material flow in complex systems by using specific tracers or probes, thereby promoting resource optimization, early disease diagnosis and experimental efficiency. For example, oil and gas field monitoring optimizes oil and gas exploitation strategies through tracer release devices; medical imaging uses fluorescent probes for cell labeling, which helps early detection and research of diseases; industrial detection uses X-ray particle tracing to improve detection accuracy; and in experimental design, the integration of tracer particles simplifies the experimental process and significantly improves experimental efficiency.

[0003] However, the existing tracing technology still has some limitations, especially in the application of tracing coal consumption process. At present, there is no effective means to directly trace various waste and residues in the coal consumption process, so as to accurately identify the corresponding coal source and its type. This not only limits the overall management of coal resources, but also affects the accuracy of environmental monitoring and the effect of resource optimization. Therefore, it is urgent to develop a method that can accurately trace the waste in the coal consumption process, so as to realize fast and accurate coal resource tracing by establishing an advanced model combined with multivariate analysis technology, big data and machine learning, thereby providing strong support for environmental monitoring and resource optimization.

[0004] In summary, fine management and tracing of waste in the coal consumption process is a problem to be solved, which can not only promote the rational allocation and utilization of resources, but also effectively improve the accuracy and efficiency of environmental monitoring. Therefore, the present application combines existing tracing technology, multivariate analysis, big data and machine learning, aiming to provide an innovative method that can accurately trace the source of various waste in the coal consumption process, to fill the gap in this technical field. SUMMARY

[0005] The purpose of the present application is to overcome the shortcomings of the prior art. In order to achieve the above purpose, a fingerprint tracing method based on coal consumption process is used to solve the problems raised in the background technology.

[0006] A fingerprint tracing method based on coal consumption process, comprising the following steps:

[0007] S1, constructing a carbon fingerprint map of coal: forming a coal carbon fingerprint vector by detecting the carbon content of organic matter components in the coal sample;

[0008] S2, build coal combustion waste fingerprint: detect solid waste, and liquid waste, form waste fingerprint vector;

[0009] S3, establish forward mapping model: take coal fingerprint vector as input matrix, waste fingerprint vector as output matrix, adopt TabPFN machine learning model to carry out multiple output regression, pretreat data through quantile transformation and field threshold bucketization, combine physical constraint check post-processing, train to obtain mapping relationship from coal fingerprint to waste fingerprint;

[0010] S4, reverse deduce coal source: based on measured waste fingerprint, utilize Bayes inversion algorithm to solve posterior distribution, and deduce original coal fingerprint through MAP estimation or MCMC sampling.

[0011] As a further scheme of the application, the specific steps in step S1 include:

[0012] A mathematical model of carbon fingerprint is built, and the carbon fingerprint of coal is defined as a multivariate function, wherein each variable represents a characteristic variable of an organic matter component of different coal rank coal;

[0013] The characteristic variables of the organic matter component are researched, and the characteristic variables include alkane carbon content, aromatic hydrocarbon carbon content, oxygen-containing functional group carbon content, total organic carbon content and alkane to aromatic hydrocarbon ratio.

[0014] For carbon fingerprint similarity measurement, Euclidean distance or Manhattan distance is used to calculate the similarity between different coal samples, so as to identify coal and build a carbon fingerprint atlas of coal.

[0015] As a further scheme of the application, the specific steps in step S2 include:

[0016] Solid waste and liquid waste generated in the coal combustion process are obtained;

[0017] The solid waste includes inorganic minerals, heavy metals, unburned carbon and slag in the ash.

[0018] The liquid waste includes tar and moisture; the main components of the tar include polycyclic aromatic hydrocarbons, alkanes, phenolic compounds, heterocyclic compounds, oxygen-containing compounds, nitrogen-containing compounds and sulfur-containing compounds.

[0019] Based on different coal ranks of coal, the difference of the organic matter component directly affects the waste composition after coal combustion, and the coal combustion waste fingerprint atlas of different coal ranks is obtained.

[0020] As a further scheme of the application, the specific steps in step S3 include:

[0021] S31. Use the carbon fingerprint spectrum of coal as the feature variable of the constructed input matrix, and the fingerprint spectrum of coal combustion waste as the feature variable of the constructed output matrix, and combine the two to establish a mapping relationship.

[0022] S32. Extreme values ​​are compressed through quantile transformation, non-uniform binning is performed based on industrial threshold, missing values ​​are filled by neighborhood interpolation, and the total ash oxides are constrained and carbon balance is verified on the output.

[0023] S33. Call the pre-trained TabPFN model, selectively unfreeze the top-level matrix for gradient fine-tuning, and combine semi-supervised distillation to enhance the adaptability of coal chemical characteristics, so as to achieve multi-output one-time prediction;

[0024] S34. Use the Attention caching mechanism to accelerate the calculation and output the mean vector of waste fingerprints and its confidence interval.

[0025] S35. Employ a tiered / plant-wide cross-validation strategy, focusing on RMSE and R... 2 The robustness of the confidence interval coverage assessment model is demonstrated, and the empirical inversion error is less than the error threshold.

[0026] As a further aspect of the present invention, step S4 specifically includes the following steps:

[0027] S41. Based on actual detection, obtain waste fingerprints and construct an inversion input benchmark;

[0028] S42. Call the TabPFN forward model and define the prediction output of any candidate coal fingerprint as a Gaussian likelihood function;

[0029] S43. Use Dirichlet distribution as the prior for coal composition, set constraints, and calibrate parameters based on historical coal quality data.

[0030] S44. Derive the logarithmic posterior using Bayes' formula, while preserving the relationships between core variables;

[0031] S45. Perform MAP estimation or MCMC sampling to output the original coal fingerprint.

[0032] Compared with the prior art, the present invention has the following technical advantages:

[0033] By employing the aforementioned technical solution and constructing a carbon fingerprint tracing model, the type and source of coal can be accurately identified, and carbon emissions and waste generation during coal combustion can be monitored. This provides a scientific basis for the efficient utilization of coal and has significant ecological, environmental, and economic value. This method can provide effective technical support for coal identification, combustion process monitoring, and waste management, contributing to the achievement of carbon reduction targets. Attached Figure Description

[0034] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings:

[0035] Figure 1 This is a schematic diagram illustrating the steps of the fingerprint tracing method according to an embodiment of this application. Detailed Implementation

[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] Please refer to Figure 1 In this embodiment of the invention, a fingerprint tracing method based on the coal consumption process includes the following steps:

[0038] S1. Constructing a carbon fingerprint of coal: By detecting the carbon content of organic matter components in a coal sample, a coal carbon fingerprint vector is formed. The specific steps include:

[0039] A mathematical model for carbon fingerprinting is constructed, defining the carbon fingerprint of coal as a multivariable function, where each variable represents a characteristic variable of the organic matter composition of coal of different ranks.

[0040] The study obtained characteristic variables of organic matter components, including alkane carbon content, aromatic hydrocarbon carbon content, oxygen-containing functional group carbon content, total organic carbon content, and the ratio of alkane to aromatic hydrocarbons.

[0041] For carbon fingerprint similarity measurement, we try to use Euclidean distance or Manhattan distance to calculate the similarity between different coal samples, thereby identifying coal and constructing a carbon fingerprint map of coal.

[0042] In this embodiment, the organic matter component of coal is an important part of its quality characteristics, directly affecting its combustion performance, gas emissions, calorific value, etc. The organic matter composition of coal changes significantly with coal rank. From lignite to anthracite, due to variations in composition and differences in physicochemical properties, the proportions of its organic matter components (such as alkanes, aromatics, and alcohols) fluctuate considerably, thus reflecting the coal rank characteristics based on its organic matter composition.

[0043] Studying the changes in organic matter composition from lignite to anthracite can reveal the changes in the content of alkanes, aromatics, and oxygen-containing functional groups.

[0044] Changes in alkane content: As the coal rank increases, the content of alkane compounds gradually decreases, and long-chain alkanes gradually decrease in high-rank coals, while low-rank coals (such as lignite) are dominated by longer-chain alkanes.

[0045] Changes in aromatic hydrocarbon composition: The content of aromatic hydrocarbons gradually increases with increasing coal rank. In anthracite, the proportion of aromatic hydrocarbon compounds increases significantly, forming complex aromatic ring structures.

[0046] Changes in oxygen-containing functional groups: In low-rank coal (such as lignite), oxygen-containing functional groups (such as phenols and alcohols) account for a relatively high proportion, while as the coal rank increases, the number of oxygen-containing functional groups gradually decreases, and the proportion of hydrocarbons relatively increases.

[0047] If the following characteristic variables are used to characterize the carbon fingerprint of coal:

[0048] C1: Carbon content of alkanes (percentage), mainly including straight-chain alkanes, branched-chain alkanes and cycloalkanes;

[0049] C2: Aromatic hydrocarbon carbon content (percentage), mainly including low-ring aromatic hydrocarbons (benzene, toluene, xylene) and high-ring aromatic hydrocarbons (naphthalene, phenanthrene, pyrene, etc.);

[0050] C3: Carbon content (percentage) of oxygen-containing functional groups (such as phenols, alcohols, carboxylic acids, ethers, and ketones);

[0051] C4: Total organic carbon content (percentage);

[0052] C5: Alkane / Aromatic hydrocarbon ratio (representing the organic matter maturity of coal);

[0053] As shown in Table 1 below:

[0054] Table 1

[0055] coal rank C1 C2 C3 C4 C5 lignite 5-15 20-30 15-30 65-70 0.2-0.5 bituminous coal 10-20 50-70 15-30 45-86 0.5-2.0 anthracite 1-5 80-90 1-5 86-97 2.0-6.0

[0056] The organic matter composition of different coal samples from a certain region was used.

[0057] To construct a mathematical model of carbon fingerprint, the carbon fingerprint of coal is defined as a multivariable function, where each variable (C1, C2, C3, C4, C5) represents the variation characteristics of organic matter composition in coal of different ranks.

[0058] Coal fingerprint model F coal (x) represents the carbon fingerprint of coal as a vector:

[0059] F coal (x) = [C1, C2, C3, C4, C5]

[0060] For lignite: F coal Lignite) = [35, 20, 25, 80, 1.75];

[0061] F coal(bituminous coal) = [20, 30, 20, 70, 0.67];

[0062] F coal (Anthracite) = [5, 50, 10, 65, 0.10];

[0063] For carbon fingerprint similarity measurement, we try to use Euclidean distance or Manhattan distance to calculate the similarity between different coal samples, thereby identifying the coal.

[0064] In this embodiment, based on Euclidean distance, the formula is:

[0065] P=√(C1 coal1 -C1 coal2 ) 2 +(C2 coal1 -C2 coal2 ) 2 +(C3 coal1 -C3 coal2 ) 2

[0066] +(C4 coal1 -C4 coal2 ) 2 +(C5 coal1 -C5 coal2 ) 2

[0067] The smaller the P-value, the closer the two are.

[0068] S2. Constructing a fingerprint map of coal combustion waste: Detecting solid and liquid waste to form waste fingerprint vectors. The specific steps include:

[0069] Acquire solid and liquid waste generated during coal combustion;

[0070] Solid waste includes inorganic minerals, heavy metals, unburned carbon, and slag in ash.

[0071] Liquid waste includes tar and water; the main components of tar include polycyclic aromatic hydrocarbons, alkanes, phenolic compounds, heterocyclic compounds, oxygen-containing compounds, nitrogen-containing compounds, and sulfur-containing compounds.

[0072] Based on the different ranks of coal, the differences in organic matter composition directly affect the composition of waste after coal combustion, resulting in fingerprint profiles of coal combustion waste of different coal ranks.

[0073] In this embodiment, solid waste and liquid waste are typically generated during the coal combustion process.

[0074] Solid waste mainly includes inorganic minerals (silicon oxide, aluminum oxide, etc.) in ash, heavy metals (lead, mercury, arsenic, etc.), unburned carbon, and slag.

[0075] Liquid waste mainly includes tar and water. Tar has a complex composition, with its main components being polycyclic aromatic hydrocarbons, alkanes, phenolic compounds, heterocyclic compounds, oxygen-containing compounds, nitrogen-containing compounds, and sulfur-containing compounds.

[0076] The differences in organic matter composition among different coal ranks directly affect the composition of coal combustion waste, as shown in Table 2 below, which presents the fingerprint profiles of coal combustion waste (tar) based on different coal ranks.

[0077] Table 2

[0078]

[0079] Table 3 below shows the fingerprint profiles of solid waste from coal combustion based on different coal ranks:

[0080] Table 3

[0081]

[0082] S3. Establish a forward mapping model: Using the coal fingerprint vector as the input matrix and the waste fingerprint vector as the output matrix, the TabPFN machine learning model is used for multi-output regression. Data is preprocessed through quantile transformation and neighborhood threshold bucketing, and then post-processed with physical constraint verification to train and obtain the mapping relationship from coal fingerprint to waste fingerprint. The specific steps include:

[0083] S31. Use the carbon fingerprint spectrum of coal as the feature variable of the constructed input matrix, and the fingerprint spectrum of coal combustion waste as the feature variable of the constructed output matrix, and combine the two to establish a mapping relationship.

[0084] S32. Extreme values ​​are compressed through quantile transformation, non-uniform binning is performed based on industrial threshold, missing values ​​are filled by neighborhood interpolation, and the total ash oxides are constrained and carbon balance is verified on the output.

[0085] S33. Call the pre-trained TabPFN model, selectively unfreeze the top-level matrix for gradient fine-tuning, and combine semi-supervised distillation to enhance the adaptability of coal chemical characteristics, so as to achieve multi-output one-time prediction;

[0086] S34. Use the Attention caching mechanism to accelerate the calculation and output the mean vector of waste fingerprints and its confidence interval.

[0087] S35. Employ a tiered / plant-wide cross-validation strategy, focusing on RMSE and R... 2The robustness of the confidence interval coverage assessment model is demonstrated, and the empirical inversion error is less than the error threshold.

[0088] In this embodiment, the carbon fingerprint of coal and the fingerprint of coal combustion waste are combined to construct a mathematical model. Machine learning algorithms can be used to optimize the model and identify the nonlinear relationships between them. The following are the detailed steps for constructing the model and matching different matrices using machine learning algorithms.

[0089] (1) Data matrix construction

[0090] Input matrix (carbon fingerprint of coal): This matrix contains 5 feature variables, each representing a different carbon composition of coal. We use these features as columns of the input matrix.

[0091] X = [C1, C2, C3, C4, C5];

[0092] in:

[0093] C1: Carbon content of alkanes (percentage), mainly including straight-chain alkanes, branched-chain alkanes and cycloalkanes.

[0094] C2: Aromatic hydrocarbon carbon content (percentage), mainly including low-ring aromatic hydrocarbons (benzene, toluene, xylene) and high-ring aromatic hydrocarbons (naphthalene, phenanthrene, pyrene, etc.).

[0095] C3: Carbon content (percentage) of oxygen-containing functional groups (such as phenols, alcohols, carboxylic acids, ethers, and ketones).

[0096] C4: Total organic carbon content (percentage)

[0097] C5: Alkane / Aromatic Hydrocarbon Ratio (represents the organic matter maturity of coal)

[0098] Output matrix (waste fingerprint after coal combustion): This matrix contains various components of solid and liquid waste. Each type of waste component (such as tar, ash, etc.) is an output variable.

[0099] Y=[T1,T2,T3,T4,T5,T6,T7,T8,T9,T10];

[0100] in:

[0101] T1: Silica content in ash;

[0102] T2: Alumina content in ash;

[0103] T3: Iron oxide content in ash;

[0104] T4: Calcium oxide content in ash;

[0105] T5: Heavy metal content in solid waste;

[0106] T6: Carbon content of alkane tar in tar;

[0107] T7: Carbon content of aromatic hydrocarbons in tar;

[0108] T8: Content of phenolic compounds in tar;

[0109] T9: Total carbon content in tar;

[0110] T10: Unburned carbon content;

[0111] (2) Machine Learning Algorithms

[0112] Machine learning algorithms can help us find the mapping relationship between the input matrix (coal fingerprint) and the output matrix (combustion waste).

[0113] TabPFN is a small-data-table "base model" based on Transformer that can predict multi-output regression or classification tasks in a single forward pass without retraining parameters for each dataset, achieving high-precision predictions by utilizing only embedded prior and context learning capabilities.

[0114] (3) Model building

[0115] 1) Data preparation and preprocessing:

[0116] Feature distribution correction and bucketing strategy

[0117] The carbon composition characteristics (C1–C6) of coal samples often exhibit a highly skewed distribution. For example, some coal samples have a high concentration of C3 (oxygen-containing functional groups) in the low value region, while a few samples have a very high value.

[0118] Before standardization, a quantile transformation is performed on each feature to compress extreme values ​​into a reasonable range, ensuring that subsequent bucketing by the Transformer will not cause most of the data to be squeezed into the same bucket due to a few outliers.

[0119] Instead of simply truncating by equal width or equal frequency, bucketing is performed based on the neighborhood threshold of each feature (such as the "20%" critical point for oxygen-containing functional groups recognized in industrial analysis). More dense buckets are added around these "critical boundaries" to allow the model to have finer resolution within that range.

[0120] a. Standardization

[0121] For all input columns C j With output column T k Perform z-score normalization separately (TabPFN will also automatically perform a round of normalization internally).

[0122] Physical constraint verification and post-processing

[0123] There are obvious mass conservation and oxide ratio constraints in combustion chemistry;

[0124] For example, the sum of SiO2+Al2O2+Fe2O3+CaO in ash should be within a certain range.

[0125] Targeted post-processing:

[0126] On the predicted values of the four oxides output by TabPFN, perform proportional normalization verification: If the sum of the four oxides exceeds 100% (extreme case), then re-normalize with the prediction confidence as the weight to ensure chemical consistency;

[0127] Add the total organic carbon in tar (T6–T9) plus the unburned carbon (T 10 ), and then perform a general carbon budget check to correct possible systematic biases.

[0128] Under normal circumstances, adopt direct acceptance + reserve margin, and calculate the margin

[0129] R = 100% - S

[0130] R represents the total proportion of "other oxides" and "volatile loss" in ash.

[0131] Do not perform forced normalization, and set upper and lower limit checks. To ensure that the model prediction falls within a reasonable range, a threshold check on S is also required: lower threshold Smin: Usually take 40-50% (based on the minimum value of the existing data); if S < Smin, it means that the model may seriously underestimate the content of the four major oxides. At this time: trigger an alarm to prompt "the content of the four oxides in ash is low, it is recommended to check the input carbon fingerprint or model parameters"; or according to domain experience, use "regional average" or "nearest neighbor sample" to correct the prediction in segments.

[0132] b. Missing value handling

[0133] Directly retain NaN, and TabPFN can automatically interpolate and add a missing indicator bit internally (after the Half-Precision layer).

[0134] Domain interpolation for missing values and noise

[0135] Domain-aware imputation is used: for missing values ​​of oxygen-containing functional groups (C3), instead of simply filling with the mean, the mean C3 value of coal samples with similar alkane / aromatic ratios from the same coal zone is used; for points in T8 that significantly exceed the common range (>0.15), a secondary correction is performed using the empirical relationship between ash and phenol under classic working conditions to prevent the model from learning extreme noise as a real signal.

[0136] 2) Model invocation and fine-tuning:

[0137] a. Direct inference

[0138] Using a pre-trained TabPFN regression model, the entire training set context and corresponding labels are input, and the model outputs the predicted distributions of all T1–T10 in the same forward pass.

[0139] b. Prior fine-tuning and a small number of gradient updates

[0140] Targeted training: In-Context Fine-Tuning: Before the one-time forward pass, the first 4 of the 10 samples can be used for "example update" (Context-only fine-tuning), that is, a small amount of internal projection matrix is ​​fine-tuned in the attention layer of the Transformer to make it more "familiar" with the chemical characteristics of the coal sample;

[0141] Specific implementation steps:

[0142] 1. Select fine-tuning parameters. Unfreeze W in all self-attention modules from layer LLL upwards in the model. Q (Query projection matrix), W K (key projection matrix), W V (Value projection matrix); at the same time, keep all other parameters (including the rest of the Transformer's projection, Feed-Forward layer, position encoding, LayerNorm weights, output header, etc.) frozen (do not participate in gradient updates).

[0143] 2. Construct fine-tuning data. From the ten sets of samples, select the top four most representative samples. As an "example update set", we perform the same preprocessing as for the forward prediction (bucketing, standardization, missing interpolation labeling, etc.) on them to build the "Context→Target" pair.

[0144] 3. Define the fine-tuning objective. Use the same mean squared error loss as in forward training: Where f θ This means that only W is affected at this time. Q W K W VDifferentiable TabPFN model.

[0145] 4. Perform fine-tuning. Choose a very small learning rate (e.g., 10⁻⁵ to 10⁻⁴) and a simple optimizer (e.g., Adam), and perform several (5–20) gradient descent steps only on the unfrozen projection matrix; at each step: input the entire Context (4 samples) into the Transformer and calculate the prediction. Calculate L fine And it propagates in reverse, updating only W. Q W K W V .

[0146] 5. Reset the frozen state. After fine-tuning, freeze these three sets of matrices again to ensure that the subsequent main prediction stage fully inherits and locks the prior adjustments after fine-tuning, and avoids further drift during batch prediction.

[0147] 6. Perform a one-time forward pass. Following the original process, feed all Contexts (either the remaining 6 out of 10 groups or all 10 groups) and Queries into the now fine-tuned model, and output the final μ in one go. Y σ Y .

[0148] Semi-supervised distillation: Hundreds of synthetic examples are generated using similar segment samples from the COALQUAL database (similar to the "TabPFN original pre-training" approach). These synthetic examples are then updated with a small number of gradients along with real samples. The prior adaptation to coal chemical characteristics is stabilized through distillation loss.

[0149] 3) Inference and Caching:

[0150] Training / test separation caching: TabPFN caches the attention keys / values ​​of training samples, avoiding repeated calculations on the same dataset and significantly accelerating batch inference.

[0151] (4) Performance evaluation and uncertainty quantification

[0152] 1) Evaluation Criteria

[0153] Multi-output regression: for each T k Calculate RMSE and R²; the overall average RMSE can be reported.

[0154] Uncertainty: Using the target distribution output by TabPFN, the 95% prediction interval is directly extracted from the prediction distribution.

[0155] 2) Cross-validation

[0156] It is recommended to use either Stratified Plants or Grouped Plants as partitioning strategies to test the robustness of the model under different coal sources and time periods.

[0157] (5) Specific model implementation

[0158] The model was applied and validated through the selection of ten sets of data. The following are the initial ten sets of data from different regions, as shown in Table 4:

[0159] Table 4

[0160] sample C1 C2 C3 C4 C5 … T10 1 0.10 0.25 0.20 0.68 0.40 … 0.1 2 0.06 0.22 0.18 0.66 0.27 … 0.12 … … … … … … … … 10 0.14 0.28 0.05 0.92 0.5 … 0.08

[0161] Detailed analysis of sample 1

[0162] Input carbon fingerprint

[0163] Alkanes (C1) = 0.10, Aromatics (C2) = 0.250, Oxygen-containing functional groups (C3) = 0.020.

[0164] Total organic carbon (C4) = 0.680, alkane / aromatic ratio (C5) = 0.400

[0165] One-time prediction

[0166] After reviewing all ten examples, TabPFN provides a 95% confidence interval of 0.242, [0.242, 0.262], for the silica content (T1) in the ash of Sample 1, with a mean of 0.252.

[0167] Error assessment

[0168] The true value is 0.250, the predicted value deviation is only +0.002 (0.8% relative error), and the true value falls entirely within the confidence interval, indicating that the model is very reliable.

[0169] Detailed analysis of sample 2

[0170] Input: C1 = 0.060, C2 = 0.220, C3 = 0.180, C4 = 0.660, C5 = 0.270

[0171] Prediction: T1 = 0.228 (0.218–0.238)

[0172] Truth value: 0.220

[0173] Error: +0.008 (+3.6% relative), confidence interval covers the true value.

[0174] Detailed analysis of sample 10

[0175] Input: C1 = 0.140, C2 = 0.280, C3 = 0.250, C4 = 0.690, C5 = 0.50

[0176] Prediction: T1 = 0.282 (0.272 – 0.392)

[0177] Truth value: 0.280

[0178] Error: +0.002 (+0.8% relative), still within the reasonable fluctuation range.

[0179] Summary of full sample performance

[0180] T1–T for all 10 samples 10 The model's average root mean square error (RMSE) is approximately 0.010, and R0 is... 2 >0.95.

[0181] The fact that the 95% confidence interval of each output indicator can cover the corresponding true value indicates that TabPFN can still provide highly reliable multi-output regression results even with small samples.

[0182] The TabPFN prediction results (mean ± 95% CI) are shown in Table 5 below.

[0183] Table 5

[0184]

[0185] Note: Pred_T k " represents the mean of the Gaussian distribution of the k-th target output by the model; the value in parentheses is "mean ± 1.96 × σ", which is the 95% confidence interval.

[0186] As can be seen from the TabPFN prediction results table, the model's average RMSE on the 10 datasets is ≈0.01, and R0 is... 2 >0.95 indicates high prediction accuracy and controllable uncertainty.

[0187] Finally, through the TabPFN process described above, we completed efficient multi-output prediction of "coal fingerprint → combustion waste fingerprint" on ten sets of samples, and the results proved that the algorithm has significant advantages in this data scenario.

[0188] S4. Reverse Coal Source Inference: Based on measured waste fingerprints, the posterior distribution is solved using a Bayesian inversion algorithm. The original coal fingerprint is then inferred through MAP estimation or MCMC sampling. The specific steps include:

[0189] S41. Based on actual detection, obtain waste fingerprints and construct an inversion input benchmark;

[0190] S42. Call the TabPFN forward model and define the prediction output of any candidate coal fingerprint as a Gaussian likelihood function;

[0191] S43. Use Dirichlet distribution as the prior for coal composition, set constraints, and calibrate parameters based on historical coal quality data.

[0192] S44. Derive the logarithmic posterior using Bayes' formula, while preserving the relationships between core variables;

[0193] S45. Perform MAP estimation or MCMC sampling to output the original coal fingerprint.

[0194] In this embodiment, since the purpose is to deduce the original coal from the waste after coal combustion, the above model is a forward derivation process, and a reverse derivation is also needed for cross-verification.

[0195] First, fit the "forward mapping" using TabPFN, which is the process described above.

[0196] f:X (coal fingerprint)→Y:9 waste fingerprint)

[0197] For any input X (normalized carbon fingerprint), TabPFN can provide a predicted distribution (mean μ) of the waste fingerprint Y in a single forward pass. Y (X) and variance

[0198] Inverse modeling (Bayesian inversion)

[0199] Treat the predicted distribution of TabPFN as a likelihood function:

[0200]

[0201] Yobs is the actual waste fingerprint detected.

[0202] 1. Preparation for observing the "waste fingerprint" Y obs ;

[0203] In the preceding forward modeling, there are ten sets of true output matrices for samples;

[0204]

[0205] For example, the observations of sample 1,

[0206] (The first four are the contents of SiO2, Al2O3, Fe2O3, and CaO in the ash content, and the last six are, in order, heavy metals, tar alkanes, tar aromatics, phenols, total carbon, and unburned carbon.)

[0207] 2. Construct the likelihood function p((Y) obs |X))

[0208] Calling the forward model

[0209] For any candidate coal fingerprint X = [C1, C2, C3, C4, C5], the TabPFNRegressor is used to obtain the corresponding predicted mean and standard deviation in a single forward pass.

[0210]

[0211] Define Gaussian likelihood;

[0212] Assume that, given X, the observed Yobs follow independent and identically distributed Gaussian errors:

[0213]

[0214] 3. Define the prior p(X) for the coal fingerprint;

[0215] Based on domain knowledge, let's establish a reasonable prior for X = [C1, C2, C3, C4, C5]:

[0216] The sum of components is 1: a Dirichlet distribution can be used;

[0217] X ~ Dirichlet(α1,…,α5), where each α j Based on the typical content of historical coalfields (such as α) j =1 indicates uniformity and no preference.

[0218] Ingredient range: Guarantee each C j ∈[0,1], and the sum is 1.

[0219] 4. Obtain the posterior distribution p((Y) obs |X));

[0220] According to Bayes' theorem:

[0221] p((Y obs |X))∝p((Y obs |X))p(X)

[0222] This can be written as the logarithmic posterior (with the constant term removed):

[0223]

[0224] 5. Solve for the posterior: MAP estimation or MCMC sampling;

[0225] A. MAP estimation (finding the most likely X) * );

[0226] Define the objective function;

[0227]

[0228] Use gradient optimization;

[0229] Start from an initial value X(0) that satisfies the prior a priori condition (such as average coal quality);

[0230] Each step is calculated using numerical gradient or automatic differentiation. Update along the negative gradient direction until convergence.

[0231] In this embodiment, for sample 1, let

[0232] The optimized result is

[0233] It is very close to the actual input [0.220, 0.250, 0.070, 0.500, 0.880, 0.010].

[0234] B. MCMC sampling (quantization uncertainty)

[0235] Define the posterior distribution using the logp((X|Y)) given above. obs ))

[0236] Run HMC or NUTS;

[0237] In each sampling iteration, TabPFN is called to calculate μ for the current X. Y (X),σ Y (X), and combined with prior evaluation of posterior density; continuous sampling yields a set of {X(s)} samples, reflecting the distribution of "all possible coal fingerprints".

[0238] result:

[0239] Statistical analysis of each C from the sample set j The mean, variance, and 95% interval are used to obtain complete information on the inverse uncertainty.

[0240] 6. Closed-loop correction combining positive and negative methods;

[0241] The reverse calculation yielded Alternatively, for the sampled set, the predicted distribution can be calculated again using the forward model f, and its relationship with the observed Y can be verified. obs The degree of matching;

[0242] If a systematic bias is found, then... By adding inverse training or optimization sets, iteratively fine-tuning priors or optimizing hyperparameters, a self-correcting closed loop is formed.

[0243] Through the above steps:

[0244] a. Using TabPFN to provide efficient and accurate positive prediction μ Y (X),σ Y (X);

[0245] b. Using Bayesian inversion, coal fingerprints are inferred within the same model uncertainty framework, and the uncertainty is quantified;

[0246] c. Combine positive and negative closed loops to continuously optimize the back-reasoning results, ensuring that the output conforms to both positive experience and domain priors.

[0247] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention. The scope of the invention is defined by the appended claims and their equivalents, all of which should be included within the scope of protection of the invention.

Claims

1. A fingerprinting method based on the coal consumption process, characterized in that, Includes the following steps: S1. Constructing a carbon fingerprint spectrum for coal: By detecting the carbon content of organic matter components in coal samples, a coal carbon fingerprint vector is formed; S2. Constructing a fingerprint map of coal combustion waste: Detecting solid waste and liquid waste to form waste fingerprint vectors; S3. Establish a forward mapping model: Using the coal fingerprint vector as the input matrix and the waste fingerprint vector as the output matrix, the TabPFN machine learning model is used for multi-output regression. The data is preprocessed by quantile transformation and neighborhood threshold bucketing, and then processed by physical constraint verification to obtain the mapping relationship from coal fingerprint to waste fingerprint. S4. Reverse Coal Source Inference: Based on measured waste fingerprints, the posterior distribution is solved using a Bayesian inversion algorithm, and the original coal fingerprint is inferred through MAP estimation or MCMC sampling.

2. The fingerprinting method based on the coal consumption process according to claim 1, characterized in that, The specific steps in step S1 include: A mathematical model for carbon fingerprinting is constructed, defining the carbon fingerprint of coal as a multivariable function, where each variable represents a characteristic variable of the organic matter composition of coal of different ranks. The study obtained characteristic variables of organic matter components, including alkane carbon content, aromatic hydrocarbon carbon content, oxygen-containing functional group carbon content, total organic carbon content, and the ratio of alkane to aromatic hydrocarbons. For carbon fingerprint similarity measurement, we try to use Euclidean distance or Manhattan distance to calculate the similarity between different coal samples, thereby identifying coal and constructing a carbon fingerprint map of coal.

3. The fingerprinting method based on the coal consumption process according to claim 1, characterized in that, The specific steps in step S2 include: Acquire solid and liquid waste generated during coal combustion; Solid waste includes inorganic minerals, heavy metals, unburned carbon, and slag in ash. Liquid waste includes tar and water; the main components of tar include polycyclic aromatic hydrocarbons, alkanes, phenolic compounds, heterocyclic compounds, oxygen-containing compounds, nitrogen-containing compounds, and sulfur-containing compounds. Based on the different ranks of coal, the differences in organic matter composition directly affect the composition of waste after coal combustion, resulting in fingerprint profiles of coal combustion waste of different coal ranks.

4. The fingerprinting method based on the coal consumption process according to claim 1, characterized in that, The specific steps in step S3 include: S31. Use the carbon fingerprint spectrum of coal as the feature variable of the constructed input matrix, and the fingerprint spectrum of coal combustion waste as the feature variable of the constructed output matrix, and combine the two to establish a mapping relationship. S32. Extreme values ​​are compressed through quantile transformation, non-uniform binning is performed based on industrial threshold, missing values ​​are filled by neighborhood interpolation, and the total ash oxides are constrained and carbon balance is verified on the output. S33. Call the pre-trained TabPFN model, selectively unfreeze the top-level matrix for gradient fine-tuning, and combine semi-supervised distillation to enhance the adaptability of coal chemical characteristics, so as to achieve multi-output one-time prediction; S34. Use the Attention caching mechanism to accelerate the calculation and output the mean vector of waste fingerprints and its confidence interval. S35. Employ a tiered / plant-wide cross-validation strategy, focusing on RMSE and R... 2 The robustness of the confidence interval coverage assessment model is demonstrated, and the empirical inversion error is less than the error threshold.

5. The fingerprinting method based on the coal consumption process according to claim 1, characterized in that, The specific steps in step S4 include: S41. Based on actual detection, obtain waste fingerprints and construct an inversion input benchmark; S42. Call the TabPFN forward model and define the prediction output of any candidate coal fingerprint as a Gaussian likelihood function; S43. Use Dirichlet distribution as the prior for coal composition, set constraints, and calibrate parameters based on historical coal quality data. S44. Derive the logarithmic posterior using Bayes' formula, while preserving the relationships between core variables; S45. Perform MAP estimation or MCMC sampling to output the original coal fingerprint.

Citation Information

Patent Citations

  • Traceability method for determining moisture of solid material

    CN104568648A

  • Multi-dimensional fusion coal quality detection method and system

    CN119046684A