Calibration method for monitoring data of fine particulate matter sensor

By transforming the Hearst exponent of the time series and the distance from the monitoring point to the pollution source into qualitative characteristics, and combining cue words and sliding window sampling, a calibration model was constructed, which solved the calibration accuracy problem of the sensor under cross-seasonal and cross-regional conditions, and achieved higher PM2.5 monitoring accuracy.

CN121786796APending Publication Date: 2026-04-03CHONGQING UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing sensor calibration methods struggle to achieve high-precision PM2.5 monitoring under cross-seasonal and cross-regional conditions. Traditional methods cannot effectively capture the nonlinear dependence and time-varying characteristics in environmental data, resulting in insufficient calibration accuracy.

Method used

The qualitative representation and inference module transforms the Hearst exponent of the time series and the distance from the monitoring point to the pollution source into qualitative characteristics. The prompt word module generates prompt words to guide the TimeCMA module. Combined with the sliding window sampling and calibration value generation modules, a calibration model is constructed to perform data calibration.

Benefits of technology

It significantly improved the calibration accuracy of PM2.5 monitoring across seasons and regions, reducing the average MSE by 28.91%, the average absolute error by 17.38%, and improving the determination coefficient R² by 208.40%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786796A_ABST
    Figure CN121786796A_ABST
Patent Text Reader

Abstract

The invention discloses a method for calibrating monitoring data of a fine particulate matter sensor, which comprises the following steps of: acquiring a time sequence obtained by monitoring particulate matter concentration by a single sensor, correcting the time sequence by a calibration model which is trained and tested to be qualified, and outputting calibration data; the calibration model is composed of a qualitative representation and reasoning module, a cue word module, a TimeCMA module and a calibration value generation module. The method is applied to monitoring data correction of cross-seasonal and cross-regional particulate matter concentration sensors, experiments based on a real cross-seasonal and cross-regional data set show that compared with an existing method, the method has the advantages in mean square error, mean absolute error and judgment coefficient R2, it is proved that the method can improve the precision of the sensor monitoring data, and the accuracy of the sensor monitoring data is improved. The problem of monitoring data correction of particulate matter concentration sensors which are arranged in a cross-season and cross-region mode can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sensor calibration technology, and in particular to a calibration method for monitoring data of fine particulate matter sensors. Background Technology

[0002] Accurate and timely fine particulate matter (such as PM2.5) 2.5 Measurement is fundamental to reliable air quality monitoring, and fine particulate matter (PM2.5) concentration monitoring is crucial for environmental managers to make informed decisions. For example, PM2.5 monitoring in hospital environments is essential. 2.5 Concentration monitoring can help hospital administrators, infection control specialists, and clinical medical staff make scientific decisions regarding hospital environmental quality control, cross-infection risk warning, and patient area management. However, PM2.5 concentration monitoring... 2.5 The reliability of data is highly dependent on the measurement accuracy of the sensors. Despite continuous advancements in sensing technology, calibration remains crucial for ensuring PM2.5 accuracy. 2.5 Key challenges in ensuring stable performance of monitoring equipment.

[0003] PM 2.5 Monitoring sensors are susceptible to measurement biases caused by various factors, including sensor aging and fluctuations in environmental conditions such as temperature and humidity. A particularly prominent issue is the need for cross-seasonal and cross-regional calibration. Seasonal changes in weather patterns and atmospheric conditions introduce systematic errors into sensor readings, requiring recalibration to maintain data accuracy. Similarly, spatial characteristics such as the distance to pollution sources increase calibration difficulty, as sensors need to adapt to different environmental scenarios to provide reliable measurements. Therefore, developing high-performance calibration methods that comprehensively consider multiple factors is crucial for improving PM2.5 accuracy. 2.5 The accuracy and reliability of the monitoring system are of great significance and are a key link in promoting the development of precision medical environment management and health protection for high-risk patients.

[0004] Traditional calibration methods (such as linear regression and random forests), while capable of modeling simple relationships between sensor and reference signals, fail to capture the inherent nonlinear dependencies and time-varying characteristics of environmental data. This leads to a significant decline in generalization performance under changing seasonal and regional conditions. To address these issues, recent research has redefined sensor calibration as a time-series modeling problem, aiming to learn the dynamic mapping between sensor signals and reference observations over time. This perspective allows calibration frameworks to utilize both short-term patterns (such as transient environmental changes) and long-term trends (such as sensor drift). In this context, deep learning architectures originally designed for time series forecasting have provided important insights. For example, convolutional neural networks (CNNs) and long short-term memory (LSTM) networks have been used to capture local and temporal dependencies in air quality time series. Recently, Transformer-based models (such as iTransformer) have demonstrated stronger global temporal correlation modeling capabilities through self-attention mechanisms. These advances indicate that calibration processes can benefit from time-series modeling principles, enabling more adaptive and generalizable sensor calibration under diverse environmental conditions. However, these models rely primarily on numerical sequences and still struggle to incorporate higher-order context or knowledge information that can guide model reasoning.

[0005] With the rapid development of Large Language Models (LLMs), researchers have begun to explore their application potential in time series modeling. However, unlike text data composed of discrete symbolic tags, time series data possess characteristics such as continuity, structure, and metric dependence, leading to a fundamental modal difference between numerical signals and linguistic representations. This difference prevents LLMs from directly capturing the temporal dependencies or physical meanings encoded in sensor data. Early research attempted to adapt LLMs to numerical sequences by designing continuous embedding interfaces or tokenization schemes, but these methods often resulted in semantic misalignment problems, where the model treats numerical values ​​as arbitrary tags without quantitative relationships. To address this issue, Time-LLM introduces a cross-modal reprogramming mechanism, projecting time series segments onto a text prototype space through fragment reprogramming, achieving effective alignment between continuous sensor signals and tag-based language embeddings. This alignment allows LLMs to infer temporal patterns using linguistic priors while preserving numerical structure. Despite these advances, frameworks like Time-LLM still face fundamental challenges in cross-modal understanding. The semantic richness of textual prompts and the quantitative accuracy of time-series signals often become entangled during training, leading to data entanglement problems. Time series with Cross-Modal Alignment (TimeCMA) addresses this issue by introducing a cross-modal alignment module. This module retrieves decoupled and robust time-series embeddings from the prompt embeddings generated by LLM based on channel similarity, thereby separating temporal features from the textual context and mitigating the entanglement effect.

[0006] However, TimeCMA still has two key limitations when applied to sensor calibration tasks: First, its cue words lack domain-specific knowledge about spatial and temporal factors, resulting in poor calibration accuracy under cross-seasonal and cross-regional conditions; second, TimeCMA is designed specifically for time series forecasting rather than calibration tasks, and its fundamental goal of predicting future values ​​is fundamentally different from the goal of aligning sensor outputs with reference measurements in calibration tasks, thus making it unsuitable for calibration tasks. Summary of the Invention

[0007] In view of this, the present invention provides a calibration method for fine particulate matter sensor monitoring data to solve the technical problem of calibrating data measured by a single particulate matter monitoring sensor when deployed across seasons and regions.

[0008] The calibration method for fine particulate matter sensor monitoring data of the present invention includes: acquiring a time series of particulate matter concentrations monitored by a single sensor, correcting the time series using a calibration model that has been trained and tested, and outputting calibration data; the calibration model consists of a qualitative representation and inference module, a prompt word module, a TimeCMA module, and a calibration value generation module, and the correction includes:

[0009] The time series and the distance of the monitoring point from the pollution source are input into the qualitative representation and inference module. The qualitative representation and inference module transforms the Hearst exponent of the time series and the distance of the monitoring point from the pollution source into qualitative characteristics, and performs qualitative inference on the time series based on the qualitative characteristics.

[0010] The qualitative characterization results and qualitative inference results are input into the prompt word module, which generates prompt words to guide the prediction of the TimeCMA module.

[0011] The sliding window sampling method is used to sample the time series. The sampled data and prompt words are input into the TimeCMA module, which then provides a predicted value for each sample in the time series.

[0012] An input sample is constructed using all predicted values ​​of a time series sample. The number of data points in each input sample is standardized to the number of sliding window samples. An input sequence is constructed using the input samples corresponding to each sample of the time series. The input sequence is then input into the calibration value generation module, which generates the calibration result for the time series.

[0013] Furthermore, the qualitative representation and inference module converts the Hearst exponent of the time series into a qualitative representation using the following method:

[0014] Let B = {mean-reverting, random, persistent} be a finite vocabulary set used to characterize behavioral relationships in a time series, and define a function h: A mapping function to map Hearst exponent values ​​to qualitative lexical terms:

[0015]

[0016] Where H is the Hearst exponent, β1 and β2 are the threshold boundaries for behavior classification, 0 < β1 < β2 < 1; qualitative descriptions using the terms mean-reverting, random, and persistent follow these rules:

[0017]

[0018] Terms that satisfy the mapping and transition rules—mean-reverting, random, and persistent time series—are called behavioral primitives.

[0019] Input h(H) into the prompt word module, which generates the prompt word: [Time Series Behavior]: This sequence presents...<h(H)> Behavioral characteristics.

[0020] Furthermore, the qualitative representation and reasoning module converts the distance from the monitoring point to the pollution source into a qualitative characterization using the following method:

[0021] Define monitoring point p i Distance d to the pollution source i For p i Let D = {close, medium, far} be a finite set of terms representing distance relationships, and let the function f be the Euclidean distance to the nearest point from the pollution source. Defined as a mapping function that maps metric distance to qualitative vocabulary:

[0022]

[0023] Where thresholds θ1 and θ2 are the boundaries for word segmentation, 0 < θ1 < θ2; qualitative descriptions using the words close, medium, and far follow these rules:

[0024] The terms close, medium, and far that satisfy conditions (3) and (4) are called distance primitives.

[0025] Furthermore, the qualitative representation and reasoning module performs qualitative spatial reasoning as follows:

[0026] Let C = {low, moderate, high} be a finite set of terms qualitatively describing particulate matter concentration. The function q: D → C is defined as a function that maps distance primitives to qualitative terms:

[0027]

[0028] When using the words low, moderate, and high for qualitative description, follow these rules:

[0029]

[0030] The words low, moderate, and high are called concentration units when they simultaneously satisfy conditions (5) and (6);

[0031] Qualitative inferences about particulate matter concentration can be made from the composite relationship between functions f and q:

[0032]

[0033] in, The distance from the monitoring point to the pollution source is represented by the symbol °, which indicates the composite operation of the function.

[0034] f(d) i )and The input prompt word module generates the following prompt word: [Expected particulate matter concentration level]: Due to the distance between the monitoring point and the main pollution source... <f(d i The expected particulate matter concentration is...

[0035] Furthermore, the calibration value generation module is a backpropagation network, which includes a first fully connected layer, a first ReLU function layer, a first Dropout layer, a second fully connected layer, a second ReLU function layer, a second Dropout layer, and a third fully connected layer connected in sequence. The number of neurons in the second fully connected layer is only half that of the first fully connected layer, and the number of neurons in the third fully connected layer is one.

[0036] Furthermore, the calibration model is trained and tested using the following method:

[0037] The sensor and reference analyzer are arranged together at m different locations P = {p1, p2, ..., p...} m Monitor particulate matter concentration;

[0038] The sensor at each monitoring point p i The observed fine particulate matter concentration values ​​∈P constitute a time series X i : Where T represents the time domain, i∈[1,m], and the time series obtained by the sensor at each monitoring point constitute the time series set X={X1,X2,…,X…} m}, X i ∈X;

[0039] Set the reference analyzer at monitoring point p i The observed fine particulate matter concentration values ​​∈P constitute a time series Y i : Where T represents the time domain, i∈[1,m], and the time series obtained by the reference analyzer at each monitoring point constitute the time series set Y={Y1,Y2,…,Y…} m}, Y i ∈Y;

[0040] Suppose that the sensor needs to be monitored at point p. k The measured values ​​were calibrated, and X was... k As test set D teThe training set D is constructed using data from the remaining monitoring points. tr D tr ={(X i (t),Y i (t))|i∈[1,m],i≠k,t∈T i}, where T i ∈T indicates that at monitoring point p i The set of monitoring time points;

[0041] Using training set D tr The calibration model is trained, and during the training process, data is transferred from the training set D. tr The calibration model was validated using a subset of data, with test set D. te The qualified calibration model is tested to obtain a qualified calibration model.

[0042] The beneficial effects of this invention are:

[0043] This invention provides a calibration method for fine particulate matter sensor monitoring data. Applying this method to the correction of particulate matter concentration sensor monitoring data deployed across seasons and regions, experiments based on real cross-seasonal and cross-regional datasets show that, compared to existing methods, the method of this invention reduces MSE by an average of 28.91% and the mean absolute error by 17.38%, while also improving the determination coefficient R0. 2 An average improvement of 208.40% confirms that the method of the present invention can improve the accuracy of sensor monitoring data and solve the problem of data correction for particulate matter concentration sensors deployed across seasons and regions. Attached Figure Description

[0044] Figure 1 A schematic diagram of the framework for calibrating the TimeCalibration model.

[0045] Figure 2 A schematic diagram illustrating the necessity of generating calibration values. (a) For the input time series x i (a) Using the sliding window sampling method, a total of N-n+1 input samples are generated. (b) Due to the overlap of data points between samples, TimeCMA may generate multiple predicted values ​​for a single data point.

[0046] Figure 3 A schematic diagram of the data processing flow for the calibration value generation module.

[0047] Figure 4 For the training set (D) tr ), Validation set (D va ) and test set (D te (Diagram showing PM levels at different times) 2.5Concentration change trend.

[0048] Figure 5 Performance comparison of the TimeCalibration model using different cue words on the AirHerit-PM2.5 dataset. (a) MSE, (b) MAE, (c) R 2 .

[0049] Figure 6 The chart shows the loss curves of the TimeCalibration model over 100 training epochs. (a) TimeCMA module loss, (b) CVG module loss.

[0050] Figure 7 Attention maps of the channel similarity matrix under different cue conditions. (a) Example input sample (sample 1 is selected, which is a segment of the input time series). (b) Attention map using cue 1. (c) Attention map using cue 6.

[0051] Figure 8 This is a Truth-Prediction graph showing the difference between the true concentration (Truth) on the test set and the calibration result (Prediction) of the method of this invention. Detailed Implementation

[0052] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0053] The calibration method for fine particulate matter sensor monitoring data in this embodiment includes: acquiring a time series of particulate matter concentrations monitored by a single sensor, correcting the time series using a trained and tested calibration model, and outputting calibration data. For example... Figure 1 As shown, the calibration model consists of a Qualitative Representation and Inference Module (QRR), a Cue Word Module, a TimeCMA Module (Time series with Cross-Modal Alignment, TimeCMA), and a Calibration Value Generation Module (CVG). The calibration includes:

[0054] The time series and the distance of the monitoring point from the pollution source are input into the qualitative representation and inference module. The qualitative representation and inference module transforms the Hearst exponent of the time series and the distance of the monitoring point from the pollution source into qualitative characteristics, and performs qualitative inference on the time series based on the qualitative characteristics.

[0055] The Hurst exponent is a key indicator characterizing the long-term memory and persistence of time series data. By quantifying the mean reversion tendency or long-term trend continuation characteristics of time series, the Hurst exponent essentially reveals the behavioral patterns of time series. The qualitative representation and inference module characterizes the behavioral features indicated by the Hurst exponent, thereby achieving the classification of time series persistence. In this embodiment, the qualitative representation and inference module converts the Hurst exponent of a time series into a qualitative representation as follows:

[0056] Let B = {mean-reverting, random, persistent} be a finite vocabulary set used to characterize behavioral relationships in a time series, and define a function h: A mapping function to map Hearst exponent values ​​to qualitative lexical terms:

[0057]

[0058] Where H is the Hearst exponent, β1 and β2 are the threshold boundaries for behavior classification, 0 < β1 < β2 < 1; qualitative descriptions using the terms mean-reverting, random, and persistent follow these rules:

[0059]

[0060] Time series that satisfy the mapping and transition rules—mean-reverting, random, and persistent—are called behavioral primitives. Mean-reverting time series indicate anti-persistence, exhibiting a tendency to revert to the mean; randomness corresponds to memoryless characteristics similar to Brownian motion; and persistence characterizes trend-persistent behavior with long-range correlation. This characterization system provides three basic classification categories for the dynamic characteristics of time series, with specific threshold parameters set as β1 = 0.4 and β2 = 0.6. This method enables the calibration model to dynamically adjust its strategy based on the mean-reverting, random, or persistent behavior exhibited by the time series, thereby achieving more environmentally conscious calibration performance.

[0061] Empirical data show a significant negative correlation between particulate matter concentration and the distance from the sensor to the pollution source. In this embodiment, the qualitative representation and inference module converts the distance from the monitoring point to the pollution source into a qualitative characterization as follows:

[0062] Define monitoring point p i Distance d to the pollution source i For p iLet D = {close, medium, far} be a finite set of terms representing distance relationships, and let the function f be the Euclidean distance to the nearest point from the pollution source. Defined as a mapping function that maps metric distance to qualitative vocabulary:

[0063]

[0064] Wherein, thresholds θ1 and θ2 are the boundaries for word division, 0 < θ1 < θ2; the qualitative descriptions using the words close (indicating that the monitoring point is close to the pollution source), medium (indicating that the monitoring point is at a moderate distance from the pollution source), and far (indicating that the monitoring point is far from the pollution source) follow the following rules:

[0065] The terms close, medium, and far that satisfy conditions (3) and (4) are called distance primitives.

[0066] Qualitative spatial reasoning is performed based on the above qualitative distance representation method to analyze the spatial correlation of fine particulate matter concentrations among different monitoring points. The method for qualitative spatial reasoning by the qualitative representation and reasoning module is as follows:

[0067] Let C = {low, moderate, high} be a finite set of terms qualitatively describing particulate matter concentration. The function q: D → C is defined as a function that maps distance primitives to qualitative terms:

[0068]

[0069] When using the words low, moderate, and high for qualitative description, follow these rules:

[0070]

[0071] The terms low, moderate, and high, when simultaneously satisfying conditions (5) and (6), are called concentration primitives. Condition (5) is based on the empirical inverse relationship between fine particulate matter concentration and distance, indicating that locations closer to the pollution source generally have higher pollutant concentrations. A qualitative inference about particulate matter concentration is made from the composite relationship between functions f and q:

[0072]

[0073] in, The symbol represents the distance between the monitoring point and the pollution source. This represents the composition of functions.

[0074] Input h(H) into the prompt word module, which generates the prompt word: [Time Series Behavior]: This sequence presents...<h(H)> Behavioral characteristics. f(d)i )and The input prompt word module generates the following prompt word: [Expected particulate matter concentration level]: Due to the distance between the monitoring point and the main pollution source... <f(d i The expected particulate matter concentration is...

[0075] A sliding window sampling method is used to sample the time series data. The sampled data and prompt words are input into the TimeCMA module, which then outputs the predicted value for each sample in the time series. Figure 2 As shown in (a), let T i ={t1,t2,…,t N} represents the sensor at position p i The set of monitoring time points, and the corresponding particulate matter concentration vector is represented as x. i =[X i (t1),X i (t2),…,X i (t N )] T = [x1,x2,…x N ] T In processing x i At that time, the TimeCMA module uses the sliding window sampling (SWS) method, which reads n consecutive data points at a time, and then slides the window forward one point to read the next n points. Therefore, for a scalar time series x of length N... i It is divided into N-n+1 overlapping samples (where 0 <n≤N)。

[0076] While the Sliding Window Sampling (SWS) method efficiently utilizes data, it can lead to multiple predictions for a single data point. Specifically, when the sample sequence length is n, the model receives a window containing n data points as input and outputs a prediction vector of length n. For the j-th sample (j∈{1,N}), its output is... When the sliding step size is 1, subsequent samples will generate predictions. The process is as follows Figure 2 As shown in (b).

[0077] Therefore, any data point x in the middle of the time series j (Excluding the beginning and end of the sequence) will be covered by multiple overlapping samples. For example... Figure 2 (b) with x j The corresponding columns show each x jMultiple predicted values ​​will be obtained. Obtaining the final calibration value based on these predicted values ​​is crucial for improving prediction accuracy; this task is performed by the calibration value generation module in this embodiment. An input sample is constructed using all predicted values ​​of a sample from the time series. The number of data points in all input samples is standardized to n. If the number of predicted values ​​is less than n, the arithmetic mean of all predicted values ​​within that input sample is used. The process involves filling in the time series data with input samples corresponding to each sample in the time series to form an input sequence. This input sequence is then fed into the calibration value generation module, which generates the calibration result for the time series.

[0078] In this embodiment, the calibration value generation module is a backpropagation network, such as... Figure 3 As shown, the backpropagation network (BP network) includes a first fully connected layer, a first ReLU function layer, a first Dropout layer, a second fully connected layer, a second ReLU function layer, a second Dropout layer, and a third fully connected layer connected in sequence. The number of neurons in the second fully connected layer is only half that of the first fully connected layer, and the number of neurons in the third fully connected layer is one.

[0079] In this embodiment, the training and inference process of the calibration value generation module (i.e., the backpropagation network) is specifically designed to handle multiple predicted values ​​generated for each data point. The calibration model is trained and tested using the following method:

[0080] The sensor and reference analyzer are arranged together at m different locations P = {p1, p2, ..., p...} m Monitor particulate matter concentration.

[0081] The sensor at each monitoring point p i The observed fine particulate matter concentration values ​​∈P constitute a time series X i : Where T represents the time domain, i∈[1,m], and the time series obtained by the sensor at each monitoring point constitute the time series set X={X1,X2,…,X…} m}, X i ∈X.

[0082] Set the reference analyzer at monitoring point p i The observed fine particulate matter concentration values ​​∈P constitute a time series Y i : Where T represents the time domain, i∈[1,m], and the time series obtained by the reference analyzer at each monitoring point constitute the time series set Y={Y1,Y2,…,Y…} m}, Y i ∈Y.

[0083] Time series X i Compared with the corresponding reference measurement value Y i Having the same length and time alignment, i.e., the domain dom(X) i ) = dom(Y i ).

[0084] Suppose that the sensor needs to be monitored at point p. k The measured values ​​were calibrated, and X was... k As test set D te The training set D is constructed using data from the remaining monitoring points. tr D tr ={(X i (t),Y i (t))|i∈[1,m],i≠k,t∈T i}, where T i ∈T indicates that at monitoring point p i The set of monitoring time points.

[0085] Using training set D tr The calibration model is trained, and during the training process, data is transferred from the training set D. tr Select a portion of the data (e.g., the training set D) tr The calibration model was validated using 30% of the data (from the test set D). te The qualified calibration model is tested to obtain a qualified calibration model.

[0086] The advantages of the calibration method for fine particulate matter sensor monitoring data in this embodiment are illustrated below through experiments.

[0087] Experimental setup

[0088] The calibration model proposed in this embodiment is named TimeCalibration.

[0089] Through ENEA MONICA TM Air quality monitoring system collects PM 2.5Concentration data: The system is equipped with electrochemical and optical sensors, supplemented by real-world concentration data provided by the LM02 benchmark analyzer, for model training and performance rating calculations. During the year-and-a-half observation period, the MONICA monitoring system and the LM02 benchmark analyzer jointly completed three observation experiments: winter 2020-2021 (February 5 to March 2, 2021, at point p1), summer 2021 (August 24 to September 14, 2021, at point p2), and winter 2021-2022 (February 9 to March 3, 2022, at point p3). During each experiment, MONICA and LM02 underwent deployment for approximately three weeks. The MONICA system collected data every 6 seconds, while the LM02 benchmark analyzer collected PM data every hour. 2.5 Numerical data. The data collected by MONICA within one hour was arithmetically averaged and then synchronized with the LM02 data stream to form the final dataset used in the experiment.

[0090] In this experiment, all summer data were used for testing, while winter data were used for training and validation. The validation set consisted of observations from the last week (7 days) of the winters of 2020-2021 and 2021-2022. Based on time-series data collected by MONICA and LM02, the training, validation, and test sets contained 774, 334, and 488 data points, respectively. Details of their temporal distribution characteristics can be found in [link to relevant documentation]. Figure 4 To ensure a sufficient sample size, a sliding window method is used. Figure 4 All time series data were segmented into sample units consisting of 24 consecutive data points, with a default sliding step size of 1. After segmentation, a training set of 751 samples, a validation set of 311 samples, and a test set of 465 samples were obtained. For consistent terminology, this dataset will be referred to as AirHerit-PM2.5 from now on. Although the range of AirHerit-PM2.5 is larger than that of a typical hospital environment, it is a typical cross-seasonal and cross-regional dataset. Therefore, this dataset was used in the experiments to test the performance of our model.

[0091] The task of this experiment is as follows: to train a model using winter data to learn PM levels monitored by MONICA. 2.5 The mapping relationship between numerical values ​​and the actual concentration measured by LM02 was established, and the trained model was then applied to MONICA summer data to calibrate PM2.5 concentration. 2.5 This allows us to obtain readings that are closer to the true concentration. This is a typical regression problem, specifically a calibration task that spans seasons and regions.

[0092] The following metrics are used to evaluate the performance of different methods: mean squared error (MSE); mean absolute error (MAE); coefficient of determination (R²). 2Training time (Tr, in seconds); testing time (Te, in milliseconds); number of parameters (P, in millions); and billions of floating-point operations per second (GFLOPS). It's important to note that MSE quantifies the variance of the squared error, penalizing larger biases more severely and highlighting the model's sensitivity to extreme values; while MAE measures the average magnitude of the absolute value of the prediction error in the raw units, reflecting the overall accuracy of the model without overemphasizing outliers. R 2 This represents the proportion of variance explained by the model. Tr is an important metric for regression models that need to be trained, as practical applications aim for shorter processing times. Te is another important metric for evaluating inference performance, covering the entire test set rather than just a single data point in this experiment. Regarding parameter count and GFLOPS, more parameters or higher GFLOPS generally indicate greater model complexity and slower processing speed. Therefore, the goal is to use models with fewer parameters and lower GFLOPS without sacrificing accuracy. Besides Tr, which measures the model's performance on the training set, the other metrics measure its performance on the test set. In the experiments, the Large Language Model (LLM)-based encoding branch in the TimeCMA module used Qwen1.5-1.8B-chat (a large language model); however, in different experiments or implementations, other types of large language models can also be used for the LLM-based encoding branch in the TimeCMA module. See Table 1 for hyperparameter configuration details. For fair comparison, we report the average results of five runs to reduce bias caused by random factors during testing.

[0093] Table 1. Hyperparameter configuration used in this model

[0094]

[0095] All experiments were implemented using Python and the PyTorch framework, and ran on an NVIDIA RTX 3090Ti graphics card with 10,752 compute cores and 24GB of global video memory.

[0096] Experimental results

[0097] 1) CVG Assessment

[0098] A comprehensive comparison was conducted to verify the performance differences of different calibration value generation algorithms. Table 2 shows the comparison results of the TimeCalibration model using four different calibration value generation modules. The study found that although the BP network slightly increases training and inference time compared to simple heuristic and statistical methods, it achieves the best calibration accuracy—with the lowest MSE and MAE scores, and R0. 2The BP network exhibits the highest coefficient of determination. This indicates that the BP network can learn the complex mapping relationship from multiple candidate predicted values ​​to the optimal calibration value through training on historical data, thereby effectively utilizing the time dependence and error patterns inherent in overlapping samples. Furthermore, while the VGG16-based method utilizes a complex pre-trained visual architecture, it incurs a higher computational load and is slower (as shown by the higher Tr and Te values ​​in Table 2). In contrast, the BP network designed in this invention achieves superior accuracy (lower MSE / MAE, higher Ri) with significantly lower computational cost. 2 This demonstrates that it is more effective in solving the tasks presented in this invention.

[0099] (ii) Prompt word design and evaluation

[0100] The effectiveness of the prompt module design in this invention was verified through comparative experiments. As shown in Table 2, six prompt templates (Prompt 1 to Prompt 6) were designed, constituting the ablation experimental system of the prompt design. Prompt 1 was used as the baseline, which only provides basic time series data and background information. Compared with Prompt 1, Prompt 2 introduced the "[Expected Particulate Matter Concentration Level]" descriptor to verify the impact of qualitative spatial reasoning based on pollution source distance on performance; Prompt 3 added the "[Hurst Index]" descriptor separately to test the independent role of qualitative temporal behavioral reasoning; Prompt 4 integrated spatiotemporal dual qualitative features to evaluate their synergistic effect; Prompt 5 further transformed qualitative features into more natural explanatory statements to test the benefit of enhanced semantic clarity to the TimeCMA module; Prompt 6 removed the explicit "[Time Period]" and "[Sampling Frequency]" descriptors based on Prompt 5 to explore whether the model can mainly rely on the rich qualitative context provided by QRR for reasoning.

[0101] Table 2 shows the different prompt designs used in this model.

[0102]

[0103]

[0104] Figure 5 The experimental results are presented. Since adjusting the prompt template primarily affects the prediction accuracy, for ease of comparison, Figure 5 List only MSE, MAE, and R 2 Three accuracy metrics. The experimental results show that the performance of TimeCalibration under different cue designs exhibits a clear and meaningful progressive pattern:

[0105] A baseline template providing only the basic timing context (Hint 1) yields MSE = 60.48, MAE = 5.53, and R... 2=0.35 performance. With the gradual introduction of qualitative reasoning elements, the system performance shows a continuous improvement: the introduction of "[expected PM 2.5 The prompt for the [Level] descriptor 2 lowered the MSE to 58.24 and the MAE to 5.46, while R 2 The value was increased to 0.37, verifying the effectiveness of qualitative spatial reasoning; examining the "[Hurst index]" descriptor alone, hint 3 yielded MSE = 57.94 and R... 2 The equivalent result of 0.37 confirms the independent value of qualitative temporal behavioral reasoning.

[0106] Hint 5's performance advantage over Hint 4 indicates that, for this task, providing semantically rich qualitative context is more impactful than simply predefined qualitative features. This demonstrates that large language models can effectively utilize high-level inference descriptors to compensate for the deficiencies of the original temporal data, thereby achieving more robust cross-modal alignment in calibration tasks.

[0107] Comprehensive suggestion 6: Remove non-critical metadata while retaining the refined qualitative descriptor, with MSE=56.67, R 2 The best overall performance was achieved with a coefficient of performance (COP) of 0.39 and an MAE of 5.47. The performance gains from COP1 to COP6 confirm the core hypothesis: injecting domain knowledge through qualitative representation and reasoning (QRR) can effectively guide the model to achieve more accurate calibration. Furthermore, the COP6 scheme demonstrates significant advantages in calibrating the model.

[0108] III) Comparison of Advanced Methods

[0109] The proposed method (model TimeCalibration) is compared with existing state-of-the-art calibration methods, including linear regression, CNN-LSTM, decision trees, random forests, iTransformer, and XGBoost. All models are initialized with pre-trained weights and trained on the same dataset, AirHerit-PM2.5, to ensure a fair comparison. Table 3 summarizes the comparison results, where indicators marked "↓" indicate that smaller values ​​are better, and indicators marked "↑" indicate that larger values ​​are better.

[0110] Table 3. Different methods in MONICA device PM 2.5 Performance comparison in calibration tasks.

[0111]

[0112] As shown in Table 3, the TimeCalibration method proposed in this invention achieves the best calibration accuracy among all comparative methods, specifically with the lowest MSE (56.67) and MAE (5.47), and the highest R... 2(0.39). This performance advantage is significant: compared with all comparison methods, MSE was reduced by an average of 15.72%, MAE by an average of 11.11%, and R... 2 The average improvement was 291.83%. Although traditional methods (such as linear regression) have the lowest computational cost, their calibration accuracy is significantly worse. Of particular note is that TimeCalibration significantly outperformed all the baseline methods compared. This indicates that despite the increased computational cost, the integration of QRR with the CVG module effectively enhances the model's performance in cross-seasonal and cross-regional sensor detection data calibration tasks.

[0113] IV) Model Analysis

[0114] First, through Figure 6 The loss curves shown visually demonstrate the model's performance. To obtain more complete information on loss changes, the training period was extended to 100 epochs. Since the TimeCMA and CVG sub-modules in the model need to be trained separately... Figure 6 The loss curves for these two sub-modules are plotted separately.

[0115] according to Figure 6 The loss curves shown indicate that the two sub-modules exhibit different training dynamics. The TimeCMA sub-module ( Figure 6 (a) During the first 13 training epochs, both training loss and validation loss showed a decreasing trend, indicating that the module could effectively learn cross-modal alignment. However, a significant divergence occurred after the 13th epoch: the validation loss began to fluctuate and increase, while the training loss continued to decrease slowly. This phenomenon suggests that the TimeCMA component has slight overfitting, that is, the model begins to memorize specific patterns of the training set rather than maintain generalization ability. In contrast, the CVG submodule ( Figure 6 (b) The module exhibits highly stable training characteristics. After 40 epochs, both training and validation losses converge smoothly, maintaining a small gap between them, indicating excellent generalization ability—crucial for its core function of generating optimal solutions from multiple candidate predictions. This differentiated training mode confirms the unique advantages of the decoupled training architecture: to prevent overfitting and optimize generalization performance, an early stopping strategy is adopted for TimeCMA, and the training epoch is set to 30. Experiments confirm that this setting effectively controls the overfitting trend of TimeCMA, while the CVG module can be trained to full convergence without worrying about a decline in generalization. The synergistic effect of the two components supports the overall calibration performance of the model.

[0116] Secondly, to explore the impact of Qualitative Representation and Reasoning (QRR) on the model's internal processing mechanism, we compared and analyzed the attention map of the channel similarity matrix in the TimeCMA cross-modal alignment module. The comparative experiment was conducted under two cue conditions: using the baseline cue template proposed in Table 2 (Cue 1), and using the QRR enhanced cue template that integrates the descriptors "[expected particulate matter concentration level]" and "[Hurst index]" (Cue 6).

[0117] Select Figure 7 The sample shown in (a) serves as a visualization input example, and attention maps generated by the matrix under different cue conditions are extracted. Figure 7 As shown in (b) and (c).

[0118] By comparing and analyzing attention maps, significant pattern differences in how QRR affects the internal processing mechanisms of the model can be observed. Under the condition of using baseline cue 1 ( Figure 7 (b) The attention distribution is relatively evenly dispersed, with values ​​mainly concentrated in the 0.12-0.18 range and no obvious focal point. This indicates that the model struggles to identify and focus on key features in the absence of qualitative guidance. However, under the QRR-enhanced cue 6 condition ( Figure 7 (c) The attention pattern exhibits more structured and discriminative characteristics: some regions in the figure receive significantly higher attention weights (up to 0.24), while other regions are strongly suppressed (down to 0.07). This polarized distribution indicates that when the model is equipped with the QRR descriptor, it can successfully focus on feature units aligned with the semantics of qualitative reasoning (such as the corresponding expected PM). 2.5 (Characteristics of horizontal and temporal behavior) while filtering secondary information. The formation of this more interpretable and targeted attention mechanism confirms how QRR guides the cross-modal alignment process, thereby achieving the calibration performance improvement observed in the experiments. Finally, Figure 8 The differences between the actual concentration of LM02 collected on the test set and the calibration results of this method are presented in a visual manner.

[0119] according to Figure 8The analysis of the truth-prediction difference between the actual concentration and the calibration results reveals several statistically significant patterns. These patterns highlight the performance advantages of the TimeCalibration method: the error distribution closely surrounds the mean value (approximately 0.31) close to zero, indicating that the model achieves extremely low systematic bias and avoids persistent overestimation or underestimation, a significant advantage over traditional methods that often exhibit significant bias; the vast majority of errors are concentrated in a narrow range, demonstrating high overall calibration accuracy and stability, which also shows that the integration of qualitative spatiotemporal inference effectively alleviates the common peak underestimation problem in air quality calibration; the weak temporal clustering of errors indicates that the model can effectively capture short-term fluctuations and long-term trends without overfitting to specific time periods. These statistical characteristics collectively validate that this method can achieve high-precision, high-stability, and highly temporally consistent calibration results. Its advantage stems from the cross-modal alignment mechanism—minimizing errors by fusing numerical time-series data with qualitative domain knowledge.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A calibration method for monitoring data from a fine particulate matter sensor, characterized in that: include: The time series of particulate matter concentrations monitored by a single sensor is collected, and the time series is corrected by a calibration model that has been trained and tested, and the calibration data is output. The calibration model consists of a qualitative representation and inference module, a prompt word module, a TimeCMA module, and a calibration value generation module. The calibration includes: The time series and the distance of the monitoring point from the pollution source are input into the qualitative representation and inference module. The qualitative representation and inference module transforms the Hearst exponent of the time series and the distance of the monitoring point from the pollution source into qualitative characteristics, and performs qualitative inference on the time series based on the qualitative characteristics. The qualitative characterization results and qualitative inference results are input into the prompt word module, which generates prompt words to guide the prediction of the TimeCMA module. The sliding window sampling method is used to sample the time series. The sampled data and prompt words are input into the TimeCMA module, which then provides a predicted value for each sample in the time series. An input sample is constructed using all predicted values ​​of a time series sample. The number of data points in each input sample is standardized to the number of sliding window samples. An input sequence is constructed using the input samples corresponding to each sample of the time series. The input sequence is then input into the calibration value generation module, which generates the calibration result for the time series.

2. The calibration method for fine particulate matter sensor monitoring data according to claim 1, characterized in that: The qualitative representation and inference module converts the Hearst exponent of the time series into a qualitative representation using the following method: Let B = {mean-reverting, random, persistent} be a finite vocabulary set used to characterize behavioral relationships in time series, and define a function h: A mapping function to map Hearst exponent values ​​to qualitative lexical terms: Where H is the Hearst exponent, β1 and β2 are the threshold boundaries for behavior classification, 0 < β1 < β2 < 1; qualitative descriptions using the terms mean-reverting, random, and persistent follow these rules: Terms that satisfy the mapping and transition rules—mean-reverting, random, and persistent time series—are called behavioral primitives. Input h(H) into the prompt word module, which generates the prompt word: [Time Series Behavior]: This sequence presents...<h(H)> Behavioral characteristics.

3. The calibration method for fine particulate matter sensor monitoring data according to claim 1 or 2, characterized in that: The qualitative representation and reasoning module converts the distance from the monitoring point to the pollution source into a qualitative characterization using the following method: Define monitoring point p i Distance d to the pollution source i For p i Let D = {close, medium, far} be a finite set of terms representing distance relationships, and let the function f be the Euclidean distance to the nearest point from the pollution source. Defined as a mapping function that maps metric distance to qualitative vocabulary: Where thresholds θ1 and θ2 are the boundaries for word segmentation, 0 < θ1 < θ2; qualitative descriptions using words closs, medium, and far follow the rules below: The terms close, medium, and far that satisfy conditions (3) and (4) are called distance primitives.

4. The calibration method for fine particulate matter sensor monitoring data according to claim 3, characterized in that: The qualitative representation and reasoning module performs qualitative spatial reasoning as follows: Let C = {low, moderate, high} be a finite set of terms qualitatively describing particulate matter concentration. The function q: D → C is defined as a function that maps distance primitives to qualitative terms: When using the words low, moderate, and high for qualitative description, follow these rules: The words low, moderate, and high are called concentration units when they simultaneously satisfy conditions (5) and (6); Qualitative inferences about particulate matter concentration can be made from the composite relationship between functions f and q: in, The distance from the monitoring point to the pollution source is represented by the symbol °, which indicates the composite operation of the function. f(d) i )and The input prompt word module generates the following prompt word: [Expected particulate matter concentration level]: Due to the distance between the monitoring point and the main pollution source... <f(d i The expected particulate matter concentration is...

5. The calibration method for fine particulate matter sensor monitoring data according to claim 4, characterized in that: The calibration value generation module is a backpropagation network, which includes a first fully connected layer, a first ReLU function layer, a first Dropout layer, a second fully connected layer, a second ReLU function layer, a second Dropout layer, and a third fully connected layer connected in sequence. The number of neurons in the second fully connected layer is only half that of the first fully connected layer, and the number of neurons in the third fully connected layer is one.

6. The calibration method for fine particulate matter sensor monitoring data according to claim 1, characterized in that: The calibration model was trained and tested using the following method: The sensor and reference analyzer are arranged together at m different locations P = {p1, p2, ..., p...} m Monitor particulate matter concentration; The sensor at each monitoring point p i The observed fine particulate matter concentration values ​​∈P constitute a time series X i : Where T represents the time domain, i∈[1,m], and the time series obtained by the sensor at each monitoring point constitute the time series set X={X1,X2,…,X…} m }, X i ∈x; Set the reference analyzer at monitoring point p i The observed fine particulate matter concentration values ​​∈P constitute a time series Y i : Where T represents the time domain, i∈[1,m], and the time series obtained by the reference analyzer at each monitoring point constitute the time series set Y={Y1,Y2,…,Y…} m }, Y i ∈Y; Suppose that the sensor needs to be monitored at point p. k The measured values ​​were calibrated, and X was... k As test set D te The training set D is constructed using data from the remaining monitoring points. tr D tr ={(C i (t),Y i (t))|i∈[1,m],i≠k,t∈T i }, where T i ∈T indicates that at monitoring point p i The set of monitoring time points; Using training set D tr The calibration model is trained, and during the training process, data is transferred from the training set D. tr The calibration model was validated using a subset of data, with test set D. te The qualified calibration model is tested to obtain a qualified calibration model.

7. The calibration method for fine particulate matter sensor monitoring data according to claim 1, characterized in that: When standardizing the number of data points in each input sample to the number of samples taken by the sliding window, if the number of predicted values ​​in the input sample is insufficient, the arithmetic mean of the predicted values ​​in the input sample is used to supplement them.

8. The calibration method for fine particulate matter sensor monitoring data according to claim 1, characterized in that: The sensor is a PM 2.5 Concentration sensor.